The Churn Ceiling
100 TB is 23 hours of bandwidth and six weeks of calendar — and the difference is not your network
Published: 2026-10-05 | jslet Research | 15 min read
📑 In This Briefing
A hundred terabytes over a 10 Gbps link is 23.3 hours of arithmetic. Multiply out the bytes, divide by the line rate, and that is the answer — assuming the link runs flat out, twenty-four hours a day, with zero protocol overhead and a source that has stopped changing.
Nobody who has run a real migration believes that number, and the reason is rarely that the network is slow. A migration is not a copy operation. It is a race between a transfer and a workload, and the workload does not pause because you have decided to move its data.
Two teams can start the same 100 TB on the same link. One finishes in three weeks. The other never finishes at all — not late, not over budget, but genuinely never reaching the point where the copy and the original agree. What separates them is a single ratio: how much data the source creates each day, divided by how much the link can move each day. Below one, the migration is finite and its calendar is predictable. At or above one, the catch-up rounds grow instead of shrinking, and no amount of retrying converges.
This is the written companion to the Data Migration Time calculator, which models the three phases and the transfer window directly. What follows is the part a phase model tends to hide: why the window is a squared penalty rather than a linear one, why the incremental rounds form a geometric series, and the five numbers that tell you in advance whether your migration has an end date.
The Bandwidth Answer
Start with the number everyone computes first. Bytes divided by line rate, no adjustments:
| Link | Raw transfer, 24/7, zero overhead | Same 100 TB, as a working estimate |
|---|---|---|
| 10 Gbps | 23.3 hours | 3–6 weeks |
| 1 Gbps | 9.7 days | 7–14 weeks |
| 100 Mbps | 97.1 days | never, at 2% daily change |
| 50 Mbps | 194.2 days | never, at 1% daily change |
The right column is not a fudge factor. Each entry comes from a specific mechanism, and the mechanisms compound rather than add. The gap between 23.3 hours and three weeks is worth decomposing, because only one of the four pieces is about bandwidth.
Three Subtractions and a Ratio
Three adjustments sit between the raw number and a realistic one, and they behave very differently.
1. Protocol overhead
Moving bytes over a network costs framing, headers, acknowledgements and per-object round trips. The penalty depends entirely on the transfer method, and the spread between methods is a factor of eight:
| Transfer method | Overhead | Effective rate on a 1 Gbps link |
|---|---|---|
| Raw TCP stream | ~3% | 970 Mbps |
| Custom UDP / Aspera / GridFTP | ~5% | 950 Mbps |
| HTTP / HTTPS | ~12% | 880 Mbps |
| Object-store multipart (S3-style) | ~18% | 820 Mbps |
| rsync / SCP / SFTP | ~25% | 750 Mbps |
The uncomfortable part is that the method most teams reach for by default — rsync, because it is incremental and resumable — is the most expensive one on the list. It buys resumability and delta detection, and pays for them with per-file round trips that the raw-stream methods do not make. That trade is often correct, but it is a trade, not a free upgrade.
2. The transfer window
Production systems rarely hand over a link for twenty-four hours. A nightly maintenance window of eight hours is a generous one, and it removes two-thirds of the day before anything else is considered.
3. Workload change during the migration
The source keeps writing. The baseline pass takes days or weeks, and every byte written during that period has to be re-copied afterwards. That catch-up pass takes time too, during which more data is written, which needs another pass. This is the term that turns a finite job into an infinite one, and it is the only one of the three that can.
Stacked on the same 100 TB over 1 Gbps:
| Adjustment applied | Timeline | Multiplier |
|---|---|---|
| Raw transfer, 24/7, no overhead | 9.7 days | 1.0× |
| + HTTP overhead (12%) | 11.0 days | 1.14× |
| + 8-hour nightly window | 33.1 days | 3.4× |
| + 1% daily change rate | 49.5 days | 5.1× |
Protocol overhead is the smallest line in the table, and it is the one teams spend the most time on. The window and the change rate are where the calendar actually goes — and the change rate decides whether there is a calendar at all.
The Churn Ceiling
Define the convergence ratio ρ as the data the source creates per day divided by the data the link can move per day:
Everything about the migration's schedule follows from that one number. Below ρ = 1 the total time is finite and equals the baseline pass divided by (1 − ρ). At ρ = 1 it diverges. Above it, the catch-up passes lengthen with each round and the migration runs permanently behind.
The ratio is uncomfortable because it is not a property of the network, and it is not a property of the data volume either. It is a property of their relationship. A migration of 10 TB can be impossible while a migration of 1 PB succeeds, if the small one has a busy enough source and the large one is cold.
Here is where the ceiling sits for a 100 TB dataset, across a range of daily change rates. The right-hand columns give the minimum effective throughput required to stay below ρ = 1:
| Daily change rate | Data created per day | Min effective rate, 24h window | Min effective rate, 8h window |
|---|---|---|---|
| 0.1% | 102 GB | 9.7 Mbps | 29.1 Mbps |
| 0.5% | 512 GB | 48.5 Mbps | 146 Mbps |
| 1% | 1,024 GB | 97 Mbps | 291 Mbps |
| 2% | 2,048 GB | 194 Mbps | 583 Mbps |
| 5% | 5,120 GB | 486 Mbps | 1,456 Mbps |
Two migrations illustrate how quickly this stops being theoretical.
Case one. 100 TB on a 1 Gbps link, transfers restricted to eight hours a night, source regenerating 2% of its volume daily. Daily capacity is 3,094 GB; daily change is 2,048 GB. ρ is 0.66 — uncomfortable, but finite. The baseline pass alone takes 33 days, and the convergence series stretches it to 97.9 days against a raw-arithmetic estimate of 9.7. A tenfold miss, from two decisions that each looked reasonable.
Case two. The same 100 TB, the same eight-hour window, the same 2% change rate, but on a 100 Mbps link instead of 1 Gbps. Daily capacity drops to 309 GB. ρ is now 6.6. The migration moves data about six times slower than the workload creates it, which means the gap widens every day it runs. This project has no completion date, and it will not acquire one by being run for longer. Nothing about the transfer will fail, which is precisely the problem: there is no error to alert on, only a delta that never closes.
The Window Is Squared
Shrinking the transfer window is the most common way to reduce production impact, and it is usually justified as a linear cost: halve the hours available, double the calendar. On a cold dataset that is exactly right. On a live one it is badly wrong, because two things change at once.
Capacity falls in proportion to the window — fewer hours per day, less data moved per day. And the calendar time the transfer occupies rises in inverse proportion, so the number of days over which new data accumulates goes up as the window goes down. The numerator of ρ grows while its denominator shrinks. The ratio scales as the inverse square of the window.
| Nightly window | Capacity/day | Baseline pass | ρ | Total | vs 24h |
|---|---|---|---|---|---|
| 24 hours | 9,281 GB | 11.0 d | 0.110 | 12.4 d | 1.0× |
| 12 hours | 4,641 GB | 22.1 d | 0.221 | 28.3 d | 2.3× |
| 8 hours | 3,094 GB | 33.1 d | 0.331 | 49.5 d | 4.0× |
| 6 hours | 2,320 GB | 44.1 d | 0.441 | 79.0 d | 6.4× |
| 4 hours | 1,547 GB | 66.2 d | 0.662 | 196 d | 15.8× |
| 3 hours | 1,160 GB | 88.3 d | 0.883 | 752 d | 60.6× |
| 2 hours | 773 GB | 132 d | 1.324 | never | — |
Read the last two rows together. Moving a 100 TB migration from an eight-hour window to a four-hour one does not double the timeline, it multiplies it by four — 49.5 days becomes 196. Move it to three hours and the estimate is 752 days, which is a little over two years for a job whose raw transfer time is under ten days. Move it to two hours and the arithmetic stops returning a date.
The reason the 1 Gbps link, the 100 TB dataset and the 1% change rate in this table are unremarkable is the point. None of them is extreme. The result is extreme because the penalty is quadratic, and quadratic penalties are invisible on a napkin.
The same table read backwards is the cheapest optimisation in this briefing. On the 8-hour row, widening the window to 24 hours takes the estimate from 49.5 days to 12.4 — a fourfold improvement with no hardware purchased, no link upgraded, and no vendor called. It is a decision about who owns the bandwidth at 2 a.m., and it is worth more than any protocol tuning on the list.
Where the Rounds Never End
The "incremental catch-up" is not one pass. Each round copies what changed during the previous round, so the round lengths form a geometric series with ratio ρ:
That closed form is the whole story of the churn term. It is mild for small ρ and violent as ρ approaches 1:
| ρ | Multiplier on the baseline pass | What it looks like in practice |
|---|---|---|
| 0.10 | 1.11× | A rounding error on the schedule |
| 0.30 | 1.43× | Six weeks becomes nine |
| 0.50 | 2.00× | The catch-up pass is as long as the baseline |
| 0.70 | 3.33× | Every round is visibly worse than the last |
| 0.90 | 10.0× | The baseline is 10% of the project |
| 0.99 | 100× | Theoretically finite, practically not |
This is also where the conventional three-phase plan quietly fails. The standard migration playbook stops after two incremental rounds — a baseline, one catch-up, one final delta — and declares the cutover. That truncation is harmless while ρ is small and increasingly dishonest as it grows:
| ρ | After two incremental rounds | True total | Understated by |
|---|---|---|---|
| 0.331 | 1.44× T₁ | 1.49× T₁ | 3.6% |
| 0.662 | 2.10× T₁ | 2.96× T₁ | 29% |
| 0.883 | 2.66× T₁ | 8.55× T₁ | 69% |
At ρ = 0.33, a three-phase plan is within 4% of the truth and there is no reason to complicate it. At 0.66 it understates the job by nearly a third. At 0.88 it reports 2.7 baseline passes for a job that will take 8.6, and the third incremental round is not the last one.
The practical consequence is that you stop planning for rounds and start planning for a threshold. Rather than asking "how many catch-up passes will it take", set a convergence target — a delta small enough to cut over on, say 0.1% of the total volume — and compute when the series reaches it. Waiting for a delta of exactly zero is not an engineering plan; it is a definition under which the migration never ends.
The Levers, Ranked
ρ has four inputs, and every real intervention moves exactly one of them. Ranking them by how much of the calendar they buy back, on the same 100 TB with a 1% daily change rate:
| Change | Which lever | ρ | Total | Saving |
|---|---|---|---|---|
| Baseline: 1 Gbps · 8h · HTTP | — | 0.331 | 49.5 d | — |
| Switch HTTP to raw TCP | overhead −9 pts | 0.300 | 42.9 d | 13% |
| Widen the window to 24 hours | window ×3 | 0.110 | 12.4 d | 75% |
| Run 4 parallel streams | concurrency ×4 | 0.083 | 9.0 d | 82% |
| Upgrade the link to 10 Gbps | bandwidth ×10 | 0.033 | 3.4 d | 93% |
| 4 streams and a 24-hour window | both | 0.028 | 2.8 d | 94% |
Protocol optimisation is the smallest lever on the list and the most popular one, because it is the only one that can be done entirely inside engineering. A 9-point overhead reduction buys 13% of the calendar. A schedule conversation buys 75%.
Why one stream cannot use one link
Parallelism is not a trick to work around a slow network; it is required by how TCP works. A single stream's throughput is bounded by the receive window divided by the round-trip time — the bandwidth-delay product. A 1 MB window on a 100 ms cross-region path caps one stream at 10 MB/s, which is about 80 Mbps on a link rated at ten thousand. The link is idle; the window is the constraint.
Loss makes the bound worse than linear. A standard approximation puts a single stream near 1.22 × MSS / (RTT × √p): at 100 ms of latency, 1,460-byte segments and 0.1% packet loss, that is roughly 4.5 Mbps. One lost segment in a thousand costs two orders of magnitude on a long-haul path. Chasing that with a bigger window is a tuning project; opening eight streams is an afternoon.
The lever that is not on the network
The largest available lever moves no bytes at all. A hundred 20 TB drives hold 2 PB, and a 24-hour drive moves them — which works out to roughly 185 Gbps of equivalent bandwidth, against 185 days for the same data on a 1 Gbps line. This is not an analogy; it is the arithmetic behind seeding phases and offline transfer appliances at every large cloud provider. The trade is latency: the van has a day of it where the network has milliseconds, which makes physical transfer a bulk-pass tool only. The incremental catch-up still runs over the wire, but by then ρ is being compared against a baseline that is already done.
The Metadata Tax
Everything above counts bytes. Storage systems also count objects, and a migration's cost per object is paid in round trips that have nothing to do with payload size. Two datasets with wildly different byte counts can invert the picture entirely:
| Dataset | Objects | Payload | Metadata at 5 ms/object | Payload at 1 Gbps |
|---|---|---|---|---|
| 4 KB files | 100 M | 381 GB | 5.79 days | 59 minutes |
| 4 MB files | 10 M | 38.1 TB | 0.58 days | 4.21 days |
The first row is the one that ruins schedules. A hundred million small objects hold less than half a terabyte, which is under an hour of transfer on a gigabit link — and about six days of open, stat and close operations. The metadata is 141 times the payload. No bandwidth upgrade touches this, because the link is not the resource being consumed.
Object stores add a second ceiling of the same kind. Request rates are provisioned per partition rather than per byte, and the widely documented working figure is on the order of 3,500 write requests per second per partitioned prefix. A hundred million objects is therefore at least eight hours of requests even in the ideal case, before any data is considered — and that ideal case assumes the prefixes are spread well enough to avoid a single hot partition, which is an application-side property, not a service-side one.
The practical test is one question asked early: what is the average object size? Above a megabyte or so, the migration is a bandwidth problem and the ratios above apply directly. Below it, the migration is an object-count problem, and the fix is compaction or batching on the source side — fewer, larger objects — rather than anything in the network path.
A Five-Number Plan
Before committing to a window, a link or a date, collect five numbers. Four of them are arithmetic; the fifth decides whether the arithmetic applies.
- Bytes to move (D). The baseline volume, measured rather than estimated. Include the data you are not planning to migrate but will end up moving anyway.
- Bytes created per day. Measure the source's write rate over a full business cycle — a week, not an hour. This is the number teams guess, and it is the one that decides everything.
- Effective throughput (C). Link rate × (1 − protocol overhead) × parallel streams. Use the overhead table above; do not use the link's headline rate.
- Window hours (w). The hours per day you are actually permitted to saturate the link, not the hours the link exists.
- Object count and average size. If the average is under a megabyte, re-derive C as requests per second and take the smaller of the two ceilings.
Then compute ρ = (D × c) ÷ C and act on the result:
- ρ ≥ 1. The migration has no end date. Adding retries, resuming better or upgrading the tooling changes nothing. Fix one of the four inputs first: more capacity, a wider window, fewer parallel sources of change, or a seeding pass that removes the bulk from the race entirely.
- 0.5 ≤ ρ < 1. Finite, but not in the time the phase plan says. Use T₁ / (1 − ρ), not the two-round model, and treat the third incremental round as unlikely to be the last.
- ρ < 0.3. A three-phase plan is accurate to within a few percent. Plan normally and spend your effort on the cutover instead.
Steps one through three are what the Data Migration Time calculator computes directly, including the window and the three phases. Step two — the change rate — is the one no calculator can measure for you, and it is the only input that can turn a three-week project into one that never finishes.
Frequently Asked Questions
How long does it take to migrate 100 TB?
Raw arithmetic gives 23.3 hours on a 10 Gbps link and 9.7 days on 1 Gbps, both assuming 24/7 transfer and zero overhead. HTTP overhead pushes the 1 Gbps case to 11.0 days; restricting transfers to an eight-hour nightly window pushes it to 33.1 days; and a source generating 1% of its own volume per day pushes it to 49.5 days. Each layer answers a different question — bandwidth, protocol, calendar, workload — and the last one can make the answer infinite.
Why does shortening the transfer window cost more than the hours it removes?
Because two things change at once. Daily capacity falls in proportion to the window, and the calendar time the transfer occupies rises in inverse proportion, so the amount of new data created during the migration grows as the capacity to catch up with it shrinks. The convergence ratio scales as the inverse square of the window. On 100 TB over 1 Gbps with 1% daily change, cutting the window from 24 hours to 8 multiplies the timeline by 4.0 rather than 3, and cutting it to 4 hours multiplies it by 15.8.
Can a data migration fail to converge?
Yes, and it is arithmetic rather than bad luck. If the source produces more data per day than the link moves per day, every round is longer than the one before it. 100 TB at a 2% daily change rate produces 2,048 GB per day; a 100 Mbps link in an eight-hour nightly window moves 309 GB per day. The second number is smaller than the first, so the migration runs at roughly one-sixth of the rate it needs, indefinitely.
Does adding bandwidth always shorten a migration?
Not until ρ is comfortably below 1, and not at all when the bottleneck is object count rather than bytes. A single TCP stream is limited by the receive window divided by the round-trip time, so a 1 MB window on a 100 ms path caps one stream near 80 Mbps whatever the link is rated, and 0.1% packet loss cuts that to single-digit Mbps. Datasets dominated by small files are usually bound by per-object metadata instead: 100 million 4 KB files carry 381 GB of payload but cost about 5.8 days of metadata handling against 59 minutes of transfer.
Methodology & Disclosure
Durations are computed as T = T₁ / (1 − ρ), where T₁ = D / C is the baseline pass, ρ = (D × c) / C is the convergence ratio, C = B × (1 − o) × n × w × 3600 / 8192 is daily capacity in GB, B is link rate in Mbps, o is protocol overhead, n is parallel streams and w is the daily window in hours. Protocol overheads are the calculator's fixed coefficients — 3% raw TCP, 5% custom UDP/Aspera/GridFTP, 12% HTTP/HTTPS, 18% object-store multipart, 25% rsync/SCP/SFTP — and reflect transport framing and per-object round trips, not congestion. Storage volumes use 1 TB = 1,024 GB to match the paired calculator; throughput uses 1 MB = 1,024 KB.
The change rate c is assumed constant and linear, which overstates the total for workloads with a diurnal peak and understates it for workloads that grow. Capacity assumes the link is fully available for every hour of the window and that disk I/O is never the bottleneck; real migrations hit source read limits before they hit link limits, so these figures are floors. The metadata estimate uses 5 ms per object for open, stat and close round trips, which is conservative for local filesystems and optimistic for high-latency object stores. The 3,500 requests-per-second prefix figure is the documented working order of magnitude for partitioned object-store prefixes and assumes well-distributed key prefixes. Ceilings from TCP windowing and loss use the standard bandwidth-delay product and Mathis-style approximations and describe single streams, not aggregate capacity. The three-phase truncation table compares the calculator's phase model against the closed-form series. This site takes no sponsorship, no affiliate commission and no vendor payment of any kind; the vendors named here have no relationship with jslet and did not review this briefing.
References & Further Reading
- Tanenbaum, A. S. Computer Networks — the observation that a station wagon full of tapes has more bandwidth than a data link, which is a load-bearing assumption in every large migration plan rather than a joke. Pearson
- Jacobson, V., Braden, R., Borman, D. TCP Extensions for High Performance (RFC 1323) and its successor RFC 7323 — window scaling and the bandwidth-delay product that caps a single stream on a long path. IETF
- Mathis, M., Semke, J., Mahdavi, J., Ott, T. The Macroscopic Behavior of the TCP Congestion Avoidance Algorithm — the throughput bound under loss used for the single-stream figures. ACM SIGCOMM Computer Communication Review
- Amazon Web Services. Amazon S3 request rate and performance guidelines — the per-prefix request rates that set an object-count ceiling separate from bandwidth. docs.aws.amazon.com
- AWS DataSync, rclone and AzCopy documentation — multi-threaded transfer defaults, verification behaviour and resumability trade-offs. aws.amazon.com · rclone.org · learn.microsoft.com
- Related: Data Migration Time Calculator · Gbps to TB per Day · Cloud Storage Cost · Log Storage & Retention · The Tail Latency Trap · Network & Bandwidth Economics
📜 Copyright & Attribution
© 2026 jslet Research. This article is an original work published on jslet (jslet.com). All rights reserved.
Sharing & Reprinting: You may share excerpts (up to 200 words) with a mandatory, do-follow link back to the original article URL. Full reproduction, translation, or adaptation requires prior written permission. Contact: support@jslet.com.
Estimation Disclaimer: Transfer durations are model-based estimates from link rate, protocol overhead, transfer window and change rate, not measurements of any specific migration. Real cutovers vary with source I/O, network contention, failure retries and validation overhead, and are typically longer than the figures here. Validate against a pilot transfer of a representative subset before committing to a date.