Discuss a project

Issue 05Data infrastructure

Three substitutions

How a backup system is actually sized: how the protected perimeter differs from storage capacity, why the vendor's stated data reduction ratio doesn't hold on an encrypted stream, and what to ask a vendor before the budget is approved.

Over the past year we’ve reviewed other people’s backup system sizing more often than we’ve built our own. And in almost every one we see the same thing: a budget calculated in spring has, by the time of deployment, grown several times over. The customer feels cheated. The vendor produces correspondence and documents where everything adds up. And both are right, because nobody actually had to lie.

The reason is that a backup sizing calculation has three points where one quantity is quietly swapped for another. The volume of protected data gets swapped for the capacity of the backup storage. Storage capacity gets swapped for the physical volume of disks. And physical volume gets swapped for a data reduction ratio that may have nothing to do with your data at all. Each substitution looks trivial on its own. Together they produce an order-of-magnitude discrepancy.

What follows is a breakdown of all three — with wording for the statement of work, and questions worth asking before the budget is signed off.

The license metric

Perimeter, not storage

A license is sized against the volume of protected sources. The capacity of the backup storage plays no part in that calculation at all.

Let’s start with the quantity that sets the license price.

A license is sized against the volume of the protected sources. Not against the volume of the copies the system will produce. Not against the capacity of the storage those copies land on. Against how much data sits in the systems you’re bringing under protection. That sounds obvious right up until you have to name a number.

Because even for a single virtual machine there are three such numbers, and all three are real. How much disk is allocated. How much of it is used. And how much of the used space falls inside the perimeter — which depends on whether you’re capturing system partitions, swap, temp directories. The gap between the first number and the third can run as high as two to one.

The budget is usually sized against the first number — it’s the easiest one to pull from the console: open it, export it, add it up. From there, two outcomes. Sized against allocated space — you overpaid, and you’ll find out a year later at renewal. Sized against used space but forgot half the systems — you underpaid, and that surfaces mid-deployment, when it’s too late to change course. The second case is worse and happens more often.

And there’s a second fork that gets confused just as often. The metric itself comes in different flavors, and a single spec typically has two of them.

A virtual machine is licensed by machine count — usually in packs of ten — and data volume plays no role: a machine is a machine, whether empty or packed to the brim. Front-end terabytes are the opposite — licensed by volume, and the number of objects doesn’t count.

The trap is that the same machine can be counted both ways. If it’s running Oracle or PostgreSQL inside, image-based backup is pointless: you’ll get either an inconsistent copy or a database outage for the duration of the snapshot. The database gets captured by an agent — with logs, with point-in-time recovery. And that’s front-end terabytes now, not “one more machine in the pack.”

The consequence is twofold. The spec needs both line items, and it needs to be known in advance which systems fall into which bucket: the list of machines running databases is drawn up separately. And those machines can’t be counted twice — once in the pack, then again in terabytes. We see double-counting in roughly one out of every three calculations, and it’s always in the vendor’s favor.

One virtual machine, three legitimate answersWhich volume counts as protected: values are for a typical machine; the proportions are typical ofbanking and telecom infrastructure8.6 TBAllocated to theVMThe sum of virtual disksizes. One click in theconsole — which iswhy it most often endsup in the budget.5.4 TBUsed by dataWhat is actuallywritten inside theguest system. Usuallynoticeably less thanallocated.3.9 TBWithin theperimeterUsed space minuswhat is excluded frombackup: systempartitions, swap,temporary directories.Up to a twofolddifferenceAll three numbers are real.Only the third goes into thelicense.Count by the first — youoverpay. Count by the thirdbut forget half the systems —you underpay, and it surfacesduring implementation.One virtual machine, three legitimateanswersWhich volume counts as protected: values are for atypical machine; the proportions are typical of bankingand telecom infrastructureAllocated to the VMThe sum of virtual disk sizes. One click in the console — which iswhy it most often ends up in the budget.8.6 TBUsed by dataWhat is actually written inside the guest system. Usuallynoticeably less than allocated.5.4 TBWithin the perimeterUsed space minus what is excluded from backup: systempartitions, swap, temporary directories.3.9 TBUp to a twofold differenceAll three numbers are real. Only the third goes intothe license.Count by the first — you overpay. Count by thethird but forget half the systems — you underpay,and it surfaces during implementation.
Fig. 1. Three legitimate answers to the question of how much data a single virtual machine holds

A separate point about backup storage — this is where the confusion gets the most expensive. Its capacity has no part in the license calculation at all: it’s driven by retention depth, the backup scheme, and immutability requirements. A system sized at several petabytes can legitimately be built for a perimeter of a few hundred terabytes. When we see a license rate multiplied by array capacity in a sizing sheet, we don’t need to look any further.

Anatomy of the estimate

How the perimeter grows in layers

The perimeter is built up in layers, and each next one gets remembered later than the one before it. The first estimate almost always accounts for one layer out of five.

Now, why the initial estimate is almost always too low. The perimeter builds up in layers, and each next layer gets remembered later than the previous one.

The first thing to land in the calculation is whatever’s visible in the virtualization console: the farm is the best-inventoried thing there is, and the number is one click away.

The second layer to arrive is databases. Some live on those same machines and are already counted; some run on physical servers and aren’t counted anywhere. And it almost always turns out that backing up a database via the hypervisor versus via an agent means a different volume, a different recovery time, and different line items in the spec.

Third comes file shares: shared folders, scan archives, exchange directories with external systems. They’re never counted upfront — there’s no single owner responsible for them.

Fourth — dev and test environments. The first instinct is: not production, exclude it. But inside sits a copy of the production database with real data, and you’ll still have to restore it after a failure.

Fifth comes mail that’s moved to the cloud, and here the metric changes completely: the license is sized by mailbox count, not by volume. The familiar “terabytes times rate” arithmetic doesn’t apply at all. Worse, mailbox count follows different laws: a company can go years without accumulating much data while still hiring people — and the license grows right along with headcount.

Ask how much space mail takes up, and you’re told how much storage the current backups occupy. That’s not a sizing metric.

That’s also where our favorite trap lives. Ask how much space mail takes up, and you get the storage occupied by the current backups. That’s not a sizing metric — it’s how much space the copies take at today’s retention depth. It has nothing to do with the mailbox count that the license actually runs on.

The perimeter builds up in layersEach next layer is remembered later than the previous one. Bar width is the conventional share of thelayer in the total volume of protected data1Virtual farmThe best inventoried. The numbercomes from the console — and thatis usually where the calculation stops.2DatabasesSome sit on the same machines andare counted; some are on physicalservers — counted nowhere.3File sharesShared folders, scan archives,exchange directories. No singleowner is responsible for them.4Development environments“It’s not production, exclude it” — yetinside is a copy of the live databasewith real data.5Cloud mailThe metric changes: the license iscounted by the number ofmailboxes, not by volume.The order in which the layers are remembered is the reverse of the order in which they should have been countedThe perimeter builds up in layersEach next layer is remembered later than the previousone. Bar width is the conventional share of the layer in thetotal volume of protected data1Virtual farmThe best inventoried. The number comes from theconsole — and that is usually where the calculationstops.2DatabasesSome sit on the same machines and are counted; someare on physical servers — counted nowhere.3File sharesShared folders, scan archives, exchange directories. Nosingle owner is responsible for them.4Development environments“It’s not production, exclude it” — yet inside is a copy ofthe live database with real data.5Cloud mailThe metric changes: the license is counted by thenumber of mailboxes, not by volume.The order in which the layers are remembered is the reverse ofthe order in which they should have been counted
Fig. 2. The layers that make up the protected perimeter, in the order they surface in a sizing exercise

In one project where we managed a backup system expansion, the final perimeter came out roughly five times larger than the initial estimate. The data hadn’t grown fivefold in that time. The initial estimate simply accounted for one layer out of five.

The reverse happens too: in another project, the perimeter that came out of an audit was slightly smaller than the volume the license had been reserved for. The gap was small but telling — it came from inventory, not from gut feel. A proper calculation protects against overpaying exactly as much as it protects against falling short.

The list of omissions

What always gets forgotten

The list of what surfaces last is boringly consistent. Inventorying sources is a project stage in its own right, not preparation for one.

The list of what surfaces last is boringly consistent. Check your own sizing against it.

Databases outside the virtual environment. Enterprise DBMSs on classic Unix platforms and on appliances. They’re invisible in the virtualization console — they never make it into the first calculation. And yet it’s usually precisely for their sake that the whole backup system was commissioned in the first place.

A copy at the second site. The system is being backed up at the primary center, and it’s counted. Its copy at the secondary site is potentially a separate line item in the license. Rules differ between vendors, and even between delivery options from the same vendor — so the question is asked directly, and the answer is obtained in writing.

Additional modules. Isolated recovery environments, malware scanning of backups, immutable storage — these are standalone line items with their own bundling rules. They’re not part of the base license, and you can’t buy them in later under the same terms.

Supporting infrastructure. Backing up cloud mail requires separate intermediaries — virtual machines that the export runs through. They’re not in the license quote. They are in the project, they consume resources, and someone has to count them.

Unused license balance. The most common source of unpleasant surprises. There are licenses sitting on the books, bought a few years ago and, as the story goes, not fully used up. They get planned in as a ready resource. Until it’s reconciled against the actual list of systems, it isn’t a resource — it’s an assumption, and it needs reconciling before the calculation, not after the contract is signed.

Where a source gets counted and where it gets forgottenThe first six rows are the protected perimeter. Below them is what almost never makes it into the initialestimateIn the firstestimateSeparate licenseitemInfrastructure, notlicenseVirtual machines at the primary siteDatabases on physical serversDatabases on appliancesFile shares and scan archivesDevelopment and test environmentsMail in a cloud serviceCopy of systems at the secondary siteIsolated recovery environmentScanning backups for malicious activityImmutable storageProxies for exporting cloud mailManagement servers and media agentsYesDepends on the terms of supplyNoWhere a source gets counted and whereit gets forgottenThe first six rows are the protected perimeter. Below themis what almost never makes it into the initial estimateInfrastructure, not licenseSeparate license itemIn the first estimateVirtual machines at theprimary siteDatabases on physicalserversDatabases on appliancesFile shares and scan archivesDevelopment and testenvironmentsMail in a cloud serviceCopy of systems at thesecondary siteIsolated recoveryenvironmentScanning backups formalicious activityImmutable storageProxies for exporting cloudmailManagement servers andmedia agentsYesDepends on the terms of supplyNo
Fig. 3. Where a source gets counted and where it gets forgotten in the initial estimate

Which leads to the point we repeat at every meeting: inventorying sources is a project stage in its own right, not preparation for one. The inputs come from two places, and they diverge: the written list assembled for procurement, and the verbal account of the engineers who run this stuff every day. Disagreement between the two is the norm.

There’s exactly one tool that works — a “what exists, what doesn’t” reconciliation checklist: source category, presence, count, volume, protection method, owner. It’s filled in jointly, signed by both sides, and only then does the sizing calculation begin.

And one more thing, from the calendar side. Count backward not from the project’s readiness date, but from the expiration date of the current support contract. The procurement process takes months: if licenses expire in December and the tender is announced in September, that leaves August for gathering the input data — and that’s already too late.

The second substitution

The target isn’t the perimeter

Four independent storage systems are not one system with combined capacity. And they’re not the quantity the software is licensed against.

This is where the second part of the conversation begins, and the second substitution.

Take a typical target architecture: four independent storage systems, several petabytes each, spread across two sites, two of them running in immutable mode. The mix can be almost anything within reason: two flash systems plus two disk-based ones, two flash plus two tape libraries, ready-made vendor backup appliances, or self-built Ceph with object access on your own hardware. That’s a question of money and performance requirements, not dogma.

But the first thing that needs to be said out loud: this is four systems, not one system with combined capacity. Independent systems don’t sum for fault tolerance — the failure of one isn’t automatically offset by the other three — nor for software licensing. And the temptation to grab the biggest number in the project and multiply it by the rate hits literally everyone.

Immutable storage is regularly confused with redundancy. A copy at the second site protects against equipment failure and a site-wide disaster. Immutability protects against a compromised administrator account. Three copies across two centers won’t save you if an attacker has admin access to both.

And one more thing that specs almost never mention: the target isn’t usually one system — it’s several tiers, each with a different job.

On-site sits a fast storage tier for instant recovery — shallow retention, but the service comes back up within hours. Next to it, a high-capacity, slow tier that holds the main retention depth. A tape library for the archive — on the same site or a remote one. A fourth tier fits nicely as object storage in a rented cloud.

The last one removes the need to build a full storage set at the second site. Instead of a second array fleet, media servers are enough there — compute without data. In a disaster at the first site, they pull data from the cloud and spin services up locally. Slower than having a ready replica next door — but you’re not paying for a second disk fleet that sits doing nothing for years.

The target is not one system but several tiersTiers differ in recovery speed, retention depth and cost per terabyte. The license metric does notdepend on how many there areSite 1four backup storage tiersFast storage systemhours to daysInstant recovery. Short retention.Capacity storage systemweeks to monthsMain retention depth. Most of the backups live here.Tape librarymonths to yearsArchive. At the same site or at a remote one.S3 object storagedisasterRented cloud. Replaces a full second copy.Site 2Media serversCompute only. There isno full copy of the datahere.In a disaster at site 1, themedia servers pull the datafrom the cloud and bring theservices up locally.Cheaper than keeping asecond full storage set.restore from the cloudin a disasterImmutability is a property of a tier, not of the whole systemThe mode in which even an administrator cannot delete a backup is enabled on selected tiers. Itprotects not against hardware failure but against a compromised account — which is why a copy atthe second site does not replace it.The target is not one system but severaltiersTiers differ in recovery speed, retention depth and cost perterabyte. The license metric does not depend on howmany there areSite 1four backup storage tiersFast storage systemhours to daysInstant recovery. Short retention.Capacity storage systemweeks to monthsMain retention depth. Most of the backups livehere.Tape librarymonths to yearsArchive. At the same site or at a remote one.S3 object storagedisasterRented cloud. Replaces a full second copy.restore from the cloud in a disasterSite 2Media serversCompute only. There is no full copy of the datahere.In a disaster at site 1, the media servers pull the datafrom the cloud and bring the services up locally.Cheaper than keeping a second full storage set.Immutability is a property of a tier, not of thewhole systemThe mode in which even an administrator cannotdelete a backup is enabled on selected tiers. Itprotects not against hardware failure but against acompromised account — which is why a copy at thesecond site does not replace it.
Fig. 4. Backup storage tiers and the role of media servers at the second site

Immutability, in this scheme, is a property of one specific tier, not of the whole construction. It’s switched on where it’s needed, and it’s a separate line item in the license.

And the thing that’s remembered last of all. A backup strategy doesn’t exist separately from a disaster recovery strategy — it’s one single construction, and it can’t be designed in halves. Backups that get taken beautifully but don’t restore in time protect nothing but the reporting.

And there’s a third one, rarely spelled out at all: the recovery strategy. Not “how do we store it” or “where do we failover to,” but “how much downtime is the business willing to accept, and what can we bring back up within that time.” That’s exactly what sets the requirements for the tiers: what has to sit on fast storage, what can go to tape, and what moves to the cloud.

An anonymized example from practice. A project built entirely on Ceph and PostgreSQL. The customer’s architect calculated a full recovery at the second site in the economy option — it came out to twenty-three hours of downtime or more. Unacceptable for the business, and the question immediately shifted from “how much does storing it cost” to “how much does a day of downtime cost.” The good news is that this got calculated at the design stage, not after an outage.

Which leads to an unpleasant corollary that both sides of the deal tend to stay quiet about. A license buys the right to protect a given volume with a given feature set — but not backup speed and not recovery time. Those are set by the architecture: stream count, network channel, disk pool throughput, exactly where you’re restoring the data from. None of those quantities appears in the license at all. You can buy the entire perimeter at maximum feature tier and still end up with a recovery time that misses every target you have.

And last. The type of target affects the licensing metric. Moving from tape to disk or object storage isn’t just a hardware swap: an auxiliary copy to object storage and object-lock mode are counted under their own rules. A calculation built for the previous architecture can’t be carried over mechanically to the new one, even if the protected volume hasn’t changed by a single terabyte.

The third substitution

The physics of ciphertext

Strongly encrypted data doesn’t compress. Not “compresses worse” — doesn’t compress at all.

The third substitution is the most technical and the most expensive one. It takes one paragraph to explain, and months to argue about.

Strongly encrypted data doesn’t compress. Not “compresses worse” — doesn’t compress. The compression limit for ciphertext is one point something in the hundredths. The reason lies in the nature of encryption: a good algorithm outputs a sequence that’s statistically indistinguishable from random, and compression works precisely on statistical redundancy. There’s nothing to compress in ciphertext.

This isn’t our own observation — it’s the vendors’ own stated position. Commvault’s documentation has a direct example: a cartridge rated at 110 gigabytes that held around 190 gigabytes with compression and no encryption fits about 124 gigabytes once encryption is turned on. The ratio drops from 1.73 to 1.13 — same media, same drive, same compression setting.

What happens to capacity with encryption enabledAn example from the Commvault documentation: a cartridge rated at 110 GB took about 190 GB withcompression and no encryption, and about 124 GB with encryptionSame media. Same drive. Same compression setting. The onlydifference is whether encryption is enabled.110 GBNominal mediacapacityno compression190 GBCompression withoutencryption1.73 : 1124 GBCompression withencryption1.13 : 1Sizing ruleThe stream is encryptedby the backup software— the ratio in thecapacity calculationequals one.What happens to capacity withencryption enabledAn example from the Commvault documentation: acartridge rated at 110 GB took about 190 GB withcompression and no encryption, and about 124 GB withencryptionSame media. Same drive. Same compression setting. Theonly difference is whether encryption is enabled.Nominal media capacityno compression110 GBCompression without encryption1.73 : 1190 GBCompression with encryption1.13 : 1124 GBSizing ruleThe stream is encrypted by the backup software —the ratio in the capacity calculation equals one.
Fig. 5. Media capacity under compression, with and without encryption

In the same place, the vendor draws a second conclusion, an even more practical one: hardware compression is not recommended on encrypted data — the volume may not shrink at all, it may grow. That’s not a paradox: attempting to compress a random sequence adds the algorithm’s own overhead structures without removing anything. A separate line states plainly that software-side encryption negates hardware compression entirely.

Not 1.7. Not 2. Not 5. One.

Which gives us a sizing rule that’s almost indecently simple. If the stream is encrypted on the backup software side, the data reduction ratio used in the capacity calculation is set to one. Not 1.7. Not 2. Not 5. One.

Breaking down the number

What the stated ratio is actually made of

The ratio in the slide deck is usually measured honestly. Just not on the stream you’ll actually have.

So where do the 1.7, 3, and 5 figures in the slide decks come from? The answer is unpleasant precisely because it’s so mundane: these numbers are usually not made up. They’re honestly measured — just not on what you assumed.

The stated ratio is almost always a composite number. It includes deduplication — eliminating repeated blocks, the biggest contributor on unencrypted data. Zero-block elimination — empty regions of allocated-but-unwritten disk space. Thin provisioning. And compression of the unencrypted parts of the stream: even when the data is encrypted, the service layer stays in the clear — metadata, job headers, catalogs — and that part compresses normally.

Add it all up on a test bench with typical data, and you’ll legitimately get one and a half to two. Carry it over to a stream encrypted on the backup software side, and you get one. Both numbers are real; the conditions are different.

What the stated data reduction ratio is made ofThe number in the slide deck is usually measured honestly — just not on the stream you will actuallyhaveDeduplicationEliminating repeated blocks across backups. The maincontributor — and only on unencrypted data.drops almost to one with aper-session keyZero-block eliminationEmpty areas of allocated but unfilled disks. Has nothing todo with the stream content.works regardless of encryptionThin provisioningSpace is reserved as data is written, not up front. Savings atthe volume level, not the data level.works regardless of encryptionCompression of the metadata partMetadata, job headers and catalogs stay in the clear evenwhen the stream is encrypted.contributes, but littleCompression of the data itselfCiphertext is statistically indistinguishable from random. Thereis nothing in it to compress.one point something in thehundredthsThe stated ratio is the sum of all rows, measured on typical data. On an encrypted stream, only what does notdepend on the content remains. The right question for the vendor: show the measurement methodologyseparately for each row.WorksPartlyDoes not work on ciphertextWhat the stated data reduction ratio ismade ofThe number in the slide deck is usually measured honestly— just not on the stream you will actually haveDeduplicationEliminating repeated blocks across backups. Themain contributor — and only on unencrypteddata.drops almost to one with a per-session keyZero-block eliminationEmpty areas of allocated but unfilled disks. Hasnothing to do with the stream content.works regardless of encryptionThin provisioningSpace is reserved as data is written, not up front.Savings at the volume level, not the data level.works regardless of encryptionCompression of the metadata partMetadata, job headers and catalogs stay in theclear even when the stream is encrypted.contributes, but littleCompression of the data itselfCiphertext is statistically indistinguishable fromrandom. There is nothing in it to compress.one point something in the hundredthsThe stated ratio is the sum of all rows, measured on typicaldata. On an encrypted stream, only what does notdepend on the content remains. The right question for thevendor: show the measurement methodology separatelyfor each row.WorksPartlyDoes not work on ciphertext
Fig. 6. The mechanisms that make up the stated data reduction ratio

There’s a second mechanism behind the discrepancy too. Deduplicating encrypted data is possible, but only with deterministic encryption: the same block, under a fixed key, always produces the same ciphertext. The moment a per-session key is used, identical blocks turn into different sequences, and deduplication collapses to nearly one. The ratio depends on an encryption setting that the slide deck never mentions a word about — and two customers with identical hardware end up with different results.

And a third case we ran into directly. Technologies for genuinely reducing host-encrypted data do exist: the array gets the key through a separate key management server, decrypts the stream, deduplicates it, and re-encrypts it with its own key. The numbers there are high and not fabricated. But a feature like that may live in only one product line, require a third-party component, and be discontinued — while the line you’re being offered has no equivalent at all. The figure is real. It’s just about a different product.

Which gives us the right way to challenge a vendor. Not “you’re inflating the ratio” — there’s always a legitimate comeback to that. But “show me the measurement methodology broken out separately: how much comes from compressing the encrypted stream itself, and how much from the other mechanisms.” They answer either with numbers or with silence, and both answers are equally useful to you.

While we’re at it, about capacity: in backup projects it has three meanings — physical, license-entitled, and usable, what’s left after the redundancy scheme. A model where the array ships fully populated but you pay for a smaller share of it is showing up more and more often. The first payment is lower — but you’re locked to that vendor for the entire growth cycle.

Checklist

What to ask before the budget is approved

Twelve questions and three pieces of wording for the statement of work.

Twelve questions. If every one of them has an answer in writing, the calculation can be defended.

Separately — two pieces of wording for the statement of work.

On capacity: usable capacity and recovery speed are to be stated as physical figures, measured on data that resists deduplication and compression; the stated data reduction ratio is to be accompanied by a description of the measurement methodology and a list of the mechanisms it accounts for.

On the form of the requirement: fix the logical volume the system is obligated to deliver, and don’t dictate the number and type of disks to the vendor. How they get there is their problem and their risk. And remember that a ratio figure in a contract, without a measurement methodology attached, means nothing: any dispute will run straight into the fact that the two sides measured different things.

Conclusion

In place of a conclusion

None of the three substitutions is fraud. Both sides cite real numbers from real documents — the numbers just refer to different quantities while going by the same name.

There’s one remedy, and it’s a boring one: for every number in the calculation, ask exactly what it measures and under what conditions it was obtained. Weeks spent on preparation save months at deployment.

If you’re sizing a backup system right now, or renewing licenses — send us your calculation, and we’ll go through it point by point against the checklist. The discrepancy is usually found in the first three.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.