Discuss a project

Issue 03Data infrastructure

Backup: counting the wrong things

Capacity, retention depth, backup window — and not a word about how long it takes to bring a service back. A breakdown of the forks where backup system sizing parts ways with reality.

Over the past year, more than a dozen technical specifications for backup systems have crossed our desk — banks, telecom, government agencies. Every one of them states the required capacity, retention depth, and backup window. A recovery time target by service class appeared in just two.

What happens next is always the same. The system goes live, the statuses are green, the reports get generated. Then someone asks to restore a one-terabyte virtual machine, and the restore takes nine hours. The question lands on our desk: how is that possible, when the array delivers twenty gigabytes per second and the network runs at twenty-five gigabits?

Our answer: because neither the array nor the network plays any part in this arithmetic. What matters is the speed of a single stream and what that stream is reading from. Everything else is a ceiling you never get close to.

Below are five forks where the sizing parts ways with reality. We go through them in the order they tend to surface in projects.

Nature of the data

What will actually compress

The data reduction ratio is determined not by the array but by what gets written to it. Deduplication wins on repetition, not on content.

The first place sizing breaks is the data reduction ratio. It comes from the array’s spec sheet, makes its way into the capacity calculation, and stays there until the pool fills up one and a half times faster than planned.

The reason is simple: the efficiency of deduplication and compression is determined not by the array but by what gets written to it. Backup software almost always compresses and deduplicates the stream on its own side — on the media server, before sending it. What arrives at the array is data that’s close to random. It can’t be compressed a second time, and that’s not a flaw in the array; it’s a property of the data.

What compresses in a backup repositoryDeduplication wins on repetition, not on content. Anything compressed or encrypted before it arrivesas a finished stream.10–20 : 1Repeated full copies of the samedataset4–8 : 1VM images of the same OS family3–6 : 1Uncompressed DBMS dumps andexports2–4 : 1File servers, documents, sharedfolders2–4 : 1Mail systems2–3.5 : 1Home directories and profiles1.1–1.4 : 1DBMS with page compression1.2–1.6 : 1Search engine indexes1.1–1.3 : 1Copies compressed by the agent atthe source1–1.1 : 1Backup software streamdeduplicated before writing1–1.05 : 1Scans, media, ready-made archives1 : 1Encryption inside the guest OS1 : 12 : 14 : 18 : 116 : 1the gain comesfrom repeatedblocksdepends on whatthe source hasalready donedata arrivesalready packedApproximate ranges; verified by measurement on the customer's data.What compresses in a backup repositoryDeduplication wins on repetition, not on content. Anythingcompressed or encrypted before it arrives as a finishedstream.Repeated full copies of the same dataset10–20 : 1VM images of the same OS family4–8 : 1Uncompressed DBMS dumps and exports3–6 : 1File servers, documents, shared folders2–4 : 1Mail systems2–4 : 1Home directories and profiles2–3.5 : 1DBMS with page compression1.1–1.4 : 1Search engine indexes1.2–1.6 : 1Copies compressed by the agent at the source1.1–1.3 : 1Backup software stream deduplicated before writing1–1.1 : 1Scans, media, ready-made archives1–1.05 : 1Encryption inside the guest OS1 : 11 : 12 : 14 : 18 : 116 : 1the gain comes from repeated blocksdepends on what the source has already donedata arrives already packedApproximate ranges; verified by measurement on thecustomer's data.
Fig. 1. Approximate data reduction ratios by type of protected data

Look at the top row. Deduplication really does deliver a multiple-fold gain — but on repetition, not on content. Twenty full copies of the same dataset shrink beautifully, because nineteen of them consist of blocks the system has already seen. That’s exactly why a ratio achieved on a deduplicated backup pool can’t be carried over to primary storage, and a primary storage ratio can’t be carried over to a backup repository. These are different numbers that come from data of a different nature.

What to do about it in practice. Calculate capacity from raw terabytes. Factor in reduction only where the backup system itself is the source of the data, and only based on measurements from a sample of the customer’s real data — a week of test backups tells you more than any presentation.

Capacity sizing

How many terabytes it really is

Logical volume isn’t “how much data we have.” It’s a full copy, plus growth multiplied by retention depth, plus every long-term retention tier.

The second fork is what the volume is actually made of. Logical storage volume isn’t “how much data we have.” It’s a full copy, plus daily growth multiplied by the depth of daily restore points, plus each long-term retention tier as a separate term.

And daily growth can’t be taken as a typical figure. Across the systems we see, it ranges from one percent to fifteen. The difference between three and eight percent on three hundred terabytes of protected data is one and a half terabytes a day, forty-five a month. At these volumes, a wrong assumption costs more than a wrong choice of vendor.

250 TB of logical volume: how the ratio drives the purchaseThe gap between an honest calculation and a plugged-in “three to one” is a threefold difference inthe capacity you buy.0100 TB200 TB300 TB62784 : 1831043 : 11251562 : 11672081.5 : 12503121 : 1data reduction ratio assumed in the calculationphysical capacitythe same plus 25% fill headroomwhat logical volume is made ofFull copyvolume of protected dataDaily restore pointsdaily growth × depthWeekly setsas a separate termMonthly setsas a separate termYearly setsas a separate termImmutable tierseparately, no shared dedupreferences250 TB of logical volume: how the ratiodrives the purchaseThe gap between an honest calculation and a plugged-in“three to one” is a threefold difference in the capacity youbuy.0100 TB200 TB300 TB62784 : 1831043 : 11251562 : 11672081.5 : 12503121 : 1data reduction ratio assumed in the calculationphysical capacitythe same plus 25% fill headroomwhat logical volume is made ofFull copyvolume of protected dataDaily restore pointsdaily growth × depthWeekly setsas a separate termMonthly setsas a separate termYearly setsas a separate termImmutable tierseparately, no shared dedup references
Fig. 2. Required repository capacity under different reduction assumptions

The chart shows what’s worth keeping in front of you when reading any commercial proposal. Between an honest calculation and a “let’s assume three to one” calculation lies a threefold difference in the capacity you buy. That works in the seller’s favor exactly once — at signing. After that, it’s no longer the seller who benefits.

A wrong assumption about daily growth costs more than a wrong choice of vendor.

The immutable tier is calculated as a separate line — more on that below, in the section on the vanished physical gap.

Recovery arithmetic

Why a restore runs in a single stream

Recovery time equals volume divided by the speed of the narrowest sequential segment of the path. Everything else is a ceiling.

Now for the arithmetic itself. Recovery time equals volume divided by the speed of the narrowest sequential segment of the data path. The array’s aggregate performance, the link bandwidth, and the number of cores on the media server don’t enter the formula.

The path looks like this: reading blocks from deduplicated storage, rehydration, decompression, transfer, writing to the target storage. The bottleneck is almost always the first link, because reading from a dedup pool isn’t a sequential read — it’s fetching blocks from all over the pool. On spinning disks, that’s tens of megabytes per second.

This is also where a misconception lives that’s worth spelling out. Backup does run in parallel: the software spreads streams across machines and disks, and adding more readers really does widen the window. On restore, parallelism is limited by the structure of the object. A single virtual disk isn’t split across streams — one disk, one stream, no matter how many readers you add. And for certain combinations of virtualization platform and agent, the limit is even stricter: a full restore of a single machine runs in one stream even when it has several disks, and no setting overrides that.

The arithmetic turns out unpleasant. At around thirty megabytes per second, a full restore stops being a usable mechanism already at a few terabytes — not because it doesn’t work, but because the result arrives later than any reasonable target. Ten terabytes is more than three days. By then, the question “when will we be back up” has stopped being a technical one.

The design takeaway: when setting a target, what you need to know isn’t how many streams the product supports in general, but how many streams it will allocate to restoring one specific object in your particular combination of versions. That’s a question worth putting to the vendor in writing, with the answer filed with the project.

Recovery time: volume divided by single-stream speedThe array's aggregate performance and the link bandwidth don't enter this formula. Both scales arelogarithmic.1 h4 h12 h1 day3 daysweek4-hour targetone daythree days0.5 TB1 TB2 TB5 TB10 TB20 TB35 MB/s120 MB/s350 MB/s800 MB/svolume of data being restored35 MB/s — single stream from a disk-based dedup pool120 MB/s — single stream from a flash dedup pool350 MB/s — non-deduplicated copy, flash800 MB/s — multi-stream read, flashRecovery time: volume divided bysingle-stream speedThe array's aggregate performance and the linkbandwidth don't enter this formula. Both scales arelogarithmic.1 h4 h12 h1 day3 daysweek4-hour targetone daythree days0.5 TB1 TB2 TB5 TB10 TB20 TB35 MB/s120 MB/s350 MB/s800 MB/svolume of data being restored35 MB/s — single stream from a disk-based deduppool120 MB/s — single stream from a flash dedup pool350 MB/s — non-deduplicated copy, flash800 MB/s — multi-stream read, flash
Fig. 3. Recovery time depending on volume and single-stream speed

Recovery tiers

There’s no single mechanism

A full restore is slow by nature. There’s no point speeding it up — it simply shouldn’t be used where a fast return is needed.

Since a full restore is slow by nature, there’s no point trying to speed it up. The right move is not to use it for anything that has to come back fast.

A mature system is three or four mechanisms with different times and different costs. Restoring from a hardware snapshot on the array doesn’t read the repository at all and gets fresh restore points back in minutes. Instant recovery brings the machine up immediately by presenting its disks straight from the backup, while the actual data transfer runs in the background. A non-deduplicated copy is read linearly and predictably. A full restore from the main pool remains the standard option for everything that doesn’t have a hard target.

The cost of these tiers rises from bottom to top, and this is where projects usually swing to one of two extremes. Either a fast tier isn’t planned at all — and then the four-hour target exists only on paper. Or it’s planned for the entire fleet, and the budget doubles for the sake of machines nobody will ever restore in emergency mode. The right size of the fast tier is determined not by a share of the total volume but by the combined data of the services that genuinely must be back within a business day. In our experience, that’s usually ten to fifteen percent of the protected volume — and this figure shouldn’t be estimated, it should be listed service by service.

Which mechanism meets which targetFull restore stays in the system but stops being the only way to bring a service back.HardwarearraysnapshotInstantrecovery fromthe flash tierNon-deduplicatedcopyFull restorefrom thededup poolTape, coldarchiveTarget: up to an hourCore banking system,processingBilling, payment gatewaysProduction databases underloadTarget: two to four hoursDocument management,CRM, portalsReporting and data martsDepartmental file servicesTarget: a day or moreInfrastructure and auxiliaryVMsTest and pre-productionenvironmentsRegulatory and long-termretentionPrimary mechanismFallback optionDoes not meet the targetWhich mechanism meets which targetFull restore stays in the system but stops being the onlyway to bring a service back.1Hardware array snapshot2Instant recovery from the flash tier3Non-deduplicated copy4Full restore from the dedup pool5Tape, cold archive12345Target: up to an hourCore banking system,processingBilling, payment gatewaysProduction databasesunder loadTarget: two to four hoursDocument management,CRM, portalsReporting and data martsDepartmental file servicesTarget: a day or moreInfrastructure and auxiliaryVMsTest and pre-productionenvironmentsRegulatory and long-termretentionPrimary mechanismFallback optionDoes not meet the target
Fig. 4. Mapping recovery mechanisms to targets by service class

Instant recovery has a catch that people step into regularly. Formally, the mechanism works from any media, but a machine launched from a slow pool starts fast and then runs unusably slow. Formally, the service is up; in practice, nobody can use it — and the report is all green. If you’re counting on instant recovery to meet a target, its tier has to be on flash; otherwise it’s not a solution but an imitation of one.

The service is up, nobody can use it, and the report is all green.

And most importantly: the “acceptable downtime” row in the table isn’t infrastructure’s to fill in. Infrastructure has no right to tell the business how long it can afford to be down. All it can do is say honestly what each option will cost.

Platform-level protection

What the platform can do on its own

Behind the words “everything’s replicated” hide four different mechanisms with different scopes and different costs.

On Nutanix — and in our projects, hyperconvergence almost always means Nutanix — half the mechanisms from the previous section are already built in, and that changes the conversation with the customer. Unfortunately, not always for the better: the phrase “everything’s replicated” regularly hides four completely different things with different scopes of application. Let’s go through them in order, from cheapest to most expensive.

A local snapshot on the same cluster. It’s taken instantly, rolls back state in minutes, and costs next to nothing in time. But it lives in the same container, on the same disks, under the same Prism Central. A cluster failure, the loss of a site, or a compromised administrator account wipes out both production and the restore points in one move. This is a rollback mechanism, not a backup — and in a report to the regulator, that’s exactly what it should be called. A second nuance people remember too late: deep snapshot chains consume the expensive flash of the production cluster — the very flash that was bought for the workload.

Replication to a neighboring cluster at the same site. Now you have a separate hardware failure domain, and that’s a fundamentally different level. Nutanix Disaster Recovery offers three modes here: asynchronous, with a restore point once an hour or less often; NearSync, built on lightweight LWS snapshots, with an interval from one to fifteen minutes; and synchronous replication with zero data loss — the latter requires a link with latency of up to five milliseconds. It covers a cluster failure. It doesn’t cover the loss of a site, and it doesn’t cover ransomware: the replica will faithfully repeat any logical corruption, just one interval later.

Replication to a remote cluster. The same, plus site protection, and here the link becomes the bottleneck. Initial synchronization runs as full copies, after that only changes, but the minute-interval modes are sensitive both to bandwidth and to the rate of data change. Starting with NCI 7.5, a single protection policy can combine up to four failure domains: synchronously to a neighboring site, asynchronously or NearSync to a region, plus long-term retention in object storage. That’s no longer “replication” but a full-fledged topology, and it needs to be designed accordingly.

Offloading snapshots to object storage — Multicloud Snapshot Technology. It takes deep chains off the production pool, puts them into any S3-compatible storage, and lets you recover to any location where Nutanix is deployed. The interval is from one hour. An important recent change: Instant Restore, which became generally available in Prism Central 7.5.1, starts a machine from metadata without waiting for all the data to be offloaded — meaning instant recovery now works from object storage too, not just from a local tier. Of all the platform mechanisms, this is the closest to a real backup. One caveat applies to all of them: without Nutanix Guest Tools, a snapshot is only crash-consistent, and that won’t be enough for a busy database.

A dedicated backup environment. A separate cluster or a separate site for backups, where backup software runs with its own repository, catalog, and immutability. It’s needed where you require independence from a compromise of the platform itself, granular recovery of objects inside applications, a single catalog across the whole fleet — including whatever sits outside hyperconvergence — and retention for years to meet regulatory requirements.

The practical rule we’ve arrived at: platform mechanisms cover failures and mistakes; backup software covers malicious intent and regulatory compliance. This isn’t competition but a division of responsibility, and in a well-designed project both are at work. Design errors begin where one area is covered with the tools of the other: compliance with snapshots, and a fifteen-minute target with a full restore from the repository.

What each protection mechanism coversPlatform mechanisms cover failures and mistakes. Malicious intent and regulatory requirements arecovered by a dedicated environment.Snapshot onthe sameclusterReplica to aneighboringclusterReplica to aremoteclusterSnapshots toobjectstorageDedicatedbackupenvironmentwith softwareand WORMHardware failureDisk or node failureFailure of the entire clusterSite failure: power,connectivity, fireError and data corruptionFile deleted, update rolledbackLogical data corruption by anapplicationGranular recovery of anapplication objectMalicious intent and regulationRansomware withadministrator rightsDeleting restore points fromthe insideRetention for years to meetregulatory requirementsRecovery to a differentplatformCoversCovers partiallyDoes not coverWhat each protection mechanism coversPlatform mechanisms cover failures and mistakes.Malicious intent and regulatory requirements are coveredby a dedicated environment.1Snapshot on the same cluster2Replica to a neighboring cluster3Replica to a remote cluster4Snapshots to object storage5Dedicated backup environment with software andWORM12345Hardware failureDisk or node failureFailure of the entire clusterSite failure: power,connectivity, fireError and data corruptionFile deleted, update rolledbackLogical data corruption byan applicationGranular recovery of anapplication objectMalicious intent andregulationRansomware withadministrator rightsDeleting restore points fromthe insideRetention for years to meetregulatory requirementsRecovery to a differentplatformCoversCovers partiallyDoes not cover
Fig. 5. Scope of platform protection mechanisms and a dedicated backup environment

Tiers and protocols

Where to put all of it

A repository is rarely homogeneous. The key question about protocols isn’t “can the array do it” but who in the architecture provides them.

A repository is rarely homogeneous. A sound architecture has up to five roles: a landing tier for the backup window, a main pool for operational retention depth, a fast tier for hard targets, an immutable tier, and an archive. Roles can be combined, but deliberately.

These roles don’t have to be filled by classic arrays alone, and in the matrix below we’ve deliberately mixed three different types of target storage — in real projects, they sit side by side.

Classic arrays. Lenovo ThinkSystem DG on a unified platform — file, block, and object access on end-to-end NVMe. Lenovo ThinkSystem DS — the same hardware and the same storage operating system, but block only. Lenovo ThinkSystem DE on SANtricity — capacity at a reasonable price, which is exactly why this class gets chosen for the main pool. More on the difference between the first two a little further down, in the part about protocols.

A hyperconverged cluster as target storage. Lenovo ThinkAgile HX650 in a hybrid configuration with Nutanix Unified Storage covers both file shares and object access straight from the platform, with no separate array. For a customer already running on Nutanix, this is usually the cheapest way to get an object tier and the fastest in terms of timeline: not a new platform, just one more cluster of a familiar one.

And right away, a caveat we’re obliged to say out loud. A separate cluster is a separate failure domain for hardware, but not for the platform: the same vendor, the same software stack, the same administrative perimeter, and often the same Prism Central. Such a repository doesn’t cover the “platform administrator compromised” scenario, and the matrix calls this out as a separate row. Plus, a hybrid disk-based configuration isn’t suitable for a fast recovery tier — only for a capacity tier.

Software-defined object storage. Cloudian HyperStore on Lenovo ThinkSystem SR650 V4 servers answers exactly two rows of the matrix: the immutable tier and the archive. Object Lock, horizontal scaling by adding nodes, and a cost per terabyte that approaches tape with incomparably faster access. Among ready-made hardware alternatives in the same role is Everpure, the former Pure Storage FlashBlade//E line: object access and immutability come out of the box, at the price of a closed platform and lock-in to a single vendor.

The general rule for all three options — and probably the most expensive one in this whole section: having an object protocol doesn’t equal certification. Cloudian, Everpure, the object tier on the virtualization platform itself — each has to be checked against the compatibility matrix for the specific version of the backup software you’re installing.

We had a project where the target storage, which honestly and fully supported S3, simply wasn’t on the list of validated targets for the chosen backup software. This came to light after the storage platform had already been decided, and the architecture had to be rebuilt. And here the story gets the sequel that’s the reason we’re telling it: about a year later, the software vendor added that storage to its list of supported targets. Today, the same choice would be the right one.

Another decision that gets made silently and paid for later is combining roles on a single system. A landing tier and a main pool on the same box look like savings right up until the first time the backup window coincides with a restore: writing the backup stream and random reads for rehydration fight over the same disks and the same controller queue. On a normal night you won’t notice, but during an incident — exactly when the system is needed — both operations slow down at once. If you do combine them, account for it in the performance calculation, not just the capacity calculation.

Keep fill level in mind as well. A deduplicated pool filled above eighty-five percent starts to degrade in speed, and garbage collection can no longer free up space as fast as new backups arrive. Fifteen to twenty percent free space isn’t headroom for growth but an operating parameter, without which the stated figures can’t be reproduced. You won’t find this line in spec sheets, but in real life it’s mandatory.

The moral isn’t that someone didn’t support someone else. The moral is that the list is alive and moves in both directions, which means you have to check it as of the project date and for the specific product version — not from memory, not from last year’s experience, and not from the “S3-compatible” line in a spec sheet. Check the model too: support for a product family doesn’t automatically mean support for every line within it.

Repository roles and types of target storageA classic array, a hyperconverged cluster, and software-defined object storage solve different rows.No single type covers them all.Lenovo DGunifiedNVMefile, block,objectLenovo DSNVMeblock onlyLenovo DEhybrid andflashblock,capacityLenovoHX650hybrid,NutanixUnifiedStorageCloudianHyperStoreonLenovoSR650 V4LenovoTS4300LTO tapeOperational environmentLanding tier: writing within thebackup windowMain dedup pool for 14–30daysFast tier for a target of up tofour hoursDatastore for instant recoveryof machinesFile share as a repositoryLong-term retention andbackup protectionImmutable object tierTiering cold sets to objectstorageRegulatory archive for yearsIndependence from theplatform failure domainRemovable offline copyPrimary optionPossible, with caveatsNot suitableRepository roles and types of targetstorageA classic array, a hyperconverged cluster, andsoftware-defined object storage solve different rows. Nosingle type covers them all.1Lenovo DG unified NVMe file, block, object2Lenovo DS NVMe block only3Lenovo DE hybrid and flash block, capacity4Lenovo HX650 hybrid, Nutanix Unified Storage5Cloudian HyperStore on Lenovo SR650 V46Lenovo TS4300 LTO tape123456OperationalenvironmentLanding tier: writingwithin the backupwindowMain dedup pool for14–30 daysFast tier for a target ofup to four hoursDatastore for instantrecovery of machinesFile share as arepositoryLong-term retentionand backupprotectionImmutable object tierTiering cold sets toobject storageRegulatory archive foryearsIndependence fromthe platform failuredomainRemovable offlinecopyPrimary optionPossible, with caveatsNot suitable
Fig. 6. Repository roles and types of target storage

A separate word on protocols, because this is the most common argument in tenders. The question “should the array support file and object protocols” is almost always asked the wrong way. The right way to frame it is: who in the architecture provides the protocol.

Backup software handles file and object access on its own. The repository can sit on a network share, a copy can go to object storage with object locking, and during instant recovery the media server itself brings up a file share from its own block volume and presents it to the hypervisor. The array doesn’t need file protocols for any of this. Functionally, a block array plus a media server covers the full set of scenarios.

Hence the test question for any proposal: show us the point in the architecture where each declared protocol is used. No such point — the requirement is excessive, and you’re paying for it. There is a point — the requirement is justified, and that’s a sound engineering decision: presenting a share directly from the array really does take load off the media servers, and an object tier on the same hardware simplifies operations. It’s just an architectural choice, not a technical necessity, and it should be presented honestly.

And one more ceiling people remember too late. It’s not in the media but in the fabric. Current-generation tape drives deliver around four hundred megabytes per second per device. Twenty-four drives come to nine and a half gigabytes per second in theory. But if the library is connected with eight-hundred-megabytes-per-second links, each link can carry exactly two drives at full speed, and the rest sit idle. The same arithmetic shows that a continuous read-out of ten petabytes at a real bandwidth of three gigabytes per second takes more than a month of nonstop work. Migrating an archive to a new platform has to be planned as a separate project with its own window, not as a “data transfer” line in the project schedule.

Immutability

The air gap is gone — what replaces it

The physical gap disappeared along with tape in the operational environment. Its replacement has two layers and only works when both are in place.

Protection of backups against deliberate deletion used to rest on physics: the cartridge is ejected from the library, and no administrative privileges can reach it. With the move to disk and object storage, that gap disappears, and something has to replace it.

The replacement consists of two layers, and it only works when both are in place. The technical layer is immutability at the storage level; object locking on a “write once, read many” model has become the standard. The organizational layer is separation of privileges: the account that manages backups must not be able to shorten the retention period, remove the lock, or delete the storage. Add to that confirmation of critical operations by a second administrator.

A detail worth checking by hand. Object locking has two modes. The soft mode allows removal by a privileged account — that is, by exactly the person you’re worried about. The strict mode allows no one to delete anything before the retention period expires. Only the strict mode protects against a targeted attack, and different storage systems enable different modes by default.

Then there’s the cost. Immutability almost always costs more than the overall ratio suggests, for two reasons. First: isolated sets, as a rule, don’t share dedup references with the main pool — they’re self-contained so that a restore remains possible even if the repository is lost entirely, and self-containment means their own full copies. Second: until the retention period expires, the space isn’t freed, even if the copy is no longer needed. An error on the long side of the retention period can only be cured by waiting.

That’s why immutable tier capacity is calculated as a separate line, from the full volume and retention depth, without the main pool’s ratio. Conservative — yes. But this is precisely the line that most often turns out to be underestimated.

Choosing a platform

Two philosophies: how the software differs

Formally, both products solve the same problem. The differences grow out of different ideas about what happens to data in an enterprise.

Our shortlist usually includes two products. Formally, they solve the same problem, but they start from different ideas about what actually happens to data in an enterprise — and that’s exactly where the differences in capabilities and price come from.

Outside these two remains a niche of platform-native products — they’re good where all virtualization lives on a single platform, and they deserve a separate discussion. We don’t compare them here: mixing general-purpose enterprise systems and platform-specific solutions in one table is a guaranteed way to get a flawed comparison.

Two products, two different views of the taskBoth cover the baseline scenario. The differences begin where a homogeneous virtual environmentends.Commvaultunderlying premise:the enterprise is heterogeneous, data lives indozens of systemsWidest range of supported sourcesDeep work with DBMSs and enterpriseapplicationsLegacy and niche systemsAdvanced retention and tiering policyIntegration with hardware arraysnapshotsIsolated recovery environmentLicense: capacity of protected data at thesource plus per-machine packs; capabilitiesare split into tiers. Requires a dedicatedadministrator.Veeamunderlying premise:the fleet is virtual; speed of recovery andsimplicity are what countStrong instant recovery mechanisms forworkloadsHardware independence of the targetstorageBackup immutability in the basepackageHardened repository on a dedicatednodeAutomated testing of recovery plansWorkload migration between sites andcloudsLicense: portable, per workload, in packs, withconversion rules for file data and workstations.Counted by peak concurrent usage.Summary based on vendors' public materials, August 2026.Two products, two different views of thetaskBoth cover the baseline scenario. The differences beginwhere a homogeneous virtual environment ends.Commvaultunderlying premise:the enterprise is heterogeneous, data lives in dozensof systemsWidest range of supported sourcesDeep work with DBMSs and enterpriseapplicationsLegacy and niche systemsAdvanced retention and tiering policyIntegration with hardware array snapshotsIsolated recovery environmentLicense: capacity of protected data at the sourceplus per-machine packs; capabilities are split intotiers. Requires a dedicated administrator.Veeamunderlying premise:the fleet is virtual; speed of recovery and simplicityare what countStrong instant recovery mechanisms forworkloadsHardware independence of the targetstorageBackup immutability in the base packageHardened repository on a dedicated nodeAutomated testing of recovery plansWorkload migration between sites and cloudsLicense: portable, per workload, in packs, withconversion rules for file data and workstations.Counted by peak concurrent usage.Summary based on vendors' public materials, August 2026.
Fig. 7. Underlying premises, strengths, and licensing models of the two platforms

Feature tables rarely determine the choice: both cover the baseline scenario. What decides it is the answers to four questions. What share of the fleet consists of non-standard sources — if your infrastructure includes legacy systems and enterprise databases with special requirements, breadth of coverage becomes decisive. How deep is the integration with your virtualization platform — native agentless operation removes an entire layer of proxy servers. How many people will operate it — a product that requires a dedicated administrator becomes a risk in a team of three rather than protection against one. And how will the system grow: out, in the number of objects, or up, in volume — that directly determines which licensing model will be cheaper three years from now.

A fifth question stands apart, because that’s where projects go down in flames. Is the chosen target array certified for the required product version? Protocol compatibility doesn’t equal certification: storage that honestly supports object access may be missing from the validated matrix for a particular piece of software. This comes to light during deployment, when the hardware is already sitting in the rack.

Checklist

What to ask before sizing the specification

Twelve questions that turn a round number into a calculation you can defend.

  1. Has acceptable downtime and acceptable data loss been set for each service class — and signed off by the system owners, not by infrastructure?
  2. Was the volume calculated from used space or from provisioned space? Have replicas, service machines, and abandoned machines been excluded?
  3. Where did the data reduction ratio come from, and exactly which data stream does it apply to?
  4. Was daily growth measured or assumed from a typical figure?
  5. Which mechanism meets the target for each class — and what confirms it, other than the aggregate throughput of the hardware?
  6. How many streams will the product allocate to restoring a single object in your combination of versions?
  7. Is immutable tier capacity calculated as a separate line, with no shared references to the main pool?
  8. For each required protocol, is the point of use in the architecture identified?
  9. Is the target storage in the compatibility matrix for the required product version?
  10. Can the account that manages backups shorten the retention period or delete the repository?
  11. Is migration of existing data set out as a separate phase, calculated against the path bandwidth?
  12. Is there a test restore procedure that measures actual recovery time — and when was it last carried out?

Bottom line

In place of a conclusion

Backup is the only infrastructure system whose value is measured at the moment everything else fails. The rest of the time it looks like a cost item, which is why it’s so easy to design it as a formality: capacity is there, retention depth is there, statuses are green.

The maturity test is simple. If the organization has a document that records acceptable downtime for each service class, and a test restore report that records the actual time, the system has been designed. If there’s no such document, it’s premature to discuss ratios, licenses, and array models: there’s nothing to calculate.

If you have a technical specification or a commercial proposal on your desk right now, run it through the checklist above. Half an hour of your time, and usually by the third question it’s already clear what you need to talk to the contractor about.

The figures in this publication are estimates that illustrate orders of magnitude; anything that affects procurement or a target is verified on the specific configuration.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.