Discuss a project

Issue 09Data infrastructure

The fork after vSphere

Five replacement options, two starting points — the classic array-based setup and a vSAN cluster — and what happens to storage and networking once the 'ESXi plus Fibre Channel array' combo disappears from the picture.

For roughly the last eight months, almost every infrastructure conversation eventually lands in the same spot: “we’re looking at where to move off VMware.” What usually follows is the name of a platform — already picked somewhere else — and a single question: how much will this cost in hardware.

That’s the wrong question. The price of hardware is not the most interesting variable here, and it’s certainly not the first one. The first is what you’re actually changing. Because you’re not changing a hypervisor.

This issue is about five replacement options and about what happens to storage and networking once the “ESXi plus Fibre Channel array” combo disappears from the picture. There are also two starting points, and they lead in different directions: the classic setup with an external array, and a vSAN cluster, where there’s no array at all. Versions, statuses, and supported configurations are current as of September 2026: the landscape moves fast, and what was true last fall is in places no longer true.

The starting point

Deadlines, not prices

The end-of-support date shapes the timeline more than any budget: the decision gets made in 2026, or it gets made under pressure.

Let’s start with the calendar, because it shapes more than any budget line. Sales of vSphere 8 stop, on the announced schedule, in October 2026; general support ends on October 11, 2027; technical support runs through October 11, 2029 — but by then it’s consultation on existing configurations, not active support. Work backward from there: a pilot on real workloads takes a quarter, migrating production takes anywhere from two quarters to a year, and stabilization plus team training takes another quarter. The decision needs to be made in 2026 or very early 2027; otherwise it gets made under deadline pressure, which is the worst possible mode for an architecture decision.

The commercial side adds three complications. Licensing moved to a per-core subscription with a minimum of 16 cores per processor — an 8-core server gets billed as if it had 16. The announced minimum order of 72 cores was withdrawn after pushback from the market, but the fact that it was even discussed shows the direction of travel. And third: holders of perpetual licenses with expired support are receiving letters questioning the legitimacy of installing updates, and the partner program for service providers has moved to an invite-only model, which has shrunk the channel in a number of countries.

Estimates of how much the bill jumps at renewal vary wildly — from one and a half times to an order of magnitude — and it depends above all on the discount in the previous contract. Run the numbers for your own core count. But the conversation about leaving doesn’t start because of price — it starts because the pricing model has become unpredictable over the horizon of the next renewal, while infrastructure gets planned in 3-year cycles, plus another 3 for the second round.

The scope of the change

It’s not the hypervisor that’s changing

Six layers change at once, and the hypervisor is the easy one. The failures happen in management, the surrounding tooling, and the network.

Here’s the main mistake we see in how the task gets framed. Leaving VMware gets described as swapping one product for another. In reality, six layers change at once, and the hypervisor is the simplest of them.

Layer one, compute: the hypervisor itself. This part is solvable — KVM, underlying all the alternatives, has worked reliably for a long time. Layer two, management. vCenter goes away, and with it the familiar objects: resource pools, roles, folders, tags, placement rules. All of that gets rebuilt from scratch in someone else’s model, and not all of it maps over one to one.

Layer three, storage, is the heaviest one — it gets two sections below. Layer four, networking, is the one people forget most often — it also gets its own section.

Layer five, the surrounding tooling: backup, disaster recovery, monitoring, integrations with security systems, agents, automation. Every item here is a separate compatibility-matrix check, and that’s exactly where things fail.

Layer six, the one nobody writes into the statement of work: guest OS licenses and the team’s know-how. OEM and boxed Windows Server licenses are tied to hardware and don’t transfer to a new platform — you need enterprise agreements. Better to sort this out before signing anything.

What changes when leaving vSphereThe hypervisor is the simplest of the six layers. Failures happen in management, the surrounding toolingand the network.Changes entirelyHypervisor and disk formatManagement layer: pools,roles, folders, tags,placement rulesSnapshot and clone modelBuilt-in backupOperating proceduresTwisted-pair accessnetworkEverything that was afunction of the platform, not aproperty of the workloadChanges partiallyPhysical servers — per thecompatibility listArray — protocol anddriverFibre Channel fabricMonitoring: agents andintegrationsDisaster recovery designSegmentation and policiesCarries over, but on its ownterms and not in everyconfigurationStays as it wasGuest systems andapplicationsDataAddressing, VLANs, routingApplication-levelintegrationsApplication-layer skillsRegulatory requirementsAnd this is exactly what has towork the same way the nextdayWhat changes when leaving vSphereThe hypervisor is the simplest of the six layers. Failureshappen in management, the surrounding tooling and thenetwork.Changes entirelyHypervisor and disk formatManagement layer: pools, roles, folders, tags,placement rulesSnapshot and clone modelBuilt-in backupOperating proceduresTwisted-pair access networkEverything that was a function of the platform, not aproperty of the workloadChanges partiallyPhysical servers — per the compatibility listArray — protocol and driverFibre Channel fabricMonitoring: agents and integrationsDisaster recovery designSegmentation and policiesCarries over, but on its own terms and not in everyconfigurationStays as it wasGuest systems and applicationsDataAddressing, VLANs, routingApplication-level integrationsApplication-layer skillsRegulatory requirementsAnd this is exactly what has to work the same waythe next day
Fig. 1. Six infrastructure layers and what happens to each one when leaving vSphere

The options

Five candidates

What’s actually on the table in September 2026 — with version numbers, not slide-deck promises.

Briefly, what’s actually available. Each one is then broken down by storage, networking, and hardware reuse.

Nutanix in its classic hyperconverged form. The current combination as of September 2026 is AOS 7.6, AHV 11.2, and Prism Central 7.6, all released in July. Storage is its own, distributed: compute and capacity grow together.

Nutanix with external storage. A separate license type, where the cluster is built only from compute nodes and capacity comes from a qualified external array. This is a fundamentally different architecture, and it’s covered in detail in the external-arrays section.

Red Hat OpenShift Virtualization. Virtual machines as Kubernetes objects; the current platform branch is 4.21. For those who don’t need the container side, there’s a separate edition scoped to virtualization only, which includes a migration tool.

SUSE Virtualization, formerly Harvester. The current line is 1.8, with 1.8.2 as the latest patch. Also Kubernetes under the hood, with its own distributed storage — SUSE Storage, aka Longhorn.

Proxmox VE. Version 9.2 shipped in May 2026: Debian 13.5, kernel 7.0, Ceph Tentacle 20.2.1 as the default storage backend, with Ceph Squid as an option. Managing multiple clusters is handled by a separate product, Proxmox Datacenter Manager, whose first stable release appeared in late 2025.

Transport

The network everyone remembers last

Storage traffic moves back from a dedicated fabric into Ethernet. What matters is bandwidth per core, not gigabits per server.

Now for something that doesn’t appear in any migration proposal.

In the classic “ESXi plus Fibre Channel array” setup, storage traffic lived on a separate fabric: physically dedicated, competing with nothing, and never part of the LAN budget at all. Every one of the five options above moves that traffic back into Ethernet — as inter-node replication in the hyperconverged schemes, or as NVMe/TCP or NFS on the front end in the external-array schemes. The network stops being transport and becomes a storage bus, and it can’t be left as-is in either case.

The network stops being transport and becomes a storage bus.

What matters is bandwidth per core, not gigabits per server. A dual-socket server from 2015 with 24 cores and a pair of 10-gigabit ports works out to roughly 0.8 gigabit per core. Today’s dual-socket server has 172 cores — two 86-core processors — and with the same pair of ports that drops to 0.12. Memory per node now reaches 2, and in places 6, terabytes; drives are sold at 15.36 terabytes each. VM density has grown many times over, while the network has stayed the same.

Bandwidth per core: what the same two ports turned intoVirtual machine density has grown many times over, the network has stayed the same. Calculation,not measurement.Dual-socket node, 201524 cores · 2 × 10G0.83 Gbit/s per coreDual-socket node, 201948 cores · 2 × 10G0.42 Gbit/s per coreDual-socket node, 202264 cores · 2 × 25G0.78 Gbit/s per coreDual-socket node, 2026172 cores (2 × 86) · 2 × 25G0.29 Gbit/s per coreThe same 2026 node172 cores · 2 × 100G1.16 Gbit/s per corenotional boundary below which the networkbecomes the density limiter00,51,01,5Total bandwidth of the node's ports divided by the number of physical cores.Bandwidth per core: what the same twoports turned intoVirtual machine density has grown many times over, thenetwork has stayed the same. Calculation, notmeasurement.Dual-socket node, 2015 · 24 cores · 2 × 10G0.83 Gbit/s per coreDual-socket node, 2019 · 48 cores · 2 × 10G0.42 Gbit/s per coreDual-socket node, 2022 · 64 cores · 2 × 25G0.78 Gbit/s per coreDual-socket node, 2026 · 172 cores (2 × 86) · 2 × 25G0.29 Gbit/s per coreThe same 2026 node · 172 cores · 2 × 100G1.16 Gbit/s per corenotional boundary below which the networkbecomes the density limiter00,51,01,5Total bandwidth of the node's ports divided by the number ofphysical cores.
Fig. 2. Network bandwidth per physical core across generations of dual-socket nodes

Hence the steps. Moving from 10 to 25 gigabit today isn’t an upgrade — it’s the baseline: any cluster with software-defined storage, or with block access over Ethernet, should be designed at 25 and up. Moving from 25 to 100 is for those where new consumers of bandwidth have shown up over the life of the old cluster. That’s the group we’d check first.

Redundancy rebuild is the biggest and most underrated consumer of bandwidth, but it needs to be counted carefully. Copies are spread across all nodes, so recovery is many-to-many: reads come from multiple nodes, writes go to multiple nodes, and the cluster’s aggregate bandwidth is in play, not a single link. And what gets rebuilt is the occupied capacity, not the drive’s rated size. On a live cluster, losing a drive gets closed out in minutes and shows up on graphs as a spike, not as a day of degraded performance.

Two other cases deserve attention: losing an entire node, where the volume involved is much larger and the available bandwidth is smaller — the node took its ports down with it; and a small cluster on a slow network, where the pattern is closer to “few-to-many.” Erasure coding is more expensive than replication here: reconstruction reads several fragments from different nodes. What’s worth checking isn’t “how many hours will this take,” but what peak the rebuild produces and what it competes with. Next: the front end to an external array — the same traffic that used to travel over a dedicated fabric. The backup window and, more importantly, the restore window: recovery time objective is bound by bandwidth, not by the backup system. Live migration when taking a node offline for maintenance — and it’s important not to misjudge the order of magnitude here. What travels over the network isn’t the memory installed in the node, but the occupied memory of the machines being evacuated, and it doesn’t travel just once: pages get copied while the machine is running, then the changed pages get resent, over and over, until what’s left is small enough for a brief pause. Allocated-but-untouched memory costs almost nothing; but on machines that are actively writing, the amount transferred noticeably exceeds the amount occupied. The lower bound for a 6-terabyte node that’s two-thirds full is 4 terabytes. With shared storage, the disks themselves don’t move — only memory and device state travel.

Then there are workloads that show up on their own, unrelated to migration. The private interconnect for an Oracle RAC cluster — though what matters there isn’t so much bandwidth as zero packet loss and stable latency. Synchronous replication of large PostgreSQL databases with a write-ahead log stream. Kafka partition rebalancing and intermediate data exchange in Spark. Index rebuilding in search clusters. Tiering out to object storage. Loading datasets and writing checkpoints, if there’s a GPU segment nearby. The morning mass logon of virtual desktop sessions. And cross-site replication, if disaster recovery has been moved from a daily cycle to something close to synchronous.

Who really takes the bandwidthWhat to check is not “is it enough now”, but what peak each of them produces and what that peakcompetes with.Constant flowSharp peakLatency-sensitiveGrows withvolumeAppearedafter leavingFCRedundancy rebuild after a nodefailureWrite replication between clusternodesBlock access to an external arrayLive migration when taking anode offlineBackup windowRestore from backupOracle RAC cluster interconnectPostgreSQL synchronousreplicationKafka partition rebalancingSpark intermediate dataIndex rebuilding in search clustersTiering out to object storageModel datasets and checkpointsMorning mass logon ofworkstationsCross-site replicationPronouncedModerateNot typicalWho really takes the bandwidthWhat to check is not “is it enough now”, but what peakeach of them produces and what that peak competeswith.1Constant flow2Sharp peak3Latency-sensitive4Grows with volume5Appeared after leaving FC12345Redundancy rebuild aftera node failureWrite replication betweencluster nodesBlock access to anexternal arrayLive migration whentaking a node offlineBackup windowRestore from backupOracle RAC clusterinterconnectPostgreSQL synchronousreplicationKafka partitionrebalancingSpark intermediate dataIndex rebuilding in searchclustersTiering out to objectstorageModel datasets andcheckpointsMorning mass logon ofworkstationsCross-site replicationPronouncedModerateNot typical
Fig. 3. Consumers of network bandwidth: traffic profile and the link to moving off a dedicated fabric

A separate note on physical layer, because it’s the quietest way to blow a budget, and it matters not to confuse two different fleets. If access is built on SFP+ — and that’s how most server installations are built — the transition looks manageable: SFP28 ports accept the old 10-gigabit modules, multimode cabling for 25 gigabit is usually adequate, though the length margin is smaller, and a site can be migrated in stages. You’ll still need to replace access cards and switches, but not necessarily the optics or the cable runs. If access runs over twisted pair, it’s much less forgiving: 25 gigabit over RJ45 doesn’t hold up in practice, there’s no intermediate step, and the whole 10GBASE-T fleet has to be replaced, cabling included. Further down the chain: uplinks, a separate fabric for storage or at least queue separation, end-to-end MTU across the whole path — the classic place where block access over Ethernet performance falls apart — and switch buffers when many replicas respond to one initiator at once. And the move to leaf-spine, because east-west traffic becomes dominant.

One caveat, so as not to overpromise: 100 gigabit doesn’t fix latency. It shortens frame transmission time and removes queuing, but round-trip latency is determined by the stack, the switching, and how the application behaves on writes.

Software-defined storage

Storage inside the platform

Giving up the array is paid for in usable capacity, network, and node resources. And file and object services are a separate question altogether.

The first of two big questions: whether to give up the array entirely and live on disks inside the servers.

For Nutanix this is a distributed file system with a replication factor of 2 or 3. The technology is mature, and the pitfalls are known: usable capacity is calculated after replication, not before it, and that’s a separate conversation with finance every single time.

For OpenShift Virtualization, the built-in storage is OpenShift Data Foundation. In internal mode, it’s Ceph, deployed inside the cluster via Rook, on the nodes’ disks. In external mode, it connects to a separately managed Ceph cluster. The second option is more honest for large installations, but it means one more system you need to know how to operate.

Proxmox has built-in Ceph, defaulting to Tentacle in version 9.2. It’s worth understanding here that Ceph isn’t a hypervisor feature — it’s an independent distributed system with its own failure model and its own requirements for node count. Three nodes is a lab minimum, not a design figure.

SUSE Virtualization defaults to SUSE Storage, aka Longhorn, first-generation engine. There’s a second-generation engine built on SPDK with NVMe-oF access — noticeably faster, and marked generally available upstream, in Longhorn 1.12. But in the SUSE Virtualization 1.8 documentation it’s still flagged as experimental and not for production use: backing images and volume encryption aren’t supported, and every node needs a dedicated core and 2 gigabytes of huge pages. We wouldn’t plan production on it as of September 2026.

And one more thing people remember at the last minute: clusters tend to end up multi-purpose. Besides running VMs, people want file shares for users and object storage for backups or archives off the same cluster. Formally, file and object access exist in almost every option: for OpenShift via ODF, that’s CephFS plus the RGW and NooBaa gateways; for Proxmox, it’s CephFS, with the object gateway set up manually, outside the interface. But this is storage for the cluster itself and for its pods — not a file server for people.

If you need enterprise-grade services for outside consumers — SMB with domain integration, multi-protocol access, quotas, access analytics, ransomware detection, object storage with versioning and object lock — all from the same console with a single support contract, then of the five options only Nutanix offers that, bundled as Unified Storage. Everywhere else it’s a kit of separate components, each with its own lifecycle. Caveats apply: it’s licensed separately, by capacity; file servers run as VMs on the same cluster and consume its resources and post-replication capacity; and on a compute-only-plus-external-array setup, this needs to be checked separately rather than assumed.

Storage inside the platform: what we pay for giving up the arrayUsable capacity is counted after redundancy. The Longhorn v2 engine is generally available upstreambut marked experimental in SUSE Virtualization.NutanixDSFCeph inOpenShift(ODF)Cephin ProxmoxLonghorn v1(SUSE)Longhorn v2(SUSE)Redundancy modelRF2 / RF33 replicas orEC3 replicas orEC2–3 replicas2–3 replicas,SPDKPractical minimum ofnodesfrom 4from 5from 5from 3from 3Node networkrequirementfrom 25Gfrom 25G,separatenetworkfrom 25G,separatenetworkfrom 25GNVMe +hugepagesVolume snapshots andclonesyesyesyesyesyesVolume encryptionyesyesyesyesnoBacking imagesyesyesyesyesnoFile access for externalconsumersFiles(separate)CephFS — forthe clusterCephFS — forthe clusternonoS3 object accessObjects(separate)RGW andNooBaa — forthe clusterRGWmanuallynonoStatus as of September2026productionproductionproductionproduction,defaultexperimentalin SUSE Virt.Minimum node count is for a real project, not a lab.Storage inside the platform: what we payfor giving up the arrayUsable capacity is counted after redundancy. TheLonghorn v2 engine is generally available upstream butmarked experimental in SUSE Virtualization.Redundancy modelNutanix DSFRF2 / RF3Ceph in OpenShift(ODF)3 replicas or ECCeph in Proxmox3 replicas or ECLonghorn v1 (SUSE)2–3 replicasLonghorn v2 (SUSE)2–3 replicas, SPDKPractical minimum of nodesNutanix DSFfrom 4Ceph in OpenShift(ODF)from 5Ceph in Proxmoxfrom 5Longhorn v1 (SUSE)from 3Longhorn v2 (SUSE)from 3Node network requirementNutanix DSFfrom 25GCeph in OpenShift(ODF)from 25G, separate networkCeph in Proxmoxfrom 25G, separate networkLonghorn v1 (SUSE)from 25GLonghorn v2 (SUSE)NVMe + hugepagesVolume snapshots and clonesNutanix DSFyesCeph in OpenShift(ODF)yesCeph in ProxmoxyesLonghorn v1 (SUSE)yesLonghorn v2 (SUSE)yesVolume encryptionNutanix DSFyesCeph in OpenShift(ODF)yesCeph in ProxmoxyesLonghorn v1 (SUSE)yesLonghorn v2 (SUSE)noBacking imagesNutanix DSFyesCeph in OpenShift(ODF)yesCeph in ProxmoxyesLonghorn v1 (SUSE)yesLonghorn v2 (SUSE)noFile access for external consumersNutanix DSFFiles (separate)Ceph in OpenShift(ODF)CephFS — for the clusterCeph in ProxmoxCephFS — for the clusterLonghorn v1 (SUSE)noLonghorn v2 (SUSE)noS3 object accessNutanix DSFObjects (separate)Ceph in OpenShift(ODF)RGW and NooBaa — for theclusterCeph in ProxmoxRGW manuallyLonghorn v1 (SUSE)noLonghorn v2 (SUSE)noStatus as of September 2026Nutanix DSFproductionCeph in OpenShift(ODF)productionCeph in ProxmoxproductionLonghorn v1 (SUSE)production, defaultLonghorn v2 (SUSE)experimental in SUSE Virt.Minimum node count is for a real project, not a lab.
Fig. 4. Platforms' built-in storage: redundancy, requirements, and status as of September 2026

Common to all four: software-defined storage is not a “free array.” You pay for it three times over. In usable capacity: replication or erasure coding eats anywhere from a third to two-thirds of raw capacity. In network, covered in the previous section. And in node CPU time and memory that don’t go toward running VMs.

External storage

The external array: what connects today

The main fork in the road is the protocol. Not everyone supports Fibre Channel, and that alone rules out some options before price ever comes up.

The second big question, and for most of our customers the decisive one, because they already have an array and nobody’s planning to write it off.

Let’s start with what usually comes as a surprise. Nutanix with external storage fundamentally does not support Fibre Channel — not now, not on the roadmap. Everything runs over Ethernet: NVMe/TCP, NFS, or the array’s own protocol. The cluster is built exclusively from compute nodes; you can’t mix them with hyperconverged nodes in the same cluster; licensing is per-core; and snapshots and clones are handled by the array itself. Qualified systems as of September 2026: Dell PowerFlex — a minimum of four storage nodes, configuration entirely on Dell hardware, and connectivity through Dell’s own SDC client rather than NVMe/TCP; FlashArray from Pure Storage, now Everpure, models //X, //XL, and the //C model added in 2026; Dell PowerStore, added in July 2026 starting with AOS 7.6 — generations one through three, excluding the X model, and only as a single-cluster appliance, with connectivity from 10 gigabit up. NetApp ONTAP — in early access, with general availability the vendor has stated for Q3 2026, connecting via NFS, with AFF A-series and some hybrid FAS systems named; support for Lenovo storage systems is promised on roughly the same timeline. Put simply, Fibre Channel here isn’t a connectivity detail — it’s a fork in the road for choosing the platform.

Fibre Channel here isn’t a connectivity detail — it’s a fork in the road for choosing the platform.

For OpenShift Virtualization, external storage connects through third-party CSI drivers, and the key selection criterion is whether the driver supports simultaneous access to a block volume from multiple nodes. Without that, there’s no live migration, and that’s the first thing to check in the driver’s specification, not in a slide deck. A classic Fibre Channel fabric doesn’t connect directly: there’s a separate product for that, IBM Fusion Access for SAN, which lets you use your existing SAN — essentially a clustered file system on top of LUNs. It works, but it’s one more license and one more system to operate.

For SUSE Virtualization, third-party CSI drivers are supported, including for root volumes, and in July 2026 a dedicated storage certification program for virtualization appeared. But there’s a landmine here worth knowing before you pick a platform: built-in VM backup only works with first-generation Longhorn volumes. For volumes on external storage, the platform can neither create backups nor restore them — that’s stated plainly in the documentation. So all backup has to move to an external product, and its compatibility needs checking separately. Plus a few details that surface at installation time: the multipath service on nodes is disabled by default, and the OS is immutable, so node preparation happens through an initial-configuration mechanism on each host.

For Proxmox, an existing fabric connects the standard way: LUNs get handed to nodes, with LVM built on top. For years the main limitation was the lack of snapshots on that setup. Version 9 closed that gap — snapshots are implemented as volume chains and work on thick LVM over both iSCSI and Fibre Channel. But even in 9.2 the feature’s status is still technology preview; graduating it out of that status is on the developers’ roadmap, but there’s no date. The platform still has no officially supported clustered file system.

External storage system: what connects as of September 2026NFS for Nutanix means ONTAP in early access. Classic hyperconverged Nutanix is not in the table: ituses only its own capacity.Direct FibreChanneliSCSINVMe/TCPNFSExternalvolumesnapshotBuilt-in VMbackupNutanix with external storageOpenShift VirtualizationSUSE VirtualizationProxmox VESupportedWith conditions or in previewNot supportedExternal storage system: what connectsas of September 2026NFS for Nutanix means ONTAP in early access. Classichyperconverged Nutanix is not in the table: it uses only itsown capacity.1Direct Fibre Channel2iSCSI3NVMe/TCP4NFS5External volume snapshot6Built-in VM backup123456Nutanix with externalstorageOpenShift VirtualizationSUSE VirtualizationProxmox VESupportedWith conditions or in previewNot supported
Fig. 5. Connecting external storage systems by platform and protocol

Geography

Two or three sites

There are almost no single-site customers left. This is where the options diverge more sharply than on any other point.

In our practice, there are almost no single-site customers left: two sites is the norm, three is common. And that’s not an implementation detail — it’s a condition for choosing the platform, because the options diverge more sharply here than on any other point.

Five things need checking, in exactly this order. A single point of management for all sites. Synchronous replication to a second site for anything that can’t tolerate data loss. Asynchronous replication to a third for everything else. Failover orchestration: not “the data made it over,” but a scenario that brings machines up in the right order and lets you test that without stopping production — and it’s usually that test that an audit ends up demanding. And live migration between sites, nice to have but not decisive.

Here’s how the options stack up. For Nutanix, all of this is bundled into a single product and managed from a single console: synchronous replication between two sites, asynchronous to a third, recovery plans with test failover. Plus file and object services replicate through the same mechanism — if the cluster ends up multi-purpose, that applies to the second site too.

If you’re staying on VMware, you already have a working multi-site setup — that’s historically one of the platform’s strongest points, and it’s the most painful thing to lose in a move.

For OpenShift, all of this is achievable, but it’s assembled from several products: a separate layer for managing multiple clusters, separate disaster recovery modes for stretched versus asynchronous setups, and latency requirements between sites. It works, but it’s noticeably more complex to design and operate.

For SUSE Virtualization, managing multiple clusters is handled natively, but there’s no native VM replication between sites: what’s left is backup to external storage and restore at the second site. For disaster recovery with a short recovery time, that’s not enough.

For Proxmox, the central manager gives a unified view and live migration between clusters, but it ties sites together loosely: each cluster stays autonomous, there’s no shared configuration, and there’s no automatic failover between sites either. Cross-site replication has to be assembled by hand from storage mechanisms and backup-server synchronization.

And a warning that applies to all of them. A cluster stretched across two sites needs a link with guaranteed latency and a third point for quorum. Two sites without a witness isn’t fault tolerance — it’s two places that can go down at the same time.

Two or three sites: what covers itWhat to check is not “is replication supported”, but whether there is a failover scenario and whether itcan be tested.UnifiedmanagementSynchronous tothe secondAsynchronousto the thirdOrchestrationand testingLive migrationbetween sitesStay on VMware (VCF 9)NutanixOpenShift VirtualizationSUSE VirtualizationProxmox VENatively, from one consoleSeparate product or manuallyNoTwo or three sites: what covers itWhat to check is not “is replication supported”, butwhether there is a failover scenario and whether it can betested.1Unified management2Synchronous to the second3Asynchronous to the third4Orchestration and testing5Live migration between sites12345Stay on VMware (VCF 9)NutanixOpenShift VirtualizationSUSE VirtualizationProxmox VENatively, from one consoleSeparate product or manuallyNo
Fig. 6. Platform capabilities for a two- or three-site connected setup

The existing fleet

What carries over from what you already have

Servers are 2–4 years old, the array is recent, the fabric is in place. But if there’s no array and everything runs on vSAN, the picture changes entirely.

The more common situation: servers are 2–4 years old, the array is recent, the fabric runs at 32 gigabit, and the customer reasonably doesn’t want to throw all of that out over someone else’s licensing policy.

If You’re Leaving vSAN

Here the picture is different, and generally simpler. There’s no array — the Fibre Channel fork disappears along with it, and the external-storage option is off the table unless the customer has also decided to move away from hyperconvergence altogether. The choice narrows down to “one hyperconverged platform instead of another.”

The fleet carries over better here than in any other case: the nodes, the local NVMe drives, the storage network (if it’s already 25 gigabit), and the topology itself were all designed for distributed storage and will move to it again. What needs to change is the software, not the hardware.

There’s one piece of bad news, but it’s a big one: there’s no in-place conversion from vSAN to another distributed storage system. Data has to be moved entirely, and, unlike the array-based setup, there’s no intermediate shared storage that can be hooked up to both platforms so machines can be moved one at a time without copying disks. The whole volume goes over the network — and the bandwidth section stops being theory and turns into a project-timeline calculation.

And two things get lost silently. Per-VM storage policies: nobody has a direct equivalent, and redundancy in other platforms is set at a coarser grain. And a stretched cluster with a witness node: it doesn’t carry over — it gets rebuilt from scratch under someone else’s rules, with different latency requirements between sites.

What of the current fleet carries over to the new platformIn the classic setup the main fork is the Fibre Channel fabric. A vSAN cluster has none at all, and itsfleet carries over best of all.Nutanix HCINutanix withexternalstorageOpenShiftVirtualizationSUSEVirtualizationProxmox VEServers 2–4 years oldLocal NVMe and SSD, including fromvSANArray with Ethernet portsArray with Fibre Channel ports onlyFibre Channel fabric and adapters10GBASE-T access switchesOEM-category guest Windows licensesCurrent backup schemevCenter roles, pools and placementrulesVM-level storage policies (SPBM)Stretched vSAN cluster with a witnessnodeCarries overCarries over with conditionsDoes not carry overWhat of the current fleet carries over tothe new platformIn the classic setup the main fork is the Fibre Channelfabric. A vSAN cluster has none at all, and its fleet carriesover best of all.1Nutanix HCI2Nutanix with external storage3OpenShift Virtualization4SUSE Virtualization5Proxmox VE12345Servers 2–4 years oldLocal NVMe and SSD,including from vSANArray with Ethernet portsArray with Fibre Channelports onlyFibre Channel fabric andadapters10GBASE-T accessswitchesOEM-category guestWindows licensesCurrent backup schemevCenter roles, pools andplacement rulesVM-level storage policies(SPBM)Stretched vSAN clusterwith a witness nodeCarries overCarries over with conditionsDoes not carry over
Fig. 7. Reuse of the current hardware and license fleet by platform

What carries over well. Servers — almost always: compatibility lists are broad, and for Nutanix compute nodes it goes back to 2017-era generations. NVMe and SSD drives — if they’re on the lists, and not unbranded. The array — if it can serve capacity over the right protocol. The team’s knowledge of guest systems and applications — completely.

What carries over with caveats. An existing Fibre Channel fabric: on a Nutanix-with-external-storage setup it doesn’t carry over at all; for OpenShift it needs a separate product; for Proxmox it works, but with snapshots still in preview status; for SUSE, it works through a third-party driver, but then you lose native backup. An array with only a Fibre Channel interface: sometimes adding Ethernet ports saves the day, if the model supports it, sometimes it doesn’t. Twisted-pair access network: doesn’t carry over — it changes.

What never carries over. Guest Windows licenses bought as OEM. Load-balancing rules tied to objects on the old management server. And the team’s habit of doing snapshots, clones, and restores in one place, one way.

And a warning about the temptation to go hybrid: keep part of the workload on the old platform, move part over. In practice, you end up running two platforms for a year, two backup systems, two monitoring stacks, and two sets of skills — and paying for licenses on both. A transition period is fine, but it needs an end date on the plan.

Migration

Migration tools and the surrounding work

The tool is half a day of work per machine. Backup, disaster recovery, and monitoring are months.

Every platform has migration tools, and they all do the same thing: read the machine’s disks from the old environment, convert the format, create the object in the new one.

For Nutanix, that’s Move, the most battle-tested of them all; it also gained no-copy conversion, where vVols volumes turn into AHV disks in place. For Red Hat, it’s the migration toolkit for virtualization, where in version 2.11 offloading the copy to the array reached general availability: the array itself moves the data, not the network. Plus warm migration, where most of the volume is copied while the machine keeps running. For speed you need VMware’s virtual disk development kit, and you’ll have to build the image with it yourself — the license doesn’t allow publishing it to a public registry. For SUSE Virtualization, it’s the vm-import-controller add-on with vSphere and OpenStack sources, with a non-obvious quirk: the machine’s name becomes the name of the Kubernetes object, and machines with names that don’t follow naming rules fail to import. For Proxmox, it’s the ESXi import wizard, also technology preview, with vSAN limitations and a speed drop when routed through the management server.

But the tool is half a day of work per machine, while the surrounding work is months. What to check before choosing a platform.

Backup: does your product support the target platform, at what level — full image or changed blocks — and what happens with restoring individual files. If the platform is SUSE and the volumes are external, there’s no native backup at all.

And a separate budget line that almost everyone skips: training. Not “read the documentation,” but a hands-on course before go-live. We run these courses ourselves, and we see the difference in support-ticket volume over the first six months.

Checklist

What to ask before choosing a platform

Eleven questions that move the platform choice out of the realm of preference and into the realm of constraints.

Bottom line

In place of a conclusion

Leaving VMware isn’t buying a different hypervisor — it’s redesigning the infrastructure from the ground up, just while keeping the workloads running. Anyone who treats this as a license swap discovers the other five layers along the way — usually at the worst possible moment.

There are several right answers. Some sites have a recent array and fabric that make Proxmox or OpenShift the obvious choice. Some have no Fibre Channel support, which rules out half the options outright. For some, what decides it isn’t the technology but the availability of local support and expertise in the country. The platform choice should fall out of your constraints, not out of a slide deck.

Three topics stayed out of scope, each deserving its own treatment: the “stay” scenario on VVF 9 or VCF 9, building from scratch, and the fate of Kubernetes — Tanzu is leaving along with vSphere.

The one thing we’d ask: count the network. In this story it isn’t a budget line item — it’s the condition for everything else to work, and it’s usually the part people assume is “already covered.” If you want to work through your own configuration, come talk to us — we’ll run the numbers together, on your data.

Discuss your project

Let’s discuss your project

Tell us about your platform or project — an engineer will reply on Telegram or by e-mail.

Message us on Telegram

Or message us on Telegram — the bot will pass your question to an engineer.