Wand AI announced the addition of StorONE to its Sovereign AI offering, giving sovereign AI programs a new way to improve storage utilization alongside the platform’s existing compute optimization. Through StorONE’s Real-Time Tiering technology, organizations can run the same AI workloads using significantly less flash storage while maintaining performance and immediate access to their data.
A sovereign AI program has two scarce assets, GPUs and storage. Both must run at high utilization for national AI to be affordable at scale.
National AI infrastructure is judged on cost per outcome, and cost per outcome is a utilization problem. Wand’s sovereign backend already attacks the compute half: by letting every data owner run on shared capacity rather than a dedicated carveout, the architecture is designed to raise utilization of national compute from roughly 25% to more than 80%. Storage is the other half of the bill, and it is still bought the old way.
However, AI storage is provisioned for its peak, so nearly all of it is flash, even though, at any given moment, most of the data sitting on it is inactive. The expensive tier runs at low utilization by construction. Capacity compounds the problem. Bought per workload and per protocol, it strands in silos that cannot be lent to the workload that actually needs it. A nation ends up buying its most expensive media to hold its coldest data, and buying it again for the next workload.
The existing storage market does not resolve this. Data reduction returns roughly 30% to 40% capacity savings while consuming 80% to 85 % of CPU and memory overhead — paying for storage efficiency with the compute the AI stack was built to use. Archive tiers lower the bill by putting cold data behind a restore step, which is the one thing a retrieval workload cannot tolerate. And proprietary arrays tie capacity to hardware only one vendor can supply, at a moment when sovereign programs need to buy the drives that are actually available to them.
The addition of StorONE closes that gap. The company’s Real-Time Tiering writes all data to flash at full performance, then continuously identifies inactive blocks and moves only those to high-capacity media inside the same volume — no policies, no bolt-on tools, no restore step, and every block immediately accessible. StorONE reports approximately 90% average flash savings at roughly 10% cost-performance and overhead impact, against the 30% to 40% that legacy data reduction returns for 80%to 85% overhead.
Because the ONE-Volume Architecture separates storage services from hardware, any server, any drive, across block, file, and object, capacity is pooled across workloads instead of stranded in per-workload arrays, and a nation can build that pool from whatever media it can source. StorONE is elective rather than required: the Wand backend runs on a program’s existing storage where those economics already work, and StorONE is brought in where flash cost, footprint, or power is the binding constraint.
“Nations are sizing AI programs around GPU utilization and then quietly losing the same money one layer down, on flash that is mostly holding cold data,” said Gal Naor, Founder and CEO of StorONE. “Real-Time Tiering moves only the inactive blocks to high-capacity media inside the same volume, so everything stays immediately accessible. One of our enterprise customers went from nine cabinets to two, resulting in an 80% smaller footprint and 78% drop in power usage. At national scale, that is capacity, megawatts, and capital returned to the AI program.”
“Sovereign AI is an economics question before it is a technology question, and the answer is utilization on both sides of the stack. We are pleased to welcome StorONE into Wand’s sovereign technology ecosystem, adding a storage efficiency capability that ministries, institutions, and agencies can draw on where flash economics are the constraint, serving the same workloads on far less flash, and spending the difference on the AI labor itself,” said Cristian Felix, Chief AI Architect of Wand AI.
How It Works
Utilization is raised on two axes. Flash utilization. Everything is written to flash at full performance; real-time algorithms continuously separate active from inactive data and relocate only the required blocks to high-capacity media within the same volume. The hot working set keeps flash economics on the data that earns them, and the cold majority stops occupying the most expensive tier while remaining directly addressable, with no external system, copy, or recall in the path. Capacity utilization. One platform serves all protocols and all media on standard hardware, with each volume carrying its own performance, protection, and retention parameters. Capacity is drawn from a shared pool rather than committed to a dedicated array per workload, so a ministry’s unused terabytes are available to the ministry that needs them.
The effect compounds beyond the storage line item. Less flash for the same workload means fewer racks, less power, and less cooling in one enterprise deployment, a reduction from nine cabinets to two, with an 80% decrease in data center footprint and a 78% reduction in power consumption. In a national AI build, those are the constraints that determine how much GPU a program can actually stand up.
The capability is complementary to the inference privacy, confidential computing, and model robustness layers already in Wand’s sovereign ecosystem. Those raise the utilization and trustworthiness of national compute; StorONE raises the utilization of the storage underneath it. The capabilities are independent, each can be adopted on its own, and none is a prerequisite for another, but they move the same number: cost per outcome for every ministry and agency the sovereign backend serves.
Related News: