Insights · Evidence-first decisions

Six times the capacity from drives you already own? Read the footnotes, then measure.

With flash memory on allocation and lead times stretching, one storage vendor is offering to find “up to 6x or more” effective capacity in SSDs customers already have. The instinct is right. The number is not a measurement. The lesson applies to every AI-readiness decision you will make this year.

Consulting News Desk27 January 20263 min readEvidence-first decisions

The offer

The pitch is well timed. Flash memory is in short supply, allocation windows are tight, and any organisation planning an AI programme is being told it needs more fast storage than it can currently buy. Into that gap comes an offer: audit your estate, qualify the servers and drives you already run, consolidate them onto a modern platform, and recover up to six times the effective capacity you thought you had.

Strip away the branding and it is an engagement rather than a product — an audit phase, a qualification phase, a consolidation phase — wrapped around techniques the vendor has shipped for years: erasure coding instead of replication, data reduction across the whole namespace instead of per volume, and a write path that absorbs bursty writes in faster memory before landing them sequentially on flash.

Two things are true at once. The underlying argument is sound. And the headline number does not mean what it appears to mean.

The instinct is right

Scarcity exposes waste that abundance hides. Most large estates carry three kinds of it: protection models that keep three full copies of everything, data stacks fragmented into silos each with its own overhead, and performance designs that rely on over-provisioned drives sitting mostly idle. When you could buy more flash next week, none of that mattered enough to fix. When you cannot, it is suddenly the cheapest capacity available.

The same logic runs through the whole of AI readiness. Before buying anything new for an AI programme, the first question should be what the existing estate can already do. Warehouses with unused headroom. Storage carrying duplicate copies of production data that should never have existed. Compute reserved for batch windows that no longer run. In most organisations we assess, the first meaningful capacity for AI comes from consolidation, not purchase.

The number is not

Now the footnotes. “Up to 6x or more” is the product of three separate multipliers stacked together: erasure coding against replication, global reduction against per-volume reduction, and reduced over-provisioning. No baseline platform is named. No dataset is described. No reduction ratio is broken out. An estate running three-way replication over highly similar data would reach the top of that range by arithmetic alone; an estate already running efficient erasure coding would not come close.

A multiplier assembled from assumptions is a forecast. A multiplier measured on your data is a result. Only one of them should appear in a business case.

The release also says customers are already benefiting, and names none, and attaches no capacity figure. And it is silent on what the engagement costs, how long qualification takes, and which hardware qualifies — which, for a proposal that amounts to migrating your data onto the vendor’s platform using your own drives, are the questions that decide whether it is a good idea.

The discipline that settles it

None of this makes the offer wrong. It makes it unevaluated. The way to evaluate it is the same way any platform claim should be evaluated: run it on a representative slice of your own corpus, with the pass mark agreed before the results come in.

For a data-reduction claim, that means a real sample of your data — not a synthetic set, and not the vendor’s demo set — measured for reduction ratio, protection overhead and performance under your access pattern. Similarity-based reduction, the technique behind the biggest of the three multipliers, swings violently by data type: it finds near-identical blocks and stores the difference, which is spectacular on some corpora and marginal on others. The only ratio that matters is yours.

The same discipline applies to the other claim buried in the release: that AI inference, as it scales, will add a second demand on the same scarce flash by persisting and reusing model state. That is plausible and worth planning for. It is also a forecast, and should be sized from measurements of your own workloads rather than accepted as a reason to buy.

Before the purchase order

Audit the estate first — copies, silos, protection overhead, idle capacity. Ask any vendor promising a multiplier for the baseline, the dataset and the breakdown. Then measure on your own data with agreed criteria, and let the number decide. It usually comes out somewhere between the vendor’s slide and the sceptic’s shrug, and it is the only number worth putting in front of a finance director.

Consulting News DeskWeekly notes on AI integration, data foundations, and agentic workflows from the IDMS consulting team — written by the people doing the integration work.