Technology

Memory Is Getting More Expensive, So AI Is Learning to Make Do With Less

NVIDIA is selling a half-memory DGX Spark for more than last year's model with twice the memory, and a seven-year-old Shield just got 50 percent pricier. The same week, a developer ran a 125-billion-parameter model on a 12 GB gaming GPU.

RAM modules laid out in shrinking rows on a table, with an expensive price tag on the last single module (illustration)

Half the memory, higher price

NVIDIA's DGX Spark is a compact desktop computer built for AI developers. It launched in October 2025 with 128 GB of unified memory and a price tag of $3,999. The idea was simple: run large AI models on your desk without needing a data center.

Two small open desktop computers; the one with less memory carries the bigger price tag (illustration)

That price didn't last long. In February 2026, NVIDIA raised it to $4,699, citing global memory supply constraints. This month, the 128 GB model climbed again to $6,950, about 75 percent above its launch price.

In the same week, NVIDIA also announced a more "accessible" version: a 64 GB DGX Spark that will go on sale on October 23 through Acer, ASUS, Dell, Gigabyte, HP and MSI. It starts at $4,999.

Put the numbers side by side and things start to look strange. The new 64 GB model is cheaper than today's 128 GB version. But it is still about 25 percent — exactly $1,000 — more expensive than the launch price of last year's model with twice the memory. And if someone wants more memory and links two 64 GB systems together, the total comes to $9,998, roughly $3,000 more than a single 128 GB machine.

According to NVIDIA, the 64 GB model can run models with around 100 billion parameters, while the 128 GB version can handle models up to around 200 billion.

A nearly seven-year-old box just got 50 percent more expensive

It would be a mistake to think the memory crunch is only hitting expensive AI computers. In the same week, NVIDIA raised the price of the Shield TV Pro, which launched in 2019 at $199.99, to $299.99 as of October 2. The standard Shield model has also been discontinued entirely.

A TV box, a remote and a receipt with the price circled in red on a coffee table (illustration)

What makes the increase unusual is that nothing inside the device has changed. The Shield TV Pro still uses the same processor, the same 3 GB of memory and the same 16 GB of storage it had nearly seven years ago. Technology products usually get cheaper as they age. This one is now 50 percent more expensive than when it launched. NVIDIA's explanation is straightforward: component costs, including memory, have risen sharply across the industry.

The Shield is not alone. Apple raised prices on its streaming devices at the end of June, Roku followed at the end of July, and Google did the same at the end of August.

What's behind the crunch?

One of the main forces behind these price increases is the seemingly bottomless appetite of AI data centers for memory. AI servers rely on a special type of high-bandwidth memory called HBM. Producing the same amount of HBM consumes more wafer capacity than conventional DRAM. Strong demand from AI accelerators is also pushing manufacturers to devote a growing share of their limited production capacity to HBM and server memory.

A vast data center campus at sunset (illustration)

The direct result is tighter supply of conventional DRAM for PCs and consumer electronics. Storage is under similar pressure, although the reason there is not HBM but the growing demand for enterprise SSDs from AI data centers. We looked at why RAM and SSD prices are rising in an earlier piece.

And that is where the irony comes in: NVIDIA, one of the world's biggest sellers of AI chips, is passing part of the cost created by the memory demand around those chips on to customers buying one of its boxes simply to watch movies in the living room.

On the other side: 125 billion parameters, 12 GB of VRAM

At almost the same time, a story arrived from the open-source world moving in the opposite direction.

A developer known as Niko1221 released an MIT-licensed project called Strata, which claims to run Qwen's 125-billion-parameter Qwen3.8-Flash-Next language model on a single gaming GPU with just 12 GB of memory. Models of this size would normally require servers equipped with hundreds of gigabytes of GPU memory.

What makes this possible is the structure of the model itself. Qwen3.8-Flash-Next is a mixture-of-experts, or MoE, model: it contains thousands of smaller "expert" subnetworks, but only a small fraction of them are activated for each token it generates. (We explained why the active parameter count alone does not set the memory requirement in our Nemotron piece.) Strata places the model in system memory, keeps the most frequently used experts in the GPU's faster memory as well, and also puts the CPU to work. The model can also be compressed down to 2- and 3-bit variants. That cuts memory usage dramatically, though it can also come with some loss in model quality, depending on the level of compression.

The project quickly picked up thousands of stars on GitHub. Reported speeds are generally in the range of several dozen tokens per second depending on the GPU and the rest of the system. But results vary significantly across different configurations, so giving Strata a single performance number would be misleading for now.

There is still a memory bill

There is one easy-to-miss detail in the Strata story: it saves GPU memory by leaning heavily on system memory instead. It requires at least 32 GB of RAM, recommends 64 GB if you want access to all model sizes, and also needs roughly 70–80 GB of disk space.

A PC with a graphics card in the middle, linked by lines of light to rows of RAM modules and SSDs on both sides (illustration)

So Strata does not eliminate the need for memory. It spreads that need across different layers. All of the experts are kept in system memory, while the few thousand used most often are also placed in the GPU's faster memory; a large lookup table lives on the SSD. With larger model variants, some weights can also be read from the SSD when system memory is not enough. And the system memory carrying most of that load is one of the same components being hit by the crunch behind NVIDIA's Shield price increase.

That is what makes the two stories so interesting when you put them next to each other. On the hardware side, memory is becoming one of the scarcest resources of the AI era, to the point where lower-memory versions of the same platforms can now cost more than more capable systems did in the past. On the software side, developers are looking for ways to run the same models using less — and cheaper — memory. The more expensive memory becomes, the more valuable it is to use it intelligently.

For anyone thinking about building a PC, the practical lesson is simple: over the coming months, memory prices may play a much bigger role in the total cost of a system than they used to.

TagsAINVIDIAHardware

Related posts

All posts