// HACKER NEWS — CYBERSECURITY
The AI Race Just Got Awkward
If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse.
Not a week goes by when Anthropic doesn’t release another article on how the Chinese are distilling their models, becoming a danger to humanity itself, etc.
It’s beneficial for them to say that because it sets the ground for these models to be restrained legally and regulatorily later on.
But it’s clear that the days of mindless distilling are over.
Not just over. The new game in town is adopting Chinese labs’ advances. Note how I call this adoption instead of the more vitriol-infused “stealing” that Anthropic tends to use.
That’s because, unlike the Western companies, the Chinese are pretty much giving away their recipes.
The latest one shamelessly copied without acknowledgement is the breakthrough in KV cache optimizations that DeepSeek has generously shared with the world.
It is a mind-blowing optimization that basically dropped the KV cache footprint for certain use cases that use a long session context, like coding, by a factor of roughly 437x compared with DeepSeek-V1.
They were the first ones to release the MLA architecture, which compressed the cache by roughly 15x, and then followed it up with ‘Compressed Sparse Attention’ and ‘Heavily Compressed Attention.’ The latest DeepSeek-V4.1-Flash pushes it even further with CSA2, cross-layer cache reuse, a causal encoder-decoder architecture, and FP4 caching, bringing the global KV cache down to 890 bytes per token.
Why do these things matter? Because for serving long-context models, one of the largest costs is the VRAM needed to hold this cache in GPU memory.