AI NewsModels & agentsReported

Opus 5.5 and GPT-6.1 Sol cut cache-read prices by 60 and 80 percent, and a blog post traces the drop to DeepSeek

A post on insufferable.dev from 29 September argues that the quiet cache-read price cuts on Anthropic Opus 5.5 (60 percent) and OpenAI GPT-6.1 Sol (80 percent) line up with DeepSeek-V4.1-Flash shrinking its KV cache from 389 GB to 890 MB per million tokens.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team paying for a long-context coding model can rebuild its cost model around the new cache-read prices, with Opus 5.5 at 0.20 dollars per million tokens and GPT-6.1 Sol at 0.10 dollars.

Someone paying for a long-context coding model watched the cache-read price fall by more than half last week and did not get a reason. A post on insufferable.dev from 29 September offers one, and ties the quiet price cuts on Anthropic Opus 5.5 and OpenAI GPT-6.1 Sol to a 437 to 1 drop in DeepSeek's KV cache footprint.

The post lays out the vendor numbers first. The author writes that Opus 5.5 cut its cache-read price by 60 percent versus Opus 5, and that GPT-6.1 Sol cut its cache-read price by 80 percent versus GPT-5.6 Sol's late-July pricing. The author also records input and cache-write cuts over the same window: Opus 5 to 5.5 moves input from 5 to 4 dollars per million tokens and cache-write from 6.25 to 5. GPT-5.6 Sol to GPT-6.1 Sol moves input from 5 to 2 and cache-write from 6.25 to 2.50.

The 437 to 1 cache drop

The post traces the cuts to a series of KV cache optimizations that DeepSeek has published over the past year. DeepSeek-V1 held 389.12 GB of KV cache per million tokens, in the author's chart. The MLA architecture came first, followed by "Compressed Sparse Attention" and a heavier variant, and now DeepSeek-V4.1-Flash holds 890 MB per million tokens, through CSA2, cross-layer cache reuse, a causal encoder-decoder architecture and FP4 caching. The ratio across the four model generations is 437, which the chart presents as 437 equal-area squares shrinking to one.

That reduction matters for a long-context model because the KV cache sits in GPU memory for the whole session, and VRAM is the budget line that caps how many users one server can hold. A smaller cache per token is more concurrent sessions per GPU, which is where a cache-read price cut comes from on a vendor that owns its hardware.

Where the attribution gets careful

The vendor prices and the DeepSeek numbers in the post are each from the source that published them. The step that joins the two, that Anthropic and OpenAI are using DeepSeek's cache work in their new models, is the author's reading, not a vendor statement. The post itself frames it that way: it points at the "silent releases" of Opus 5.5 and GPT-6.1 Sol, the stellar user reviews on routine use, and the small gap to each lab's flagship (Claude Fable 5.1 and GPT-6 Astra). The post then concludes that the Western labs adopted the recipes DeepSeek has given away.

Treat the cost math as a hypothesis, and treat the vendor prices as the fact that matters this week. The cache-read cut is the one line on each pricing table that moved the most, and a team that serves long-context coding workloads can rework its own cost model around the new numbers without needing the attribution to be right.

Source

Reported byinsufferable.dev

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX