DeepSeek just made 1M-token context much less absurd to run. V4.1-Flash uses only 890 bytes of KV cache per token, 4x less than V4-Flash. This is the kind of boring-sounding breakthrough that could make long-context agents actually practical.
@HuggingPapersDeepSeek just released DeepSeek-V4.1-Flash > A 552B MoE multimodal model with 1M-token context that compresses the KV cache to 890 bytes per token, slashing deployment costs for long-context agents.