Firehose

Filtered to Papers, tagged “persistent KV cache” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

17 SEP 2026 · Paper

This paper introduces DeepSeek-V4.1-Flash, a more efficient model that reduces the computational cost of long-horizon agents by optimizing its prefill process and cache compression. Practitioners can benefit from this model's improved performance and reduced storage needs for agentic workloads.