This paper introduces DeepSeek-V4.1-Flash, a more efficient model that reduces the computational cost of long-horizon agents by optimizing its prefill process and cache compression. Practitioners can benefit from this model's improved performance and reduced storage needs for agentic workloads.
Firehose
Filtered to Papers, tagged “multimodal Mixture-of-Experts (MoE)” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives