Firehose
Everything worth reading, newest first — posts from tracked people and companies, plus new papers. For raw numbers (models, repos, benchmarks) see Dashboard.
Great first-party data alone does not guarantee effective marketing, as it often falls short of activating customer signals into real-time campaigns due to integration bottlenecks and siloed tools in the martech stack. A unified data founda…
Simulation for physical AI systems relies on generating large amounts of physically grounded data, which is challenging to collect in the real world due to safety, cost, and practicality concerns. Simulation engines like MuJoCo, Isaac Sim, …
Dow built a Carbon Footprint Ledger (CFL) on the Databricks Data Intelligence Platform to unify enterprise data and calculate cradle-to-gate Product Carbon Footprints (PCFs) for its entire portfolio, reducing processing time from weeks to a…
Databricks' Lakehouse architecture is being used to integrate scattered R&D data from various source systems into a unified, AI-ready product, enabling industrial AI in industries like heavy-duty applications, where data from multiple sourc…
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
Xaira Therapeutics' X-Cell model for drug discovery relies on a large dataset, X-Atlas, which contains information-rich data on gene expression in human cells, enabling the model to predict changes to gene expression and understand the rela…
Databricks has announced the Public Preview of Discover and Domains, powered by Unity Catalog, which provides an internal marketplace for data and AI assets, enabling users to find trusted, relevant data and AI assets through business-align…
OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.
Canvases turn AI into interactive workspaces where you can visualize information, explore workflows, and take action across complex tasks. The post How to build interactive experiences with canvases appeared first on The GitHub Blog .
Google introduces Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, models designed to improve efficiency, latency, and reliability for building AI agents at scale, with Gemini 3.6 Flash offering 17% reduced output token usage compared…
Nativ: Run AI models locally on your Mac Prince Canuma is the developer behind the excellent MLX-VLM Python library for running vision-LLMs using MLX on a Mac. I'm really excited about his new project, which wraps MLX in a full macOS deskto…
With this post, I’ll wrap up my notes from the second Future of Software Development Retreat . But before I do, I should note that the full Thoughtworks report on the retreat is now available . They have five headline findings: Code generat…
We analyzed global HTTP traffic to explore how kickoff times, streaming habits, and hydration breaks reshaped online activity worldwide. From late-night traffic surges to halftime browsing spikes, here is how the world connected during the …
Earlier this month I hosted a fireside chat session at the AI Engineer World's Fair with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design…
Netflix's Q2 earnings were in line with expectations, with revenue growth slowing down to 7.5% year-over-year, indicating the company's most exciting days may be behind it, and it's likely in a mature phase. The company's subscriber growth …
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
Today we're expanding Vercel Agent . It started by triaging alerts and reviewing your pull requests. Now it has a home in your dashboard, where it can investigate production, answer questions about your projects, and take action once you ap…
Searchable on Vercel 5x increase in development velocity 100+ billion tokens processed Customer-requested features shipped in as little as 30 minutes Zero model SDK implementation or API key rotation with AI Gateway Searchable helps brands …
The top AI news of the week includes the announcement of the AIE Security track and the release of Sonar CEO Tariq Shaukat's emphasis on verification for safety/security/correctness. Meanwhile, US debate over restricting Chinese open models…
Hello! I’m on a funny journey right now where I’m trying to learn how to make websites in a sort of 2010 style, where I have an SQL database and render some HTML on the backend. It’s kind of an interesting journey because …
AI Gateway now supports service tiering. Service tiers let you optimize for latency, throughput, and cost per request to match your use case. Pick a faster tier for interactive workloads (less queueing, higher token throughput), or a lower …
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway. Gemini 3.6 Flash improves quality across coding, agentic tasks, and web development while consuming fewer tokens and making fewer model calls. It produces cleaner w…
Laguna S 2.1 from Poolside is now available on AI Gateway. There are 2 versions of the model available: Free version (256K context window): poolside/laguna-s-2.1-free Paid version (1M context window): poolside/laguna-s-2.1 Laguna S 2.1 is a…
Vercel MCP now supports purchasing Vercel products. You can: Upgrade your team to the Pro plan Add prepaid credits for v0 (requires a paid v0 plan) or AI Gateway Purchase the SIEM add-on (requires an Enterprise plan) Purchase and register a…
Vercel Connect now includes preset connectors for 90+ services, including Shopify, Okta, Workday, Jira, and Sanity. Preset connectors are predefined configurations for supported services. They reduce manual setup by pre-populating the brand…
Vercel now compiles Python functions to bytecode at build time. In our benchmarks, cold starts for the median-sized function dropped from 2.8s to 1.3s . When Python imports a module without cached bytecode, it parses and compiles the source…
Grabette is an open-source system for recording robot-manipulation data, allowing users to capture demonstrations with a handheld gripper and a camera, and then process the data into a robot-ready dataset. The system consists of a handheld …
David Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBC, bringing global leadership in finance, technology, and governance.
Cloudflare Internal DNS brings authoritative and recursive DNS for private networks to the same global network and control plane that runs Cloudflare's Zero Trust, networking, and public DNS.
Glaspoort's Lakebase branching setup branches every environment from production, creating ephemeral per-PR databases for CI/CD, and treats migrations as the single source of truth, avoiding the "reset-from-parent trap" in development and ac…
I keep hearing anecdotes from people who used coding agents to reverse-engineer and automate devices in their homes. I think this is an interesting illustration of the impact of the reduced cost of writing code. Prior to agents, it was enti…
Researchers at Databricks developed a solution to map freeform text to large taxonomies of 100k+ labels, outperforming existing frontier models in accuracy and cost. By combining vector search with the Databricks AI Classify function, they …
Seven options, considered
Who’s Afraid of Chinese Models? Interesting proposal from Ben Thompson that both addresses the hypocrisy of labs outlawing distillation against their models despite training on unlicensed data, and could help US open models compete more eff…
Celebrating $100 million contributed by the community to the people who build and sustain open source every day. The post $100 million for open source: A milestone built by the community appeared first on The GitHub Blog .
The global implications on the AI ecosystem.
Kimi K3 is a very good model with excellent benchmarks.
It is 5:45 on a Friday morning, and a store manager is standing in the back office...
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now UK government: Gap between open and closed weight models o…
Everyone is worried about Chinese models, but the frontier labs will be fine; we need to enable open U.S. alternatives.
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
AI transparency is the practice of making an artificial intelligence system's data, model behavior...
We have been having extensive discussions around open source strategy. We will discuss it more at our next board meeting, but one thing we’d like to do soon is to create a language model with the approximate capability of GPT-3 that can run…
Team Owners can now clear the team's Remote Cache of all artifacts in one click. This is useful when you believe there are poisoned artifacts in your cache. In your team's Build and Deployment settings, visit the Remote Caching section and …
Vercel Workflows now keeps each run's state, queue dispatch, and output streams in a single home region: the region where the run starts by default, or any target region you choose. A run keeps its home region for its lifetime, so for agent…
Apply for Anthropic’s AI for Science rare disease research grants
20 JUL 2026 · Paper
The prevailing inference framework for diffusion models formulates generation fundamentally as a problem of numerical integration. This perspective casts the model as an exact estimator, neglecting the inherent statistic
20 JUL 2026 · Paper
This paper proposes a new way for large language models to learn from feedback, allowing them to retain more detailed information about the quality of their responses and learn from it in a more nuanced way. Practitioners might care because this approach could lead to better performance on tasks where the model doesn't have a clear way to evaluate its own output.
20 JUL 2026 · Paper
This paper introduces RynnBrain 1.1, a family of large-scale embodied foundation models that can perform tasks like spatial reasoning, localization, and planning, and shows promising results in real-world robot experiments. Practitioners may care about the potential of these models for robot manipulation and control.
20 JUL 2026 · Paper
This paper proposes a new pruning method for coding agents that prunes tool outputs directly inside the agent, rather than relying on a separate code classifier, and shows it can save up to 39% of tokens while preserving task quality.
20 JUL 2026 · Paper
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement,
20 JUL 2026 · Paper
This paper develops a video personalization method that focuses on human-object interactions, aiming to improve the accuracy of video generation by better understanding human-object relationships and incorporating intra-subject references. Practitioners may care about this research as it could lead to more realistic and engaging video content.
20 JUL 2026 · Paper
This paper proposes a new method for training language models to generate coherent and faithful responses, even when the input data has changed significantly. Practitioners might care about this research because it could lead to more robust and adaptable language models that can handle real-world scenarios where data distributions shift.
20 JUL 2026 · Paper
This paper creates a new AI model that can edit and generate videos without needing masks, and can also learn to mimic image editing capabilities. Practitioners might care about this because it could lead to more diverse and realistic video editing data, and enable AI models to understand and generate human-like video editing instructions.
20 JUL 2026 · Paper
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. W
20 JUL 2026 · Paper
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic be
20 JUL 2026 · Paper
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot
20 JUL 2026 · Paper
This paper evaluates whether language models can be used to generate molecules that fit specific 3D constraints, such as protein pockets and spatial relationships between molecules. Practitioners in the field of drug design might care about this research because it could lead to more efficient and effective methods for generating candidate molecules.
20 JUL 2026 · Paper
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal
20 JUL 2026 · Paper
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or
Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.
AI Mania Is Eviscerating Global Decision-Making Here's an entertaining perspective from Nik Suresh on the AI mania that is overwhelming the large companies that he consults with. It's crammed with spicy anecdotes from anonymous sources. In …
In Rewriting Bun in Rust Jarred Sumner made the following claim: Claude Code v2.1.181 (released June 17th) and later use the Rust port of Bun. Startup got 10% faster on Linux but otherwise, barely anyone noticed. Boring is good. I decided t…
19 JUL 2026 · Paper
This paper proposes a new method called Distilled Reinforcement Learning that improves large language model post-training by providing fine-grained guidance to transfer new knowledge from a teacher model to a student model. Practitioners might care because it outperforms standard reinforcement learning and on-policy distillation methods in terms of knowledge transfer and model performance.
19 JUL 2026 · Paper
This paper introduces a framework called EvolvingWorld that allows characters and worlds to evolve together over time in interactive literary worlds, enabling more realistic and coherent simulations. Practitioners interested in developing more immersive and dynamic interactive stories might care about this approach.
19 JUL 2026 · Paper
This paper develops a new approach to understanding videos by predicting when specific events or evidence occur within the video. Practitioners working on video analysis and AI models might care about this research because it could lead to more accurate and robust video understanding systems.
19 JUL 2026 · Paper
This paper develops a new method for generating realistic hand-object interactions in 3D animation, combining appearance and motion to create smooth and believable movements. Practitioners in the field of computer animation and AI may care about this research for its potential to improve the realism and consistency of interactive simulations.
19 JUL 2026 · Paper
This paper develops a mathematical framework to analyze the Transformer architecture, using differential geometry to model its core components. Practitioners may care about this work because it provides new insights into the stability and optimization dynamics of Large Language Models.
Tool: SQLite Query Explainer Julia Evan's, in Learning a few things about running SQLite : Maybe one day I’ll learn to read a query plan. Big same.... which inspired me to have Fable build this interactive explain t
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
Claude make Fable 5 permanent An update from the @claudeai account on Twitter: Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access …
nascheme/quixote A certain vintage of Python web nerd might be delighted to learn that the most recent commit to the Quixote web framework was six hours ago . The oldest commit in that repo is from 21 years ago, and that w
a quiet day
A delayed education on the trenches and scales of suffering
18 JUL 2026 · Paper
This paper proposes a method to generate synthetic data for training API-calling agents without the need for an actual environment. This is useful for scalability, as collecting high-quality data at scale can be difficult. Practitioners might care about this approach because it could speed up the development of AI agents that can interact with APIs.
18 JUL 2026 · Paper
This paper proposes a new method for reinforcement learning in large language models, called Group Entropy-Controlled Policy Optimization (GEPO), which helps balance exploration and exploitation by controlling entropy levels across different tasks. Practitioners might care about GEPO because it can lead to more balanced and task-specific exploration levels.
18 JUL 2026 · Paper
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image ana
Cloudflare has deployed two WAF rules in response to high-severity vulnerabilities disclosed to us by the WordPress security team. The new rules protect all Cloudflare customers using affected WordPress versions, but customers should still …
Responsible AI: Governance, Principles, and Practical GuideResponsible AI is the...
Ask a tech company CFO where the quarter's margin is landing, and you will get a...
The best Stratechery content from the week of July 13, 2026, including the end of the mainframe, the continuing adventures of OpenAI, and answering the question, "Is Netflix Washed?".
The cost of writing code dropped; the cost of owning it didn't. A framework for deciding which changes are actually cheap in the AI era. The post The cost of saying yes has changed appeared first on The GitHub Blog .
Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.
As usual, part 2 of the weekly deals with speculative, regulatory, political and alignment questions.
Sarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability, and return on compute.
a great week for open models continues.
Hello! I’ve been working on a Django site recently, and I decided to use SQLite as the database. When I was getting started with using SQLite as database for a website I read a bunch of blog posts about how it is totally fine to use S…
17 JUL 2026 · Paper
This paper introduces S1-Omni, a unified AI model that can reason about scientific data, generate predictions, and create new scientific content, which could be useful for researchers and practitioners who need to analyze and understand complex scientific information.
17 JUL 2026 · Paper
This paper investigates the use of the Muon optimizer in reinforcement learning (RL) post-training and finds that it can significantly improve the success rate of RL agents, especially when combined with other techniques like policy optimization and advantage estimation. Practitioners in RL may care about this research to explore new ways to improve the performance of their agents.
17 JUL 2026 · Paper
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completion. Such measurements
17 JUL 2026 · Paper
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objectives. Existing compute services, whether offered as direct accelerator rental or
17 JUL 2026 · Paper
This paper introduces a benchmark to evaluate video generation models' ability to reason about physical laws, which is crucial for creating reliable world simulators. Practitioners caring about developing more realistic and physically intelligent AI models will find this research valuable.
To a sceptic, spending $165K to migrate Bun from Zig to Rust sounds very expensive. But to a realist, shortening a 1-2 year migration down to 11 days opens amazing new opportunities for devs. However, a thoroughly-tested project is required…
Gemini Omni and personal avatars in Google Vids make video creation easier than ever.
You’ll be able to securely link and interact with your go-to services directly in AI Mode.
Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.
This week saw the releases of, among other things:
Lila is betting that science, not the internet, is the last untapped source of training data. We went to find out what that actually looks like in a room full of robots.
When people think of legacy modernization, most folks aren't imagining the target environment will be Java 8. But this was the challenge facing Nik Malykhin when he needed to run a Java 1.5 codebase on today's hardware. His early use of LLM…
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
Thinky's first full LLM release is a banger and bonus: it's open weights!
Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company.
16 JUL 2026 · Paper
This paper introduces a vision-language-action model that can perform mobile manipulation tasks in unseen environments with minimal training data, and how it can be scaled up to achieve better performance. Practitioners might care about this model for building robots that can adapt to new tasks with minimal fine-tuning.
16 JUL 2026 · Paper
This paper proposes a new method for advantage shaping in reinforcement learning called Contrastive Policy Optimization (CPO), which uses contrastive disagreement between reference-guided and vanilla generation distributions to indicate correctness. Practitioners might care because it can improve the effectiveness of reinforcement learning methods in generating correct responses.
16 JUL 2026 · Paper
This paper develops a system to extract and represent skills from human-created resources like videos, code, and articles, allowing software agents to learn from these multimodal inputs. Practitioners might care about this work because it could enable more effective training of agents in various domains.
New to GitHub? This beginner's guide explains version control, repositories, and pull requests—plus everything else you need to start working confidently on GitHub. The post GitHub for Beginners: Your roadmap to mastering the GitHub essenti…
Hierarchical Interest Representation is a research area for Meta Ads. We’re exploring an upstream representation layer over the universe of Ads entities – users, advertisers, products, services – learning unified embeddings that connect use…
It’s a quiet week so let’s do the monthly right on schedule.
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
IBM announced preliminary results that spooked the software market generally; this is a story, however, specifically about IBM and its mainframe franchise.
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
15 JUL 2026 · Paper
This paper develops a method to transform video foundation models' representations into compact, reconstruction-capable, and generation-friendly video latents, which can be used in various generative modeling tasks. Practitioners can use VideoRAE to improve the performance of their models by leveraging the semantic and spatio-temporal structure captured by the frozen video foundation encoder.
15 JUL 2026 · Paper
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if k verifier calls all accept it. Under conditionally independent gates, the recent Odds Law (arXiv:2606.1
15 JUL 2026 · Paper
This paper develops a new method for generating high-fidelity 3D images of thin-shell objects, like garments, by learning a continuous surface representation. Practitioners caring about 3D generation and object modeling might care about this approach because it achieves better results with fewer computational resources.
15 JUL 2026 · Paper
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one teacher can specialize
15 JUL 2026 · Paper
This paper introduces Open-AoE, an open dataset and toolchain for egocentric manipulation learning, providing a scalable and structured platform for training embodied models. Practitioners can use Open-AoE to improve their robot learning models, especially those focused on human-robot interaction and embodied intelligence.
a continuation: Codex adding 1M users a day now.
At this year's AIE World’s Fair, AI engineering entered a new phase: building systems around agents, rather than just building with agents.
I previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have changed again.
Google Images is turning 25. Here’s a look back at some major milestones — and new ways to explore and create visual content.
Good news, for once.
When a failed DNSSEC key rollover took down the .al TLD, we deployed a Negative Trust Anchor to restore resolution. This time, though, clients didn't have to take our word for it: 1.1.1.1 returned EDE 33, a new DNS error code that signals d…
LLMs generate code incredibly fast, but to ensure they generate exactly what is intended, they need clear boundaries. Abstractions and Domain-Specific Languages (DSLs) provide a strong harness that guides LLMs right from the start. Unmesh J…
OpenAI has refashioned Codex as the new ChatGPT; is the company abandoning the chat category they pioneered?
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
a quiet day lets us fact check some numbers against the sound of silence of Claude Code reporting...
Introducing Claude for Teachers
Anthropic commits $10 million to Canadian AI research
14 JUL 2026 · Paper
This paper proposes a new type of AI system that can watch and remember things over long periods, and use that memory to reason about the world. Practitioners might care because this could lead to more capable assistants that can learn and adapt over time.
OpenAI’s GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna.
TL; DR At Meta’s scale, a few milliseconds of latency degradation can have a significant negative impact on ads performance. When a Linux kernel upgrade risked regressing latency across Meta’s ad serving fleet, we turned to sche…
Precursor, our new continuous behavioral validation engine for bot management, offers visibility into how humans and bots actually interact across the full user journey. By turning session-level behavior into bot detection signals, it ident…
Some more of my notes from Thoughtworks Future of Software Development Retreat . When we had our first retreat in Utah early this year, nobody had heard of Harness Engineering . This time we had a whole session on it. When comes to the guid…
Google and AIM launched ATL Saathi, a Gemini-powered AI tool empowering Indian educators in robotics labs.
Look at the past history of this blog. There are many blog posts about programming with AI, a few of them date back to January 2024 (like this: https://antirez.com/news/140). I’m a relatively well regarded programmer, after all. I don’t hav…
Apple is suing AI for stealing trade secrets; there is one guilty employee, but this mostly feels like lashing out.
I feel that some vibecoded software changes somewhat randomly and unexpectedly. That made me think about Bruegel’s “The Tower of Babel” which shows an already quite chaotic depiction of the Tower of Babel. The story is usu…
13 JUL 2026 · Paper
This paper introduces Qwen-Music, a music generation model that can produce highly musical and high-fidelity songs from text descriptions and existing songs. Practitioners in music generation and AI may care about this research for its potential to create realistic and engaging music.
13 JUL 2026 · Paper
This paper introduces RAGU, an open-source GraphRAG engine that improves large language models with structured knowledge by separating extraction and consolidation, and trains a compact extractor that outperforms larger models on knowledge-graph construction and GraphRAG tasks. Practitioners might care because RAGU can efficiently generate more accurate and complete context for tasks like factoid-level evidence recall and multi-hop question answering.