Firehose

Filtered to Papers · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

17 SEP 2026 · Paper

This paper investigates how different components of coding harnesses, such as planning, action space, and context management, impact the performance of autonomous coding agents in software engineering tasks. Practitioners might care about understanding how to design harnesses that effectively utilize these components to improve agent performance.

17 SEP 2026 · Paper

This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.

17 SEP 2026 · Paper

This paper optimizes the performance of large language models (LLMs) on fine-grained visual perception tasks by learning to selectively focus on relevant regions of the image, rather than relying on high-resolution visual encoding. By doing so, it can improve accuracy with fewer visual tokens, making it more efficient and effective for real-world applications.

17 SEP 2026 · Paper

This paper evaluates the ability of general-purpose models to understand and act on spatial intelligence through visual demonstrations, active perception, and metric control. Practitioners might care about this research because it can help develop models that can effectively navigate and interact with their environment.

17 SEP 2026 · Paper

This paper develops a new framework, JEPA-Anything, that enables predictive models to work across different domains and systems, allowing for world modeling and learning from interaction. Practitioners might care about this because it could lead to more generalizable and versatile AI models.

17 SEP 2026 · Paper

This paper introduces a new method for video generation, called Video DeltaNet, which combines attention mechanisms to improve efficiency and quality in livestream video generation. Practitioners may care about this paper if they work on video generation tasks and want to explore more efficient and effective methods.

17 SEP 2026 · Paper

This paper proposes a framework called UFO to evaluate the alignment of multi-modal image generation models, which is crucial for achieving consistency with human judgments. Practitioners might care about this paper because it offers a more comprehensive approach to evaluating multi-modal image generation models.

17 SEP 2026 · Paper

This paper introduces DeepSeek-V4.1-Flash, a more efficient model that reduces the computational cost of long-horizon agents by optimizing its prefill process and cache compression. Practitioners can benefit from this model's improved performance and reduced storage needs for agentic workloads.

17 SEP 2026 · Paper

This paper investigates whether giving a language model extra information, such as a worked solution, improves its learning through on-policy self-distillation. A practitioner might care about how to optimize this technique for better performance.

17 SEP 2026 · Paper

This paper proposes a method to control the length of large reasoning models to make them more efficient, by learning when to use less computation for easy problems and more for hard ones. Practitioners might care about this because it can help improve the accuracy-efficiency trade-offs in complex tasks.

17 SEP 2026 · Paper

This paper proposes a new method for self-retiring on-policy distillation in agentic reinforcement learning, which allows agents to learn more effectively by switching between teacher-student training and reinforcement learning alone. Practitioners might care about this approach because it can improve the performance of agents in complex tasks.

17 SEP 2026 · Paper

This paper develops a new framework called WeVisDoc to improve the performance of document parsing systems by addressing their weaknesses in diverse layouts and acquisition conditions. Practitioners may care about this research if they want to build robust document parsing systems that can handle real-world challenges.

17 SEP 2026 · Paper

This paper investigates why some AI models can generate excessively long responses and proposes a solution to mitigate this issue by aligning the models' termination tokens. Practitioners might care about this problem because it can lead to wasted generation budgets and inefficient AI applications.

17 SEP 2026 · Paper

This paper develops a feed-forward model that can predict the articulation of objects from sparse, unordered point cloud observations, allowing it to learn from multiple views and generalize to new inputs. Practitioners in computer vision and robotics may care about this model as it addresses the challenge of modeling articulated objects from limited and incomplete observations.

17 SEP 2026 · Paper

This paper explores whether supervising both agent actions and environment observations during reinforcement learning improves agent exploration and performance. Practitioners might care because it could lead to better initialization for reinforcement learning tasks.

17 SEP 2026 · Paper

This paper creates a benchmark for Telugu spoken question answering, allowing researchers to evaluate models that can understand and respond to questions in Telugu, a high-resource but understudied language. Practitioners working on natural language processing for low-resource languages might care about this work to develop more accurate and culturally sensitive models.

17 SEP 2026 · Paper

This paper proposes a framework called SELF-INDEX that allows an index to automatically improve its performance without human intervention, leading to better information retrieval for complex tasks and benefiting applications such as search agents and agent memory systems.

16 SEP 2026 · Paper

This paper evaluates the trustworthiness of enterprise AI assistants in high-pressure situations, such as hiring, healthcare, and finance, where compliance with rules is crucial. Practitioners should care about this research to ensure their AI assistants are reliable and transparent in complex decision-making scenarios.

16 SEP 2026 · Paper

This paper evaluates whether a multimodal generative model can reason about the physical world by testing its ability to integrate information from different modalities, such as text, images, video, and audio. Practitioners might care about this research because it can help develop more advanced models that can better understand and generate complex scenarios.

16 SEP 2026 · Paper

This paper investigates how the way large language models generate multiple candidate responses affects their performance and energy consumption. Practitioners might care because optimizing test-time scaling can lead to significant improvements in model accuracy and efficiency.

16 SEP 2026 · Paper

This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.

16 SEP 2026 · Paper

This paper investigates whether people's gaze patterns can reveal how they understand each other in collaborative tasks, and whether this understanding is related to the success of the task. Practitioners working on human-robot collaboration or other tasks with asymmetric information might care about this research because it could help them design better interfaces that take into account how people communicate with each other.

16 SEP 2026 · Paper

This paper creates a system called ScienceIDE that converts scientific code into environments that can be used to train artificial agents to perform scientific tasks. Practitioners might care because this could lead to more efficient and effective ways to develop scientific intelligence.

16 SEP 2026 · Paper

This paper introduces Agora, a system that uses Git to enable collective auto-research by sharing and versioning research results among multiple agents, allowing them to build upon each other's work and avoid duplicated search. Practitioners might care about this because it could lead to more efficient and effective research in areas like AI and machine learning.

16 SEP 2026 · Paper

This paper proposes a new method for aligning large language models with human preferences, called Comparison-based Preference Optimization (ComPO), which is more efficient than existing methods and can mitigate a problem called likelihood displacement. Practitioners might care about this paper because it offers a new approach to aligning LLMs with human preferences, which is essential for developing more reliable and trustworthy AI models.

16 SEP 2026 · Paper

This paper improves autoregressive vision-language-action models by creating a new method for action tokenization that better preserves the relationships between actions, allowing the model to perform more accurately in different contexts. Practitioners might care about this because it could lead to more reliable and generalizable vision-language-action models.

16 SEP 2026 · Paper

This paper investigates a common problem in reinforcement learning for language models called Value Flattening, where critics fail to accurately estimate state values, and proposes a new method, SP^3O, to mitigate this issue by supervising only a few well-separated states per response.

16 SEP 2026 · Paper

This paper develops a new method for image captioning that also grounds each phrase with a specific region of the image, allowing for more accurate and detailed descriptions. Practitioners might care about this work if they're building AI systems that need to understand and interact with the physical world.

16 SEP 2026 · Paper

This paper presents a new technique to reduce memory usage and speed up inference for large neural networks, allowing them to run on consumer hardware with limited memory. A practitioner might care about this because it enables the deployment of large models in edge devices and reduces the need for expensive storage.

16 SEP 2026 · Paper

This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.

16 SEP 2026 · Paper

This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.

15 SEP 2026 · Paper

This paper introduces a benchmark for restoring obfuscated platform messages and investigating associated websites to combat online abuse, with the goal of improving the accuracy of risk reports and user safety. Practitioners might care about this research if they want to develop more effective tools to detect and mitigate online threats.

15 SEP 2026 · Paper

This paper proposes a framework for training-free skill evolution in GUI agents, allowing them to adapt to dynamic graphical user interfaces without additional training. Practitioners may care about this research if they develop GUI agents that need to handle changing environments.

15 SEP 2026 · Paper

This paper develops a framework to evaluate the social reasoning of large language models (LLMs) in a more realistic setting, by simulating interactions between the LLM and users who provide feedback on the LLM's predictions. Practitioners might care about this research because it helps improve the social reasoning of LLMs, which are increasingly used for advice and decision-making.

15 SEP 2026 · Paper

This paper proposes a new method for 3D hand mesh reconstruction from egocentric event-based cameras, which can handle low-light conditions and motion blur, and provides more accurate hand information and inter-hand relationships than previous approaches.

15 SEP 2026 · Paper

This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.

15 SEP 2026 · Paper

This paper creates a new type of AI model that can generate interactive worlds, allowing users to explore, control events, and provide feedback through text and keyboard input. Practitioners in AI development might care about this research because it could lead to more engaging and interactive AI experiences.

15 SEP 2026 · Paper

This paper proposes a new method for estimating confidence in language models, called XConf, which uses the model's past experiences to inform its confidence, rather than just relying on the current inference process. Practitioners might care about this because it could lead to more reliable and trustworthy deployment of language models.

15 SEP 2026 · Paper

This paper introduces LimiX-2, a new AI model that uses a new paradigm called Contextual Mechanism Networks (CMNs) to learn from structured data. Practitioners might care about this because it could lead to more accurate and causal AI models.

15 SEP 2026 · Paper

This paper introduces Fathom, a technique to speed up decoding in large language models by selectively reading only the relevant parts of the key-value cache, reducing the computational cost and memory access. Practitioners in the field of natural language processing and deep learning may care about optimizing decoding efficiency for large models.

15 SEP 2026 · Paper

This paper teaches a robotic hand to walk, support itself, and interact with its environment using its fingers, without needing a separate locomotion system. A practitioner might care about this research because it could lead to more compact and versatile robots that can perform multiple tasks.

15 SEP 2026 · Paper

This paper introduces ScienceBuddy, a tool that helps researchers work with intelligent agents that can learn and improve on their own, and how this can lead to new discoveries and advancements in scientific research. Practitioners might care because it could revolutionize the way scientists work with AI.

15 SEP 2026 · Paper

This paper explores how AI can be applied across different stages of game development, from playing games to designing and testing them, and how to reuse capabilities across these stages. Practitioners might care about how to apply AI to improve game development efficiency and effectiveness.

15 SEP 2026 · Paper

PhysStream is a video generation model that can control and manipulate dynamic scenes in a physically meaningful way, allowing for fine-grained control over motion and object placement. This can be useful for interactive applications where the generated video needs to be adjusted in real-time.

15 SEP 2026 · Paper

This paper develops a new approach to world-action models that can effectively combine multiple visual modalities, such as depth and point tracks, to improve performance. Practitioners in robotics and AI might care about this research because it could lead to more accurate and robust models for tasks like grasping and manipulation.

15 SEP 2026 · Paper

This paper tests how well AI agents can withstand prolonged interactions and unexpected events, and finds that even seemingly safe agents can fail in complex, long-term scenarios. Practitioners should care because it highlights the need to design more resilient autonomous systems that can handle unexpected failures.

15 SEP 2026 · Paper

This paper tests the robustness of rubrics generated by language models as reward signals in reinforcement learning, finding that even generic rubrics can be exploited 64% of the time, while tailored rubrics can be used to create fake answers. Practitioners should care because this can lead to biased grading and evaluation.

15 SEP 2026 · Paper

This paper proposes a new framework for joint multimodal representation learning and generation, allowing for flexible-length aligned transmodal tokens that can be used for both retrieval and generation tasks. Practitioners might care about this paper because it shows how to improve generative performance by training a shared multimodal encoder alongside downstream models.

14 SEP 2026 · Paper

This paper investigates whether diffusion language models can continue reasoning across generation chunks without keeping earlier text in context, and whether using a fixed-size "register" can improve performance. Practitioners might care about this because it could lead to more efficient and flexible language generation models.

14 SEP 2026 · Paper

This paper improves the performance and efficiency of diffusion transformers, a type of AI model used for video generation, by reducing the computational cost of attention mechanisms. Practitioners caring about accelerating AI models on hardware can benefit from this research.

14 SEP 2026 · Paper

This paper introduces HypoEvolve, a framework that uses genetic algorithms to enable multi-agent LLMs to discover scientific hypotheses by collaborating on hypothesis synthesis, evaluation, and revision. Practitioners might care about this because it could lead to more effective AI systems for scientific discovery and drug repurposing.

14 SEP 2026 · Paper

This paper evaluates the performance of a neural network-based tumor segmentation model on a diverse dataset of brain tumor images. Practitioners in the field of medical imaging may care about the findings as they could inform the development of more robust and generalizable models for tumor segmentation.

14 SEP 2026 · Paper

This paper proposes a method to automatically select skills for a large language model (LLM) without requiring explicit skill text in the context, allowing for more efficient and accurate skill routing. Practitioners may care about this approach as it could lead to improved performance and reduced model size in applications where skill selection is critical.

14 SEP 2026 · Paper

This paper explores how transformer representations change over time and how these changes can be understood and manipulated. Practitioners might care about this research because it could lead to more robust and efficient transformer models, especially in applications where model updates need to be edited or compressed.

14 SEP 2026 · Paper

This paper proposes a way to make language models understand and respond to users' mental states, so they can better collaborate with humans in the long term. Practitioners might care about this because it could lead to more effective AI assistants that can support people's goals and needs.

14 SEP 2026 · Paper

This paper introduces a fast and efficient post-hoc defense against a type of attack that can bypass safety features in language models, allowing the model to continue functioning but with compromised security. Practitioners caring about model security may be interested in this approach as it can provide an additional layer of protection without requiring significant computational resources.

14 SEP 2026 · Paper

This paper introduces HarnessVLN, a training-free framework for embodied navigation that uses a unified tool interface to validate proposed actions against spatial evidence and task progress, allowing agents to generalize and learn from multimodal large language models.

14 SEP 2026 · Paper

This paper proposes a framework for generalizable recursive self-improvement (RSI) of agent harnesses, which can improve execution mechanisms without being specific to a particular task or benchmark. Practitioners can care about this work because it aims to create more adaptable and transferable AI agents.

14 SEP 2026 · Paper

This paper proposes a new approach to handling streaming omni-modal models, called Omni-Streaming Thinking (OST), which helps prevent models from prematurely committing to interpretations based on incomplete audio or visual information. A practitioner might care about this because it can lead to more accurate and reliable responses in real-time applications.

14 SEP 2026 · Paper

This paper introduces Dream-RSI, a framework for recursive self-improvement in exploration, which helps autonomous AI agents discover high-value solutions more efficiently by using a replay simulator to provide low-cost feedback. Practitioners might care because effective exploration is crucial for AI progress, and Dream-RSI can improve discovery quality and reduce costs.