Firehose

Filtered to Papers, tagged “Continual learning” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

17 SEP 2026 · Paper

This paper investigates how different components of coding harnesses, such as planning, action space, and context management, impact the performance of autonomous coding agents in software engineering tasks. Practitioners might care about understanding how to design harnesses that effectively utilize these components to improve agent performance.

17 SEP 2026 · Paper

This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.

17 SEP 2026 · Paper

This paper optimizes the performance of large language models (LLMs) on fine-grained visual perception tasks by learning to selectively focus on relevant regions of the image, rather than relying on high-resolution visual encoding. By doing so, it can improve accuracy with fewer visual tokens, making it more efficient and effective for real-world applications.

17 SEP 2026 · Paper

This paper develops a new framework, JEPA-Anything, that enables predictive models to work across different domains and systems, allowing for world modeling and learning from interaction. Practitioners might care about this because it could lead to more generalizable and versatile AI models.

17 SEP 2026 · Paper

This paper proposes a method to control the length of large reasoning models to make them more efficient, by learning when to use less computation for easy problems and more for hard ones. Practitioners might care about this because it can help improve the accuracy-efficiency trade-offs in complex tasks.

17 SEP 2026 · Paper

This paper proposes a new method for self-retiring on-policy distillation in agentic reinforcement learning, which allows agents to learn more effectively by switching between teacher-student training and reinforcement learning alone. Practitioners might care about this approach because it can improve the performance of agents in complex tasks.

17 SEP 2026 · Paper

This paper develops a new framework called WeVisDoc to improve the performance of document parsing systems by addressing their weaknesses in diverse layouts and acquisition conditions. Practitioners may care about this research if they want to build robust document parsing systems that can handle real-world challenges.

17 SEP 2026 · Paper

This paper investigates why some AI models can generate excessively long responses and proposes a solution to mitigate this issue by aligning the models' termination tokens. Practitioners might care about this problem because it can lead to wasted generation budgets and inefficient AI applications.

17 SEP 2026 · Paper

This paper proposes a framework called SELF-INDEX that allows an index to automatically improve its performance without human intervention, leading to better information retrieval for complex tasks and benefiting applications such as search agents and agent memory systems.

16 SEP 2026 · Paper

This paper evaluates whether a multimodal generative model can reason about the physical world by testing its ability to integrate information from different modalities, such as text, images, video, and audio. Practitioners might care about this research because it can help develop more advanced models that can better understand and generate complex scenarios.

16 SEP 2026 · Paper

This paper investigates how the way large language models generate multiple candidate responses affects their performance and energy consumption. Practitioners might care because optimizing test-time scaling can lead to significant improvements in model accuracy and efficiency.

16 SEP 2026 · Paper

This paper introduces Agora, a system that uses Git to enable collective auto-research by sharing and versioning research results among multiple agents, allowing them to build upon each other's work and avoid duplicated search. Practitioners might care about this because it could lead to more efficient and effective research in areas like AI and machine learning.

16 SEP 2026 · Paper

This paper develops a framework for robots to learn from context without relying on pre-programmed demonstrations, allowing them to adapt to new environments. Practitioners might care because this technology could enable robots to perform tasks more efficiently and effectively in real-world situations.

16 SEP 2026 · Paper

This paper proposes a new framework for Mixture-of-Agents that allows query routing and agent fine-tuning to evolve together, improving the ability of agents to adapt to changing capabilities. Practitioners might care about this approach because it can lead to more efficient and effective data-driven specialization in complex tasks.

15 SEP 2026 · Paper

This paper introduces a benchmark for restoring obfuscated platform messages and investigating associated websites to combat online abuse, with the goal of improving the accuracy of risk reports and user safety. Practitioners might care about this research if they want to develop more effective tools to detect and mitigate online threats.

15 SEP 2026 · Paper

This paper proposes a framework for training-free skill evolution in GUI agents, allowing them to adapt to dynamic graphical user interfaces without additional training. Practitioners may care about this research if they develop GUI agents that need to handle changing environments.

15 SEP 2026 · Paper

This paper proposes a new method for 3D hand mesh reconstruction from egocentric event-based cameras, which can handle low-light conditions and motion blur, and provides more accurate hand information and inter-hand relationships than previous approaches.

15 SEP 2026 · Paper

This paper introduces EvolveTrade, a self-evolving framework that allows large language model trading agents to refine their policies over time, enabling them to adapt to changing market regimes and improve their performance. Practitioners in finance and AI may care about this research as it provides a way to build more robust and adaptive trading agents.

15 SEP 2026 · Paper

This paper proposes a new method for estimating confidence in language models, called XConf, which uses the model's past experiences to inform its confidence, rather than just relying on the current inference process. Practitioners might care about this because it could lead to more reliable and trustworthy deployment of language models.

15 SEP 2026 · Paper

This paper introduces ScienceBuddy, a tool that helps researchers work with intelligent agents that can learn and improve on their own, and how this can lead to new discoveries and advancements in scientific research. Practitioners might care because it could revolutionize the way scientists work with AI.

15 SEP 2026 · Paper

This paper explores how AI can be applied across different stages of game development, from playing games to designing and testing them, and how to reuse capabilities across these stages. Practitioners might care about how to apply AI to improve game development efficiency and effectiveness.

15 SEP 2026 · Paper

This paper develops a new approach to world-action models that can effectively combine multiple visual modalities, such as depth and point tracks, to improve performance. Practitioners in robotics and AI might care about this research because it could lead to more accurate and robust models for tasks like grasping and manipulation.

15 SEP 2026 · Paper

This paper tests how well AI agents can withstand prolonged interactions and unexpected events, and finds that even seemingly safe agents can fail in complex, long-term scenarios. Practitioners should care because it highlights the need to design more resilient autonomous systems that can handle unexpected failures.

15 SEP 2026 · Paper

This paper proposes a new framework for joint multimodal representation learning and generation, allowing for flexible-length aligned transmodal tokens that can be used for both retrieval and generation tasks. Practitioners might care about this paper because it shows how to improve generative performance by training a shared multimodal encoder alongside downstream models.

14 SEP 2026 · Paper

This paper investigates whether diffusion language models can continue reasoning across generation chunks without keeping earlier text in context, and whether using a fixed-size "register" can improve performance. Practitioners might care about this because it could lead to more efficient and flexible language generation models.

14 SEP 2026 · Paper

This paper introduces HypoEvolve, a framework that uses genetic algorithms to enable multi-agent LLMs to discover scientific hypotheses by collaborating on hypothesis synthesis, evaluation, and revision. Practitioners might care about this because it could lead to more effective AI systems for scientific discovery and drug repurposing.

14 SEP 2026 · Paper

This paper proposes a method to automatically select skills for a large language model (LLM) without requiring explicit skill text in the context, allowing for more efficient and accurate skill routing. Practitioners may care about this approach as it could lead to improved performance and reduced model size in applications where skill selection is critical.

14 SEP 2026 · Paper

This paper proposes a way to make language models understand and respond to users' mental states, so they can better collaborate with humans in the long term. Practitioners might care about this because it could lead to more effective AI assistants that can support people's goals and needs.

14 SEP 2026 · Paper

This paper introduces a fast and efficient post-hoc defense against a type of attack that can bypass safety features in language models, allowing the model to continue functioning but with compromised security. Practitioners caring about model security may be interested in this approach as it can provide an additional layer of protection without requiring significant computational resources.

14 SEP 2026 · Paper

This paper proposes a framework for generalizable recursive self-improvement (RSI) of agent harnesses, which can improve execution mechanisms without being specific to a particular task or benchmark. Practitioners can care about this work because it aims to create more adaptable and transferable AI agents.

14 SEP 2026 · Paper

This paper proposes a new approach to handling streaming omni-modal models, called Omni-Streaming Thinking (OST), which helps prevent models from prematurely committing to interpretations based on incomplete audio or visual information. A practitioner might care about this because it can lead to more accurate and reliable responses in real-time applications.