Firehose

Filtered to tagged “benchmarking” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

17 SEP 2026 · Paper

This paper develops a method to efficiently scale agent research loops, allowing for more effective self-improvement and reusable improvements across diverse environments. Practitioners might care about this research because it could lead to significant cost savings and improved performance in automated code completion and generation tasks.

16 SEP 2026 · Paper

This paper introduces ProgramDistill, a benchmark that evaluates coding agents on their ability to infer behavior from working software and implement it in an incomplete application. Practitioners in AI/ML and web development might care about this work because it provides a scalable and controlled benchmark for evaluating and training coding agents.