Firehose

Filtered to tagged “enterprise” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

16 SEP 2026 · Paper

This paper evaluates the trustworthiness of enterprise AI assistants in high-pressure situations, such as hiring, healthcare, and finance, where compliance with rules is crucial. Practitioners should care about this research to ensure their AI assistants are reliable and transparent in complex decision-making scenarios.

12 SEP 2026 · Hacker News · 274 pts · 156 comments ↗

The Real-SWE benchmark evaluates AI models on private, real-world, enterprise codebases, challenging their ability to navigate complex business logic and company-specific coding patterns. The benchmark features 8 tasks from private production codebases, each with 8 independent runs per model, resulting in a resolution rate of 15% or lower for most models, with Fable 5.1 achieving the highest resolution rate of 38.8%. The cost of running the benchmark varies from $2.50 to $6.96 per rollout, with Gemini 3.8 Flash and GPT-5.6 Sol being the most cost-effective models. AI summary