This paper evaluates the trustworthiness of enterprise AI assistants in high-pressure situations, such as hiring, healthcare, and finance, where compliance with rules is crucial. Practitioners should care about this research to ensure their AI assistants are reliable and transparent in complex decision-making scenarios.
Firehose
Filtered to tagged “enterprise” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
artificial intelligence 87continual learning 32AI 24reinforcement learning 14agentic coding 13AI safety 13open-weight models 13AI agents 10existential risk 9AI ethics 8cybersecurity 8ethics 7language models 7machine learning 7natural language processing 6open-source 6Reinforcement learning 6security 6artificial general intelligence 5Diffusion models 5recursive self-improvement 5robotics 5software development 5Agentic AI 4large language models 4mathematics 4multi-agent systems 4Recursive self-improvement 4agentic AI 3agents 3
The Real-SWE benchmark evaluates AI models on private, real-world, enterprise codebases, challenging their ability to navigate complex business logic and company-specific coding patterns. The benchmark features 8 tasks from private production codebases, each with 8 independent runs per model, resulting in a resolution rate of 15% or lower for most models, with Fable 5.1 achieving the highest resolution rate of 38.8%. The cost of running the benchmark varies from $2.50 to $6.96 per rollout, with Gemini 3.8 Flash and GPT-5.6 Sol being the most cost-effective models. AI summary