This paper develops a framework to evaluate the social reasoning of large language models (LLMs) in a more realistic setting, by simulating interactions between the LLM and users who provide feedback on the LLM's predictions. Practitioners might care about this research because it helps improve the social reasoning of LLMs, which are increasingly used for advice and decision-making.
Firehose
Filtered to Papers, tagged “verifiable ground truth” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives