This paper investigates why some AI models can generate excessively long responses and proposes a solution to mitigate this issue by aligning the models' termination tokens. Practitioners might care about this problem because it can lead to wasted generation budgets and inefficient AI applications.
Firehose
Filtered to Papers, tagged “termination-token mismatch” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives