AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
This paper proposes a method called AgentOPSD for turn-level credit assignment in agentic reinforcement learning, which helps to identify pivotal decisions that determine outcomes in long-horizon tasks. Practitioners may care about this because it can improve the performance of reinforcement learning agents in complex tasks.