Learning from Teacher Continuations at Student States Paper • 2609.36246 • Published 13 days ago • 41
Selecting Diverse SFT Traces Improves Post-RL Generalization Paper • 2609.33780 • Published 14 days ago • 38
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published Sep 3 • 114
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78 • 3
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78