JevAny · Adaptive Decision Systems Collection Calibration-aware RL for typed decisions, agents, native multimodal evidence, test-time adaptation, and symbolic harnesses. • 2 items • Updated 2 days ago • 1
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 4 days ago • 16
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 4 days ago • 16
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 4 days ago • 16
💻 Qwopus-Coder Collection Reasoning-distilled coding models optimized for specialized domains like agentic workflows. • 10 items • Updated 21 days ago • 55
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116
To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models Paper • 2602.12566 • Published Feb 13 • 1 • 1
To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Models Paper • 2602.12566 • Published Feb 13 • 1
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use Paper • 2510.05592 • Published Oct 7, 2025 • 113
Nemotron-Post-Training-v3 Collection Collection of datasets used in the post-training phase of Nemotron Nano, Super, and Ultra v3. • 50 items • Updated Aug 11 • 202
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116 • 6
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation Paper • 2604.13010 • Published Apr 14 • 20 • 7
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116 • 6
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation Paper • 2604.13010 • Published Apr 14 • 20 • 7