MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers Paper • 2610.06801 • Published 7 days ago • 37
GRACE: Generation-aware latent compression for efficient video generation Paper • 2610.10524 • Published 5 days ago • 76
SGF+: Decoupling Gradient Flows for Autoregressive Video Generation Paper • 2610.10429 • Published 5 days ago • 55
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 13 days ago • 140
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 13 days ago • 553
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 17 days ago • 145
facebook/dinov3-vitb16-pretrain-lvd1689m Image Feature Extraction • 85.7M • Updated Aug 19, 2025 • 484k • 414
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 25 days ago • 139
What Does Privileged Information Add to On-Policy Self-Distillation? Paper • 2609.20612 • Published 25 days ago • 36
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published Sep 10 • 657
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 28 days ago • 33
EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents Paper • 2609.01281 • Published Sep 1 • 15
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation Paper • 2608.05879 • Published Aug 29 • 9
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience Paper • 2609.03241 • Published Sep 3 • 51
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published Sep 3 • 71