JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper ⢠2608.03974 ⢠Published 4 days ago ⢠84
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models Paper ⢠2606.17539 ⢠Published Jun 16 ⢠15
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics Paper ⢠2606.09826 ⢠Published Jun 8 ⢠19
Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing Paper ⢠2606.05172 ⢠Published Apr 16 ⢠1
LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation Paper ⢠2606.02553 ⢠Published Jun 1 ⢠20
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation Paper ⢠2605.18739 ⢠Published May 18 ⢠116
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Paper ⢠2604.24954 ⢠Published Apr 27 ⢠26
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper ⢠2604.22748 ⢠Published Apr 24 ⢠232
Efficient-Large-Model/Sana_Sprint_1.6B_1024px_teacher Text-to-Image ⢠Updated 7 days ago ⢠35 ⢠1
Efficient-Large-Model/Sana_Sprint_1.6B_1024px_diffusers Text-to-Image ⢠Updated 7 days ago ⢠⢠26