FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Paper • 2605.27284 • Published May 26 • 9
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Paper • 2510.10396 • Published Oct 12, 2025
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 10 days ago • 155
FineVLA: Fine-Grained Instruction Alignment For VLA Collection This is the collection of FineVLA, including the RoboFine-Bench RoboFine-VLM and FineVLA-policyLA • 3 items • Updated Jun 4 • 1
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies Paper • 2605.27284 • Published May 26 • 9
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Paper • 2605.30993 • Published May 29 • 65
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146
FineVLA: Fine-Grained Instruction Alignment For VLA Collection This is the collection of FineVLA, including the RoboFine-Bench RoboFine-VLM and FineVLA-policyLA • 3 items • Updated Jun 4 • 1
FineVLA: Fine-Grained Instruction Alignment For VLA Collection This is the collection of FineVLA, including the RoboFine-Bench RoboFine-VLM and FineVLA-policyLA • 3 items • Updated Jun 4 • 1