SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 6 days ago • 66
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 4 days ago • 133
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation Paper • 2607.09362 • Published 12 days ago • 12
Latent-Identity Tuning in Text-to-Image Personalization Models Paper • 2607.11885 • Published 9 days ago • 14
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Paper • 2607.04461 • Published 17 days ago • 10
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published 9 days ago • 43
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published 9 days ago • 22
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 9 days ago • 83
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published 13 days ago • 86
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 8 days ago • 102
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence Paper • 2607.06838 • Published 15 days ago • 14
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Paper • 2606.30124 • Published 23 days ago • 8
Multiplayer Interactive World Models with Representation Autoencoders Paper • 2607.05352 • Published 16 days ago • 28
Perceptual Flow Matching for Few-Step Generative Modeling Paper • 2607.03524 • Published 19 days ago • 18
InstanceControl: Controllable Complex Image Generation without Instance Labeling Paper • 2606.31924 • Published 22 days ago • 15
Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment Paper • 2607.02920 • Published 19 days ago • 4
CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training Paper • 2607.02998 • Published 15 days ago • 6
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling Paper • 2607.02980 • Published 19 days ago • 79