Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published 25 days ago • 29
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence Paper • 2608.10720 • Published 22 days ago • 16
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published Jul 30 • 73
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published Jul 7 • 38
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments Paper • 2607.02440 • Published Jul 2 • 51
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects Paper • 2605.19587 • Published May 19 • 10
World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning Paper • 2606.03603 • Published Jun 2 • 31
World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning Paper • 2606.03603 • Published Jun 2 • 31
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling Paper • 2605.13301 • Published May 13 • 167
Running Agents Featured 58 Hy3-preview ⚡ 58 Hy3-preview multi-turn streaming chat with function calling
GEMS: Agent-Native Multimodal Generation with Memory and Skills Paper • 2603.28088 • Published Mar 30 • 88
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Paper • 2602.10604 • Published Feb 11 • 201