Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision Paper • 2609.07099 • Published 16 days ago • 3
Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model Paper • 2609.07154 • Published 16 days ago • 3
Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision Paper • 2609.07099 • Published 16 days ago • 3
Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model Paper • 2609.07154 • Published 16 days ago • 3
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published Aug 13 • 16
Building Bridges: A Dataset for Evaluating Gender-Fair Machine Translation into German Paper • 2406.06131 • Published Jun 10, 2024
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding Paper • 2506.06275 • Published Jun 6, 2025
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios Paper • 2606.06177 • Published Jun 4
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese Paper • 2603.26511 • Published Mar 27
BabyBabelLM: A Multilingual Benchmark of Developmentally Plausible Training Data Paper • 2510.10159 • Published Oct 11, 2025 • 3
Measuring what Matters: Construct Validity in Large Language Model Benchmarks Paper • 2511.04703 • Published Nov 3, 2025 • 8
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources Paper • 2509.25531 • Published Sep 29, 2025 • 11
BigCodeArena: Unveiling More Reliable Human Preferences in Code Generation via Execution Paper • 2510.08697 • Published Oct 9, 2025 • 41
BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications Paper • 2509.24908 • Published Sep 29, 2025 • 3
EmbeddingGemma: Powerful and Lightweight Text Representations Paper • 2509.20354 • Published Sep 24, 2025 • 51
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings Paper • 2509.14405 • Published Sep 17, 2025 • 2