Datasets:
benchmark_id stringlengths 12 103 | slug stringlengths 3 103 | name stringlengths 2 118 | source stringclasses 5
values | source_url stringlengths 19 110 | description stringlengths 0 995 | categories listlengths 0 21 | languages listlengths 0 13 | modality stringclasses 5
values | publisher stringclasses 315
values | released_at timestamp[s]date 2010-02-19 00:00:00 2026-09-13 00:00:00 ⌀ | openness stringclasses 3
values | paper_url stringlengths 31 62 ⌀ | repo_url stringlengths 28 117 ⌀ | dataset_url stringclasses 395
values | document_count int64 1 29 ⌀ | model_count int64 1 586 ⌀ | score_count int64 0 586 | highest_score float64 0.01 2.1M ⌀ | score_unit stringclasses 3
values |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
claire-radar:2609.04611 | claire-radar-2609-04611 | $\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction | claire_radar | https://arxiv.org/abs/2609.04611 | A benchmark where a developer agent builds a complete customer-service agent from business records, client requirements, a production API, inherited codebase, and cost/model limits, scored by deployment against held-out simulated users across 53 tasks in four domains. | [
"Agents",
"Agents & Tool Use",
"General AI"
] | [] | null | null | 2026-09-04T00:00:00 | unknown | https://arxiv.org/abs/2609.04611 | null | null | 1 | null | 0 | null | null |
claire-radar:2608.08814 | claire-radar-2608-08814 | 360CityArena | claire_radar | https://arxiv.org/abs/2608.08814 | 360CityArena evaluates embodied agents in a photorealistic virtual urban environment built from 360-degree video of Tokyo's Akihabara district. It includes 175 tasks across environment understanding, path reasoning, and spatial reasoning, testing localization, landmark search, path planning, and relational spatial reas... | [
"Robotics & Embodied AI",
"cs.CV"
] | [] | null | null | 2026-08-09T00:00:00 | unknown | https://arxiv.org/abs/2608.08814 | https://github.com/360MM-Team/360CityArena | null | 1 | null | 0 | null | null |
claire-radar:2606.01057 | claire-radar-2606-01057 | 3DCodeBench | claire_radar | https://arxiv.org/abs/2606.01057 | 3DCodeBench evaluates vision-language model agents on procedural 3D modeling by converting text and image references into Blender Python code. It includes 212 object categories with ground-truth scripts, and scores outputs on executability, image similarity (SigLIP-2/DINOv3), 3D shape distance (Chamfer/Uni3D), and LLM-... | [
"General AI",
"Geometric reasoning",
"Multimodal"
] | [] | null | null | 2026-05-31T00:00:00 | unknown | https://arxiv.org/abs/2606.01057 | https://github.com/gaoypeng/3dcodebench | null | 1 | null | 0 | null | null |
claire-radar:2608.26947 | claire-radar-2608-26947 | 4DSynth-Nav | claire_radar | https://arxiv.org/abs/2608.26947 | Evaluates embodied agents on interactive navigation tasks in procedurally generated 4D environments with independently tunable difficulty axes. | [
"Robotics & Embodied AI",
"cs.RO"
] | [] | null | null | 2026-08-27T00:00:00 | unknown | https://arxiv.org/abs/2608.26947 | null | null | 1 | null | 0 | null | null |
claire-radar:2609.09964 | claire-radar-2609-09964 | 5-Dialects-BN | claire_radar | https://arxiv.org/abs/2609.09964 | Evaluation object not clearly defined; abstract mentions alignment of Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five Bangla regional varieties, but no explicit task or scoring setup is provided. | [
"General AI",
"Language & Knowledge",
"cs.CL"
] | [] | null | null | 2026-09-09T00:00:00 | unknown | https://arxiv.org/abs/2609.09964 | null | null | 1 | null | 0 | null | null |
claire-radar:2605.24045 | claire-radar-2605-24045 | A Large-Scale Dataset and Benchmark | claire_radar | https://arxiv.org/abs/2605.24045 | InteractBind provides a large-scale dataset of ~100k protein-ligand pairs with fine-grained binding-site localization tasks and interaction maps for six non-covalent interaction types, plus affinity and similarity-controlled splits. | [
"Biology & Drug Discovery",
"Language & Knowledge",
"cs.LG"
] | [] | null | null | 2026-05-21T00:00:00 | unknown | https://arxiv.org/abs/2605.24045 | null | null | 1 | null | 0 | null | null |
opencompass:1580 | opencompass-1580-a-bench | A-Bench | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/A-Bench | A-Bench is a benchmark designed to diagnose whether LMMs are masters at evaluating AIGIs. 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts, and tested across 18 leading LMMs. A-Bench是一个旨在诊断 LMMs 是否擅长评估 AIGIs 的基准,从 16 个文本到图像模型中采样了 2,864 个 AIGIs,每个都与由人类专家标... | [
"多模态",
"Multimodal",
"多模态模型",
"VLM",
"视觉生成",
"Visual Generation",
"图像理解",
"Image Understanding",
"不支持",
"Unsupported"
] | [] | multimodal | SJTU, NTU. | 2024-06-05T00:00:00 | unknown | https://arxiv.org/abs/2406.03070 | https://github.com/Q-Future/A-Bench | https://huggingface.co/datasets/q-future/A-Bench | 1 | null | 0 | null | null |
opencompass:1367 | opencompass-1367-a-okvqa | A-OKVQA | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/A-OKVQA | A-OKVQA assesses commonsense reasoning abilities. It is a crowdsourced dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer. A-OKVQA用于评估多模态大模型的常识及推理能力,由25K个不同的问题组成,需要对图像中描述的场景进行某种形式的常识性推理来回答。 | [
"多模态",
"Multimodal",
"推理",
"Reasoning",
"知识",
"Knowledge",
"VQA",
"多模态模型",
"VLM",
"逻辑推理",
"知识储备",
"不支持",
"Unsupported"
] | [] | multimodal | Allen Institute for AI | 2022-06-03T00:00:00 | unknown | https://arxiv.org/abs/2206.01718 | https://github.com/allenai/aokvqa | null | 1 | null | 0 | null | null |
claire-radar:2605.22321 | claire-radar-2605-22321 | A3S-Bench | claire_radar | https://arxiv.org/abs/2605.22321 | A3S-Bench is an executable security test suite for evaluating autonomous agents across multi-turn scenarios and agent-level risks using an action-grounded oracle. | [
"Agents",
"Agents & Tool Use",
"General AI"
] | [] | null | null | 2026-05-21T00:00:00 | unknown | https://arxiv.org/abs/2605.22321 | https://github.com/antgroup/Agent3Sigma-Stage | null | 1 | null | 0 | null | null |
artificial-analysis:aa-analystagent | artificial-analysis-aa-analystagent | AA-AnalystAgent | artificial_analysis | https://artificialanalysis.ai/evaluations/aa-analyst-agent | Quantitative analysis on spreadsheets & documents | [
"agentic",
"business",
"reasoning"
] | [] | null | null | null | unknown | null | null | null | 1 | 30 | 30 | 0.6 | null |
artificial-analysis:aa-briefcase | artificial-analysis-aa-briefcase | AA-Briefcase | artificial_analysis | https://artificialanalysis.ai/evaluations/aa-briefcase | Agentic knowledge work, Elo | [
"agentic",
"business"
] | [] | null | null | 2026-06-18T00:00:00 | unknown | null | null | null | 1 | 65 | 65 | 1,710.26 | null |
llm-stats:aa-briefcase | llm-stats-aa-briefcase | AA-Briefcase | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/aa-briefcase?top_n=500 | AA-Briefcase is an Artificial Analysis evaluation of AI systems on professional knowledge-work tasks, reported as an Elo score. | [
"productivity",
"reasoning",
"agents"
] | [] | text | null | null | unknown | null | null | null | 1 | 3 | 3 | 1,577 | null |
llm-stats:aa-index | llm-stats-aa-index | AA-Index | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/aa-index?top_n=500 | No official academic documentation found for this benchmark. Extensive research through ArXiv, IEEE/ACL/NeurIPS papers, and university research sites yielded no peer-reviewed sources for an 'aa-index' benchmark. This entry requires verification from official academic sources. | [
"general"
] | [] | text | null | null | unknown | null | null | null | 1 | 4 | 4 | 0.677 | null |
artificial-analysis:aa-lcr | artificial-analysis-aa-lcr | AA-LCR | artificial_analysis | https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning | Long context reasoning | [
"intelligence-index",
"long-context",
"reasoning"
] | [] | null | null | 2025-08-05T00:00:00 | unknown | null | null | null | 1 | 510 | 510 | 0.833333 | null |
llm-stats:aa-lcr | llm-stats-aa-lcr | AA-LCR | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/aa-lcr?top_n=500 | Agent Arena Long Context Reasoning benchmark | [
"long_context",
"reasoning"
] | [] | text | null | null | unknown | null | null | null | 1 | 18 | 18 | 0.8 | null |
model-reports:aa_lcr | aa_lcr | AA-LCR | model_reports | https://artificialanalysis.ai/evaluations/aa-lcr | Run by Artificial Analysis rather than the vendor. Third-party execution is the point, but it also means the vendor did not control the setup. | [
"long_context"
] | [] | null | null | 2025-09-16T00:00:00 | unknown | null | null | null | 2 | 2 | 2 | 74.7 | percent |
artificial-analysis:aa-omniscience-accuracy | artificial-analysis-aa-omniscience-accuracy | AA-Omniscience Accuracy | artificial_analysis | https://artificialanalysis.ai/evaluations/omniscience | Knowledge | [
"intelligence-index",
"knowledge"
] | [] | null | null | 2025-11-16T00:00:00 | unknown | null | null | null | 1 | 489 | 489 | 0.6535 | null |
llm-stats:aa-omniscience-index | llm-stats-aa-omniscience-index | AA-Omniscience Index | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/aa-omniscience-index?top_n=500 | AA-Omniscience Index is Artificial Analysis's knowledge-reliability metric. It rewards correct answers, penalizes hallucinations, and does not penalize abstention. Scores range from -100 to 100, where 0 means as many correct as incorrect answers. | [
"reasoning",
"science",
"knowledge"
] | [] | text | null | 2025-11-16T00:00:00 | unknown | null | null | null | 1 | 2 | 2 | 126 | null |
artificial-analysis:aa-omniscience-non-hallucination | artificial-analysis-aa-omniscience-non-hallucination | AA-Omniscience Non-Hallucination Rate | artificial_analysis | https://artificialanalysis.ai/evaluations/omniscience | 1 - hallucination rate | [
"intelligence-index",
"knowledge",
"faithfulness"
] | [] | null | null | 2025-11-16T00:00:00 | unknown | null | null | null | 1 | 489 | 489 | 0.990991 | null |
claire-radar:2606.07462 | claire-radar-2606-07462 | AARRI-Bench | claire_radar | https://arxiv.org/abs/2606.07462 | AARRI-Bench evaluates LLM agents on entry-level research intern tasks, measuring success rate in containerized environments with fixed tasks and scoring. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-06-05T00:00:00 | unknown | https://arxiv.org/abs/2606.07462 | https://github.com/AARR-bench/AARRI-bench | null | 1 | null | 0 | null | null |
claire-radar:2606.11150 | claire-radar-2606-11150 | ABC-Bench | claire_radar | https://arxiv.org/abs/2606.11150 | ABC-Bench evaluates LLM agents on biosecurity-relevant tasks including liquid handling robot code generation, DNA fragment design, and DNA synthesis screening evasion, with wet-lab validation. | [
"Biology & Drug Discovery",
"Language & Knowledge",
"Robotics & Embodied AI",
"cs.AI"
] | [] | null | null | 2026-06-09T00:00:00 | unknown | https://arxiv.org/abs/2606.11150 | null | null | 1 | null | 0 | null | null |
opencompass:1148 | opencompass-1148-abspyramid | AbsPyramid | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AbsPyramid | ABSPYRAMID is a unified entailment graph of 221K textual descriptions of abstraction knowledge. ABSPYRAMID collects abstract knowledge for three components
of diverse events to comprehensively evaluate the abstraction ability of language
models in the open domain. ABSPYRAMID 是包含 221,000 条文本描述的抽象知识,收集了多种事件的三个组成部分的抽象知识,以... | [
"知识",
"Knowledge",
"NAACL 2024",
"大语言模型",
"LLM",
"知识储备",
"不支持",
"Unsupported"
] | [] | null | Tencent AI Lab | 2024-06-16T00:00:00 | unknown | null | https://github.com/HKUST-KnowComp/AbsPyramid | null | 1 | null | 0 | null | null |
llm-stats:acebench | llm-stats-acebench | ACEBench | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/acebench?top_n=500 | ACEBench is a comprehensive benchmark for evaluating Large Language Models' tool usage capabilities across three primary evaluation types: Normal (basic tool usage scenarios), Special (tool usage with ambiguous or incomplete instructions), and Agent (multi-agent interactions simulating real-world dialogues). The benchm... | [
"reasoning",
"finance",
"general",
"healthcare",
"tool_calling"
] | [] | text | null | null | unknown | null | null | null | 1 | 2 | 2 | 0.765 | null |
claire-radar:2608.09476 | claire-radar-2608-09476 | ActBench | claire_radar | https://arxiv.org/abs/2608.09476 | Self-evolving benchmark of behavioral safety in cowork agents, evaluating risk from execution trajectories across 600 cases, 213 scenarios, 15 risk behaviors, and 48 web-service APIs. | [
"General AI",
"Safety",
"Safety & Trustworthiness"
] | [] | null | null | 2026-08-10T00:00:00 | unknown | https://arxiv.org/abs/2608.09476 | https://github.com/zjuicsr/ActBench | null | 1 | null | 0 | null | null |
opencompass:1319 | opencompass-1319-actionatlas | ActionAtlas | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/ActionAtlas | ActionAtlas is a multiple-choice video question answering benchmark, including 934 videos showcasing 580 unique actions across 56 sports, with a total of 1896 actions within choices. ActionAtlas是一个多项选择视频问答基准测试,包括934个视频,展示了56项运动中的580个独特动作,选项共包含1896个动作。 | [
"多模态",
"Multimodal",
"NeurIPS 2024",
"多模态模型",
"VLM",
"视频理解",
"Video Understanding",
"不支持",
"Unsupported"
] | [] | multimodal | University of Washington | 2024-10-08T00:00:00 | unknown | https://arxiv.org/abs/2410.05774 | https://github.com/mrsalehi/action-atlas | https://huggingface.co/datasets/mrsalehi/ActionAtlas-v1.0 | 1 | null | 0 | null | null |
claire-radar:2608.04682 | claire-radar-2608-04682 | Active-SWE | claire_radar | https://arxiv.org/abs/2608.04682 | Active-SWE evaluates coding agents on proactive bug fixing: detecting and fixing multiple bugs without issue reports. It includes 1,663 tasks across six bug categories and eight languages, with stages for recorded bugs, potential bugs, and judge validation. | [
"Code & Software",
"Software & AI Compute",
"cs.SE"
] | [] | null | null | 2026-08-05T00:00:00 | unknown | https://arxiv.org/abs/2608.04682 | https://github.com/XLearning-SCU/Active-SWE | null | 1 | null | 0 | null | null |
claire-radar:2607.10180 | claire-radar-2607-10180 | ActiveFly-Bench | claire_radar | https://arxiv.org/abs/2607.10180 | ActiveFly-Bench is a benchmark for UAV embodied perception, decomposing active perception into three tasks: Aerial Embodied Question Answering (Air-EQA), Observation Behavior Planning (OBP), and Fine-grained Language-guided UAV Control (FLUC). It includes datasets from real-world and simulated outdoor environments. | [
"Multimodal",
"Reasoning",
"Robotics & Embodied AI"
] | [] | null | null | 2026-07-11T00:00:00 | unknown | https://arxiv.org/abs/2607.10180 | null | null | 1 | null | 0 | null | null |
claire-radar:2605.24043 | claire-radar-2605-24043 | ActiveSciBench | claire_radar | https://arxiv.org/abs/2605.24043 | ActiveSciBench evaluates budget-constrained, closed-loop scientific discovery across enzyme-kinetics and gene-regulatory-network tasks. | [
"AI Scientist",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-05-21T00:00:00 | unknown | https://arxiv.org/abs/2605.24043 | https://github.com/scientific-discovery/LLM-AutoSciLab | null | 1 | null | 0 | null | null |
claire-radar:2607.16165 | claire-radar-2607-16165 | ActiveVision | claire_radar | https://arxiv.org/abs/2607.16165 | ActiveVision evaluates whether multimodal large language models can perform repeated, active visual perception across 17 tasks in three categories. | [
"General AI",
"Vision & 3D",
"cs.CV"
] | [] | null | null | 2026-07-17T00:00:00 | unknown | https://arxiv.org/abs/2607.16165 | null | null | 1 | null | 0 | null | null |
llm-stats:activitynet | llm-stats-activitynet | ActivityNet | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/activitynet?top_n=500 | A large-scale video benchmark for human activity understanding. Provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours. The benchmark covers a wide range of complex human activities that are of interest to people... | [
"video",
"vision"
] | [] | video | null | null | unknown | null | null | null | 1 | 1 | 1 | 0.619 | null |
claire-radar:2609.03677 | claire-radar-2609-03677 | AD-Diff Bench | claire_radar | https://arxiv.org/abs/2609.03677 | Evaluates set-difference captioning on autonomous driving image subsets, with three splits based on annotation filtering, CLIP filtering, and web scraping, yielding natural-language descriptions of differences between target and reference sets. | [
"Autonomous Driving",
"General AI",
"Vision & 3D",
"cs.CV"
] | [] | null | null | 2026-09-03T00:00:00 | unknown | https://arxiv.org/abs/2609.03677 | https://github.com/KIT-MRT/AD-Diff | null | 1 | null | 0 | null | null |
opencompass:1155 | opencompass-1155-ada-leval | Ada-LEval | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/Ada-LEval | Ada-LEval is a length-adaptable benchmark for evaluating the long-context understanding
of LLMs. Ada-LEval includes two challenging subsets, TSort and BestAnswer, which enable
a more reliable evaluation of LLMs’ long context capabilities. Ada-LEval 用于评估大型语言模型(LLMs)对长上下文的理解能力。Ada-LEval 包含两个具有挑战性的子集,TSort 和 BestAnswer,能够... | [
"长文本",
"Long-Context",
"NAACL 2024",
"大语言模型",
"LLM",
"长上下文",
"Long Context",
"官方自建",
"Official",
"不支持",
"Unsupported"
] | [] | null | Shanghai AI Laboratory | 2024-06-16T00:00:00 | unknown | null | https://github.com/open-compass/Ada-LEval | null | 1 | null | 0 | null | null |
claire-radar:2606.21144 | claire-radar-2606-21144 | AdaMem-Bench | claire_radar | https://arxiv.org/abs/2606.21144 | AdaMem-Bench simulates weeks of interaction with week-by-week question answering to evaluate memory policies for personalized long-horizon LLM agents. It measures QA accuracy and memory volume across different models. | [
"General AI",
"Language & Knowledge",
"cs.CL"
] | [] | null | null | 2026-06-19T00:00:00 | unknown | https://arxiv.org/abs/2606.21144 | null | null | 1 | null | 0 | null | null |
claire-radar:2606.05622 | claire-radar-2606-05622 | AdaPlanBench | claire_radar | https://arxiv.org/abs/2606.05622 | Evaluates adaptive planning of LLM agents under progressively disclosed world and user constraints across 307 household tasks. | [
"Agents & Tool Use",
"General AI",
"Planning",
"cs.CL"
] | [] | null | null | 2026-06-04T00:00:00 | unknown | https://arxiv.org/abs/2606.05622 | https://github.com/JiayuJeff/AdaPlanBench | null | 1 | null | 0 | null | null |
claire-radar:2607.18063 | claire-radar-2607-18063 | Adaptive Adversaries | claire_radar | https://arxiv.org/abs/2607.18063 | A 21-scenario benchmark evaluating memoryless LLM defenders against adaptive attackers that observe earlier responses and change tactics across multiple rounds. | [
"Agents",
"Agents & Tool Use",
"Robotics & Embodied AI"
] | [] | null | null | 2026-07-20T00:00:00 | unknown | https://arxiv.org/abs/2607.18063 | null | null | 1 | null | 0 | null | null |
claire-radar:2608.26204 | claire-radar-2608-26204 | ADeptS-Bench | claire_radar | https://arxiv.org/abs/2608.26204 | Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can safely interact with visual interfaces while handling ambiguou… | [
"Agents & Tool Use",
"Manufacturing & Process Control",
"cs.CR"
] | [] | null | null | 2026-08-25T00:00:00 | unknown | https://arxiv.org/abs/2608.26204 | null | null | 1 | null | 0 | null | null |
llm-stats:advancedif | llm-stats-advancedif | AdvancedIF | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/advancedif?top_n=500 | AdvancedIF is a rubric-based benchmark measuring complex, multi-turn, and system-prompted instruction following ability, scored with a calibrated LLM judge against per-instruction rubrics. | [
"reasoning",
"instruction_following",
"general"
] | [] | text | null | null | unknown | null | null | null | 1 | 2 | 2 | 0.85 | null |
claire-radar:2607.11849 | claire-radar-2607-11849 | AdvancedMathBench | claire_radar | https://arxiv.org/abs/2607.11849 | AdvancedMathBench is a benchmark suite for advanced mathematical reasoning, containing ProverBench (296 proof problems) and VerifierBench (888 proof trajectories with expert labels), with an automatic verification pipeline. | [
"Language & Knowledge",
"Mathematics & Formal Science",
"Reasoning"
] | [] | null | null | 2026-07-13T00:00:00 | unknown | https://arxiv.org/abs/2607.11849 | null | null | 1 | null | 0 | null | null |
claire-radar:2606.24589 | claire-radar-2606-24589 | AdversaBench | claire_radar | https://arxiv.org/abs/2606.24589 | AdversaBench is an automated LLM red-teaming evaluation that mutates seed prompts and confirms resulting failures with multiple judges and a tiebreaker. | [
"Cybersecurity",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-06-23T00:00:00 | unknown | https://arxiv.org/abs/2606.24589 | https://github.com/khanak0509/AdversaBench | null | 1 | null | 0 | null | null |
claire-radar:2608.00832 | claire-radar-2608-00832 | AdvPlan-Bench | claire_radar | https://arxiv.org/abs/2608.00832 | AdvPlan-Bench is an offline benchmark for adversarial evaluation of structured plan-generation agents. It uses typed action chains, adversarial response sets, selector diagnostics, and metrics like BLUE-vs-RED advantage and Nash-gap. Includes 150 synthetic scenarios across five planning templates. | [
"General AI",
"Language & Knowledge",
"cs.LG"
] | [] | null | null | 2026-08-01T00:00:00 | unknown | https://arxiv.org/abs/2608.00832 | null | null | 1 | null | 0 | null | null |
claire-radar:2607.14726 | claire-radar-2607-14726 | AE-UAV | claire_radar | https://arxiv.org/abs/2607.14726 | AE-UAV is an air-to-air event-based UAV tracking benchmark with 178 flight sequences and continuous-time cubic B-spline annotations, supporting evaluation at arbitrary temporal resolutions. It includes multimodal auxiliary data and predefined train/validation/test splits. | [
"General AI",
"Vision & 3D",
"cs.CV"
] | [] | null | null | 2026-07-16T00:00:00 | unknown | https://arxiv.org/abs/2607.14726 | https://github.com/MSP-xEN/AE-UAV | null | 1 | null | 0 | null | null |
opencompass:2452 | opencompass-2452-aecbench | AECBench | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AECBench | AECBench is an open-source benchmark for evaluating LLMs in architecture, engineering, and construction (AEC), covering 23 tasks and about 4,800 samples across five cognitive levels. AECBench 是面向建筑、工程与施工(AEC)领域的大语言模型评测基准,覆盖 5 个认知层级、23 类任务和约 4,800 个样本,用于评估模型在知识记忆、理解、推理、计算与应用方面的能力。 | [
"科学",
"Science",
"学科",
"Examination",
"知识",
"Knowledge",
"科学智能",
"AI for Science",
"科学推理",
"Scientific Reasoning",
"知识储备",
"不支持",
"Unsupported"
] | [
"Chinese"
] | null | 华东建筑设计研究院有限公司、同济大学 | 2026-04-10T00:00:00 | restricted | https://arxiv.org/pdf/2509.18776 | https://github.com/ArchiAI-LAB/AECBench | https://huggingface.co/datasets/jackluoluo/AECBench | 1 | null | 0 | null | null |
claire-radar:2607.04383 | claire-radar-2607-04383 | AEGBench | claire_radar | https://arxiv.org/abs/2607.04383 | AEGBench is a difficulty-stratified evaluation for open-vocabulary audio event grounding, where systems localize queried sound events in time. | [
"General AI",
"Speech & Audio",
"cs.SD"
] | [] | null | null | 2026-07-05T00:00:00 | unknown | https://arxiv.org/abs/2607.04383 | null | https://huggingface.co/datasets/zihan-audio/AEGBench | 1 | null | 0 | null | null |
llm-stats:community:5f95f778-c521-43fa-b80e-6a55465601e3 | llm-stats-community-5f95f778-c521-43fa-b80e-6a55465601e3 | ael_gate_benchmark_cases_template | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/community%3A5f95f778-c521-43fa-b80e-6a55465601e3?top_n=500 | [] | [] | null | null | null | unknown | null | null | null | 1 | null | 0 | null | null | |
claire-radar:2608.16349 | claire-radar-2608-16349 | AeroCopilotBench | claire_radar | https://arxiv.org/abs/2608.16349 | AeroCopilotBench is a two-tier benchmark for evaluating LLM agents as aviation copilots. Tier-1 uses 1,200 multiple-choice questions for knowledge assessment, while Tier-2 includes 73 procedural tasks in an interactive virtual cockpit environment, with safety-gated evaluation. | [
"Agents",
"Agents & Tool Use",
"General AI"
] | [] | null | null | 2026-08-17T00:00:00 | unknown | https://arxiv.org/abs/2608.16349 | null | null | 1 | null | 0 | null | null |
claire-radar:2608.14721 | claire-radar-2608-14721 | AeroGround | claire_radar | https://arxiv.org/abs/2608.14721 | AeroGround evaluates vision-language models on aerial-ground collaborative reasoning using a simulated dataset of ~29,000 multimodal observation groups and 2,250 QA instances covering cross-view correspondence, spatial understanding, and reasoning. | [
"Reasoning",
"Robotics & Embodied AI",
"Vision & 3D"
] | [] | null | null | 2026-08-12T00:00:00 | unknown | https://arxiv.org/abs/2608.14721 | null | null | 1 | null | 0 | null | null |
llm-stats:aethercode | llm-stats-aethercode | AetherCode | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/aethercode?top_n=500 | AetherCode is a competitive-programming benchmark of olympiad-level algorithmic coding problems. | [
"reasoning",
"coding"
] | [] | text | null | null | unknown | null | null | null | 1 | 2 | 2 | 0.679 | null |
claire-radar:2608.24954 | claire-radar-2608-24954 | AFDBench | claire_radar | https://arxiv.org/abs/2608.24954 | Evaluates generative meteorological reasoning through 7,732 expert-written forecast discussions paired with AI weather inputs, using metrics for numerical accuracy, style, and grounding. | [
"AI Scientist",
"General AI",
"Language & Knowledge",
"Reasoning"
] | [] | null | null | 2026-08-25T00:00:00 | unknown | https://arxiv.org/abs/2608.24954 | null | null | 1 | null | 0 | null | null |
claire-radar:2606.14240 | claire-radar-2606-14240 | AFFORDANCE20Q | claire_radar | https://arxiv.org/abs/2606.14240 | Affordance20Q is a benchmark for evaluating affordance reasoning in LLMs using a 20-questions game. It comprises 1,009 games over 454 objects and 59 affordances, where models identify a hidden object's affordance by asking yes/no questions about physical properties. | [
"General AI",
"Language & Knowledge",
"Reasoning"
] | [] | null | null | 2026-06-12T00:00:00 | unknown | https://arxiv.org/abs/2606.14240 | https://github.com/1171-jpg/Affordance20Q.git | null | 1 | null | 0 | null | null |
opencompass:506 | opencompass-506-afqmc | AFQMC | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AFQMC | AFQMC is an Ant Financial chinese semantic similarity task, which requires to judge whether two sentences have the same meaning or not. AFQMC一个蚂蚁金服中文语义相似度任务,要求判断两个句子是否具有相同的语义。 | [
"语言",
"Language",
"大语言模型",
"LLM",
"语言理解",
"Comprehension",
"开源收录",
"Open-Source",
"不支持",
"Unsupported"
] | [
"Chinese"
] | null | null | 2020-04-13T00:00:00 | unknown | https://arxiv.org/abs/2209.02970 | https://github.com/IDEA-CCNL/Fengshenbang-LM | null | 1 | null | 0 | null | null |
claire-radar:2608.23628 | claire-radar-2608-23628 | AFT-Bench | claire_radar | https://arxiv.org/abs/2608.23628 | Holds task, backend, initial state, injected failure, agent, and language model fixed while varying the tool interface to measure callability versus operability. | [
"Agents",
"Code & Software",
"General AI"
] | [] | null | null | 2026-08-23T00:00:00 | unknown | https://arxiv.org/abs/2608.23628 | null | null | 1 | null | 0 | null | null |
claire-radar:2607.01152 | claire-radar-2607-01152 | AGC-Bench | claire_radar | https://arxiv.org/abs/2607.01152 | AGC-Bench evaluates artificial general creativity across 78 datasets covering brainstorming, problem solving, STEM, narrative, figurative language, and humor. It uses an agentic harness and a public leaderboard with an open-weight judge model. | [
"General AI",
"Language & Knowledge",
"cs.CL"
] | [] | null | null | 2026-07-01T00:00:00 | unknown | https://arxiv.org/abs/2607.01152 | null | null | 1 | null | 0 | null | null |
claire-radar:github:kyashrathore/agent-app-benchmark | claire-radar-github-kyashrathore-agent-app-benchmark | Agent App Benchmark | claire_radar | https://github.com/kyashrathore/agent-app-benchmark | Evaluates GUI application performance for coding agents using deterministic historical session workloads, measuring app start, session switching, memory, and CPU metrics. | [
"Agents",
"Language & Knowledge",
"Software & AI Compute"
] | [] | null | null | 2026-08-22T00:00:00 | unknown | null | https://github.com/kyashrathore/agent-app-benchmark | null | 1 | null | 0 | null | null |
claire-radar:github:datapace-ai/agent-memory-benchmark | claire-radar-github-datapace-ai-agent-memory-benchmark | Agent Memory Benchmark | claire_radar | https://github.com/datapace-ai/agent-memory-benchmark | Evaluates agent memory systems on a 10-question pilot using LongMemEval, measuring answer accuracy under three judge rules plus token usage, latency, ingestion cost, forgetting curve, and superseded-value handling. | [
"Agents",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-09-08T00:00:00 | unknown | null | https://github.com/datapace-ai/agent-memory-benchmark | null | 1 | null | 0 | null | null |
claire-radar:2606.04874 | claire-radar-2606-04874 | Agent Planning Benchmark | claire_radar | https://arxiv.org/abs/2606.04874 | Agent Planning Benchmark (APB) is a diagnostic benchmark with 4,209 multimodal cases across 22 domains, evaluating planning capabilities in five settings including tool noise and unsolvable tasks. | [
"Agents",
"General AI",
"Multimodal",
"Planning",
"Robustness",
"Safety & Trustworthiness"
] | [] | null | null | 2026-06-03T00:00:00 | unknown | https://arxiv.org/abs/2606.04874 | https://github.com/Mikivishy/AgentPlanningBenchmark | null | 1 | null | 0 | null | null |
claire-radar:2607.24882 | claire-radar-2607-24882 | Agent Retrieval Bench | claire_radar | https://arxiv.org/abs/2607.24882 | File-level retrieval benchmark for coding agents, covering four positive tasks (code2test, comment2context, trace2code, edit2ripple) and a selective-retrieval subset with natural no-gold and counterfactual controls across 25 repositories, 427 samples, with frozen base-commit corpora. | [
"Agents",
"Information retrieval",
"Language & Knowledge",
"Software & AI Compute"
] | [] | null | null | 2026-07-27T00:00:00 | unknown | https://arxiv.org/abs/2607.24882 | https://github.com/eyuansu62/agent-retrieval-bench | null | 1 | null | 0 | null | null |
llm-stats:agent-startup-bench | llm-stats-agent-startup-bench | Agent Startup Bench | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/agent-startup-bench?top_n=500 | Agent Startup Bench measures AI agents on high-economic-value, startup-style tasks that require autonomous planning and execution to deliver practical, verifiable results. | [
"reasoning",
"general",
"agents"
] | [] | text | null | null | unknown | null | null | null | 1 | 2 | 2 | 0.688 | null |
claire-radar:github:martin-beck/agent-systems-benchmark | claire-radar-github-martin-beck-agent-systems-benchmark | Agent Systems Benchmark (ASB) | claire_radar | https://github.com/martin-beck/agent-systems-benchmark | A Linux terminal framework for measuring how AI coding agents scale under concurrent sessions while maintaining task quality, latency, and resource consumption within declared bounds. | [
"Agents",
"Language & Knowledge",
"Software & AI Compute"
] | [] | null | null | 2026-09-06T00:00:00 | unknown | null | https://github.com/martin-beck/agent-systems-benchmark | null | 1 | null | 0 | null | null |
claire-radar:github:sentinelden/agent-injection-bench | claire-radar-github-sentinelden-agent-injection-bench | agent-injection-bench | claire_radar | https://github.com/sentinelden/agent-injection-bench | Evaluates prompt-injection robustness of tool-calling agent stacks using a 24-scenario corpus (21 adversarial, 3 control) delivered through channels such as tool results, documents, web pages, filenames and multi-turn dialogue, with deterministic canary, tool-call, exfiltration and refusal success conditions; results r... | [
"Agents",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-09-11T00:00:00 | unknown | null | https://github.com/sentinelden/agent-injection-bench | null | 1 | null | 0 | null | null |
claire-radar:github:ajbermudezh22/agent-model-bench | claire-radar-github-ajbermudezh22-agent-model-bench | agent-model-bench | claire_radar | https://github.com/ajbermudezh22/agent-model-bench | Evaluates language models on tool calling and strict JSON adherence using 32 cases, scoring exact tool and argument matches and JSON parsing with strict and content-based criteria. | [
"Agents",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-08-21T00:00:00 | unknown | null | https://github.com/ajbermudezh22/agent-model-bench | null | 1 | null | 0 | null | null |
claire-radar:2607.10059 | claire-radar-2607-10059 | AgentAbstain | claire_radar | https://arxiv.org/abs/2607.10059 | AgentAbstain is a paired-task benchmark for evaluating LLM agents' ability to abstain from acting in scenarios such as ambiguity, conflicting constraints, or tool failures. It includes 263 paired tasks across 42 sandbox environments, with a proposed pipeline for generating fresh task instances. | [
"Agents",
"General AI",
"Language & Knowledge",
"Reasoning"
] | [] | null | null | 2026-07-11T00:00:00 | unknown | https://arxiv.org/abs/2607.10059 | null | null | 1 | null | 0 | null | null |
opencompass:1242 | opencompass-1242-agentboard | AgentBoard | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AgentBoard | AgentBoard is tailored to analytical evaluation of LLM agents. It offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. AgentBoard专用于LLM Agent的分... | [
"智能体",
"Agent",
"NeurIPS 2024",
"任务执行",
"Task Execution",
"不支持",
"Unsupported"
] | [] | null | The University of Hong Kong | 2024-06-24T00:00:00 | restricted | https://arxiv.org/abs/2401.13178 | https://github.com/hkust-nlp/AgentBoard | https://huggingface.co/datasets/hkust-nlp/agentboard | 1 | null | 0 | null | null |
claire-radar:2608.14680 | claire-radar-2608-14680 | AGENTCHAOSBENCH | claire_radar | https://arxiv.org/abs/2608.14680 | AGENTCHAOSBENCH is a dataset of sanitized execution traces from five agentic applications with injected runtime faults, used to evaluate fault detection and localization from telemetry. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-08-04T00:00:00 | unknown | https://arxiv.org/abs/2608.14680 | null | null | 1 | null | 0 | null | null |
claire-radar:2606.23189 | claire-radar-2606-23189 | AgentCIBench | claire_radar | https://arxiv.org/abs/2606.23189 | AgentCIBench is an executable evaluation harness for testing whether computer-use agents improperly disclose information across application contexts. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-06-22T00:00:00 | unknown | https://arxiv.org/abs/2606.23189 | null | null | 1 | null | 0 | null | null |
claire-radar:2609.06972 | claire-radar-2609-06972 | AgentDrift | claire_radar | https://arxiv.org/abs/2609.06972 | Step-labeled corpus of 12,536 synthetic LLM agent tool-call trajectories in five domains, labeled per step for benign, injection point, hijacked, or failed injection, supporting detection, localization, and attempt-vs-success evaluation. | [
"Agents",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-09-07T00:00:00 | unknown | https://arxiv.org/abs/2609.06972 | https://github.com/Asif-0209/AgentDrift | null | 1 | null | 0 | null | null |
claire-radar:2606.16723 | claire-radar-2606-16723 | AgentFairBench | claire_radar | https://arxiv.org/abs/2606.16723 | AgentFairBench evaluates demographic disparity in the actions of LLM agents across hiring, lending, and medical triage. It uses synthetic, demographic-neutral profiles in counterfactual matched sets varying name-coded race/gender. Metrics include counterfactual flip rate, mean absolute score difference, action-rate dis... | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-06-15T00:00:00 | unknown | https://arxiv.org/abs/2606.16723 | null | null | 1 | null | 0 | null | null |
opencompass:1351 | opencompass-1351-agentharm | AgentHarm | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AgentHarm | AgentHarm tests the robustness of LLMs to jailbreak attacks. It includes a diverse set of 110 explicitly malicious agent tasks (440 with augmentations), covering 11 harm categories including fraud, cybercrime, and harassment. AgentHarm用于评估LLM智能体对越狱攻击的鲁棒性,包括110套恶意智能体任务(其中有440个强化任务),涵盖欺诈、网络犯罪和骚扰等11个危害类别。 | [
"安全",
"Safety",
"智能体",
"Agent",
"任务执行",
"Task Execution",
"安全对齐",
"Safety Alignment",
"不支持",
"Unsupported"
] | [] | null | Gray Swan AI | 2024-10-11T00:00:00 | restricted | https://arxiv.org/abs/2404.02151 | https://github.com/UKGovernmentBEIS/inspect_evals | https://huggingface.co/datasets/ai-safety-institute/AgentHarm | 1 | null | 0 | null | null |
opencompass:2061 | opencompass-2061-agenthazard | AgentHazard | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AgentHazard | 移动端 GUI Agent 通过与设备环境的交互来完成任务,在完成任务的过程中会遇到一些未知或不可信的信息来源,这些信息可能含有攻击性的内容,致使 Agent 无法正常完成任务,甚至对用户的隐私和财产带来危害。本评测集兼具动态执行环境和静态评测数据集,旨在为移动端 GUI Agent 提供一个仿真度高的模拟环境,以评估其在真实场景下执行的行为和安全性。 移动端 GUI Agent 通过与设备环境的交互来完成任务,在完成任务的过程中会遇到一些未知或不可信的信息来源,这些信息可能含有攻击性的内容,致使 Agent 无法正常完成任务,甚至对用户的隐私和财产带来危害。本评测集兼具动态执行环境和静态评测数据集,旨在为移动端 GUI Agent... | [
"多模态",
"Multimodal",
"安全",
"Safety",
"智能体",
"Agent",
"任务执行",
"Task Execution",
"安全对齐",
"Safety Alignment",
"跨模态推理",
"Cross-modal Reasoning",
"不支持",
"Unsupported"
] | [] | multimodal | Institute for AI Industry Research, Tsinghua University | 2025-07-16T00:00:00 | unknown | https://arxiv.org/abs/2507.04227 | https://github.com/Zsbyqx20/AgentHazard | null | 1 | null | 0 | null | null |
claire-radar:2605.25707 | claire-radar-2605-25707 | AgentHijack | claire_radar | https://arxiv.org/abs/2605.25707 | Evaluates the robustness of computer use agents under common environment corruptions such as pop-ups, resolution changes, and competing applications. The benchmark introduces 9 configurable corruptions and evaluates agent performance on desktop tasks using multimodal LLM-based agents, measuring task completion rates. | [
"Agents",
"Agents & Tool Use",
"General AI",
"Robustness"
] | [] | null | null | 2026-05-25T00:00:00 | unknown | https://arxiv.org/abs/2605.25707 | https://github.com/tmlr-group/AgentHijack | null | 1 | null | 0 | null | null |
claire-radar:2607.29626 | claire-radar-2607-29626 | AgentHPOBench | claire_radar | https://arxiv.org/abs/2607.29626 | AgentHPOBench evaluates LLM agents as sequential hyperparameter optimizers across 30 executable ML tasks. Agents observe accumulated configurations, metrics, and logs, then propose the next configuration. Scoring compares agents and conventional HPO baselines under a unified protocol. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-07-31T00:00:00 | unknown | https://arxiv.org/abs/2607.29626 | null | null | 1 | null | 0 | null | null |
claire-radar:github:bluebunnyanon/agent-benchmark-2d-maze | claire-radar-github-bluebunnyanon-agent-benchmark-2d-maze | Agentic Evaluation in 2D Mazes | claire_radar | https://github.com/bluebunnyanon/agent-benchmark-2d-maze | Evaluates interactive multimodal agents in structured 2D maze environments requiring long-horizon planning, mechanism interaction, and error recovery, scored by mechanism-aware progress. | [
"Agents",
"Agents & Tool Use",
"Multimodal",
"Software & AI Compute"
] | [] | null | null | 2026-08-22T00:00:00 | unknown | null | https://github.com/bluebunnyanon/agent-benchmark-2d-maze | null | 1 | null | 0 | null | null |
claire-radar:github:aah20/agentic-infrastructure-change-benchmark | claire-radar-github-aah20-agentic-infrastructure-change-benchmark | Agentic Infrastructure ChangeBench | claire_radar | https://github.com/AAH20/agentic-infrastructure-change-benchmark | Evaluates AI agent proposals for cloud, Kubernetes, and network infrastructure changes across correctness, safety, reliability, economics, efficiency, and evidence using deterministic scenario-based checks. | [
"Finance",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-09-05T00:00:00 | unknown | null | https://github.com/AAH20/agentic-infrastructure-change-benchmark | null | 1 | null | 0 | null | null |
claire-radar:github:andersballegaard/agentic-netdevops-benchmark | claire-radar-github-andersballegaard-agentic-netdevops-benchmark | Agentic NetDevOps Benchmark | claire_radar | https://github.com/AndersBallegaard/agentic-netdevops-benchmark | A multi-stage environment with a Linux management host and five VyOS routers, requiring agents to complete netdevops tasks from information gathering to evaluation. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-09-08T00:00:00 | unknown | null | https://github.com/AndersBallegaard/agentic-netdevops-benchmark | null | 1 | null | 0 | null | null |
claire-radar:github:gitayg/agentic-security-benchmark | claire-radar-github-gitayg-agentic-security-benchmark | Agentic Security Benchmark | claire_radar | https://github.com/gitayg/agentic-security-benchmark | Measures detection and prevention of agentic-AI security products on 286 attack and 875 benign samples across five AMTSO attack vectors via a scoring adapter harness. | [
"Language & Knowledge",
"Software & AI Compute",
"cs.AI"
] | [] | null | null | 2026-09-07T00:00:00 | unknown | null | https://github.com/gitayg/agentic-security-benchmark | null | 1 | null | 0 | null | null |
claire-radar:2607.01647 | claire-radar-2607-01647 | AgenticDataBench | claire_radar | https://arxiv.org/abs/2607.01647 | Evaluates LLM-based data agents on realistic data science workflows across 15 domains, with fine-grained ground-truth labels and skill-level scoring. | [
"General AI",
"Language & Knowledge",
"cs.DB"
] | [] | null | null | 2026-07-02T00:00:00 | unknown | https://arxiv.org/abs/2607.01647 | https://github.com/AgenticDataBench/AgenticDataBench | null | 1 | null | 0 | null | null |
claire-radar:2606.24026 | claire-radar-2606-24026 | AgenticInterpBench | claire_radar | https://arxiv.org/abs/2606.24026 | Evaluates language model agents on explaining components of transformer circuits, with 84 semi-synthetic circuits and 163 component-level annotations. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-06-23T00:00:00 | unknown | https://arxiv.org/abs/2606.24026 | null | null | 1 | null | 0 | null | null |
claire-radar:2607.02255 | claire-radar-2607-02255 | AgenticSTS | claire_radar | https://arxiv.org/abs/2607.02255 | AgenticSTS is a reproducible long-horizon agent testbed that evaluates bounded, typed-retrieval memory configurations through game outcomes and controlled ablations. | [
"Agents",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-07-02T00:00:00 | unknown | https://arxiv.org/abs/2607.02255 | null | null | 1 | null | 0 | null | null |
claire-radar:github:aah20/ai-agent-infrastructure-benchmark | claire-radar-github-aah20-ai-agent-infrastructure-benchmark | AgentInfraBench | claire_radar | https://github.com/AAH20/ai-agent-infrastructure-benchmark | Evaluates AI agent infrastructure across isolation, identity, network, authority, durability, observability, and unit economics using weighted scores with hard safety gates. | [
"Agents",
"Language & Knowledge",
"Software & AI Compute"
] | [] | null | null | 2026-09-05T00:00:00 | unknown | null | https://github.com/AAH20/ai-agent-infrastructure-benchmark | null | 1 | null | 0 | null | null |
claire-radar:2608.26623 | claire-radar-2608-26623 | AgentJudgeBench | claire_radar | https://arxiv.org/abs/2608.26623 | LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-08-27T00:00:00 | unknown | https://arxiv.org/abs/2608.26623 | null | null | 1 | null | 0 | null | null |
claire-radar:2607.06624 | claire-radar-2607-06624 | AgentLens | claire_radar | https://arxiv.org/abs/2607.06624 | AgentLens evaluates interactive coding agents across entire trajectories, combining formal verification with LLM-written trajectory reviews and side-by-side comparisons to score dimensions such as instruction compliance, tool use, and interaction style. | [
"Agents",
"Agents & Tool Use",
"General AI"
] | [] | null | null | 2026-07-07T00:00:00 | unknown | https://arxiv.org/abs/2607.06624 | https://github.com/agent-lens/agent-lens-bench | null | 1 | null | 0 | null | null |
claire-radar:2608.00009 | claire-radar-2608-00009 | AgentMemBench | claire_radar | https://arxiv.org/abs/2608.00009 | AgentMemBench evaluates five long-term memory management strategies for conversational AI agents across three public datasets (LoCoMo, MultiDoc2Dial, MSC), covering multi-session dialogue, document grounding, and persona-grounded chat. Scoring uses retrieval metrics (Recall@k, MRR, nDCG@k), Answer F1, LLM-judge faithfu... | [
"General AI",
"Language & Knowledge",
"cs.CL"
] | [] | null | null | 2026-06-16T00:00:00 | unknown | https://arxiv.org/abs/2608.00009 | null | null | 1 | null | 0 | null | null |
claire-radar:2606.02240 | claire-radar-2606-02240 | AgentRedBench | claire_radar | https://arxiv.org/abs/2606.02240 | AgentRedBench evaluates LLM agents against indirect prompt injection and underspecified-authorization attacks across 24 enterprise SaaS integrations. It defines 215 attack scenarios with immutable versioning, and tracks attack success rate (ASR) for models and defenses. Open-source codebase and schemas enable replay. | [
"General AI",
"Language & Knowledge",
"cs.CR"
] | [] | null | null | 2026-06-01T00:00:00 | unknown | https://arxiv.org/abs/2606.02240 | null | null | 1 | null | 0 | null | null |
claire-radar:2608.15286 | claire-radar-2608-15286 | AgentRelBench | claire_radar | https://arxiv.org/abs/2608.15286 | AgentRelBench measures repeated agent runs in a fixed task suite, computing severity-weighted damage from database state diffs with no LLM judging. | [
"Agents",
"Agents & Tool Use",
"General AI"
] | [] | null | null | 2026-08-15T00:00:00 | unknown | https://arxiv.org/abs/2608.15286 | https://github.com/shivenkk/agentrelbench | null | 1 | null | 0 | null | null |
opencompass:1778 | opencompass-1778-agentrewardbench | AgentRewardBench | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AgentRewardBench | AgentRewardBench, the first benchmark to assess the effectiveness of LLM judges for evaluating web agents. AgentRewardBench contains 1302 trajectories across 5 benchmarks and 4 LLMs. AgentRewardBench 是首个用于评估大型语言模型(LLM)评判者评估网络代理有效性的基准测试。AgentRewardBench 包含来自 5 个基准测试和 4 个大型语言模型的 1302 条轨迹。 | [
"推理",
"Reasoning",
"智能体",
"Agent",
"任务执行",
"Task Execution",
"不支持",
"Unsupported"
] | [] | null | McGill University,Mila Quebec AI Institute,etc. | 2025-04-11T00:00:00 | restricted | https://arxiv.org/abs/2504.08942 | null | https://huggingface.co/datasets/McGill-NLP/agent-reward-bench | 1 | null | 0 | null | null |
llm-stats:agents-last-exam | llm-stats-agents-last-exam | Agents' Last Exam | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/agents-last-exam?top_n=500 | Agents' Last Exam is a challenging benchmark for AI agents on hard, long-horizon tasks that test sustained reasoning, planning, and tool use, reported with and without tool access. | [
"reasoning",
"agents",
"tool_calling"
] | [] | text | null | null | unknown | null | null | null | 1 | 10 | 10 | 0.527 | null |
model-reports:agents_last_exam | agents_last_exam | Agents' Last Exam | model_reports | https://lastexam.ai/ | Multi-step agentic benchmark using Claude Code harness. Tool Search disabled. Scores depend on harness, reasoning effort, context length, and timeout settings. | [
"coding_agent"
] | [] | null | null | 2025-01-23T00:00:00 | unknown | null | null | null | 3 | 3 | 3 | 31.8 | percent |
claire-radar:2606.05405 | claire-radar-2606-05405 | Agents' Last Exam (ALE) | claire_radar | https://arxiv.org/abs/2606.05405 | A living benchmark of 1K+ long-horizon professional tasks with verifiable outcomes, covering 55 sub-fields in 13 industry clusters based on O*NET / SOC 2018. | [
"General AI",
"Language & Knowledge",
"cs.AI"
] | [] | null | null | 2026-06-03T00:00:00 | unknown | https://arxiv.org/abs/2606.05405 | https://github.com/rdi-berkeley/agents-last-exam | null | 1 | null | 0 | null | null |
claire-radar:2607.27294 | claire-radar-2607-27294 | AgentS4D | claire_radar | https://arxiv.org/abs/2607.27294 | AgentS4D evaluates runtime safety of LLM-based workspace agents across a four-dimensional framework, with 328 risk-injected cases and seven lifecycle checkpoints, measuring unsafe behavior and evidence across six risk-entry sources and nine harms. | [
"General AI",
"Safety",
"Safety & Trustworthiness"
] | [] | null | null | 2026-07-29T00:00:00 | unknown | https://arxiv.org/abs/2607.27294 | null | null | 1 | null | 0 | null | null |
claire-radar:2608.00805 | claire-radar-2608-00805 | AgentSLABench | claire_radar | https://arxiv.org/abs/2608.00805 | A resource-aware agent benchmark with 16 Dockerized task environments that profiles correctness alongside latency, cost, compute, memory, and network use under declared budgets. | [
"Agents & Tool Use",
"Software & AI Compute",
"cs.AI"
] | [] | null | null | 2026-08-01T00:00:00 | unknown | https://arxiv.org/abs/2608.00805 | https://github.com/MeherBhaskar/agentslabench | null | 1 | null | 0 | null | null |
claire-radar:2608.15127 | claire-radar-2608-15127 | AgentSysBench | claire_radar | https://arxiv.org/abs/2608.15127 | AgentSysBench is a benchmark suite and measurement toolkit for characterizing agentic workloads on LLM serving systems. It includes ten representative agentic applications and unified instrumentation, identifying six properties that distinguish agentic workloads from conventional LLM inference. | [
"General AI",
"Language & Knowledge",
"cs.OS"
] | [] | null | null | 2026-08-15T00:00:00 | unknown | https://arxiv.org/abs/2608.15127 | null | null | 1 | null | 0 | null | null |
claire-radar:2606.24597 | claire-radar-2606-24597 | AgentWorldBench | claire_radar | https://arxiv.org/abs/2606.24597 | AgentWorldBench evaluates language world models on simulation fidelity across 7 domains (MCP, Search, Terminal, SWE, Android, Web, OS) using real-world trajectories and rubric-based scoring. | [
"General AI",
"Language & Knowledge",
"cs.CL"
] | [] | null | null | 2026-06-23T00:00:00 | unknown | https://arxiv.org/abs/2606.24597 | https://github.com/QwenLM/Qwen-AgentWorld | https://huggingface.co/datasets/Qwen/AgentWorldBench | 1 | null | 0 | null | null |
claire-radar:2609.02339 | claire-radar-2609-02339 | AGI Maze Prediction Datasets and Benchmark | claire_radar | https://arxiv.org/abs/2609.02339 | Evaluates predictive models on maze prediction tasks to study world dynamics learning. | [
"General AI",
"Language & Knowledge",
"cs.LG"
] | [] | null | null | 2026-09-02T00:00:00 | unknown | https://arxiv.org/abs/2609.02339 | null | null | 1 | null | 0 | null | null |
llm-stats:agieval | llm-stats-agieval | AGIEval | llm_stats | https://api.zeroeval.com/leaderboard/benchmarks/agieval?top_n=500 | A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission tests (LSAT), math competitions, lawyer qualification tests, and civil service exams. Contains 20 tasks (18 multiple-choice, 2 cloze) designed to assess understanding, kn... | [
"legal",
"math",
"reasoning",
"general"
] | [] | text | null | null | unknown | null | null | null | 1 | 10 | 10 | 0.658 | null |
model-reports:agieval | agieval | AGIEval | model_reports | https://github.com/ruixiangcui/AGIEval | Human-exam derived; overlaps heavily with MMLU-style coverage. | [
"knowledge"
] | [] | null | null | 2023-04-13T00:00:00 | unknown | null | https://github.com/ruixiangcui/AGIEval | null | 4 | null | 0 | null | null |
opencompass:497 | opencompass-497-agieval | AGIEval | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AGIEval | AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving. This benchmark is derived from 20 official, public, and high-standard admission and qualification exams intended for general human test-takers, suc... | [
"学科",
"Examination",
"大语言模型",
"LLM",
"知识储备",
"Knowledge",
"开源收录",
"Open-Source",
"不支持",
"Unsupported"
] | [
"Chinese"
] | null | null | 2023-09-18T00:00:00 | unknown | https://arxiv.org/pdf/2304.06364 | https://github.com/ruixiangcui/AGIEval | null | 1 | null | 0 | null | null |
claire-radar:2605.26302 | claire-radar-2605-26302 | AgingBench | claire_radar | https://arxiv.org/abs/2605.26302 | Evaluates the reliability of deployed AI agents over extended operational lifetimes. The benchmark organizes agent aging into four mechanisms—compression, interference, revision, and maintenance—and uses temporal dependency graphs and paired counterfactual probes to produce diagnostic profiles of the memory pipeline's ... | [
"Agents",
"General AI",
"Language & Knowledge"
] | [] | null | null | 2026-05-25T00:00:00 | unknown | https://arxiv.org/abs/2605.26302 | https://github.com/VITA-Group/AgingBench | null | 1 | null | 0 | null | null |
opencompass:1753 | opencompass-1753-agmmu | AgMMU | opencompass_hub | https://hub.opencompass.org.cn/dataset-detail/AgMMU | A Comprehensive Agricultural Multimodal Understanding and Reasoning Benchmark 农业综合多模态理解和推理基准。 | [
"多模态",
"Multimodal",
"农业",
"多模态模型",
"VLM",
"跨模态推理",
"Cross-modal Reasoning",
"不支持",
"Unsupported"
] | [] | multimodal | Rice University,etc. | 2025-04-14T00:00:00 | unknown | https://arxiv.org/abs/2504.10568 | https://github.com/AgMMU/AgMMU | null | 1 | null | 0 | null | null |
claire-radar:2606.24526 | claire-radar-2606-24526 | AGORA | claire_radar | https://arxiv.org/abs/2606.24526 | Agora evaluates agentic document reasoning across eight domain collections of 9,664 authentic workplace documents. It includes 362 questions requiring location of sparse evidence and reconciliation of terminology, units, and time conventions. The benchmark is designed to exceed model context windows, necessitating deli... | [
"General AI",
"Language & Knowledge",
"Reasoning"
] | [] | null | null | 2026-06-23T00:00:00 | unknown | https://arxiv.org/abs/2606.24526 | null | null | 1 | null | 0 | null | null |
claire-radar:2609.08402 | claire-radar-2609-08402 | AGOS-Bench | claire_radar | https://arxiv.org/abs/2609.08402 | AGOS-Bench evaluates VLMs on the urban Air-Ground Object Search task where a UAV and UGV coordinate to locate a target vehicle from multi-view visual references, reporting success rate and success weighted by path length. | [
"Multimodal",
"Robotics & Embodied AI"
] | [] | null | null | 2026-09-08T00:00:00 | unknown | https://arxiv.org/abs/2609.08402 | null | null | 1 | null | 0 | null | null |
claire-radar:2605.22366 | claire-radar-2605-22366 | AgroTools | claire_radar | https://arxiv.org/abs/2605.22366 | AgroTools evaluates tool-augmented multimodal agents in agriculture, with 539 QA instances, 1,097 images, 14 executable tools, and structured tool-use traces for process and outcome evaluation. | [
"General AI",
"Multimodal"
] | [] | null | null | 2026-05-21T00:00:00 | unknown | https://arxiv.org/abs/2605.22366 | null | https://huggingface.co/datasets/AgroTools/AgroTools | 1 | null | 0 | null | null |
Benchmark Radar Dataset
Overview
Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).
As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:
- Catalog & Scores: A curated, normalized cross-registry archive of 3,212 benchmark suites and 13,022 reported model score observations across LLM Stats, OpenCompass Hub, Artificial Analysis, and premier model technical reports.
- Radar Discoveries: A 24/7 automated intelligence pipeline tracking emerging papers, code repositories, datasets, and community attention across 37+ sources, capturing 13,275 research artifacts and 23,142 daily observations.
Dataset Structure & Usage
The dataset is partitioned into 4 configs (subsets), easily loaded via
Hugging Face datasets:
from datasets import load_dataset
# 1. Load benchmark catalog (default)
catalog = load_dataset("ktwu01/benchmark-radar", "catalog")
# 2. Load model evaluation scores
scores = load_dataset("ktwu01/benchmark-radar", "scores")
# 3. Load emerging radar artifacts (papers, repositories, datasets)
artifacts = load_dataset("ktwu01/benchmark-radar", "radar_artifacts")
# 4. Load daily discovery observations (attention, downloads, events)
observations = load_dataset("ktwu01/benchmark-radar", "radar_observations")
1. catalog (Default Config)
Contains 3,212 normalized benchmark records.
benchmark_id: Canonical unique identifier (e.g.llm-stats:mmlu-pro,opencompass:1580)slug: URL-safe unique identifiername: Benchmark display namesource: Source provider (llm_stats,opencompass_hub,artificial_analysis,model_reports)source_url: URL to original source entrydescription: Plaintext description and task summarycategories: Assigned category tags (e.g.Reasoning,Code,Multimodal)languages: Languages evaluated (e.g.en,zh)modality: Evaluated modality (text,multimodal,code, etc.)publisher: Creator organization or research labreleased_at: Official release date (YYYY-MM-DD proxy)openness: Data and code openness tierpaper_url: Associated paper URL (arXiv or DOI)repo_url: Associated code repository URL (GitHub)dataset_url: Direct dataset download or Hub URLdocument_count: Number of verified source documents citing this benchmarkmodel_count: Number of distinct models evaluatedscore_count: Number of numeric scores recordedhighest_score: Maximum observed score across all evaluated modelsscore_unit: Score metric / unit (e.g.%,accuracy,Elo)
2. scores
Contains 13,022 reported evaluation score observations across 870+ frontier models.
obs_id: Unique observation identifierkey: Benchmark keymodel_id: Source-specific stable model identifier, or null when the source does not provide onemodel_name: Model display nameorganization: Model creator/lab (e.g.DeepSeek,OpenAI,Anthropic,Google,Meta)value: Numeric reported scoreraw_value: Original score string from sourcevalue_kind: Value data type (number,percentage, etc.)reported_date: Model announcement or report publication datedate_precision: Precision level of reported datereported_by: Source reporting modality (self_reported,third_party)source: Score ingestion sourcesource_url: URL of source leaderboard or technical reportdocument_id: Source document citation ID
3. radar_artifacts
Contains 13,275 academic artifacts surfaced by daily radar.
id: Artifact identifier (e.g.artifact:arxiv:2203.17257)type: Entity type (artifact)label: Title or repository nameurl: Primary external URLcategories: Extracted capability categoriessources: Discovery sources that observed this artifactfirst_seen_at: First discovery date (YYYY-MM-DD)last_seen_at: Latest discovery date (YYYY-MM-DD)observation_count: Total appearances across snapshotslatest_score: Attention and composite radar scoremetrics: Dictionary of observed signals (e.g. stars, citations, downloads)
4. radar_observations
Contains 23,142 discrete daily discovery events.
id: Observation identifierentity_id: Referenced artifact IDsnapshot_date: Radar snapshot date (YYYY-MM-DD)source: Discovery source platformsource_id: Upstream source identifierurl: Surfaced event URLdiscovered_at: Timestamp of discoverypublished_at: Original creation / publication timestamptotal_score: Radar composite relevance scoreevent_kind: Event category (discovered,released,updated)organizations: Identified affiliated organizationsmetrics: Point-in-time metrics (stars, likes, downloads)
Data Provenance and Principles
- Full-Corpus Coverage: All 3,212 benchmarks across all sources are preserved. Unscored benchmarks are retained with verified paper, code, and dataset links.
- Strict Evidence Citation: Every score links directly to its source document or technical report.
- Daily Automated Sync: Radar discoveries are refreshed daily at 05:00 UTC via GitHub Actions.
Licensing
The technical report and Benchmark Radar's original editorial content are available under CC BY-NC-SA 4.0. Commercial dataset packaging or product integration requires prior written permission. Third-party benchmark metadata and source material retain their original terms. Review the repository licensing notice before reuse, especially for commercial dataset packaging.
Citation
@article{benchmark_radar_2026,
title={Benchmark Radar: Daily Discovery and Full-Corpus Search
Across the AI Evaluation Landscape},
author={Wu, Koutian and Contributors},
journal={arXiv preprint arXiv:2609.11115},
year={2026}
}
- Downloads last month
- 433