Dataset Viewer
Auto-converted to Parquet Duplicate
benchmark_id
stringlengths
12
103
slug
stringlengths
3
103
name
stringlengths
2
118
source
stringclasses
5 values
source_url
stringlengths
19
110
description
stringlengths
0
995
categories
listlengths
0
21
languages
listlengths
0
13
modality
stringclasses
5 values
publisher
stringclasses
315 values
released_at
timestamp[s]date
2010-02-19 00:00:00
2026-09-13 00:00:00
⌀
openness
stringclasses
3 values
paper_url
stringlengths
31
62
⌀
repo_url
stringlengths
28
117
⌀
dataset_url
stringclasses
395 values
document_count
int64
1
29
⌀
model_count
int64
1
586
⌀
score_count
int64
0
586
highest_score
float64
0.01
2.1M
⌀
score_unit
stringclasses
3 values
claire-radar:2609.04611
claire-radar-2609-04611
$\tau^\tau$-Bench: An Environment for End-To-End, Realistic Agent Construction
claire_radar
https://arxiv.org/abs/2609.04611
A benchmark where a developer agent builds a complete customer-service agent from business records, client requirements, a production API, inherited codebase, and cost/model limits, scored by deployment against held-out simulated users across 53 tasks in four domains.
[ "Agents", "Agents & Tool Use", "General AI" ]
[]
null
null
2026-09-04T00:00:00
unknown
https://arxiv.org/abs/2609.04611
null
null
1
null
0
null
null
claire-radar:2608.08814
claire-radar-2608-08814
360CityArena
claire_radar
https://arxiv.org/abs/2608.08814
360CityArena evaluates embodied agents in a photorealistic virtual urban environment built from 360-degree video of Tokyo's Akihabara district. It includes 175 tasks across environment understanding, path reasoning, and spatial reasoning, testing localization, landmark search, path planning, and relational spatial reas...
[ "Robotics & Embodied AI", "cs.CV" ]
[]
null
null
2026-08-09T00:00:00
unknown
https://arxiv.org/abs/2608.08814
https://github.com/360MM-Team/360CityArena
null
1
null
0
null
null
claire-radar:2606.01057
claire-radar-2606-01057
3DCodeBench
claire_radar
https://arxiv.org/abs/2606.01057
3DCodeBench evaluates vision-language model agents on procedural 3D modeling by converting text and image references into Blender Python code. It includes 212 object categories with ground-truth scripts, and scores outputs on executability, image similarity (SigLIP-2/DINOv3), 3D shape distance (Chamfer/Uni3D), and LLM-...
[ "General AI", "Geometric reasoning", "Multimodal" ]
[]
null
null
2026-05-31T00:00:00
unknown
https://arxiv.org/abs/2606.01057
https://github.com/gaoypeng/3dcodebench
null
1
null
0
null
null
claire-radar:2608.26947
claire-radar-2608-26947
4DSynth-Nav
claire_radar
https://arxiv.org/abs/2608.26947
Evaluates embodied agents on interactive navigation tasks in procedurally generated 4D environments with independently tunable difficulty axes.
[ "Robotics & Embodied AI", "cs.RO" ]
[]
null
null
2026-08-27T00:00:00
unknown
https://arxiv.org/abs/2608.26947
null
null
1
null
0
null
null
claire-radar:2609.09964
claire-radar-2609-09964
5-Dialects-BN
claire_radar
https://arxiv.org/abs/2609.09964
Evaluation object not clearly defined; abstract mentions alignment of Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five Bangla regional varieties, but no explicit task or scoring setup is provided.
[ "General AI", "Language & Knowledge", "cs.CL" ]
[]
null
null
2026-09-09T00:00:00
unknown
https://arxiv.org/abs/2609.09964
null
null
1
null
0
null
null
claire-radar:2605.24045
claire-radar-2605-24045
A Large-Scale Dataset and Benchmark
claire_radar
https://arxiv.org/abs/2605.24045
InteractBind provides a large-scale dataset of ~100k protein-ligand pairs with fine-grained binding-site localization tasks and interaction maps for six non-covalent interaction types, plus affinity and similarity-controlled splits.
[ "Biology & Drug Discovery", "Language & Knowledge", "cs.LG" ]
[]
null
null
2026-05-21T00:00:00
unknown
https://arxiv.org/abs/2605.24045
null
null
1
null
0
null
null
opencompass:1580
opencompass-1580-a-bench
A-Bench
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/A-Bench
A-Bench is a benchmark designed to diagnose whether LMMs are masters at evaluating AIGIs. 2,864 AIGIs from 16 text-to-image models are sampled, each paired with question-answers annotated by human experts, and tested across 18 leading LMMs. A-Bench是一个旨在诊断 LMMs 是否擅长评估 AIGIs 的基准,从 16 个文本到图像模型中采样了 2,864 个 AIGIs,每个都与由人类专家标...
[ "多模态", "Multimodal", "多模态模型", "VLM", "视觉生成", "Visual Generation", "图像理解", "Image Understanding", "不支持", "Unsupported" ]
[]
multimodal
SJTU, NTU.
2024-06-05T00:00:00
unknown
https://arxiv.org/abs/2406.03070
https://github.com/Q-Future/A-Bench
https://huggingface.co/datasets/q-future/A-Bench
1
null
0
null
null
opencompass:1367
opencompass-1367-a-okvqa
A-OKVQA
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/A-OKVQA
A-OKVQA assesses commonsense reasoning abilities. It is a crowdsourced dataset composed of a diverse set of about 25K questions requiring a broad base of commonsense and world knowledge to answer. A-OKVQA用于评估多模态大模型的常识及推理能力,由25K个不同的问题组成,需要对图像中描述的场景进行某种形式的常识性推理来回答。
[ "多模态", "Multimodal", "推理", "Reasoning", "知识", "Knowledge", "VQA", "多模态模型", "VLM", "逻辑推理", "知识储备", "不支持", "Unsupported" ]
[]
multimodal
Allen Institute for AI
2022-06-03T00:00:00
unknown
https://arxiv.org/abs/2206.01718
https://github.com/allenai/aokvqa
null
1
null
0
null
null
claire-radar:2605.22321
claire-radar-2605-22321
A3S-Bench
claire_radar
https://arxiv.org/abs/2605.22321
A3S-Bench is an executable security test suite for evaluating autonomous agents across multi-turn scenarios and agent-level risks using an action-grounded oracle.
[ "Agents", "Agents & Tool Use", "General AI" ]
[]
null
null
2026-05-21T00:00:00
unknown
https://arxiv.org/abs/2605.22321
https://github.com/antgroup/Agent3Sigma-Stage
null
1
null
0
null
null
artificial-analysis:aa-analystagent
artificial-analysis-aa-analystagent
AA-AnalystAgent
artificial_analysis
https://artificialanalysis.ai/evaluations/aa-analyst-agent
Quantitative analysis on spreadsheets & documents
[ "agentic", "business", "reasoning" ]
[]
null
null
null
unknown
null
null
null
1
30
30
0.6
null
artificial-analysis:aa-briefcase
artificial-analysis-aa-briefcase
AA-Briefcase
artificial_analysis
https://artificialanalysis.ai/evaluations/aa-briefcase
Agentic knowledge work, Elo
[ "agentic", "business" ]
[]
null
null
2026-06-18T00:00:00
unknown
null
null
null
1
65
65
1,710.26
null
llm-stats:aa-briefcase
llm-stats-aa-briefcase
AA-Briefcase
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/aa-briefcase?top_n=500
AA-Briefcase is an Artificial Analysis evaluation of AI systems on professional knowledge-work tasks, reported as an Elo score.
[ "productivity", "reasoning", "agents" ]
[]
text
null
null
unknown
null
null
null
1
3
3
1,577
null
llm-stats:aa-index
llm-stats-aa-index
AA-Index
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/aa-index?top_n=500
No official academic documentation found for this benchmark. Extensive research through ArXiv, IEEE/ACL/NeurIPS papers, and university research sites yielded no peer-reviewed sources for an 'aa-index' benchmark. This entry requires verification from official academic sources.
[ "general" ]
[]
text
null
null
unknown
null
null
null
1
4
4
0.677
null
artificial-analysis:aa-lcr
artificial-analysis-aa-lcr
AA-LCR
artificial_analysis
https://artificialanalysis.ai/evaluations/artificial-analysis-long-context-reasoning
Long context reasoning
[ "intelligence-index", "long-context", "reasoning" ]
[]
null
null
2025-08-05T00:00:00
unknown
null
null
null
1
510
510
0.833333
null
llm-stats:aa-lcr
llm-stats-aa-lcr
AA-LCR
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/aa-lcr?top_n=500
Agent Arena Long Context Reasoning benchmark
[ "long_context", "reasoning" ]
[]
text
null
null
unknown
null
null
null
1
18
18
0.8
null
model-reports:aa_lcr
aa_lcr
AA-LCR
model_reports
https://artificialanalysis.ai/evaluations/aa-lcr
Run by Artificial Analysis rather than the vendor. Third-party execution is the point, but it also means the vendor did not control the setup.
[ "long_context" ]
[]
null
null
2025-09-16T00:00:00
unknown
null
null
null
2
2
2
74.7
percent
artificial-analysis:aa-omniscience-accuracy
artificial-analysis-aa-omniscience-accuracy
AA-Omniscience Accuracy
artificial_analysis
https://artificialanalysis.ai/evaluations/omniscience
Knowledge
[ "intelligence-index", "knowledge" ]
[]
null
null
2025-11-16T00:00:00
unknown
null
null
null
1
489
489
0.6535
null
llm-stats:aa-omniscience-index
llm-stats-aa-omniscience-index
AA-Omniscience Index
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/aa-omniscience-index?top_n=500
AA-Omniscience Index is Artificial Analysis's knowledge-reliability metric. It rewards correct answers, penalizes hallucinations, and does not penalize abstention. Scores range from -100 to 100, where 0 means as many correct as incorrect answers.
[ "reasoning", "science", "knowledge" ]
[]
text
null
2025-11-16T00:00:00
unknown
null
null
null
1
2
2
126
null
artificial-analysis:aa-omniscience-non-hallucination
artificial-analysis-aa-omniscience-non-hallucination
AA-Omniscience Non-Hallucination Rate
artificial_analysis
https://artificialanalysis.ai/evaluations/omniscience
1 - hallucination rate
[ "intelligence-index", "knowledge", "faithfulness" ]
[]
null
null
2025-11-16T00:00:00
unknown
null
null
null
1
489
489
0.990991
null
claire-radar:2606.07462
claire-radar-2606-07462
AARRI-Bench
claire_radar
https://arxiv.org/abs/2606.07462
AARRI-Bench evaluates LLM agents on entry-level research intern tasks, measuring success rate in containerized environments with fixed tasks and scoring.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-06-05T00:00:00
unknown
https://arxiv.org/abs/2606.07462
https://github.com/AARR-bench/AARRI-bench
null
1
null
0
null
null
claire-radar:2606.11150
claire-radar-2606-11150
ABC-Bench
claire_radar
https://arxiv.org/abs/2606.11150
ABC-Bench evaluates LLM agents on biosecurity-relevant tasks including liquid handling robot code generation, DNA fragment design, and DNA synthesis screening evasion, with wet-lab validation.
[ "Biology & Drug Discovery", "Language & Knowledge", "Robotics & Embodied AI", "cs.AI" ]
[]
null
null
2026-06-09T00:00:00
unknown
https://arxiv.org/abs/2606.11150
null
null
1
null
0
null
null
opencompass:1148
opencompass-1148-abspyramid
AbsPyramid
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AbsPyramid
ABSPYRAMID is a unified entailment graph of 221K textual descriptions of abstraction knowledge. ABSPYRAMID collects abstract knowledge for three components of diverse events to comprehensively evaluate the abstraction ability of language models in the open domain. ABSPYRAMID 是包含 221,000 条文本描述的抽象知识,收集了多种事件的三个组成部分的抽象知识,以...
[ "知识", "Knowledge", "NAACL 2024", "大语言模型", "LLM", "知识储备", "不支持", "Unsupported" ]
[]
null
Tencent AI Lab
2024-06-16T00:00:00
unknown
null
https://github.com/HKUST-KnowComp/AbsPyramid
null
1
null
0
null
null
llm-stats:acebench
llm-stats-acebench
ACEBench
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/acebench?top_n=500
ACEBench is a comprehensive benchmark for evaluating Large Language Models' tool usage capabilities across three primary evaluation types: Normal (basic tool usage scenarios), Special (tool usage with ambiguous or incomplete instructions), and Agent (multi-agent interactions simulating real-world dialogues). The benchm...
[ "reasoning", "finance", "general", "healthcare", "tool_calling" ]
[]
text
null
null
unknown
null
null
null
1
2
2
0.765
null
claire-radar:2608.09476
claire-radar-2608-09476
ActBench
claire_radar
https://arxiv.org/abs/2608.09476
Self-evolving benchmark of behavioral safety in cowork agents, evaluating risk from execution trajectories across 600 cases, 213 scenarios, 15 risk behaviors, and 48 web-service APIs.
[ "General AI", "Safety", "Safety & Trustworthiness" ]
[]
null
null
2026-08-10T00:00:00
unknown
https://arxiv.org/abs/2608.09476
https://github.com/zjuicsr/ActBench
null
1
null
0
null
null
opencompass:1319
opencompass-1319-actionatlas
ActionAtlas
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/ActionAtlas
ActionAtlas is a multiple-choice video question answering benchmark, including 934 videos showcasing 580 unique actions across 56 sports, with a total of 1896 actions within choices. ActionAtlas是一个多项选择视频问答基准测试,包括934个视频,展示了56项运动中的580个独特动作,选项共包含1896个动作。
[ "多模态", "Multimodal", "NeurIPS 2024", "多模态模型", "VLM", "视频理解", "Video Understanding", "不支持", "Unsupported" ]
[]
multimodal
University of Washington
2024-10-08T00:00:00
unknown
https://arxiv.org/abs/2410.05774
https://github.com/mrsalehi/action-atlas
https://huggingface.co/datasets/mrsalehi/ActionAtlas-v1.0
1
null
0
null
null
claire-radar:2608.04682
claire-radar-2608-04682
Active-SWE
claire_radar
https://arxiv.org/abs/2608.04682
Active-SWE evaluates coding agents on proactive bug fixing: detecting and fixing multiple bugs without issue reports. It includes 1,663 tasks across six bug categories and eight languages, with stages for recorded bugs, potential bugs, and judge validation.
[ "Code & Software", "Software & AI Compute", "cs.SE" ]
[]
null
null
2026-08-05T00:00:00
unknown
https://arxiv.org/abs/2608.04682
https://github.com/XLearning-SCU/Active-SWE
null
1
null
0
null
null
claire-radar:2607.10180
claire-radar-2607-10180
ActiveFly-Bench
claire_radar
https://arxiv.org/abs/2607.10180
ActiveFly-Bench is a benchmark for UAV embodied perception, decomposing active perception into three tasks: Aerial Embodied Question Answering (Air-EQA), Observation Behavior Planning (OBP), and Fine-grained Language-guided UAV Control (FLUC). It includes datasets from real-world and simulated outdoor environments.
[ "Multimodal", "Reasoning", "Robotics & Embodied AI" ]
[]
null
null
2026-07-11T00:00:00
unknown
https://arxiv.org/abs/2607.10180
null
null
1
null
0
null
null
claire-radar:2605.24043
claire-radar-2605-24043
ActiveSciBench
claire_radar
https://arxiv.org/abs/2605.24043
ActiveSciBench evaluates budget-constrained, closed-loop scientific discovery across enzyme-kinetics and gene-regulatory-network tasks.
[ "AI Scientist", "General AI", "Language & Knowledge" ]
[]
null
null
2026-05-21T00:00:00
unknown
https://arxiv.org/abs/2605.24043
https://github.com/scientific-discovery/LLM-AutoSciLab
null
1
null
0
null
null
claire-radar:2607.16165
claire-radar-2607-16165
ActiveVision
claire_radar
https://arxiv.org/abs/2607.16165
ActiveVision evaluates whether multimodal large language models can perform repeated, active visual perception across 17 tasks in three categories.
[ "General AI", "Vision & 3D", "cs.CV" ]
[]
null
null
2026-07-17T00:00:00
unknown
https://arxiv.org/abs/2607.16165
null
null
1
null
0
null
null
llm-stats:activitynet
llm-stats-activitynet
ActivityNet
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/activitynet?top_n=500
A large-scale video benchmark for human activity understanding. Provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours. The benchmark covers a wide range of complex human activities that are of interest to people...
[ "video", "vision" ]
[]
video
null
null
unknown
null
null
null
1
1
1
0.619
null
claire-radar:2609.03677
claire-radar-2609-03677
AD-Diff Bench
claire_radar
https://arxiv.org/abs/2609.03677
Evaluates set-difference captioning on autonomous driving image subsets, with three splits based on annotation filtering, CLIP filtering, and web scraping, yielding natural-language descriptions of differences between target and reference sets.
[ "Autonomous Driving", "General AI", "Vision & 3D", "cs.CV" ]
[]
null
null
2026-09-03T00:00:00
unknown
https://arxiv.org/abs/2609.03677
https://github.com/KIT-MRT/AD-Diff
null
1
null
0
null
null
opencompass:1155
opencompass-1155-ada-leval
Ada-LEval
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/Ada-LEval
Ada-LEval is a length-adaptable benchmark for evaluating the long-context understanding of LLMs. Ada-LEval includes two challenging subsets, TSort and BestAnswer, which enable a more reliable evaluation of LLMs’ long context capabilities. Ada-LEval 用于评估大型语言模型(LLMs)对长上下文的理解能力。Ada-LEval 包含两个具有挑战性的子集,TSort 和 BestAnswer,能够...
[ "长文本", "Long-Context", "NAACL 2024", "大语言模型", "LLM", "长上下文", "Long Context", "官方自建", "Official", "不支持", "Unsupported" ]
[]
null
Shanghai AI Laboratory
2024-06-16T00:00:00
unknown
null
https://github.com/open-compass/Ada-LEval
null
1
null
0
null
null
claire-radar:2606.21144
claire-radar-2606-21144
AdaMem-Bench
claire_radar
https://arxiv.org/abs/2606.21144
AdaMem-Bench simulates weeks of interaction with week-by-week question answering to evaluate memory policies for personalized long-horizon LLM agents. It measures QA accuracy and memory volume across different models.
[ "General AI", "Language & Knowledge", "cs.CL" ]
[]
null
null
2026-06-19T00:00:00
unknown
https://arxiv.org/abs/2606.21144
null
null
1
null
0
null
null
claire-radar:2606.05622
claire-radar-2606-05622
AdaPlanBench
claire_radar
https://arxiv.org/abs/2606.05622
Evaluates adaptive planning of LLM agents under progressively disclosed world and user constraints across 307 household tasks.
[ "Agents & Tool Use", "General AI", "Planning", "cs.CL" ]
[]
null
null
2026-06-04T00:00:00
unknown
https://arxiv.org/abs/2606.05622
https://github.com/JiayuJeff/AdaPlanBench
null
1
null
0
null
null
claire-radar:2607.18063
claire-radar-2607-18063
Adaptive Adversaries
claire_radar
https://arxiv.org/abs/2607.18063
A 21-scenario benchmark evaluating memoryless LLM defenders against adaptive attackers that observe earlier responses and change tactics across multiple rounds.
[ "Agents", "Agents & Tool Use", "Robotics & Embodied AI" ]
[]
null
null
2026-07-20T00:00:00
unknown
https://arxiv.org/abs/2607.18063
null
null
1
null
0
null
null
claire-radar:2608.26204
claire-radar-2608-26204
ADeptS-Bench
claire_radar
https://arxiv.org/abs/2608.26204
Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark comprehensively evaluates whether they can safely interact with visual interfaces while handling ambiguou…
[ "Agents & Tool Use", "Manufacturing & Process Control", "cs.CR" ]
[]
null
null
2026-08-25T00:00:00
unknown
https://arxiv.org/abs/2608.26204
null
null
1
null
0
null
null
llm-stats:advancedif
llm-stats-advancedif
AdvancedIF
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/advancedif?top_n=500
AdvancedIF is a rubric-based benchmark measuring complex, multi-turn, and system-prompted instruction following ability, scored with a calibrated LLM judge against per-instruction rubrics.
[ "reasoning", "instruction_following", "general" ]
[]
text
null
null
unknown
null
null
null
1
2
2
0.85
null
claire-radar:2607.11849
claire-radar-2607-11849
AdvancedMathBench
claire_radar
https://arxiv.org/abs/2607.11849
AdvancedMathBench is a benchmark suite for advanced mathematical reasoning, containing ProverBench (296 proof problems) and VerifierBench (888 proof trajectories with expert labels), with an automatic verification pipeline.
[ "Language & Knowledge", "Mathematics & Formal Science", "Reasoning" ]
[]
null
null
2026-07-13T00:00:00
unknown
https://arxiv.org/abs/2607.11849
null
null
1
null
0
null
null
claire-radar:2606.24589
claire-radar-2606-24589
AdversaBench
claire_radar
https://arxiv.org/abs/2606.24589
AdversaBench is an automated LLM red-teaming evaluation that mutates seed prompts and confirms resulting failures with multiple judges and a tiebreaker.
[ "Cybersecurity", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-06-23T00:00:00
unknown
https://arxiv.org/abs/2606.24589
https://github.com/khanak0509/AdversaBench
null
1
null
0
null
null
claire-radar:2608.00832
claire-radar-2608-00832
AdvPlan-Bench
claire_radar
https://arxiv.org/abs/2608.00832
AdvPlan-Bench is an offline benchmark for adversarial evaluation of structured plan-generation agents. It uses typed action chains, adversarial response sets, selector diagnostics, and metrics like BLUE-vs-RED advantage and Nash-gap. Includes 150 synthetic scenarios across five planning templates.
[ "General AI", "Language & Knowledge", "cs.LG" ]
[]
null
null
2026-08-01T00:00:00
unknown
https://arxiv.org/abs/2608.00832
null
null
1
null
0
null
null
claire-radar:2607.14726
claire-radar-2607-14726
AE-UAV
claire_radar
https://arxiv.org/abs/2607.14726
AE-UAV is an air-to-air event-based UAV tracking benchmark with 178 flight sequences and continuous-time cubic B-spline annotations, supporting evaluation at arbitrary temporal resolutions. It includes multimodal auxiliary data and predefined train/validation/test splits.
[ "General AI", "Vision & 3D", "cs.CV" ]
[]
null
null
2026-07-16T00:00:00
unknown
https://arxiv.org/abs/2607.14726
https://github.com/MSP-xEN/AE-UAV
null
1
null
0
null
null
opencompass:2452
opencompass-2452-aecbench
AECBench
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AECBench
AECBench is an open-source benchmark for evaluating LLMs in architecture, engineering, and construction (AEC), covering 23 tasks and about 4,800 samples across five cognitive levels. AECBench 是面向建筑、工程与施工(AEC)领域的大语言模型评测基准,覆盖 5 个认知层级、23 类任务和约 4,800 个样本,用于评估模型在知识记忆、理解、推理、计算与应用方面的能力。
[ "科学", "Science", "学科", "Examination", "知识", "Knowledge", "科学智能", "AI for Science", "科学推理", "Scientific Reasoning", "知识储备", "不支持", "Unsupported" ]
[ "Chinese" ]
null
华东建筑设计研究院有限公司、同济大学
2026-04-10T00:00:00
restricted
https://arxiv.org/pdf/2509.18776
https://github.com/ArchiAI-LAB/AECBench
https://huggingface.co/datasets/jackluoluo/AECBench
1
null
0
null
null
claire-radar:2607.04383
claire-radar-2607-04383
AEGBench
claire_radar
https://arxiv.org/abs/2607.04383
AEGBench is a difficulty-stratified evaluation for open-vocabulary audio event grounding, where systems localize queried sound events in time.
[ "General AI", "Speech & Audio", "cs.SD" ]
[]
null
null
2026-07-05T00:00:00
unknown
https://arxiv.org/abs/2607.04383
null
https://huggingface.co/datasets/zihan-audio/AEGBench
1
null
0
null
null
llm-stats:community:5f95f778-c521-43fa-b80e-6a55465601e3
llm-stats-community-5f95f778-c521-43fa-b80e-6a55465601e3
ael_gate_benchmark_cases_template
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/community%3A5f95f778-c521-43fa-b80e-6a55465601e3?top_n=500
[]
[]
null
null
null
unknown
null
null
null
1
null
0
null
null
claire-radar:2608.16349
claire-radar-2608-16349
AeroCopilotBench
claire_radar
https://arxiv.org/abs/2608.16349
AeroCopilotBench is a two-tier benchmark for evaluating LLM agents as aviation copilots. Tier-1 uses 1,200 multiple-choice questions for knowledge assessment, while Tier-2 includes 73 procedural tasks in an interactive virtual cockpit environment, with safety-gated evaluation.
[ "Agents", "Agents & Tool Use", "General AI" ]
[]
null
null
2026-08-17T00:00:00
unknown
https://arxiv.org/abs/2608.16349
null
null
1
null
0
null
null
claire-radar:2608.14721
claire-radar-2608-14721
AeroGround
claire_radar
https://arxiv.org/abs/2608.14721
AeroGround evaluates vision-language models on aerial-ground collaborative reasoning using a simulated dataset of ~29,000 multimodal observation groups and 2,250 QA instances covering cross-view correspondence, spatial understanding, and reasoning.
[ "Reasoning", "Robotics & Embodied AI", "Vision & 3D" ]
[]
null
null
2026-08-12T00:00:00
unknown
https://arxiv.org/abs/2608.14721
null
null
1
null
0
null
null
llm-stats:aethercode
llm-stats-aethercode
AetherCode
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/aethercode?top_n=500
AetherCode is a competitive-programming benchmark of olympiad-level algorithmic coding problems.
[ "reasoning", "coding" ]
[]
text
null
null
unknown
null
null
null
1
2
2
0.679
null
claire-radar:2608.24954
claire-radar-2608-24954
AFDBench
claire_radar
https://arxiv.org/abs/2608.24954
Evaluates generative meteorological reasoning through 7,732 expert-written forecast discussions paired with AI weather inputs, using metrics for numerical accuracy, style, and grounding.
[ "AI Scientist", "General AI", "Language & Knowledge", "Reasoning" ]
[]
null
null
2026-08-25T00:00:00
unknown
https://arxiv.org/abs/2608.24954
null
null
1
null
0
null
null
claire-radar:2606.14240
claire-radar-2606-14240
AFFORDANCE20Q
claire_radar
https://arxiv.org/abs/2606.14240
Affordance20Q is a benchmark for evaluating affordance reasoning in LLMs using a 20-questions game. It comprises 1,009 games over 454 objects and 59 affordances, where models identify a hidden object's affordance by asking yes/no questions about physical properties.
[ "General AI", "Language & Knowledge", "Reasoning" ]
[]
null
null
2026-06-12T00:00:00
unknown
https://arxiv.org/abs/2606.14240
https://github.com/1171-jpg/Affordance20Q.git
null
1
null
0
null
null
opencompass:506
opencompass-506-afqmc
AFQMC
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AFQMC
AFQMC is an Ant Financial chinese semantic similarity task, which requires to judge whether two sentences have the same meaning or not. AFQMC一个蚂蚁金服中文语义相似度任务,要求判断两个句子是否具有相同的语义。
[ "语言", "Language", "大语言模型", "LLM", "语言理解", "Comprehension", "开源收录", "Open-Source", "不支持", "Unsupported" ]
[ "Chinese" ]
null
null
2020-04-13T00:00:00
unknown
https://arxiv.org/abs/2209.02970
https://github.com/IDEA-CCNL/Fengshenbang-LM
null
1
null
0
null
null
claire-radar:2608.23628
claire-radar-2608-23628
AFT-Bench
claire_radar
https://arxiv.org/abs/2608.23628
Holds task, backend, initial state, injected failure, agent, and language model fixed while varying the tool interface to measure callability versus operability.
[ "Agents", "Code & Software", "General AI" ]
[]
null
null
2026-08-23T00:00:00
unknown
https://arxiv.org/abs/2608.23628
null
null
1
null
0
null
null
claire-radar:2607.01152
claire-radar-2607-01152
AGC-Bench
claire_radar
https://arxiv.org/abs/2607.01152
AGC-Bench evaluates artificial general creativity across 78 datasets covering brainstorming, problem solving, STEM, narrative, figurative language, and humor. It uses an agentic harness and a public leaderboard with an open-weight judge model.
[ "General AI", "Language & Knowledge", "cs.CL" ]
[]
null
null
2026-07-01T00:00:00
unknown
https://arxiv.org/abs/2607.01152
null
null
1
null
0
null
null
claire-radar:github:kyashrathore/agent-app-benchmark
claire-radar-github-kyashrathore-agent-app-benchmark
Agent App Benchmark
claire_radar
https://github.com/kyashrathore/agent-app-benchmark
Evaluates GUI application performance for coding agents using deterministic historical session workloads, measuring app start, session switching, memory, and CPU metrics.
[ "Agents", "Language & Knowledge", "Software & AI Compute" ]
[]
null
null
2026-08-22T00:00:00
unknown
null
https://github.com/kyashrathore/agent-app-benchmark
null
1
null
0
null
null
claire-radar:github:datapace-ai/agent-memory-benchmark
claire-radar-github-datapace-ai-agent-memory-benchmark
Agent Memory Benchmark
claire_radar
https://github.com/datapace-ai/agent-memory-benchmark
Evaluates agent memory systems on a 10-question pilot using LongMemEval, measuring answer accuracy under three judge rules plus token usage, latency, ingestion cost, forgetting curve, and superseded-value handling.
[ "Agents", "General AI", "Language & Knowledge" ]
[]
null
null
2026-09-08T00:00:00
unknown
null
https://github.com/datapace-ai/agent-memory-benchmark
null
1
null
0
null
null
claire-radar:2606.04874
claire-radar-2606-04874
Agent Planning Benchmark
claire_radar
https://arxiv.org/abs/2606.04874
Agent Planning Benchmark (APB) is a diagnostic benchmark with 4,209 multimodal cases across 22 domains, evaluating planning capabilities in five settings including tool noise and unsolvable tasks.
[ "Agents", "General AI", "Multimodal", "Planning", "Robustness", "Safety & Trustworthiness" ]
[]
null
null
2026-06-03T00:00:00
unknown
https://arxiv.org/abs/2606.04874
https://github.com/Mikivishy/AgentPlanningBenchmark
null
1
null
0
null
null
claire-radar:2607.24882
claire-radar-2607-24882
Agent Retrieval Bench
claire_radar
https://arxiv.org/abs/2607.24882
File-level retrieval benchmark for coding agents, covering four positive tasks (code2test, comment2context, trace2code, edit2ripple) and a selective-retrieval subset with natural no-gold and counterfactual controls across 25 repositories, 427 samples, with frozen base-commit corpora.
[ "Agents", "Information retrieval", "Language & Knowledge", "Software & AI Compute" ]
[]
null
null
2026-07-27T00:00:00
unknown
https://arxiv.org/abs/2607.24882
https://github.com/eyuansu62/agent-retrieval-bench
null
1
null
0
null
null
llm-stats:agent-startup-bench
llm-stats-agent-startup-bench
Agent Startup Bench
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/agent-startup-bench?top_n=500
Agent Startup Bench measures AI agents on high-economic-value, startup-style tasks that require autonomous planning and execution to deliver practical, verifiable results.
[ "reasoning", "general", "agents" ]
[]
text
null
null
unknown
null
null
null
1
2
2
0.688
null
claire-radar:github:martin-beck/agent-systems-benchmark
claire-radar-github-martin-beck-agent-systems-benchmark
Agent Systems Benchmark (ASB)
claire_radar
https://github.com/martin-beck/agent-systems-benchmark
A Linux terminal framework for measuring how AI coding agents scale under concurrent sessions while maintaining task quality, latency, and resource consumption within declared bounds.
[ "Agents", "Language & Knowledge", "Software & AI Compute" ]
[]
null
null
2026-09-06T00:00:00
unknown
null
https://github.com/martin-beck/agent-systems-benchmark
null
1
null
0
null
null
claire-radar:github:sentinelden/agent-injection-bench
claire-radar-github-sentinelden-agent-injection-bench
agent-injection-bench
claire_radar
https://github.com/sentinelden/agent-injection-bench
Evaluates prompt-injection robustness of tool-calling agent stacks using a 24-scenario corpus (21 adversarial, 3 control) delivered through channels such as tool results, documents, web pages, filenames and multi-turn dialogue, with deterministic canary, tool-call, exfiltration and refusal success conditions; results r...
[ "Agents", "General AI", "Language & Knowledge" ]
[]
null
null
2026-09-11T00:00:00
unknown
null
https://github.com/sentinelden/agent-injection-bench
null
1
null
0
null
null
claire-radar:github:ajbermudezh22/agent-model-bench
claire-radar-github-ajbermudezh22-agent-model-bench
agent-model-bench
claire_radar
https://github.com/ajbermudezh22/agent-model-bench
Evaluates language models on tool calling and strict JSON adherence using 32 cases, scoring exact tool and argument matches and JSON parsing with strict and content-based criteria.
[ "Agents", "General AI", "Language & Knowledge" ]
[]
null
null
2026-08-21T00:00:00
unknown
null
https://github.com/ajbermudezh22/agent-model-bench
null
1
null
0
null
null
claire-radar:2607.10059
claire-radar-2607-10059
AgentAbstain
claire_radar
https://arxiv.org/abs/2607.10059
AgentAbstain is a paired-task benchmark for evaluating LLM agents' ability to abstain from acting in scenarios such as ambiguity, conflicting constraints, or tool failures. It includes 263 paired tasks across 42 sandbox environments, with a proposed pipeline for generating fresh task instances.
[ "Agents", "General AI", "Language & Knowledge", "Reasoning" ]
[]
null
null
2026-07-11T00:00:00
unknown
https://arxiv.org/abs/2607.10059
null
null
1
null
0
null
null
opencompass:1242
opencompass-1242-agentboard
AgentBoard
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AgentBoard
AgentBoard is tailored to analytical evaluation of LLM agents. It offers a fine-grained progress rate metric that captures incremental advancements as well as a comprehensive evaluation toolkit that features easy assessment of agents for multi-faceted analysis through interactive visualization. AgentBoard专用于LLM Agent的分...
[ "智能体", "Agent", "NeurIPS 2024", "任务执行", "Task Execution", "不支持", "Unsupported" ]
[]
null
The University of Hong Kong
2024-06-24T00:00:00
restricted
https://arxiv.org/abs/2401.13178
https://github.com/hkust-nlp/AgentBoard
https://huggingface.co/datasets/hkust-nlp/agentboard
1
null
0
null
null
claire-radar:2608.14680
claire-radar-2608-14680
AGENTCHAOSBENCH
claire_radar
https://arxiv.org/abs/2608.14680
AGENTCHAOSBENCH is a dataset of sanitized execution traces from five agentic applications with injected runtime faults, used to evaluate fault detection and localization from telemetry.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-08-04T00:00:00
unknown
https://arxiv.org/abs/2608.14680
null
null
1
null
0
null
null
claire-radar:2606.23189
claire-radar-2606-23189
AgentCIBench
claire_radar
https://arxiv.org/abs/2606.23189
AgentCIBench is an executable evaluation harness for testing whether computer-use agents improperly disclose information across application contexts.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-06-22T00:00:00
unknown
https://arxiv.org/abs/2606.23189
null
null
1
null
0
null
null
claire-radar:2609.06972
claire-radar-2609-06972
AgentDrift
claire_radar
https://arxiv.org/abs/2609.06972
Step-labeled corpus of 12,536 synthetic LLM agent tool-call trajectories in five domains, labeled per step for benign, injection point, hijacked, or failed injection, supporting detection, localization, and attempt-vs-success evaluation.
[ "Agents", "General AI", "Language & Knowledge" ]
[]
null
null
2026-09-07T00:00:00
unknown
https://arxiv.org/abs/2609.06972
https://github.com/Asif-0209/AgentDrift
null
1
null
0
null
null
claire-radar:2606.16723
claire-radar-2606-16723
AgentFairBench
claire_radar
https://arxiv.org/abs/2606.16723
AgentFairBench evaluates demographic disparity in the actions of LLM agents across hiring, lending, and medical triage. It uses synthetic, demographic-neutral profiles in counterfactual matched sets varying name-coded race/gender. Metrics include counterfactual flip rate, mean absolute score difference, action-rate dis...
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-06-15T00:00:00
unknown
https://arxiv.org/abs/2606.16723
null
null
1
null
0
null
null
opencompass:1351
opencompass-1351-agentharm
AgentHarm
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AgentHarm
AgentHarm tests the robustness of LLMs to jailbreak attacks. It includes a diverse set of 110 explicitly malicious agent tasks (440 with augmentations), covering 11 harm categories including fraud, cybercrime, and harassment. AgentHarm用于评估LLM智能体对越狱攻击的鲁棒性,包括110套恶意智能体任务(其中有440个强化任务),涵盖欺诈、网络犯罪和骚扰等11个危害类别。
[ "安全", "Safety", "智能体", "Agent", "任务执行", "Task Execution", "安全对齐", "Safety Alignment", "不支持", "Unsupported" ]
[]
null
Gray Swan AI
2024-10-11T00:00:00
restricted
https://arxiv.org/abs/2404.02151
https://github.com/UKGovernmentBEIS/inspect_evals
https://huggingface.co/datasets/ai-safety-institute/AgentHarm
1
null
0
null
null
opencompass:2061
opencompass-2061-agenthazard
AgentHazard
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AgentHazard
移动端 GUI Agent 通过与设备环境的交互来完成任务,在完成任务的过程中会遇到一些未知或不可信的信息来源,这些信息可能含有攻击性的内容,致使 Agent 无法正常完成任务,甚至对用户的隐私和财产带来危害。本评测集兼具动态执行环境和静态评测数据集,旨在为移动端 GUI Agent 提供一个仿真度高的模拟环境,以评估其在真实场景下执行的行为和安全性。 移动端 GUI Agent 通过与设备环境的交互来完成任务,在完成任务的过程中会遇到一些未知或不可信的信息来源,这些信息可能含有攻击性的内容,致使 Agent 无法正常完成任务,甚至对用户的隐私和财产带来危害。本评测集兼具动态执行环境和静态评测数据集,旨在为移动端 GUI Agent...
[ "多模态", "Multimodal", "安全", "Safety", "智能体", "Agent", "任务执行", "Task Execution", "安全对齐", "Safety Alignment", "跨模态推理", "Cross-modal Reasoning", "不支持", "Unsupported" ]
[]
multimodal
Institute for AI Industry Research, Tsinghua University
2025-07-16T00:00:00
unknown
https://arxiv.org/abs/2507.04227
https://github.com/Zsbyqx20/AgentHazard
null
1
null
0
null
null
claire-radar:2605.25707
claire-radar-2605-25707
AgentHijack
claire_radar
https://arxiv.org/abs/2605.25707
Evaluates the robustness of computer use agents under common environment corruptions such as pop-ups, resolution changes, and competing applications. The benchmark introduces 9 configurable corruptions and evaluates agent performance on desktop tasks using multimodal LLM-based agents, measuring task completion rates.
[ "Agents", "Agents & Tool Use", "General AI", "Robustness" ]
[]
null
null
2026-05-25T00:00:00
unknown
https://arxiv.org/abs/2605.25707
https://github.com/tmlr-group/AgentHijack
null
1
null
0
null
null
claire-radar:2607.29626
claire-radar-2607-29626
AgentHPOBench
claire_radar
https://arxiv.org/abs/2607.29626
AgentHPOBench evaluates LLM agents as sequential hyperparameter optimizers across 30 executable ML tasks. Agents observe accumulated configurations, metrics, and logs, then propose the next configuration. Scoring compares agents and conventional HPO baselines under a unified protocol.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-07-31T00:00:00
unknown
https://arxiv.org/abs/2607.29626
null
null
1
null
0
null
null
claire-radar:github:bluebunnyanon/agent-benchmark-2d-maze
claire-radar-github-bluebunnyanon-agent-benchmark-2d-maze
Agentic Evaluation in 2D Mazes
claire_radar
https://github.com/bluebunnyanon/agent-benchmark-2d-maze
Evaluates interactive multimodal agents in structured 2D maze environments requiring long-horizon planning, mechanism interaction, and error recovery, scored by mechanism-aware progress.
[ "Agents", "Agents & Tool Use", "Multimodal", "Software & AI Compute" ]
[]
null
null
2026-08-22T00:00:00
unknown
null
https://github.com/bluebunnyanon/agent-benchmark-2d-maze
null
1
null
0
null
null
claire-radar:github:aah20/agentic-infrastructure-change-benchmark
claire-radar-github-aah20-agentic-infrastructure-change-benchmark
Agentic Infrastructure ChangeBench
claire_radar
https://github.com/AAH20/agentic-infrastructure-change-benchmark
Evaluates AI agent proposals for cloud, Kubernetes, and network infrastructure changes across correctness, safety, reliability, economics, efficiency, and evidence using deterministic scenario-based checks.
[ "Finance", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-09-05T00:00:00
unknown
null
https://github.com/AAH20/agentic-infrastructure-change-benchmark
null
1
null
0
null
null
claire-radar:github:andersballegaard/agentic-netdevops-benchmark
claire-radar-github-andersballegaard-agentic-netdevops-benchmark
Agentic NetDevOps Benchmark
claire_radar
https://github.com/AndersBallegaard/agentic-netdevops-benchmark
A multi-stage environment with a Linux management host and five VyOS routers, requiring agents to complete netdevops tasks from information gathering to evaluation.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-09-08T00:00:00
unknown
null
https://github.com/AndersBallegaard/agentic-netdevops-benchmark
null
1
null
0
null
null
claire-radar:github:gitayg/agentic-security-benchmark
claire-radar-github-gitayg-agentic-security-benchmark
Agentic Security Benchmark
claire_radar
https://github.com/gitayg/agentic-security-benchmark
Measures detection and prevention of agentic-AI security products on 286 attack and 875 benign samples across five AMTSO attack vectors via a scoring adapter harness.
[ "Language & Knowledge", "Software & AI Compute", "cs.AI" ]
[]
null
null
2026-09-07T00:00:00
unknown
null
https://github.com/gitayg/agentic-security-benchmark
null
1
null
0
null
null
claire-radar:2607.01647
claire-radar-2607-01647
AgenticDataBench
claire_radar
https://arxiv.org/abs/2607.01647
Evaluates LLM-based data agents on realistic data science workflows across 15 domains, with fine-grained ground-truth labels and skill-level scoring.
[ "General AI", "Language & Knowledge", "cs.DB" ]
[]
null
null
2026-07-02T00:00:00
unknown
https://arxiv.org/abs/2607.01647
https://github.com/AgenticDataBench/AgenticDataBench
null
1
null
0
null
null
claire-radar:2606.24026
claire-radar-2606-24026
AgenticInterpBench
claire_radar
https://arxiv.org/abs/2606.24026
Evaluates language model agents on explaining components of transformer circuits, with 84 semi-synthetic circuits and 163 component-level annotations.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-06-23T00:00:00
unknown
https://arxiv.org/abs/2606.24026
null
null
1
null
0
null
null
claire-radar:2607.02255
claire-radar-2607-02255
AgenticSTS
claire_radar
https://arxiv.org/abs/2607.02255
AgenticSTS is a reproducible long-horizon agent testbed that evaluates bounded, typed-retrieval memory configurations through game outcomes and controlled ablations.
[ "Agents", "General AI", "Language & Knowledge" ]
[]
null
null
2026-07-02T00:00:00
unknown
https://arxiv.org/abs/2607.02255
null
null
1
null
0
null
null
claire-radar:github:aah20/ai-agent-infrastructure-benchmark
claire-radar-github-aah20-ai-agent-infrastructure-benchmark
AgentInfraBench
claire_radar
https://github.com/AAH20/ai-agent-infrastructure-benchmark
Evaluates AI agent infrastructure across isolation, identity, network, authority, durability, observability, and unit economics using weighted scores with hard safety gates.
[ "Agents", "Language & Knowledge", "Software & AI Compute" ]
[]
null
null
2026-09-05T00:00:00
unknown
null
https://github.com/AAH20/ai-agent-infrastructure-benchmark
null
1
null
0
null
null
claire-radar:2608.26623
claire-radar-2608-26623
AgentJudgeBench
claire_radar
https://arxiv.org/abs/2608.26623
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-08-27T00:00:00
unknown
https://arxiv.org/abs/2608.26623
null
null
1
null
0
null
null
claire-radar:2607.06624
claire-radar-2607-06624
AgentLens
claire_radar
https://arxiv.org/abs/2607.06624
AgentLens evaluates interactive coding agents across entire trajectories, combining formal verification with LLM-written trajectory reviews and side-by-side comparisons to score dimensions such as instruction compliance, tool use, and interaction style.
[ "Agents", "Agents & Tool Use", "General AI" ]
[]
null
null
2026-07-07T00:00:00
unknown
https://arxiv.org/abs/2607.06624
https://github.com/agent-lens/agent-lens-bench
null
1
null
0
null
null
claire-radar:2608.00009
claire-radar-2608-00009
AgentMemBench
claire_radar
https://arxiv.org/abs/2608.00009
AgentMemBench evaluates five long-term memory management strategies for conversational AI agents across three public datasets (LoCoMo, MultiDoc2Dial, MSC), covering multi-session dialogue, document grounding, and persona-grounded chat. Scoring uses retrieval metrics (Recall@k, MRR, nDCG@k), Answer F1, LLM-judge faithfu...
[ "General AI", "Language & Knowledge", "cs.CL" ]
[]
null
null
2026-06-16T00:00:00
unknown
https://arxiv.org/abs/2608.00009
null
null
1
null
0
null
null
claire-radar:2606.02240
claire-radar-2606-02240
AgentRedBench
claire_radar
https://arxiv.org/abs/2606.02240
AgentRedBench evaluates LLM agents against indirect prompt injection and underspecified-authorization attacks across 24 enterprise SaaS integrations. It defines 215 attack scenarios with immutable versioning, and tracks attack success rate (ASR) for models and defenses. Open-source codebase and schemas enable replay.
[ "General AI", "Language & Knowledge", "cs.CR" ]
[]
null
null
2026-06-01T00:00:00
unknown
https://arxiv.org/abs/2606.02240
null
null
1
null
0
null
null
claire-radar:2608.15286
claire-radar-2608-15286
AgentRelBench
claire_radar
https://arxiv.org/abs/2608.15286
AgentRelBench measures repeated agent runs in a fixed task suite, computing severity-weighted damage from database state diffs with no LLM judging.
[ "Agents", "Agents & Tool Use", "General AI" ]
[]
null
null
2026-08-15T00:00:00
unknown
https://arxiv.org/abs/2608.15286
https://github.com/shivenkk/agentrelbench
null
1
null
0
null
null
opencompass:1778
opencompass-1778-agentrewardbench
AgentRewardBench
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AgentRewardBench
AgentRewardBench, the first benchmark to assess the effectiveness of LLM judges for evaluating web agents. AgentRewardBench contains 1302 trajectories across 5 benchmarks and 4 LLMs. AgentRewardBench 是首个用于评估大型语言模型(LLM)评判者评估网络代理有效性的基准测试。AgentRewardBench 包含来自 5 个基准测试和 4 个大型语言模型的 1302 条轨迹。
[ "推理", "Reasoning", "智能体", "Agent", "任务执行", "Task Execution", "不支持", "Unsupported" ]
[]
null
McGill University,Mila Quebec AI Institute,etc.
2025-04-11T00:00:00
restricted
https://arxiv.org/abs/2504.08942
null
https://huggingface.co/datasets/McGill-NLP/agent-reward-bench
1
null
0
null
null
llm-stats:agents-last-exam
llm-stats-agents-last-exam
Agents' Last Exam
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/agents-last-exam?top_n=500
Agents' Last Exam is a challenging benchmark for AI agents on hard, long-horizon tasks that test sustained reasoning, planning, and tool use, reported with and without tool access.
[ "reasoning", "agents", "tool_calling" ]
[]
text
null
null
unknown
null
null
null
1
10
10
0.527
null
model-reports:agents_last_exam
agents_last_exam
Agents' Last Exam
model_reports
https://lastexam.ai/
Multi-step agentic benchmark using Claude Code harness. Tool Search disabled. Scores depend on harness, reasoning effort, context length, and timeout settings.
[ "coding_agent" ]
[]
null
null
2025-01-23T00:00:00
unknown
null
null
null
3
3
3
31.8
percent
claire-radar:2606.05405
claire-radar-2606-05405
Agents' Last Exam (ALE)
claire_radar
https://arxiv.org/abs/2606.05405
A living benchmark of 1K+ long-horizon professional tasks with verifiable outcomes, covering 55 sub-fields in 13 industry clusters based on O*NET / SOC 2018.
[ "General AI", "Language & Knowledge", "cs.AI" ]
[]
null
null
2026-06-03T00:00:00
unknown
https://arxiv.org/abs/2606.05405
https://github.com/rdi-berkeley/agents-last-exam
null
1
null
0
null
null
claire-radar:2607.27294
claire-radar-2607-27294
AgentS4D
claire_radar
https://arxiv.org/abs/2607.27294
AgentS4D evaluates runtime safety of LLM-based workspace agents across a four-dimensional framework, with 328 risk-injected cases and seven lifecycle checkpoints, measuring unsafe behavior and evidence across six risk-entry sources and nine harms.
[ "General AI", "Safety", "Safety & Trustworthiness" ]
[]
null
null
2026-07-29T00:00:00
unknown
https://arxiv.org/abs/2607.27294
null
null
1
null
0
null
null
claire-radar:2608.00805
claire-radar-2608-00805
AgentSLABench
claire_radar
https://arxiv.org/abs/2608.00805
A resource-aware agent benchmark with 16 Dockerized task environments that profiles correctness alongside latency, cost, compute, memory, and network use under declared budgets.
[ "Agents & Tool Use", "Software & AI Compute", "cs.AI" ]
[]
null
null
2026-08-01T00:00:00
unknown
https://arxiv.org/abs/2608.00805
https://github.com/MeherBhaskar/agentslabench
null
1
null
0
null
null
claire-radar:2608.15127
claire-radar-2608-15127
AgentSysBench
claire_radar
https://arxiv.org/abs/2608.15127
AgentSysBench is a benchmark suite and measurement toolkit for characterizing agentic workloads on LLM serving systems. It includes ten representative agentic applications and unified instrumentation, identifying six properties that distinguish agentic workloads from conventional LLM inference.
[ "General AI", "Language & Knowledge", "cs.OS" ]
[]
null
null
2026-08-15T00:00:00
unknown
https://arxiv.org/abs/2608.15127
null
null
1
null
0
null
null
claire-radar:2606.24597
claire-radar-2606-24597
AgentWorldBench
claire_radar
https://arxiv.org/abs/2606.24597
AgentWorldBench evaluates language world models on simulation fidelity across 7 domains (MCP, Search, Terminal, SWE, Android, Web, OS) using real-world trajectories and rubric-based scoring.
[ "General AI", "Language & Knowledge", "cs.CL" ]
[]
null
null
2026-06-23T00:00:00
unknown
https://arxiv.org/abs/2606.24597
https://github.com/QwenLM/Qwen-AgentWorld
https://huggingface.co/datasets/Qwen/AgentWorldBench
1
null
0
null
null
claire-radar:2609.02339
claire-radar-2609-02339
AGI Maze Prediction Datasets and Benchmark
claire_radar
https://arxiv.org/abs/2609.02339
Evaluates predictive models on maze prediction tasks to study world dynamics learning.
[ "General AI", "Language & Knowledge", "cs.LG" ]
[]
null
null
2026-09-02T00:00:00
unknown
https://arxiv.org/abs/2609.02339
null
null
1
null
0
null
null
llm-stats:agieval
llm-stats-agieval
AGIEval
llm_stats
https://api.zeroeval.com/leaderboard/benchmarks/agieval?top_n=500
A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission tests (LSAT), math competitions, lawyer qualification tests, and civil service exams. Contains 20 tasks (18 multiple-choice, 2 cloze) designed to assess understanding, kn...
[ "legal", "math", "reasoning", "general" ]
[]
text
null
null
unknown
null
null
null
1
10
10
0.658
null
model-reports:agieval
agieval
AGIEval
model_reports
https://github.com/ruixiangcui/AGIEval
Human-exam derived; overlaps heavily with MMLU-style coverage.
[ "knowledge" ]
[]
null
null
2023-04-13T00:00:00
unknown
null
https://github.com/ruixiangcui/AGIEval
null
4
null
0
null
null
opencompass:497
opencompass-497-agieval
AGIEval
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AGIEval
AGIEval is a human-centric benchmark specifically designed to evaluate the general abilities of foundation models in tasks pertinent to human cognition and problem-solving. This benchmark is derived from 20 official, public, and high-standard admission and qualification exams intended for general human test-takers, suc...
[ "学科", "Examination", "大语言模型", "LLM", "知识储备", "Knowledge", "开源收录", "Open-Source", "不支持", "Unsupported" ]
[ "Chinese" ]
null
null
2023-09-18T00:00:00
unknown
https://arxiv.org/pdf/2304.06364
https://github.com/ruixiangcui/AGIEval
null
1
null
0
null
null
claire-radar:2605.26302
claire-radar-2605-26302
AgingBench
claire_radar
https://arxiv.org/abs/2605.26302
Evaluates the reliability of deployed AI agents over extended operational lifetimes. The benchmark organizes agent aging into four mechanisms—compression, interference, revision, and maintenance—and uses temporal dependency graphs and paired counterfactual probes to produce diagnostic profiles of the memory pipeline's ...
[ "Agents", "General AI", "Language & Knowledge" ]
[]
null
null
2026-05-25T00:00:00
unknown
https://arxiv.org/abs/2605.26302
https://github.com/VITA-Group/AgingBench
null
1
null
0
null
null
opencompass:1753
opencompass-1753-agmmu
AgMMU
opencompass_hub
https://hub.opencompass.org.cn/dataset-detail/AgMMU
A Comprehensive Agricultural Multimodal Understanding and Reasoning Benchmark 农业综合多模态理解和推理基准。
[ "多模态", "Multimodal", "农业", "多模态模型", "VLM", "跨模态推理", "Cross-modal Reasoning", "不支持", "Unsupported" ]
[]
multimodal
Rice University,etc.
2025-04-14T00:00:00
unknown
https://arxiv.org/abs/2504.10568
https://github.com/AgMMU/AgMMU
null
1
null
0
null
null
claire-radar:2606.24526
claire-radar-2606-24526
AGORA
claire_radar
https://arxiv.org/abs/2606.24526
Agora evaluates agentic document reasoning across eight domain collections of 9,664 authentic workplace documents. It includes 362 questions requiring location of sparse evidence and reconciliation of terminology, units, and time conventions. The benchmark is designed to exceed model context windows, necessitating deli...
[ "General AI", "Language & Knowledge", "Reasoning" ]
[]
null
null
2026-06-23T00:00:00
unknown
https://arxiv.org/abs/2606.24526
null
null
1
null
0
null
null
claire-radar:2609.08402
claire-radar-2609-08402
AGOS-Bench
claire_radar
https://arxiv.org/abs/2609.08402
AGOS-Bench evaluates VLMs on the urban Air-Ground Object Search task where a UAV and UGV coordinate to locate a target vehicle from multi-view visual references, reporting success rate and success weighted by path length.
[ "Multimodal", "Robotics & Embodied AI" ]
[]
null
null
2026-09-08T00:00:00
unknown
https://arxiv.org/abs/2609.08402
null
null
1
null
0
null
null
claire-radar:2605.22366
claire-radar-2605-22366
AgroTools
claire_radar
https://arxiv.org/abs/2605.22366
AgroTools evaluates tool-augmented multimodal agents in agriculture, with 539 QA instances, 1,097 images, 14 executable tools, and structured tool-use traces for process and outcome evaluation.
[ "General AI", "Multimodal" ]
[]
null
null
2026-05-21T00:00:00
unknown
https://arxiv.org/abs/2605.22366
null
https://huggingface.co/datasets/AgroTools/AgroTools
1
null
0
null
null
End of preview. Expand in Data Studio

Benchmark Radar Dataset

Paper Website GitHub HF Dataset

Overview

Benchmark Radar is a living registry, search engine, and discovery pipeline for AI evaluation benchmarks. This dataset mirrors the full-corpus findings of the Benchmark Radar technical report ("Benchmark Radar: Daily Discovery and Full-Corpus Search Across the AI Evaluation Landscape", arXiv:2609.11115).

As described in the paper's Two Input Paths framework, Benchmark Radar combines two complementary systems:

  1. Catalog & Scores: A curated, normalized cross-registry archive of 3,212 benchmark suites and 13,022 reported model score observations across LLM Stats, OpenCompass Hub, Artificial Analysis, and premier model technical reports.
  2. Radar Discoveries: A 24/7 automated intelligence pipeline tracking emerging papers, code repositories, datasets, and community attention across 37+ sources, capturing 13,275 research artifacts and 23,142 daily observations.

Dataset Structure & Usage

The dataset is partitioned into 4 configs (subsets), easily loaded via Hugging Face datasets:

from datasets import load_dataset

# 1. Load benchmark catalog (default)
catalog = load_dataset("ktwu01/benchmark-radar", "catalog")

# 2. Load model evaluation scores
scores = load_dataset("ktwu01/benchmark-radar", "scores")

# 3. Load emerging radar artifacts (papers, repositories, datasets)
artifacts = load_dataset("ktwu01/benchmark-radar", "radar_artifacts")

# 4. Load daily discovery observations (attention, downloads, events)
observations = load_dataset("ktwu01/benchmark-radar", "radar_observations")

1. catalog (Default Config)

Contains 3,212 normalized benchmark records.

  • benchmark_id: Canonical unique identifier (e.g. llm-stats:mmlu-pro, opencompass:1580)
  • slug: URL-safe unique identifier
  • name: Benchmark display name
  • source: Source provider (llm_stats, opencompass_hub, artificial_analysis, model_reports)
  • source_url: URL to original source entry
  • description: Plaintext description and task summary
  • categories: Assigned category tags (e.g. Reasoning, Code, Multimodal)
  • languages: Languages evaluated (e.g. en, zh)
  • modality: Evaluated modality (text, multimodal, code, etc.)
  • publisher: Creator organization or research lab
  • released_at: Official release date (YYYY-MM-DD proxy)
  • openness: Data and code openness tier
  • paper_url: Associated paper URL (arXiv or DOI)
  • repo_url: Associated code repository URL (GitHub)
  • dataset_url: Direct dataset download or Hub URL
  • document_count: Number of verified source documents citing this benchmark
  • model_count: Number of distinct models evaluated
  • score_count: Number of numeric scores recorded
  • highest_score: Maximum observed score across all evaluated models
  • score_unit: Score metric / unit (e.g. %, accuracy, Elo)

2. scores

Contains 13,022 reported evaluation score observations across 870+ frontier models.

  • obs_id: Unique observation identifier
  • key: Benchmark key
  • model_id: Source-specific stable model identifier, or null when the source does not provide one
  • model_name: Model display name
  • organization: Model creator/lab (e.g. DeepSeek, OpenAI, Anthropic, Google, Meta)
  • value: Numeric reported score
  • raw_value: Original score string from source
  • value_kind: Value data type (number, percentage, etc.)
  • reported_date: Model announcement or report publication date
  • date_precision: Precision level of reported date
  • reported_by: Source reporting modality (self_reported, third_party)
  • source: Score ingestion source
  • source_url: URL of source leaderboard or technical report
  • document_id: Source document citation ID

3. radar_artifacts

Contains 13,275 academic artifacts surfaced by daily radar.

  • id: Artifact identifier (e.g. artifact:arxiv:2203.17257)
  • type: Entity type (artifact)
  • label: Title or repository name
  • url: Primary external URL
  • categories: Extracted capability categories
  • sources: Discovery sources that observed this artifact
  • first_seen_at: First discovery date (YYYY-MM-DD)
  • last_seen_at: Latest discovery date (YYYY-MM-DD)
  • observation_count: Total appearances across snapshots
  • latest_score: Attention and composite radar score
  • metrics: Dictionary of observed signals (e.g. stars, citations, downloads)

4. radar_observations

Contains 23,142 discrete daily discovery events.

  • id: Observation identifier
  • entity_id: Referenced artifact ID
  • snapshot_date: Radar snapshot date (YYYY-MM-DD)
  • source: Discovery source platform
  • source_id: Upstream source identifier
  • url: Surfaced event URL
  • discovered_at: Timestamp of discovery
  • published_at: Original creation / publication timestamp
  • total_score: Radar composite relevance score
  • event_kind: Event category (discovered, released, updated)
  • organizations: Identified affiliated organizations
  • metrics: Point-in-time metrics (stars, likes, downloads)

Data Provenance and Principles

  • Full-Corpus Coverage: All 3,212 benchmarks across all sources are preserved. Unscored benchmarks are retained with verified paper, code, and dataset links.
  • Strict Evidence Citation: Every score links directly to its source document or technical report.
  • Daily Automated Sync: Radar discoveries are refreshed daily at 05:00 UTC via GitHub Actions.

Licensing

The technical report and Benchmark Radar's original editorial content are available under CC BY-NC-SA 4.0. Commercial dataset packaging or product integration requires prior written permission. Third-party benchmark metadata and source material retain their original terms. Review the repository licensing notice before reuse, especially for commercial dataset packaging.

Citation

@article{benchmark_radar_2026,
  title={Benchmark Radar: Daily Discovery and Full-Corpus Search
          Across the AI Evaluation Landscape},
  author={Wu, Koutian and Contributors},
  journal={arXiv preprint arXiv:2609.11115},
  year={2026}
}
Downloads last month
433

Papers for ktwu01/benchmark-radar