The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
Error code: DatasetGenerationError
Exception: ArrowInvalid
Message: Schema at index 2 was different:
enabled: bool
reason: string
stdout_bytes_omitted: int64
stderr_bytes_omitted: int64
parsed_event_count_omitted: int64
vs
submission_exists: double
build_success: double
task_score: double
overall_score: double
reward: double
error_message: string
Traceback: Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
writer.write_table(table)
~~~~~~~~~~~~~~~~~~^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 764, in write_table
self.write_rows_on_file() # in case there are buffered rows to write first
~~~~~~~~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 662, in write_rows_on_file
table = pa.concat_tables(self.current_rows)
File "pyarrow/table.pxi", line 6321, in pyarrow.lib.concat_tables
File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
return check_status(status)
File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
raise convert_status(status)
pyarrow.lib.ArrowInvalid: Schema at index 2 was different:
enabled: bool
reason: string
stdout_bytes_omitted: int64
stderr_bytes_omitted: int64
parsed_event_count_omitted: int64
vs
submission_exists: double
build_success: double
task_score: double
overall_score: double
reward: double
error_message: string
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1869, in _prepare_split_single
num_examples, num_bytes = writer.finalize()
~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 781, in finalize
self.write_rows_on_file()
~~~~~~~~~~~~~~~~~~~~~~~^^
File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 662, in write_rows_on_file
table = pa.concat_tables(self.current_rows)
File "pyarrow/table.pxi", line 6321, in pyarrow.lib.concat_tables
File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
raise convert_status(status)
pyarrow.lib.ArrowInvalid: Schema at index 2 was different:
enabled: bool
reason: string
stdout_bytes_omitted: int64
stderr_bytes_omitted: int64
parsed_event_count_omitted: int64
vs
submission_exists: double
build_success: double
task_score: double
overall_score: double
reward: double
error_message: string
The above exception was the direct cause of the following exception:
Traceback (most recent call last):
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
~~~~~~~~~~~~~~~~~~~~~~~~~^
builder, max_dataset_size_bytes=max_dataset_size_bytes
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
)
^
File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
for job_id, done, content in self._prepare_split_single(
~~~~~~~~~~~~~~~~~~~~~~~~~~^
gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
):
^
File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1880, in _prepare_split_single
raise DatasetGenerationError("An error occurred while generating the dataset") from e
datasets.exceptions.DatasetGenerationError: An error occurred while generating the datasetNeed help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.
text string |
|---|
docker |
run |
--rm |
--init |
--network |
bridge |
--workdir |
/workspace |
--entrypoint |
codex |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/mixed_signal_stm32_dev_board/workspace,dst=/workspace |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/mixed_signal_stm32_dev_board/submission/final_project,dst=/workspace/final_project |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/mixed_signal_stm32_dev_board/runtime_task,dst=/task,readonly |
--mount |
type=bind,src=/home/d/coding/eda-bench/.tmp/runtime_homes/eda_bench_codex_rjr35gkd/.codex,dst=/root/.codex |
eda-bench-agent |
--dangerously-bypass-approvals-and-sandbox |
--enable |
multi_agent |
exec |
--json |
-m |
gpt-5.4-mini |
--skip-git-repo-check |
--ephemeral |
-C |
/workspace |
-c |
model_reasoning_summary="auto" |
-c |
model_reasoning_effort="high" |
You are the root owner of this Linux container.
Subagents are enabled.
Work only inside /workspace.
Read the task assets in /task.
Use the Codex /goals feature if it is available: create a goal for completing this KiCad task, keep working until the final project is actually complete, and mark the goal complete before y... |
null |
null |
null |
null |
null |
null |
docker |
run |
--rm |
--init |
--network |
bridge |
--workdir |
/workspace |
--entrypoint |
codex |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/rp2350b_dev_board/workspace,dst=/workspace |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/rp2350b_dev_board/submission/final_project,dst=/workspace/final_project |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/rp2350b_dev_board/runtime_task,dst=/task,readonly |
--mount |
type=bind,src=/home/d/coding/eda-bench/.tmp/runtime_homes/eda_bench_codex_rxukdl2b/.codex,dst=/root/.codex |
eda-bench-agent |
--dangerously-bypass-approvals-and-sandbox |
--enable |
multi_agent |
exec |
--json |
-m |
gpt-5.4-mini |
--skip-git-repo-check |
--ephemeral |
-C |
/workspace |
-c |
model_reasoning_summary="auto" |
-c |
model_reasoning_effort="high" |
You are the root owner of this Linux container.
Subagents are enabled.
Work only inside /workspace.
Read the task assets in /task.
Use the Codex /goals feature if it is available: create a goal for completing this KiCad task, keep working until the final project is actually complete, and mark the goal complete before y... |
null |
null |
null |
null |
null |
null |
docker |
run |
--rm |
--init |
--network |
bridge |
--workdir |
/workspace |
--entrypoint |
codex |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_p26w6lmm/tasks/mixed_signal_stm32_dev_board/workspace,dst=/workspace |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_p26w6lmm/tasks/mixed_signal_stm32_dev_board/submission/final_project,dst=/workspace/final_project |
--mount |
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_p26w6lmm/tasks/mixed_signal_stm32_dev_board/runtime_task,dst=/task,readonly |
--mount |
type=bind,src=/home/d/coding/eda-bench/.tmp/runtime_homes/eda_bench_codex_ozj7mgoz/.codex,dst=/root/.codex |
EDA Bench Evaluation Provenance
This dataset contains published evaluation records for EDA Bench. Use it to audit reported runs and inspect the recorded execution environment.
Contents
The repository includes records such as:
- harness and tool version metadata
- task prompts and runtime task copies
- model responses and command records
- standard output and error logs
- raw and normalized grading metrics
- run-level reports
- source snapshots used to identify the evaluated harness
A record can be empty when a run produced no output. Pin a repository commit when reproducing a specific result.
Related Task Data
Task definitions, reference projects, source provenance, grader contracts, and canaries are in eda-bench-neurips-2026/eda-bench-tasks.
The default benchmark is web-enabled and uses public source projects. Results measure task solving under the disclosed prompts, tools, access policy, and evaluation contract. They do not measure closed-book PCB invention.
Limits
The records can include failed, incomplete, or intentionally damaged benchmark artifacts. Do not treat them as fabrication-ready hardware.
Scores follow the released benchmark oracle. They do not certify electrical safety, manufacturability, regulatory compliance, thermal behavior, electromagnetic compatibility, or production readiness.
Licenses can differ across source snapshots and task artifacts. Review the task dataset licensing notes and each included source notice before redistributing a subset.
- Downloads last month
- 626