Dataset Preview
Duplicate
The full dataset viewer is not available (click to read why). Only showing a preview of the rows.
The dataset generation failed
Error code:   DatasetGenerationError
Exception:    ArrowInvalid
Message:      Schema at index 2 was different: 
enabled: bool
reason: string
stdout_bytes_omitted: int64
stderr_bytes_omitted: int64
parsed_event_count_omitted: int64
vs
submission_exists: double
build_success: double
task_score: double
overall_score: double
reward: double
error_message: string
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1848, in _prepare_split_single
                  writer.write_table(table)
                  ~~~~~~~~~~~~~~~~~~^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 764, in write_table
                  self.write_rows_on_file()  # in case there are buffered rows to write first
                  ~~~~~~~~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 662, in write_rows_on_file
                  table = pa.concat_tables(self.current_rows)
                File "pyarrow/table.pxi", line 6321, in pyarrow.lib.concat_tables
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                  return check_status(status)
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: Schema at index 2 was different: 
              enabled: bool
              reason: string
              stdout_bytes_omitted: int64
              stderr_bytes_omitted: int64
              parsed_event_count_omitted: int64
              vs
              submission_exists: double
              build_success: double
              task_score: double
              overall_score: double
              reward: double
              error_message: string
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1869, in _prepare_split_single
                  num_examples, num_bytes = writer.finalize()
                                            ~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 781, in finalize
                  self.write_rows_on_file()
                  ~~~~~~~~~~~~~~~~~~~~~~~^^
                File "/usr/local/lib/python3.14/site-packages/datasets/arrow_writer.py", line 662, in write_rows_on_file
                  table = pa.concat_tables(self.current_rows)
                File "pyarrow/table.pxi", line 6321, in pyarrow.lib.concat_tables
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: Schema at index 2 was different: 
              enabled: bool
              reason: string
              stdout_bytes_omitted: int64
              stderr_bytes_omitted: int64
              parsed_event_count_omitted: int64
              vs
              submission_exists: double
              build_success: double
              task_score: double
              overall_score: double
              reward: double
              error_message: string
              
              The above exception was the direct cause of the following exception:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 1369, in compute_config_parquet_and_info_response
                  parquet_operations, partial, estimated_dataset_info = stream_convert_to_parquet(
                                                                        ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      builder, max_dataset_size_bytes=max_dataset_size_bytes
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/src/services/worker/src/worker/job_runners/config/parquet_and_info.py", line 948, in stream_convert_to_parquet
                  builder._prepare_split(split_generator=splits_generators[split], file_format="parquet")
                  ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1694, in _prepare_split
                  for job_id, done, content in self._prepare_split_single(
                                               ~~~~~~~~~~~~~~~~~~~~~~~~~~^
                      gen_kwargs=gen_kwargs, job_id=job_id, **_prepare_split_args
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  ):
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/builder.py", line 1880, in _prepare_split_single
                  raise DatasetGenerationError("An error occurred while generating the dataset") from e
              datasets.exceptions.DatasetGenerationError: An error occurred while generating the dataset

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

text
string
docker
run
--rm
--init
--network
bridge
--workdir
/workspace
--entrypoint
codex
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/mixed_signal_stm32_dev_board/workspace,dst=/workspace
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/mixed_signal_stm32_dev_board/submission/final_project,dst=/workspace/final_project
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/mixed_signal_stm32_dev_board/runtime_task,dst=/task,readonly
--mount
type=bind,src=/home/d/coding/eda-bench/.tmp/runtime_homes/eda_bench_codex_rjr35gkd/.codex,dst=/root/.codex
eda-bench-agent
--dangerously-bypass-approvals-and-sandbox
--enable
multi_agent
exec
--json
-m
gpt-5.4-mini
--skip-git-repo-check
--ephemeral
-C
/workspace
-c
model_reasoning_summary="auto"
-c
model_reasoning_effort="high"
You are the root owner of this Linux container. Subagents are enabled. Work only inside /workspace. Read the task assets in /task. Use the Codex /goals feature if it is available: create a goal for completing this KiCad task, keep working until the final project is actually complete, and mark the goal complete before y...
null
null
null
null
null
null
docker
run
--rm
--init
--network
bridge
--workdir
/workspace
--entrypoint
codex
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/rp2350b_dev_board/workspace,dst=/workspace
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/rp2350b_dev_board/submission/final_project,dst=/workspace/final_project
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_t_abbnn1/tasks/rp2350b_dev_board/runtime_task,dst=/task,readonly
--mount
type=bind,src=/home/d/coding/eda-bench/.tmp/runtime_homes/eda_bench_codex_rxukdl2b/.codex,dst=/root/.codex
eda-bench-agent
--dangerously-bypass-approvals-and-sandbox
--enable
multi_agent
exec
--json
-m
gpt-5.4-mini
--skip-git-repo-check
--ephemeral
-C
/workspace
-c
model_reasoning_summary="auto"
-c
model_reasoning_effort="high"
You are the root owner of this Linux container. Subagents are enabled. Work only inside /workspace. Read the task assets in /task. Use the Codex /goals feature if it is available: create a goal for completing this KiCad task, keep working until the final project is actually complete, and mark the goal complete before y...
null
null
null
null
null
null
docker
run
--rm
--init
--network
bridge
--workdir
/workspace
--entrypoint
codex
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_p26w6lmm/tasks/mixed_signal_stm32_dev_board/workspace,dst=/workspace
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_p26w6lmm/tasks/mixed_signal_stm32_dev_board/submission/final_project,dst=/workspace/final_project
--mount
type=bind,src=/mnt/eda-bench-runs/tmp/eda_bench_run_p26w6lmm/tasks/mixed_signal_stm32_dev_board/runtime_task,dst=/task,readonly
--mount
type=bind,src=/home/d/coding/eda-bench/.tmp/runtime_homes/eda_bench_codex_ozj7mgoz/.codex,dst=/root/.codex
End of preview.

EDA Bench Evaluation Provenance

This dataset contains published evaluation records for EDA Bench. Use it to audit reported runs and inspect the recorded execution environment.

Contents

The repository includes records such as:

  • harness and tool version metadata
  • task prompts and runtime task copies
  • model responses and command records
  • standard output and error logs
  • raw and normalized grading metrics
  • run-level reports
  • source snapshots used to identify the evaluated harness

A record can be empty when a run produced no output. Pin a repository commit when reproducing a specific result.

Related Task Data

Task definitions, reference projects, source provenance, grader contracts, and canaries are in eda-bench-neurips-2026/eda-bench-tasks.

The default benchmark is web-enabled and uses public source projects. Results measure task solving under the disclosed prompts, tools, access policy, and evaluation contract. They do not measure closed-book PCB invention.

Limits

The records can include failed, incomplete, or intentionally damaged benchmark artifacts. Do not treat them as fabrication-ready hardware.

Scores follow the released benchmark oracle. They do not certify electrical safety, manufacturability, regulatory compliance, thermal behavior, electromagnetic compatibility, or production readiness.

Licenses can differ across source snapshots and task artifacts. Review the task dataset licensing notes and each included source notice before redistributing a subset.

Downloads last month
626