How do terminal benchmarks construct executable tasks and check outputs, filesystem changes and final container state?
-
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Paper • 2306.14898 • Published -
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Paper • 2601.11868 • Published • 38 -
Learning CLI Agents with Structured Action Credit under Selective Observation
Paper • 2605.08013 • Published • 1 -
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
Paper • 2605.22535 • Published • 8