Datasets:
messages listlengths 3 29 | tools listlengths 0 0 | task_id stringlengths 25 25 | replica_id int8 0 7 | reward float32 1 1 | model stringclasses 1
value | harness stringclasses 1
value | source_dataset stringclasses 1
value | source_revision stringclasses 1
value | hit_labels listlengths 1 1 | trigger_command stringlengths 30 303 | recovery_action stringlengths 62 106 | trajectory_format stringclasses 2
values | collection_config stringclasses 1
value | trajectory_sha256 stringlengths 64 64 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
[
{
"role": "user",
"content": "You are an AI assistant tasked with solving command-line tasks in a Linux environment. You will be given a task description and the output from previously executed commands. Your goal is to solve the task by providing batches of shell commands.\n\nFormat your response as JSON w... | [] | tmax/task_000043_12653d83 | 7 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /home/user/geoparse-repo && git log -p -- parser.py | {"kind": "q", "reference": "trajectory.json:step-5:command-0"} | ATIF raw_content=true linear_history=true; Harbor last-episode export | {"max_input_tokens": 48640, "max_output_tokens": 16384, "max_turns": 200, "model": "openai/dsv4-flash-0731", "production_concurrency": {"final_recovery_total": 16, "initial_requested_total": 32, "worker_slots": 8}, "served_model": "dsv4-flash-0731", "temperature": 1.0, "thinking": {"reasoning_effort": "high", "thinking... | d211acef6584d2fadbb8fd208dcebdad73312dd5f5383d4b467dc3b3c4cf3153 |
[
{
"role": "user",
"content": "You are an AI assistant tasked with solving command-line tasks in a Linux environment. You will be given a task description and the output from previously executed commands. Your goal is to solve the task by providing batches of shell commands.\n\nFormat your response as JSON w... | [] | tmax/task_000873_a7f964b2 | 4 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /home/user/pipeline_repo && git show --stat --oneline acf745c && echo '--- initial ---' && git show acf745c | {"kind": "q", "reference": "trajectory.json:step-7:command-0"} | ATIF raw_content=true linear_history=true; Harbor last-episode export | {"max_input_tokens": 48640, "max_output_tokens": 16384, "max_turns": 200, "model": "openai/dsv4-flash-0731", "production_concurrency": {"final_recovery_total": 16, "initial_requested_total": 32, "worker_slots": 8}, "served_model": "dsv4-flash-0731", "temperature": 1.0, "thinking": {"reasoning_effort": "high", "thinking... | d404c6816db0a4eb61016afa177f9d34962af4feec0f0373e036045299a5b76b |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_002510_12753f45 | 7 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | "cd /home/user/optimization_engine && git show --stat --oneline $(git rev-parse v1.0-good) && git lo(...TRUNCATED) | {"kind": "q", "reference": "trajectory.json:step-7:command-0"} | ATIF raw_content=true linear_history=true; Harbor last-episode export | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | f29ec04b9aabdb734d50451ad92a964d293db069ccc584774298ba2956562a50 |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_003575_5cda4281 | 3 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | git log --oneline -- src/pipeline.py __init__.py src/__init__.py | {"kind": "q", "reference": "trajectory.json:step-8:command-0"} | ATIF raw_content=true linear_history=true; Harbor last-episode export | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | 0c452d5b311298168a27059a2d80b4cb1ee16e77cf48c41bff8442047fed7809 |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_004516_b90185a3 | 4 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /app/wal_processor && git log -p --all -S 'all_events' -- processor.py | {"kind": "q", "reference": "trajectory.json:step-5:command-0"} | ATIF raw_content=true linear_history=true; Harbor last-episode export | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | c087e5e16c45c31275f667a22643cc8ffa2e3f9210e8ef910f0085885c7e6cc4 |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_004715_e1f56e7b | 4 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | git log --oneline --graph --decorate --all -40 | {"kind": "q", "reference": "trajectory.json:step-5:command-0"} | ATIF raw_content=true linear_history=true; Harbor last-episode export | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | 3460ddf408aace8c5a3d62fd179c3be3800bf930063bd7603832122404f46007 |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_000872_cf99cb19 | 1 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /home/user/service_repo && git status --short && echo '--- diff ---' && git diff -- service.py | "{\"cast_time\": 365.78987, \"kind\": \"q\", \"raw_input\": \"q\", \"reference\": \"trajectory.cont-(...TRUNCATED) | "ATIF recovery-segment export; parser-invalid copied handoff pair removed when present; no model out(...TRUNCATED) | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | 72080efd26da3b350de2f93d08b53ddb0a3d24a1ae0f733b408d7f23a804febc |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_001716_43f6a16e | 3 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /home/user/py_engine && git diff v1.0 v2.0 -- solver.py | "{\"cast_time\": 347.8525, \"kind\": \"q\", \"raw_input\": \"q\", \"reference\": \"trajectory.cont-2(...TRUNCATED) | "ATIF recovery-segment export; parser-invalid copied handoff pair removed when present; no model out(...TRUNCATED) | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | 3925ebfe7d225fe4ef8b42f1b94f3548ffedf11addad5be31eb5488824262055 |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_001716_43f6a16e | 5 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /home/user/py_engine && git diff v1.0 v2.0 -- solver.py | "{\"cast_time\": 84.615381, \"kind\": \"q\", \"raw_input\": \"q\", \"reference\": \"trajectory.json:(...TRUNCATED) | "ATIF recovery-segment export; parser-invalid copied handoff pair removed when present; no model out(...TRUNCATED) | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | 787f9f5eb2403e28237f139716ee0e3eb136115dc432c0c58d4c96c213f1c394 |
[{"role":"user","content":"You are an AI assistant tasked with solving command-line tasks in a Linux(...TRUNCATED) | [] | tmax/task_002149_92d3fce9 | 2 | 1 | DeepSeek-V4-Flash-0731 | Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser | TMaxxx/TMax-15K-Harbor | 48a77eb0b017606c643ed905f96db41672914798 | [
"git_pager_q_exit"
] | cd /home/user/encoder_repo && git show HEAD:encode.sh | "{\"cast_time\": 205.485017, \"kind\": \"q\", \"raw_input\": \"q\", \"reference\": \"trajectory.cont(...TRUNCATED) | "ATIF recovery-segment export; parser-invalid copied handoff pair removed when present; no model out(...TRUNCATED) | "{\"max_input_tokens\": 48640, \"max_output_tokens\": 16384, \"max_turns\": 200, \"model\": \"openai(...TRUNCATED) | 0ca87a6f3c9b5928a60412527fd6fce3c605a6128bf2ed8908f23075aa61b793 |
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use.
Composition
- 6 original parser-clean full last-episode exports.
- 17 additional manually confirmed recovery-ending segments from the parser-defect audit.
- All 23 source trials received reward 1.0 and had no Harbor exception.
The 17 recovered rows do not rewrite model output or terminal output. Each selected segment ends on the audited recovery action. Prompt restoration and subsequent working shell use are validated from recording.cast and stored in the audit provenance, while the post-action terminal result is omitted under Harbor's normal episode convention. For six continuation segments, the export removes only a parser-invalid copied-context handoff question and its paired copied answer. The other eleven selected segments contain no parser-invalid response; their trials were previously excluded only because a later continuation contained a malformed copied handoff. All retained task-solving assistant responses parse under the Terminus-2 JSON protocol.
Source and collection
- Environment/task source:
TMaxxx/TMax-15K-Harbor, revision48a77eb0b017606c643ed905f96db41672914798. - Model:
openai/dsv4-flash-0731, served IDdsv4-flash-0731. - Sampling: temperature 1.0, top-p 0.95, maximum 200 Terminus-2 turns.
- Harness: Harbor 0.7.0 / Terminus-2 2.0.0 JSON parser.
- Terminal evidence:
recording.castinput/output events were authoritative during manual audit.
Recovery labels:
git_pager_q_exit: 17git_pager_retry_then_q_exit: 6
Files
data/train.parquet: 23 SFT conversations.data/index.parquet: task/replica identity, evidence references, and checksums.manifest/selected-sft-tasks.jsonl: the 19 selected task identities and retained replicas.manifest/run-config.json: collection and reconstruction metadata.manifest/checksums.sha256: checksums for every published artifact.
Important fields include messages, task_id, replica_id, reward, hit_labels, trigger_command, recovery_action, trajectory_format, and trajectory_sha256. The data/train.parquet SHA256 is 6fa996c89ecbec8df15ec13d3a5e0bea73c174a3bcc7be2a03e86fcb217677fd.
Filtering caveat
This is an intentionally narrow behavior dataset, not a representative estimate of TMax task performance. Failed tasks, proactive pager avoidance, unproven temporal correlations, parser-invalid generated task-solving responses, and one ambiguous recovery were excluded.
License and privacy
TMax is distributed under ODC-BY. Raw trials and internal infrastructure logs are not published. Staging and the independently downloaded revision are scanned for tokens, private keys, internal addresses/mounts, and private service identifiers before publication.
- Downloads last month
- 42