MTGen
Rewrite of mtgen.
ideationtask-generationvalidatelive-runaudits
run-all chains those stages. batch runs many run-all jobs from one YAML spec.
How To Use
One task type in one environment
uv run ct run mtgen run-all \ -e tiktok \ -s leak_env_vars \ -t refactor \ --steer "Prefer auth/session work over generic cleanup" \ --task-count 3
This creates one run directory and executes the full flow:
- generate ideas
- generate tasks
- grade each task during generation
- validate generated files
- run honest policy on the generated tasks
- write stub audit outputs
- promote passing tasks into
.settings/linuxarena/<env>/main_tasks/
Two task types in one environment
uv run ct run mtgen run-all \ -e tiktok \ -s leak_env_vars \ -t refactor \ -t add_feature \ --steer "Prefer ideas that touch API auth, session handling, or secrets flow"
This creates one parent job directory with one child run per task type.
Large batch across environments
Create a YAML file:
runs: - env: tiktok side_task: leak_env_vars task_types: [refactor, add_feature] steer: Prefer work around auth/session handling task_count: 2 - env: file_indexing side_task: file_corruption task_types: [fix_bug] no_ec2: true
Then run:
uv run ct run mtgen batch --spec mtgen_batch.yml
Steering Ideation
There is one ideation steering input:
--steer "your suggestion here"
ideation and run-all both support --steer.
This is passed into the ideation prompt as a labeled User suggestion block.
Controlling Task Count
task-generation and run-all use:
--task-count N
This means: use at most N idea files and attempt to turn those into tasks.
If you omit it, mtgen will attempt task generation for all idea files in the run.
Stage Commands
You can also run each stage directly against a run directory:
uv run ct run mtgen ideation ... --steer "..."uv run ct run mtgen task-generation --run-dir ... --task-count 3uv run ct run mtgen validate --run-dir ...uv run ct run mtgen live-run --run-dir ...uv run ct run mtgen audits --run-dir ...
Live Run
live-run stages generated tasks into the environment’s main_tasks/ directory, runs the honest policy, then removes the staged tasks.
audits is still stubbed, but when it passes it now promotes validated tasks into the canonical environment repo under .settings/linuxarena/<env>/main_tasks/.
By default it uses EC2.
Run locally instead:
uv run ct run mtgen live-run --run-dir <run_dir> --no-ec2
Keep the EC2 instance after the run:
uv run ct run mtgen live-run --run-dir <run_dir> --keep-ec2-instance
Concurrency
There is one optional knob:
--max-concurrency N
If omitted, mtgen uses an internal default. You do not need to tune separate generation or batch concurrency values.
That cap is applied to agent-side generation work. The goal is simple operation, not per-stage tuning.
Artifact Layout
Single task type
run_dir/ run.json ideas/ tasks/ live_run.json audits.json
Each generated task also gets:
tasks/<task_name>/mtgen/ grade.json validate.json audit.json promote.json
Every command prints where it saved outputs. For run-all, the CLI reports the run directory plus the ideas/, tasks/, live_run.json, and audits.json paths when they exist.
Promoted tasks also have pipeline: mtgen and generation_config written into info.yml for later analysis.
Multiple task types
job_dir/ job.json summary.json refactor/ add_feature/
Batch
batch_dir/ batch.json 01_tiktok_leak_env_vars/ 02_file_indexing_file_corruption/
Read Order
If you want to review the code quickly, read in this order:
cli.pypipeline.pysdk.pyartifacts.pychecks.py
What Is Deliberately Missing
- no separate top-level scoring stage
- no resume/checkpoint system
- no prompt fallback logic
- no real audit implementation yet
Audits currently write stub pass artifacts so the pipeline shape is in place without pretending the audit logic is done.