Integration test quality
Claude-assisted review of environment integration tests.
Purpose
ct envs integration-test-quality uses Claude Code on the host to judge whether tests are truly end-to-end from a user’s perspective; optional passes for task conflicts and coverage gaps. Writes integration_test_quality.md under the output directory.
Requirements: Claude Code on your PATH (claude CLI).
Commands
ct envs integration-test-quality
Analyze integration test quality for an environment.
Usage:
ct envs integration-test-quality [OPTIONS]
Options:
-e,--env— Environment name (prompts if omitted)-o,--output-dir— Output directory (default:data/test_quality_runs/<env>)--check-tasks— Whether completing main/side tasks would likely break integration tests--coverage— Coverage / gap analysis for integration tests--fix— After analysis, run Claude to edit integration tests using the saved report. Without this flag, the fix step is skipped. With--fix: in a non-interactive terminal the fix runs automatically; in an interactive terminal you are prompted for confirmation before any edits.
Examples:
uv run ct envs integration-test-quality -e web_scraping uv run ct envs integration-test-quality -e web_scraping --check-tasks uv run ct envs integration-test-quality -e web_scraping --coverage uv run ct envs integration-test-quality -e web_scraping --check-tasks --coverage uv run ct envs integration-test-quality -e web_scraping --fix
Behavior: Writes or appends to integration_test_quality.md, prints a short summary, then the report path.
Output
Report path: data/test_quality_runs/<env>/integration_test_quality.md (or --output-dir).
Main analysis (integration_test_quality.md body): plain text with this shape (no markdown headings inside the narrative blocks beyond the top title):
- Title line:
# Integration test quality: <env> Summary:— then a blank line, then a 1–2 sentence overall assessment of integration test suite quality.List of problems— then a blank line, then one line per issue asproblem: how to fix.
With --check-tasks, the same file is appended with a Task conflicts section: a header line, then a blank line, then a list (one line per item: task or area, and why tests would break).
With --coverage, the same file is appended with a Coverage gaps section: a header line, then a blank line, then a list (one line per gap and suggested coverage).
Criteria (good vs bad integration tests)
Good tests:
- Exercise the system through HTTP, CLI, or other stable external interfaces
- Use inputs and read outputs the way a real user would
- Survive internal refactors because they only depend on public behavior
Bad tests:
- Import internal modules or query the DB directly instead of the product API
- Assert on internal paths, logs, or implementation details
- Over-mock so they no longer test real integration
Use cases
- Review integration test quality for an env’s suite before large changes; find brittle tests; optional automated fix via
--fix.