開始使用
Test Suite Integrity Audit

Test Suite Integrity Audit

Audits your test suite's output for silently vanished tests, unexplained skips, and suspiciously fast runs — turns green CI into a real signal.
#生產力
評分
需要更多評價
已售出
0
使用方式
下載


name: test-suite-integrity-audit
description: Audit a test suite's run output (or its config/CI log) for the three ways a test suite quietly stops proving anything — planned tests silently vanishing from the executed count, tests newly skipped/disabled without a tracked reason, and suspiciously fast runs that suggest a suite didn't actually execute. Turns "all green" into a real signal instead of a false one. Use when asked to audit test coverage, check if CI is actually testing what it claims, or investigate why a bug shipped despite passing tests.
trigger: The user pastes test run output, a CI log, or a test suite's config/file list and asks whether the tests are actually running, whether coverage silently dropped, or why something broke despite "all tests passing."
version: 1.0.0
compatible_agents: [claude-code, codex, openclaw]

Test Suite Integrity Audit — catching tests that quietly stopped counting

A green CI run tells you tests passed — it doesn't tell you the same tests ran as last time. A
suite can silently lose coverage (a file fails to load, a bundle crashes before registering its
tests, a flag disables a block) and still report "all passing," because fewer tests just means
fewer tests, not a failure. This Skill makes that invisible drop visible.

Anti-sycophancy rule (applies throughout)

Do not treat "all green" as evidence of health on its own — that is exactly the blind spot this
Skill exists to catch. Ground every finding in the actual numbers and file/config content given to
you; never invent a baseline test count you weren't given. If no prior run's numbers are available
for comparison, say so explicitly and run only the checks that don't require a baseline (see Check
3 below).

What to check

Work from whatever the user gives you: raw test runner output, a CI log, a test file listing, or a
description of the pipeline. Run every check that the available material supports.

1. Planned vs. executed vs. passed count mismatch

Most test runners report (or can be made to report) how many tests were planned/collected versus
executed versus passed. Compare all three:

  • Planned > Executed → tests were collected but never ran (a file errored during collection, a
    bundle crashed silently, an import failed and the runner swallowed it). This is the single
    most important thing to catch
    — it means the total went down and nothing complained.
  • Executed = Passed but Executed < a known prior baseline → the same silent-drop pattern, just
    without an explicit "planned" number to compare against. Ask for or infer the prior count from
    git history, a previous CI run, or the test file count if none is given.

2. Newly skipped/disabled tests without a tracked reason

Scan for skip/disable markers (.skip, xit, @pytest.mark.skip, #[ignore], commented-out test
blocks, feature-flag-gated test blocks, etc.). For each one found:

  • Is there a comment, linked issue, or commit message explaining why? If yes, note it as
    accounted-for.
  • If no explanation exists, flag it — an unexplained skip is functionally identical to a silently
    deleted test, just more discoverable if someone thinks to grep for it (most people don't).

3. Suspiciously fast or suspiciously uniform run times

A test suite that normally takes minutes finishing in seconds, or every test in a suite reporting
an identical (often 0ms or 1ms) duration, is a strong signal the tests didn't actually execute their
logic — a mocked-out runner, a config pointing at the wrong directory, or a CI step that's silently
a no-op. Flag this even without a specific count mismatch, since it's often the first symptom.

Output format

## Test Suite Integrity Audit

### 1. Count mismatch
Planned: N | Executed: N | Passed: N
[FINDING or "No mismatch found"] — if a mismatch: which tests are missing (by file/name if
identifiable), and the most likely cause based on what's visible (collection error, import failure,
etc.)

### 2. Unexplained skips/disables
- `path/to/test.js:42` — `.skip` on "should validate email format", no comment/issue reference found
- ... (or "No unexplained skips found")

### 3. Timing anomalies
[FINDING or "No anomalies found"] — e.g. "Suite reports 847 tests in 0.3s (previously ~40s) — likely
not actually executing"

## Verdict
**[CONFIRMED SILENT DROP / CLEAN / INCONCLUSIVE — need X]** — one line naming what to fix first and
why it matters (e.g. "8 of 179 tests silently stopped executing after a bundler config change —
fix the import path in X before trusting this suite's green status again").

Handling incomplete input

If the user gives you only current output with no baseline or history, say plainly which checks
you could and couldn't run, and what additional input (a prior CI run, git blame on the test
directory, the actual test files) would let you finish the audit — don't guess at a baseline number
to make the report look more complete than it is.