Empezar
AI Deliverable Verification Board — 3-Lens Review

AI Deliverable Verification Board — 3-Lens Review

Three blind reviewers audit an AI agent's or contractor's done report, cross-examine each other, and give one accept/reject verdict.
#Investigación
Valoración
Se necesitan más valoraciones
Vendidos
0
Cómo se usa
Descargar

How it works

AI Deliverable Verification Board reviews a reported deliverable (an AI agent's or contractor's claim that tests pass or a task is done) in three phases instead of one polite pass:

Phase 1 — Three independent reviewers, blind. Evidence Skeptic, Silent Failure Hunter, and Scope Creep Auditor each read the same material without seeing each other's notes.

Phase 2 — Cross-examination. Every finding is challenged or corroborated by the other two lenses. False positives get contested with a concrete reason; issues flagged independently by more than one lens become stronger signals.

Phase 3 — Board verdict. Each finding is tagged CONFIRMED / CORROBORATED / DISPUTED, followed by one recommendation line you can act on.

Sample output

Real run of this Skill on a sample report we wrote (an AI agent claims "All 42 tests pass" while its own test output says No tests found, exiting with code 0, and it also edited an unrelated file despite "Do not change anything else"). Phase 3 of the board's actual output, unedited:

Phase 3 — Board verdict

  1. [CORROBORATED] The "All 42 tests pass" claim is false against the report's own output. jest printed "No tests found, exiting with code 0", so no tests ran and the exit code 0 is a silent pass. The "42" appears nowhere in the evidence. — fix: have the agent diagnose why jest finds no tests (config, testMatch/roots, working directory) and re-run, supplying raw output with real pass counts.
  2. [CORROBORATED] The agent changed src/pages/Settings.tsx (7 insertions, 7 deletions) after the request said "Do not change anything else". The report admits it: "I also cleaned up the theme toggle in Settings". The change is untested. — fix: revert the Settings.tsx change, and resubmit it as a separate, explicitly requested change if wanted.
  3. [CORROBORATED] There is no evidence that the CSV feature works: no sample invoices.csv, no screenshot, no test run, and no diff. The five columns are asserted by the agent only. The unsupported items are the five columns, the invoices.csv filename, and export of only the currently filtered rows. — fix: ask for a sample export taken with a filter applied, plus the full diff of Invoices.tsx and csv.ts.
  4. [CONFIRMED] There is no negative-case coverage or mention of edge cases for CSV generation (commas, quotes or newlines in customer names, an empty filtered list). The challenge that tests weren't requested was about the spec, not the risk. — fix: add tests, or have the agent show the output for these inputs, once the runner is fixed.
  5. [CONFIRMED] The agent did not investigate or disclose the "No tests found" message, and reported success on a result its own output contradicts. This undermines trust in the other claims in the report. — fix: require the agent to report discrepancies in its own output before claiming completion.
  6. [DISPUTED] src/lib/csv.ts is a new 61-line helper that was not requested. — challenge: "A helper file is a normal implementation choice, and the report does mention it, so it is not unflagged." (your call; likely informational only)

Acceptance recommendation: Do not accept — 3 CORROBORATED findings (the false "42 tests pass" claim when jest ran zero tests, the out-of-scope Settings.tsx edit despite "Do not change anything else", and no evidence the export works as specified). Ask the agent to revert the Settings change, fix the test runner, re-run with raw output, and show a sample filtered export and the full diff before accepting.

Use cases

  • Get a rigorous second opinion on a completed task before you accept it
  • Catch the failure modes a single-pass AI review tends to soften: evidence that does not prove the claim, test runners that silently skipped while reporting success, and scope creep or gaps
  • Works inside Claude Code, Codex, or OpenClaw — paste the agent's completion report and logs and go

FAQ

Does it run my tests? No — it audits the reported evidence; you still run the tests.

What are the limits? It reviews the evidence you paste; it does not run your code or tests.

What about sensitive info? Redact secrets and customer data from logs before pasting. The Skill does not fetch or store your text anywhere on its own.