
Audits your test suite's output for silently vanished tests, unexplained skips, and suspiciously fast runs — turns green CI into a real signal.

Hands-on prompt engineering and LLM evaluation. Built around a diagnose-then-ship loop: forensic transcript critique ▎ that names the exact failure mode, paired with a fixer that produces complete, deployment-ready rewrites — not ▎ vague advice.

Point it at your repo (or a specific change) and it scaffolds a unit-test suite for you — auto-detects your runner (Jest/Vitest, pytest, Go, XCTest, JUnit/Kotest), writes idiomatic AAA tests matching your conventions, then RUNS them and fixes test-side failures until green. It never edits your source — a real bug is reported, not papered over. Zero config.

Turn prompts or workflows into testable, launch-ready AI Skill products.

Academic sanity-check with interactive HTML dashboard. Detects look-ahead bias, over-parameterization, and benchmarks against 10 canonical blueprints via animated SVG gauges and color-coded KPI cards. Parses natural language, code, or CSV uploads. Output auto-adapts (EN/ZH-TW/ZH-CN). Built on De Prado (2018). For students and strategy developers. Not financial advice.

Point it at your native iOS repo and it scaffolds a complete XCUITest E2E UI-test suite for you — discovers your screens (SwiftUI + UIKit), plans the accessibilityIdentifier tagging, and generates a base test case, ready-to-run per-screen specs, server-side error-state specs, and the scheme/test-plan steps. Zero config.

Stop guessing whether Claude Code can do something. This skill lists the features that look like they work but quietly do not - session-only cron jobs, connectors that show 'connected' with zero usable tools, skills that load empty because of one quote mark in the YAML. Every claim was verified by running it, and each section ends with the exact command that settles the question on your own machine.
