
This agent helps Python teams using the Anthropic or OpenAI SDK safely cut LLM spend without risking regressions. With zero code changes, it maps call sites and runs your test scenarios to rank token waste from missed prompt caching or heavy model tiers. It tests fixes in isolated git worktrees, enforcing a strict "do-no-harm" gate that reverts any patch failing to preserve behavior. It delivers a clear savings audit that converts into verified, behavior-preserving code patches upon opt-in.

Three AI reviewers cross-examine a backlog item for thin user value, hidden effort, and roadmap misfit before you commit a sprint to it.

Analysez logs, diagnostics, erreurs et éléments techniques fournis afin d’identifier des causes probables, prioriser les anomalies et préparer un plan de correction.

Three independent AI executives — Strategy, Tech, and Growth — blind-review your business decision and cross-examine each other before delivering one verdict.

Hands-on prompt engineering and LLM evaluation. Built around a diagnose-then-ship loop: forensic transcript critique ▎ that names the exact failure mode, paired with a fixer that produces complete, deployment-ready rewrites — not ▎ vague advice.

Four independent AI reviewers — Security Skeptic, Reliability Realist, Maintainability Pragmatist, and Performance Pessimist — each review your diff blind to the others' verdicts, so real bugs don't get averaged away into one polite summary.

Paid ads advisor for performance marketers who need a diagnosis when ROAS drops, CPA spikes, or CVR falls — ranked hypotheses with confidence levels, not open-ended analysis. Routes to one of five task flows (Build / Diagnose / Decide / Report / Advise), applies the matching framework, and delivers structured output with stated assumptions. Unlike generic AI that asks 10 intake questions first, this outputs an immediate concrete deliverable — then refines with follow-up, not before.

Built on clinical epidemiology, biostatistics, causal inference, and evidence-based medicine. TrialReviewer analyzes RCTs, observational studies, meta-analyses, and diagnostic research using effect sizes, survival analysis, bias detection, GRADE, and benefit-harm frameworks to judge whether findings are valid, meaningful, and practice-changing.

marketingskills/skills/ab-testing. Corey Haines`s expert-level skill. It checks detectable lift, traffic limits, safe metrics, QA risk, and locked decision rules before launch. When you provide baseline rate, MDE, power, and traffic, it runs Python sample-size math to estimate required sample size and test duration, so you avoid wasted traffic, broken tracking, peeking, and fake wins. Support Corey Haines on https://buymeacoffee.com/coreyhaines