
Hook Spec — test the axis, not the score
What this does
Hook tools give you a score. This one refuses to. It measures which axis your own numbers actually support, and throws out the ones they do not.
Most hook advice is a list of styles somebody believes in. The real question is whether a style label explains the spread in your data. Usually it does not.
A worked case
Fourteen ad creatives, three hook types labelled by hand:
| Hook type | Creatives | Pooled CTR | Spread inside the group |
|---|---|---|---|
| moment (names when it happens) | 7 | 7.43% | 10.52% → 2.96% |
| announce | 5 | 4.04% | 8.15% → 2.17% |
| problem | 2 | 3.73% | — |
Between groups, "moment" beat "announce" by 1.84× with p ≈ 0. Settled, apparently. But the spread inside each group was larger than the gap between them. The label was not the explanation, and the hypothesis was thrown out.
A second axis, built afterward, held up far better: does the hook call the thing by its specific name, or by its category?
| Axis | Creatives | CTR | CPC | Ratio |
|---|---|---|---|---|
| Named (the actual event or product name) | 5 | 9.72% | ₩38 | CTR 2.44× · CPC 4.31× |
| Unnamed (category words) | 9 | 3.98% | ₩162 |
That axis ships with this skill as a prior to test first, not a law. It was built from the data it was tested on, and the skill will tell you the same thing about any axis you build.
What you paste in
Your hooks — the actual opening text, verbatim — and real numbers for each: impressions and clicks, views and retained, sends and opens. Four or more hooks makes this useful; fewer and it will say so.
What you get back
- 2 to 4 candidate axes built from your own copy, shown as a table for you to correct before anything is computed. Your reading of your own hooks beats the tool's
- Within-group homogeneity first. An axis whose groups are not internally consistent is rejected, not crowned
- Between-group comparison only for the axes that survived
- Axes ranked by how much they reduce within-group spread, not by p-value
- A verdict per axis:
HOLDS·STRAINED·REJECTED·TOO FEW - For the surviving axis, a short production spec: the rule in one sentence, three of your own hooks that follow it and three that break it, with numbers
If every axis is rejected
That is a result, not a failure. It means your hooks differ on something none of these axes capture. The skill names the two creatives with the widest gap inside the same group and asks you what is different about them. That question is the deliverable.
One more thing it will tell you
When the hook is weak, adding more variants of it recovers only half the loss. Measured: three variants of a strong hook landed at 10.52 / 10.35 / 9.68% — indistinguishable, p = 0.21. A weak hook's variants split 5.85% vs 2.96%, p = 5.4e-06. Fix the hook before multiplying creatives.
What it refuses to do
Score or predict a hook that has not run. Drop the homogeneity check because it rejected your favourite axis. Report a rate without its denominator. Treat a post-hoc axis as confirmed.


