Olivia Brooks·
Subject line testing was pure vibes on my team until this matrix. Twelve lines across four mechanisms, each with a prediction to check after send — we finally learn something per test.
Generates twelve subject lines across four psychological mechanisms with falsifiable predictions for post-send learning.
Subject Line Test Matrix
You are an email marketing scientist. Build a subject line test matrix — the goal is not twelve nice lines, it is twelve lines that teach us something when the results come in.
Email content and goal: {{email_summary}}
Audience and their relationship to us: {{audience_context}} (cold list, engaged customers, lapsed — and what they last heard from us)
Brand voice guardrails: {{voice_guardrails}}
Historical winners/losers if known: {{history}}
Generate exactly 12 subject lines, 3 per mechanism:
1. **Curiosity gap** — opens a loop the email closes. Rule: the gap must be honest; the email must actually deliver the reveal, no bait.
2. **Specific value** — names the concrete thing inside (number, timeframe, artifact). No adjectives doing the work a fact should do.
3. **Self-interest question** — a question the reader answers with "...actually, yes?" Rule: never a question with an obvious no.
4. **Pattern break** — violates inbox convention for this audience (lowercase, unexpected brevity, admission). Must still respect the voice guardrails.
For every line: character count (flag anything over 45 for mobile truncation), a preview-text pairing that extends rather than repeats the line, and a spam-trigger check (all caps, excessive punctuation, trigger words).
Then the science part: for each MECHANISM write one falsifiable prediction ("if curiosity beats specific value with this lapsed audience, it suggests X about their state"). Recommend which 2 mechanisms to test against each other first given my audience context and history, and the minimum sends per arm for a meaningful read at my list size {{list_size}} (state your assumed baseline open rate).
End with the one line you would bet on and one sentence why.
Ergebnisse
## Subject Line Matrix: Q3 pricing-guide launch → engaged customer list (8,200)
### Curiosity gap
1. "The pricing mistake in 9 of 10 proposals we reviewed" (49 — over 45, front-load: "The pricing mistake in 9 of 10 proposals") · Preview: "We kept a tally. Here's the fix." · Spam: clean
2. "We changed our minds about hourly rates" (39) · Preview: "The data that did it, inside." · Spam: clean
3. "What your quietest client is comparing you to" (44) · Preview: "It's not who you think." · Spam: clean
### Specific value
4. "Your 2026 pricing guide: 14 pages, 3 templates" (45) · Preview: "Benchmarks from 400 firms included." · Spam: clean
5. "Raise rates in 30 days — the exact email script" (46 — trim to "The exact email script for raising rates") · Preview: "Copy, adapt, send." · Spam: clean
6. "Benchmark: what firms your size charge in 2026" (45) · Preview: "Median, top quartile, and the gap." · Spam: clean
### Self-interest question
7. "Are you the cheapest one in the room?" (37) · Preview: "There's a fast way to check." · Spam: clean
8. "When did you last raise your rates?" (35) · Preview: "If you paused, this is for you." · Spam: clean
9. "Is one client 40% of your revenue?" (34) · Preview: "The pricing fix for that." · Spam: clean
### Pattern break
10. "we underpriced for three years" (30) · Preview: "What it cost us, in numbers." · Spam: clean; lowercase within voice guardrails ("human, plainspoken")
11. "Pricing." (8) · Preview: "Everything we know, one guide." · Spam: clean
12. "Please don't open this before quoting a client" (46 — borderline, keep: negation is load-bearing) · Preview: "Bad timing costs money." · Spam: "don't open" is safe; no caps/punctuation flags
### Predictions per mechanism
- **Curiosity:** wins → the list still reads us for insight, not just offers; content emails deserve more slots.
- **Specific value:** wins → the list is utility-driven; future subject lines should lead with artifacts, not ideas.
- **Question:** wins → pricing anxiety is active, not latent; a nurture sequence on pricing confidence would land.
- **Pattern break:** wins → format fatigue; our regular templates have gone invisible and need rotation regardless.
### Test design
First test: **Specific value vs. Curiosity** (your history shows emoji/urgency losers but no clean mechanism test; these two are the plausible ceiling for an engaged list). At 8,200 subscribers, assumed 34% baseline open rate: ~2,050 per arm detects a ~4pt lift at 80% power — run 4 arms of 2,050 (lines 2, 4, 6, 8) and read mechanism-level, not line-level.
**The bet: line 4.** Engaged customers open for the artifact — "14 pages, 3 templates" makes the value auditable before the click, and auditable beats intriguing on warm lists in nearly every test I have run.
Modell: Claude Sonnet 4.6
2 Likes1 SavesScore: 2
