Short answer: test one variable at a time, decide the metric and the sample size before you post, and repeat a test before acting on it. Short-form results are noisy: a single video's reach varies for reasons you cannot see. Use several posts per variant, compare like with like, and keep a written log so the team learns from the tests instead of from memory.
What to test
| Variable | Example variants | Primary metric |
|---|---|---|
| Hook | Result-first vs mistake-first | Retention in the first seconds |
| Format | Screen demo vs slideshow | Saves, profile visits |
| Length | 20 s vs 45 s | Completion, average watch time |
| Presenter | Founder vs faceless | Follows, comments |
| Posting time | Morning vs evening | Views in first 24 h |
| Call to action | "Link in profile" vs spoken URL | Tagged clicks, vanity URL visits |
| Caption | Search phrase vs statement | Views from search, where reported |
For hook ideas, see B2B video hooks.
Designing a fair test
- One variable. Keep topic, length and account type the same.
- Pick the metric first. Write it down before posting.
- Several posts per variant. One post each is anecdote. Choose a number in advance, for example five per variant (an assumption, not a statistical rule).
- Comparable slots. Same days and times, or alternate them.
- Comparable accounts. Test across accounts of similar age and size, or within one account over time.
- Fixed window. Read results at the same age for every post, such as 72 hours.
Test log template
| Field | Example |
|---|---|
| Test ID | T-014 |
| Hypothesis | Result-first hooks keep more viewers past 3 s |
| Variable | Hook |
| Variants | A: result-first, B: mistake-first |
| Accounts | @example.sheets, @example.invoicing |
| Posts per variant | 5 |
| Metric | Viewers remaining at 3 s |
| Read at | 72 h |
| Result | Fill in |
| Decision | Adopt, reject, or rerun |
Reading results honestly
- Look at the spread, not just the average. One viral post can distort it.
- If the difference is small relative to the spread, call it inconclusive.
- Rerun a winning test before changing your whole calendar.
- Don't cherry-pick: log every post in the test, including flops.
- Don't count a post that did not publish; confirm status first.
Turning results into decisions
Each finished test should end in one of three decisions: adopt, reject or rerun. When you adopt a change, update the content calendar and the SOP so it becomes the default, and note the test ID next to the rule. Review the test log quarterly. Patterns across several tests, such as a hook style that keeps winning, are more reliable than any single result.
Across many accounts
Running the same test across several accounts gives more posts per variant in less time. Avoid posting the identical clip twice to one account; use matched clips instead. Tie signups back to variants with UTM links and the approach in measuring signups.
Running tests on iPhones with Farmero
Farmero helps with the "did it publish" part: each post is confirmed on the profile or marked uncertain, so the test log only counts posts that went live. Fan-out creates a separate card per account, so variant A and B can be assigned and timed per account. Its MCP server lets an AI agent read the verification rate and queue posts while publishing stays behind your switch. It does not measure views; read those in each platform's own insights. See MCP server.
FAQ
How many posts per variant do I need?
More than one. Choose a number before starting and be cautious with small differences.
Can I test two things at once?
Only with a much larger number of posts. For small teams, test one variable.
How long should I wait before reading results?
Pick a fixed age, such as 72 hours, and use it for every post.
Should I delete losing variants?
No need. Log them and move on.
