GenGrowth
Methodology

How to Run Your First Growth Experiment: A Step-by-Step Playbook

GenGrowth Team·4 min read·Updated August 5, 2026
Technical line illustration: an open ring binder lying flat with five tabbed dividers fanned out and a row of five small blank step tiles laid out beside it

Most growth experiments fail because they skip the boring parts: hypothesis formation, sample sizing, and measurement criteria. This playbook covers the full process from idea to iteration, with templates you can use today.

Why Most Growth Experiments Fail

According to data from Reforge and GrowthHackers, roughly 70-80% of growth experiments fail to produce statistically significant results. That number sounds discouraging, but it is actually expected -- the goal is not to win every experiment, but to run enough experiments that the winners compound. The problem is that most teams do not run enough experiments, and the experiments they do run are poorly designed.

The three most common failure modes are:

  • No clear hypothesis. "Let us try posting on Reddit" is not a hypothesis. "Posting value-first content on r/SaaS will drive 50 qualified visitors per post within 7 days" is a hypothesis.
  • No success criteria defined upfront. If you do not decide what "success" looks like before you start, you will rationalize any result as a win or dismiss any result as inconclusive.
  • Insufficient sample size. Running an A/B test with 200 visitors and declaring a winner is statistically meaningless. You need to calculate minimum sample sizes before launching.

Step 1: Form a Testable Hypothesis

Every experiment starts with a hypothesis that follows this structure:

"If we [take this action], then [this metric] will [change in this direction] by [this amount], because [this reasoning]."

Examples:

  • "If we add a product comparison table to our pricing page, then our pricing-to-signup conversion rate will increase by 15%, because visitors currently leave to compare us with competitors on third-party sites."
  • "If we publish 10 glossary pages targeting long-tail keywords, then organic traffic from informational queries will increase by 2,000 monthly visits within 60 days, because we currently have zero coverage for these terms and competition is low."

Notice that each hypothesis includes a specific metric, a directional prediction, a magnitude estimate, and a causal reasoning. This specificity forces you to think clearly about what you are testing and why.

Step 2: Define Success and Failure Criteria

Before running any experiment, write down three things:

  1. Primary metric: The one number that determines success or failure.
  2. Minimum detectable effect (MDE): The smallest change that would be practically meaningful.
  3. Guardrail metrics: Secondary metrics that must not degrade.

Step 3: Calculate Sample Size

For A/B tests and conversion experiments, sample size determines how long you need to run the experiment. For a baseline conversion rate of 3% and an MDE of 20% relative improvement (to 3.6%), you need approximately 14,500 visitors per variation at 80% power and 95% significance.

For content and SEO experiments, you typically need 60-90 days of data to see organic traffic effects.

Step 4: Design the Experiment

Keep experiments as simple as possible. Test one variable at a time. Document the experiment design using this template:

  • Hypothesis: [from Step 1]
  • Primary metric: [from Step 2]
  • Guardrail metrics: [from Step 2]
  • Duration: [from Step 3]
  • Control: What the current experience looks like
  • Treatment: What the new experience looks like
  • Rollback plan: How to revert if something breaks

Step 5: Execute with Tracking

Every experiment needs clean attribution. Use a consistent UTM naming convention for each treatment, record the exact values in the experiment plan, and preserve the comparison window so another reviewer can reproduce the result.

Step 6: Measure and Analyze

When the experiment reaches its planned duration, analyze results using this checklist:

  1. Did the primary metric change in the predicted direction?
  2. Is the change statistically significant (p < 0.05)?
  3. Is the change practically significant (exceeds your MDE)?
  4. Did any guardrail metrics degrade?
  5. Are there segment-level differences?

Step 7: Iterate

  • Winner: Ship the treatment and design a follow-up experiment.
  • Loser: Document why the hypothesis was wrong. Revise and test again.
  • Inconclusive: Increase sample size or extend duration. If still inconclusive, move on.

Experiment Velocity Benchmarks

The best growth teams run 8-12 experiments per month. Early-stage startups should focus on 3-4 per month. The key metric is experiment velocity, not win rate.

For more on measurement infrastructure, see our marketing attribution guide. For a real-world example, read our Week 1 experiment report.

Put the method to work

Start with one verifiable SEO signal

The SEO and Tech Agents require a verified account and inspect public HTML without Search Console or site-ownership access. The marketing run is not saved to an app project.

GT

GenGrowth Team

Growth Automation Engineers

We build tools that help product teams automate growth experiments.