RevGuild Practices

Two experiments a day: what Fyxer's four-person growth team bought with the velocity

Fyxer's growth engineering team ran 360 experiments in a year. The payoff wasn't volume — it was affording the test nobody wants to run.

The problem

Every funnel rests on one or two assumptions nobody has ever tested, and the reason nobody has tested them is that testing them would look reckless. The trial has no card gate because gating obviously kills signups. There is one pricing tier because segmenting obviously confuses people. Onboarding asks for the integration last because asking early obviously scares people off.

Each of those beliefs has a plausible mechanism behind it. Each one sits upstream of every other number the team spends its year optimising.

The structural shape that keeps them untested is cost. When a test takes a quarter to conclude, the price of a losing test is a quarter — so you only run tests you expect to win. That filter is rational for each individual decision and ruinous in aggregate, because it systematically excludes the tests whose outcome you cannot predict, which are the only tests carrying real information. What you get instead is a year of confident, incremental work with a respectable win rate and a flat funnel.

The two obvious fixes both make it worse. Adding a process step — an experiment review board, a hypothesis template, a quarterly test roadmap — raises the cost of each test, which tightens the filter that caused the problem. Buying a testing tool removes the instrumentation cost but not the political one: no tool makes it survivable to be publicly wrong about the company’s founding assumption.

How it works

Reported by Kyle Poyar in Growth Unhinged on 4 March 2026, from interviews with Kameron Tanseli. Except inside a quote block, the wording is Poyar’s narration. Tanseli’s role is recorded as of that date; his own site now says he leads growth engineering at Flora AI.

One gap: the article’s section on how the team operates sits behind a paywall we did not buy. Everything below is from the free portion — the cadence figures and the experiment log.

The cadence. Over 12 months Fyxer ran 514 experiments, more than two per working day. The growth engineering team launched 360 of those on its own, and that team is four engineers strong — 90 experiments per engineer per year. Fyxer, an AI email assistant, went from $1 million to $30 million ARR across 2025. The experiments below run in the order the article dates them.

February — the card gate. Fyxer had a 7-day free trial with free-to-paid conversion of about 5%. Tanseli ran an experiment asking new signups to add a credit card upfront. Conversion jumped from 5% to 35%, and Fyxer called it a winner after only 8 days. Poyar asked what happened to signups: they dipped, and paying customers doubled anyway. Fyxer’s own in-product paywall was optional during the test, which the article says made this essentially free money on new-user traffic. Tanseli on why the moment was right:

It was around the time when a lot of AI apps required a credit card upfront

Kameron Tanseli, on the market norm at the time source ↗

There were changing perceptions around the willingness to do this. We also had a high intent rate since people were connecting to their email.

Kameron Tanseli, on perceptions and intent source ↗

February to March — the annual shift. Most signups were choosing month-to-month. Against a control that defaulted to monthly, the winner defaulted to the yearly plan, offered a 25% yearly discount and communicated the effective price per month. It 2.3x’ed the share of new trials going annual; Fyxer now sees 50% of paying customers sign up for annual plans.

March — pricing. Fyxer had a single package at $30 per user per month. Tanseli added a larger Pro tier at $50, across three variations, running 3–10 March. The winner raised the price and defaulted to annual communicated as $45.83 per month: month 0 revenue per trial rose 67%, checkout dipped by only 6%, trial starts were relatively unaffected.

March — trial lengths, and the losses. Tanseli first tried gamifying the trial — a longer trial for inviting teammates — across multiple versions over weeks. Nothing worked. He then tested lengths from 3 to 28 days. The 3-day trials failed outright:

People would cancel immediately and we had terrible trial start rates.

Kameron Tanseli, on 3-day trials source ↗

Seven days worked best for overall conversion, though personal-email signups converted better at 14. Segmenting the two lifted the personal-user trial start rate from 13.4% to 22.1%.

April to June — smaller surfaces. Adding social proof to the “Connect your email” screen, plus an explanation of what Fyxer’s assistant would do, moved email connection from 57.1% to 59.8% — a +4.7% uplift reported as significant at p=97.5%. That is the only significance figure in the readable text; the source states no threshold and no rule for calling a winner.

One account from outside. GrowthBook, the A/B testing vendor Fyxer runs on, puts Fyxer’s win rate at 25% — three quarters of their ideas failed. It confirms the four-person team and its 360 experiments, and separately reports a plan to scale that team from 6 to 13 while targeting 1,000 experiments. Two figures clash with Growth Unhinged: $35M ARR rather than $30M, and 541 experiments rather than 514, against an identical 360. One is a transposition we cannot resolve.

Why it works — our read

What the practice changes at the input is not the number of experiments. It is the price of a wrong one. At a quarter per test, only hypotheses you expect to win clear the bar. At eight days, being publicly wrong costs about a week — and the eligibility rule for what may be tested quietly changes underneath you. That, and not the throughput, is what a team is buying here. “Run more experiments” is the wrong lesson to take from it.

The second link in the chain is the denominator. A card gate is guaranteed to reduce trial starts — that part was never in doubt. A team whose scoreboard is trial starts would have read the identical experiment as a catastrophe and reverted it. The outcome is only legible because it was measured in free-to-paid conversion and total paying customers, where a volume metric is allowed to fall. Choosing what you are measured on is upstream of choosing what you test.

The third link is the filter. Velocity without one is just faster noise: two tests a day with a loose bar ships a steady stream of false positives that never get unshipped, because nobody re-runs a winner. The load-bearing number in this whole story is not 514, it is the 25% win rate — a figure you can only compute if losers are recorded as losers rather than quietly abandoned.

The enabling structure is engineering rather than marketing. The loop closes in days because the people forming the hypothesis can ship it. The counterfactual is the same four people as growth marketers with an engineering queue in between: the cadence collapses to the queue’s latency, each test is expensive again, and the safe-tests-only filter closes back over the funnel.

Velocity is not an output metric. It is what buys you the right to test the assumption your funnel is built on.

Where it breaks

Ours, not the sources’. None of the below is claimed or denied by Growth Unhinged or GrowthBook.

  • The speed advantage is intrinsically tied to effect size. A 5% → 35% result resolves in eight days because it is a sevenfold swing. Trigger condition: an expected effect under roughly 20% relative. Earliest symptom: your winners stop replicating when you re-run them.
  • Survivorship. This is a company that went from $1M to $30M ARR in a year, and hypergrowth traffic makes almost anything measurable. Run the same playbook on a slow-growing base and you will mostly measure noise, with the same confidence attached. Symptom: test durations quietly stretch from days to weeks and nobody re-plans around it.
  • The gate is contingent on demand you may not have. A card gate only survives where demand is abundant and intent at the point of the gate is already high. Shrinking the top of the funnel is affordable in that position and expensive outside it. Symptom: signups dip and paying customers do not double.

If you cannot detect a 20% relative change in your primary conversion metric inside two weeks, do not adopt this cadence. Run fewer, longer, bigger-swing tests instead.

What has to be true at your company

  • Someone who can form a hypothesis and ship it to production traffic without a queue. One person with both halves beats two people with one each.
  • Enough traffic that your primary conversion metric moves detectably in under two weeks. Run the power calculation before adopting the cadence, not after the first surprising result.
  • Losers recorded as losers, in one place, against a metric named before the test started. Without a loss log you have no win rate, and without a win rate no way to know whether your velocity produces knowledge or noise.
  • A money denominator agreed in advance. If the number on the dashboard is a volume metric, the org will revert your best test before you can explain it.
  • Standing sign-off, granted once. Whoever owns pricing and the funnel has to pre-authorise the class of test, not each instance — per-test approval reintroduces the cost that caused the problem.

Translating for a different shape: at eight people you probably already have someone who can form and ship a hypothesis, so your binding constraint is traffic, not structure — the adaptation is fewer, bigger, properly powered tests on the same assumption-first principle. At 800 people, or in an enterprise motion with a months-long sales cycle, the cadence is not available on the revenue metric at all. What transfers is the question, plus a faster proxy surface — pricing page, docs, trial, onboarding — where the belief can be checked at web speed rather than deal speed.

Try it this week

One hour, one person, no approval needed. Write down the three things about your funnel that are true “because everyone knows”. Next to each, write two numbers: the metric that would have to move for the belief to be wrong, and the metric you would actually be judged on if you ran the test. If those two are different, you have found the reason the test has never been run.

Then do one calculation. At your current traffic on that surface, how long would a test need to run to detect a 20% relative change? Any sample-size calculator will do.

It is working if at least one of the three turns out to be untested rather than tested, and the duration comes back in weeks. That is your first real experiment.

It is not working if all three durations come back in months. The cadence is not available to you, and copying it anyway will produce confident nonsense. Pick the single most load-bearing assumption and run it for the full duration instead.

Sources

  1. The AI native growth team Growth Unhinged (Kyle Poyar) · article · published Mar 4, 2026 · accessed Aug 3, 2026 · primary
  2. How a Team of 4 Used A/B Testing to Help Fyxer Grow from $1M to $35M ARR in 1 Year GrowthBook · post · published Apr 4, 2026 · accessed Aug 3, 2026
  3. Kameron Tanseli | Home kamrn.com · post · accessed Aug 3, 2026

Kameron Tanseli — lead of the growth engineering team, Fyxer — role and organisation as stated in the cited source at the time it was published. Quotes are verbatim from the linked source; everything under “Why it works — our read” and “Where it breaks” is RevGuild analysis, not a claim by Kameron Tanseli. Spotted something wrong? hello@revguild.org — we correct in place and say so.

Frequently asked questions

Isn't this just 'run more A/B tests'?

No, and reading it that way is the main failure mode. Volume on its own gives you a faster version of the same safe roadmap. The thing worth copying is what the speed unlocks: when a test resolves in days rather than a quarter, the cost of being wrong in public collapses, and you can finally run the test whose result you cannot predict. Fyxer's headline win — gating the trial behind a credit card — is exactly the test a slow org never gets to.

We don't have Fyxer's traffic. Is any of this usable?

The cadence isn't, and copying it at low volume is worse than not testing, because you get the same confidence from underpowered results. What transfers at any size is the question: which number in our funnel is a belief we have never actually checked? Run one properly powered test on that, for however many weeks it takes, instead of four two-week tests on things you already believe.

Won't a credit card gate kill our signups?

Almost certainly, and that is the wrong number to defend. At Fyxer signups did dip — Growth Unhinged reports that plainly — and total paying customers still doubled. Before you argue about the gate, settle which of those two your team would be judged on, because that decides the answer before the test runs. Note the specifics you would be borrowing: one company, in early 2025, selling an AI email assistant to people who had just connected their inbox, at a moment when Tanseli says AI apps commonly asked for a card. None of the sources claim it generalises, and neither do we.

Eight days sounds far too short to call a test. Isn't that p-hacking?

Our read is that eight days was readable because the effect was a sevenfold swing on a high-traffic surface, not because eight days is a rule. The article gives a significance figure for exactly one, much smaller test (p=97.5% on a +4.7% uplift) and states no general threshold and no stopping rule, so we cannot tell you what theirs was. Run the standard sample-size arithmetic yourself and the gap is stark: on a 5% baseline, a sevenfold lift is detectable in well under a hundred conversions per arm, while a 5% relative lift needs six figures. Same eight days, completely different question.

Do we need growth engineers, or can our marketing team do this?

Ownership matters, the job title does not. The concrete test: can the person who forms the hypothesis merge to production traffic without filing a ticket for someone else to pick up? If your marketers can, call them growth engineers and move on. If they cannot, the reporting line is the thing to fix — a team with a standing engineering dependency runs at that dependency's latency however the org chart is drawn, and no tool closes that gap.

Apply →