Skip to content

Peek all you want.

Anytime-valid confidence intervals: statistically valid at every moment of the experiment, so watching the truth narrow in real time is the product, not cheating.

No hosted signup yet — we’re pre-launch. Early access means you’re first in, and you get a say.

Effect on primary metric · 95% confidence sequenceConclusive · stopDrag the day, every peek is valid
Day
28 of 28
Visitors per arm
100,800
Estimate
+24.1%
Anytime-valid interval
+5.0% to +46.8%
Conventional p if you peek now
< 0.001“winner”
Illustration of the mechanism — how anytime-valid intervals behave under continuous peeking. Not customer data; we have no customers yet. The relative interval is unbounded until day 3, when the control interval clears zero and the band begins. The vertical axis is a symmetric log scale so the first bounded interval fits the frame. Synthetic run: 3,600 visitors per arm per day, +3.0% baseline, true lift +20%, seeded. The interval uses Splitmill’s boundary; the conventional line is a two-proportion z-test at α = 0.05, which called a “winner” on day 2, retracted it on day 4, and flipped 5 times before the honest interval concluded on day 21.

Every conventional A/B test quietly assumes you’ll look at the results exactly once. You look every day.

That mismatch manufactures false winners — and the industry built its dashboards on top of it.

Splitmill runs on anytime-valid confidence sequences: the interval is statistically valid at every moment of the experiment. Watching the truth narrow in real time isn’t cheating anymore. It’s the product.

In memoriamGoogle Optimize2012 – September 30, 2023

On September 30, 2023, Google shut down Optimize. Hundreds of thousands of websites lost their A/B testing tool overnight. The endorsed replacements started around $10,000 a year. A lot of teams just… stopped testing.

If that was you: welcome. You’re the reason this exists. Read the obituary

Your winner was noise.

Not all of them. But more than anyone admits — and the worst part is you can’t tell which. Peek-until-green isn’t analysis, it’s a slot machine with a quarterly review. Splitmill removes the machine’s ability to pay out on noise.

  1. Illustrations of the product; not customer data.

    The plan is sealed.

    Primary metric, minimum detectable effect, direction, minimum runtime — locked when the experiment starts and stamped with a cryptographic hash. Nobody gets to discover, three weeks in, that the “real” goal was a different number.

  2. Illustrations of the product; not customer data.

    The interval never lies about when you looked.

    Anytime-valid confidence sequences stay honest under continuous monitoring. Check hourly. Check obsessively. The math already assumed you would.

  3. Illustrations of the product; not customer data.

    Broken data gets a locked door.

    If the traffic split doesn’t match what you configured — sample ratio mismatch, the classic silent experiment-killer — Splitmill doesn’t put a warning triangle next to a green number. It blocks the decision panel. No verdict from poisoned data.

    break it yourself — live demo (coming soon)
  4. Illustrations of the product; not customer data.

    Guardrails can’t be quietly skipped.

    Skipping a guardrail metric requires an explicit, recorded waiver. Your future self can see exactly what you chose to ignore.

  5. Illustrations of the product; not customer data.

    The denominator is honest.

    Results are computed over visitors who actually saw the variant, and the dashboard states plainly what fraction of assigned traffic that is.

Illustrations of the product; not customer data.

An instrument, not a hype-man.

The only A/B testing tool that will talk you out of an A/B test.

Here’s a table nobody in this industry likes to show you. At 1,000 weekly visitors, “detectable” doesn’t mean button color. It means completely different page. Most tools happily let you run that button-color test for six weeks and hand you noise with a bow on it.

Splitmill’s feasibility advisor runs this math before your experiment starts — green, yellow, red — and when the answer is red, it says so and proposes methods that actually work at your traffic: painted-door tests, preference tests, a much bolder variant.

A tool that’s honest about when not to use itself is a tool you can trust when it says go.

ask the advisor yourself — live demo (coming soon)

What a conventional 4-week test can detectAdvisor: check first
Weekly visitorsSmallest detectable liftAdvisor
500~84% relativeRed
1,000~57%Red
5,000~24%Yellow
10,000~17%Green
  1. Assumptions: conventional fixed-horizon test · 3% baseline conversion · ~4-week runtime · 80% power · two-sided α = 0.05 · 50/50 split. Our math — check it.
  2. Splitmill’s anytime-valid math is stricter still; the advisor tells you the real number.

It’s Tuesday. You have an idea. The test is live by lunch.

If you used Optimize, you already know where everything is. The operator surface is deliberately Optimize-shaped — the workflow your team already had muscle memory for, rebuilt on an engine that deserves it. (Shaped like it. Not affiliated with the company that shot it.)

  1. 09:40

    Say it.

    “I want to test whether a shorter headline on the pricing page increases signups.” The AI turns that sentence into a complete draft experiment — hypothesis, primary metric, guardrails, feasibility check.

  2. 10:15

    Review it.

    You review, adjust, and confirm. Point, click, rewrite the headline in the visual editor — an overlay in your own browser tab. No Chrome extension, no developer ticket, no staging-environment séance.

  3. 11:50

    Seal it. Sealed

    The plan locks, the hash is stamped, the test goes live. When results come in, the AI explains them in plain English — constrained by the actual statistics. Drafts are the AI’s job. Decisions are yours.

delivery today: client snippet + browser SDK · minified bundle · DOM delivery fails open to your control page — never a blank screen. Server, edge & framework SDKs for a zero-flicker first paint are planned. Times above are narrative, not a benchmark.

No one sunsets what you host yourself.

The lesson of September 2023 wasn’t “pick a better vendor.” It was: an experimentation program is institutional memory, and you don’t build institutional memory on someone else’s whim.

Splitmill is self-hostable end to end. Your experiments live in your database. Your events flow to a first-party endpoint on your own domain. Your historical results outlive every vendor’s strategy meeting, including ours.

And consent isn’t a banner bolted on top — it’s inside the assignment logic. Before consent: computed in memory, no cookie written, no event sent. Denied: the buffer is discarded. Granted: everything flushes, cleanly. Built for the regulatory world the old tools died avoiding.

Test like nobody can take it away. Nobody can.

Your infrastructure · your domain
Ingestion endpoint
first-party, yourdomain.com
Event store
ClickHouse, dedup-by-design
Control plane
experiments, plans, seals
Stats engine
anytime-valid sequences, gates
SDK — browser available · node + edge planned · consent-aware by design
outside this line: nothing you depend on.

Check our math. Please.

stats/
anytime-valid confidence sequences (Waudby-Smith et al. lineage), computed per experiment against your sealed plan.
bucketing/
deterministic: MurmurHash3, 10,000-bucket integer space, two-hash eligibility/assignment — traffic changes never silently reshuffle existing users. Cross-SDK conformance fixtures.
gates/
SRM: chi-squared, p < 0.001 → decision panel blocked. Exposure-first analysis; attribution windowed from first exposure. Secondary metrics labeled exploratory — never “winners.”
diagnostics/
flight recorder: trace any visitor’s evaluation → delivery → attribution as a live timeline. “Is it even working?” takes seconds, not a support ticket.

This section is deliberately unglamorous. You’re the audience that reads footnotes for fun.

Doors open soon. The grudge is already warm.

We’re pre-launch: the engine runs end to end today, and we’re choosing early users deliberately. Tell us what you lost in 2023, or what you never got to start — first in line, and a real say in what ships.

No pricing page. No login button. Pricing is unannounced — anyone who tells you otherwise is a different company.

Early-access registerIntake closed

Intake coming soon — there’s no waitlist endpoint wired up in this preview.

Splitmill — Peek all you want.