Experiments
Feature Flags

Best free feature flagging tools worth trying in 2026

A graphic of a bar chart with an arrow pointing upward.

A free feature flag tool is only useful if it makes production releases safer without creating a second problem: stale flags, brittle SDK calls, hidden pricing meters, or experiment results nobody trusts.

Feature flags sound simple. Wrap code in a conditional, turn it on for a few users, and roll it out when the change looks safe. In practice, a flag platform becomes part of your release system. It touches deploys, incident response, QA, environment management, permissions, product analytics, and A/B testing.

That is why "free" needs a careful definition. A free open-source flag system may be the right answer when you need control. A hosted free tier may be better when you need speed. A lightweight OpenFeature backend may be enough for a small service. A broader platform may be better when flags need to connect directly to experiment metrics.

This guide compares free feature flagging tools worth trying in 2026, with a developer-first lens.

Quick comparison

ToolFree pathBest forMain watchout
GrowthBookFree Cloud Starter plan and free self-hosted optionTeams that want feature flags, experiments, and product analytics togetherAdvanced governance and statistics vary by plan
UnleashOpen-source self-hosting and 14-day Enterprise trialTeams that want mature open-source feature managementHosted Enterprise pricing is seat-based and not a free long-term SaaS plan
FlagsmithFree cloud plan and open-source/self-hosted optionsTeams that want flags, remote config, segments, and deployment flexibilityFree cloud limits can be tight for teams beyond one operator
ConfigCatForever Free hosted planTeams that want simple hosted flags with broad SDK coverageFree plan has flag, environment, targeting, and config download limits
DevCycleFree hosted planDevelopers that want a modern hosted flag platform with OpenFeature supportFree plan has MAU, request, and event limits
PostHogMonthly free allowance for feature flag requestsStartups that want flags inside an analytics and experimentation suiteCosts can spread across events, replays, flags, and other product meters
StatsigFree Developer tierTeams that want gates, configs, experiments, and analytics in one managed platformScale moves into metered events and paid packaging
FliptFree open-source self-hostingTeams that want Git-native feature flag workflowsMore self-hosting and workflow assembly than hosted tools
GO Feature FlagFree open-source, OpenFeature-native systemTeams that want lightweight flags without a databaseBetter for developers comfortable with config-driven operations

This list intentionally excludes tools that are strong but not really free in a durable way for most teams. It also keeps OpenFeature separate from the product list. OpenFeature is an open specification for vendor-agnostic feature flag APIs. It is important because it reduces code-level lock-in, but it is not a complete flag management platform by itself.

How to evaluate free feature flagging tools

The best free feature flag platform depends on the job you need flags to do.

Release control is the baseline

At minimum, a feature flag tool should let you decouple deploy from release. You should be able to deploy code hidden behind a flag, turn it on for internal users, expand to a small percentage of production traffic, and roll back without redeploying.

That baseline requires more than a boolean toggle. Look for environments, targeting rules, percentage rollouts, deterministic assignment, SDK fit, audit history, and a way to see which flags exist across services.

Experimentation changes the requirement

If a flag controls a product change, the next question is usually whether the change worked. That is where many free feature flag tools split into two groups.

Some tools control exposure and leave measurement to another analytics platform. Unleash, Flagsmith, ConfigCat, Flipt, and GO Feature Flag can work this way. That is fine when your team already has a strong analytics and statistics workflow.

Other tools connect flags to experiment readouts. GrowthBook, PostHog, Statsig, and DevCycle all put feature flagging closer to experimentation. GrowthBook is the clearest free choice when the experiment analysis should use warehouse-native metrics and stay transparent to the data team.

Pricing meters matter earlier than you think

Free feature flag tools can be limited by users, requests, CDN downloads, client-side monthly active users, server config requests, events, environments, projects, products, permission groups, or advanced governance.

The wrong meter can make a free pilot misleading. A tool may be free for one developer testing one service, then expensive once every product surface evaluates flags frequently.

Model the tool at today's usage, 3x usage, and 10x usage. Include both read path and write path: SDK evaluations, config downloads, server requests, flag updates, audit needs, and who needs access to the UI.

1. GrowthBook

GrowthBook is the strongest free feature flagging tool for technical teams that also care about experimentation, product analytics, and data ownership.

Best for

GrowthBook fits teams that do not want feature flags to live in a release silo. A flag should control who sees a change. It should also be able to become an experiment, connect to metrics, and help the team decide whether to keep rolling out.

That makes GrowthBook a good default for SaaS teams, AI product teams, growth teams, and engineering-led product organizations. You can start with feature flags, then add A/B testing and product analytics without switching platforms.

The current GrowthBook pricing page lists a free Cloud Starter plan with up to three users, unlimited feature flags, unlimited experiments, and unlimited traffic. It also lists a free self-hosted open-source plan with unlimited feature flags, unlimited experiments, and unlimited traffic.

Key strengths

GrowthBook's first strength is that feature flags and experiments are built to work together. The feature flag experiments docs show how teams can use experiment rules on flags to randomly assign users and track exposure. That avoids the common handoff where engineering manages rollout in one tool and product measures the result somewhere else.

The second strength is deployment flexibility. Teams can use GrowthBook Cloud for speed or self-host when they need more infrastructure control. That is useful for organizations that start small but later need stricter data residency, network, or audit requirements.

The third strength is local flag evaluation in common SDK patterns. GrowthBook's feature flag architecture is designed so SDKs can evaluate rules from cached feature definitions instead of making a blocking network request for each normal flag check. That matters in latency-sensitive code paths, but the right claim is practical: evaluate your own runtime path and SDK configuration rather than assuming any flag tool is automatically free from performance tradeoffs.

Watchouts

GrowthBook's free Cloud plan is intentionally small by user count. If many PMs, engineers, analysts, and QA teammates need platform access, paid plans may become relevant quickly.

Some advanced feature management, governance, and experimentation capabilities also vary by plan. The pricing page currently shows items like advanced permissioning, release plans, scheduled flags, prerequisites, approval flows, and advanced statistics in paid or enterprise areas. Verify the exact plan fit before making GrowthBook the central release control plane for a large organization.

Pricing and implementation notes

Start with one real feature flag and one experiment candidate. Test internal targeting, percentage rollout, SDK behavior, rollback, exposure tracking, and metric readout. If the team can control release and measure impact with the same platform, GrowthBook has validated the job that matters most.

GrowthBook is the best free default when you want flags to grow into a disciplined experimentation program.

2. Unleash

Unleash is one of the strongest open-source feature flagging platforms for teams that want self-hosted feature management and mature release controls.

Best for

Unleash fits engineering teams that want to own their feature flag platform, run it in their infrastructure, and focus on release management. It is especially relevant for platform teams, regulated environments, and organizations that want open-source feature management without building an in-house system.

Unleash also has a clear enterprise path. The current Unleash pricing page positions Pay-As-You-Go Enterprise as a hosted 14-day trial with a $75 per-seat monthly price and a self-hosted minimum. That means the free path is mainly open-source self-hosting or short hosted evaluation, not a permanent hosted free tier.

Key strengths

Unleash has mature feature-management concepts: projects, environments, activation strategies, variants, stickiness, gradual rollouts, feature flag metrics, and lifecycle controls. Its pricing page lists unlimited feature flags, projects, environments, experiments, A/B/n testing with flag variants, custom activation strategies, targeting, user segmentation, and 25+ SDKs on Enterprise packaging.

Unleash is also a strong fit when the organization wants governance around flags: approvals, audit logs, naming conventions, stale flag tracking, and role-based access. Those capabilities matter once flags move from a developer convenience to a release control system.

Watchouts

If you want hosted free forever, Unleash is not the cleanest fit. It is best understood as a strong open-source self-hosted option with paid enterprise packaging.

The second watchout is experimentation depth. Unleash supports A/B/n testing with flag variants, but teams should verify how they will calculate outcomes. In many deployments, Unleash controls assignment while analytics or experimentation analysis happens somewhere else.

Pricing and implementation notes

Use Unleash when the main problem is open-source feature management. For a proof of concept, self-host it, connect one service, create one gradual rollout, test variant stickiness, and define your flag cleanup process.

If your team also wants experiment statistics and warehouse-native metrics in the same platform, compare Unleash against GrowthBook before standardizing.

3. Flagsmith

Flagsmith is a strong free and open-source feature flagging tool for teams that want deployment flexibility and a simple cloud starting point.

Best for

Flagsmith fits teams that want feature flags, remote configuration, segmentation, multivariate flags, and the option to run the platform as cloud, private cloud, or self-hosted. It is a good option when feature management is the primary need and the team wants open-source control.

The current Flagsmith pricing page lists a free cloud plan with up to 50,000 requests per month, one team member, unlimited feature flags, unlimited environments, unlimited identities and segments under fair-use terms, and API access.

Key strengths

Flagsmith's open-source posture is clear. The open-source feature flags page says core functionality includes flags, segments, identities, remote configuration, user targeting, multivariate flags, and local evaluation across supported languages. It also says enterprise governance features such as SSO, audit logs, role-based access, change requests, and support are paid rather than fully open source.

That is a useful boundary. Teams can get real feature flagging for free, then decide whether enterprise governance is worth paying for later.

Flagsmith also works with OpenFeature, which can reduce code-level lock-in for teams that want a standardized API in application code.

Watchouts

The free cloud plan is small for team collaboration because it includes one team member. That can be fine for evaluation or a solo project, but it is not enough for a product team where engineers, PMs, QA, and support need access.

The pricing page also shows that A/B and MVT testing appear in paid Start-Up packaging. Teams using the free plan should verify exactly what experimentation workflow they need and whether they plan to analyze results in another analytics tool.

Pricing and implementation notes

Flagsmith is worth trying when you want a credible open-source feature flag platform with a hosted free start. For a proof of concept, test both cloud and self-hosting assumptions. A team that starts cloud because it is easy may later choose self-hosting for traffic, cost, or governance reasons.

4. ConfigCat

ConfigCat is a simple hosted feature flag tool with a useful Forever Free plan and a pricing model based on config downloads rather than flag reads or MAUs.

Best for

ConfigCat fits teams that want managed feature flags without operating infrastructure and without buying a large product-development suite. It is especially useful for small teams that need predictable hosted flags, broad SDK coverage, and straightforward rollout controls.

Current ConfigCat pricing says the Forever Free plan has the same feature set as paid plans, requires no card, and includes 5 million config JSON downloads per month, 20 GB network traffic, 10 feature flags, two environments, two products, two segments, four targeting rules per flag, and four percentage options per flag.

Key strengths

ConfigCat's pricing meter is easy to reason about. The pricing page explains that feature flags are stored in config JSON files on ConfigCat's CDN, SDKs download and cache config locally, and ConfigCat counts config JSON downloads rather than feature flag reads or evaluations. For many teams, that is easier to forecast than per-user or per-evaluation pricing.

The free plan also supports unlimited seats, service connections, MAUs, contexts, and feature flag reads, according to the current pricing table. That is useful when many developers need access but the actual flag footprint is small.

Watchouts

The free plan's 10-flag limit is real. It is enough for evaluation and small products, but it will not support a mature feature flag program without cleanup discipline or a paid plan.

ConfigCat is also a feature flag and remote configuration product, not a full experimentation and product analytics platform. If your team wants feature flags to feed experiment results and warehouse metrics, GrowthBook will usually be a stronger fit.

Pricing and implementation notes

Use ConfigCat when you want a hosted flag service that is easy to understand and quick to adopt. For a proof of concept, test SDK caching, config download behavior, targeting, and rollback in a production-like environment.

If the first few flags become permanent application configuration, write a cleanup policy early. A small free flag limit can become a helpful forcing function.

5. DevCycle

DevCycle is a modern hosted feature flag platform with a free plan, strong developer ergonomics, and OpenFeature support across its SDKs.

Best for

DevCycle fits developers who want to try a managed feature flag platform quickly. It is especially relevant when OpenFeature compatibility matters or when the team wants a hosted product with A/B testing, debugging tools, integrations, flag schemas, and MCP capabilities.

Current DevCycle pricing lists a free plan with unlimited seats, unlimited flags, all integrations, debugging tools, A/B testing, MCP Server, flag schemas, custom property schemas, 1,000 client-side MAUs, 10,000 cloud config requests, 100,000 server config requests, and 5,000 events per month. The page also says DevCycle is now part of Dynatrace, which is a buyer due-diligence point for roadmap and packaging.

Key strengths

DevCycle's free plan is broad for evaluation. Unlimited seats and flags reduce friction when a whole development team wants to try the product. OpenFeature support is also meaningful because it can reduce code-level lock-in if the team later changes providers.

The product is oriented toward developers rather than only release managers. Schemas, debugging tools, APIs, CLI, integrations, and feature opt-in flows make it appealing when engineering teams want flags to fit daily development habits.

Watchouts

The free plan has meaningful usage limits. Client-side MAUs, cloud config requests, server config requests, and events all matter. A developer proof of concept can fit easily; a production rollout with many users and services may move into Business quickly.

DevCycle is also a managed SaaS path, not the obvious choice when the requirement is self-hosting or open-source infrastructure control.

Pricing and implementation notes

Use DevCycle when developer experience and hosted simplicity matter more than self-hosting. For a proof of concept, test both client-side and server-side evaluation because the free plan meters them differently.

Also ask roadmap questions. A product being part of Dynatrace may be positive for enterprise support, but buyers should verify current packaging and product direction before standardizing.

6. PostHog

PostHog is a useful free feature flagging option when flags should live inside a broader product analytics and experimentation suite.

Best for

PostHog fits startups and product teams that want analytics, session replay, feature flags, experiments, surveys, and product data workflows in one place. Feature flags are not isolated release toggles in PostHog. They sit beside the event data and product analysis that explain how users respond.

Current PostHog pricing lists a monthly free allowance including 1 million feature flag requests and experiments billed with feature flags, alongside free allowances for analytics events and session recordings.

Key strengths

The main strength is context. If a flag controls a feature and the same tool tracks product events, cohorts, funnels, and replays, teams can investigate a rollout without switching across many systems.

PostHog is also useful for early-stage teams that want one product suite. A small team can instrument events, create flags, run simple experiments, and watch user behavior without buying three tools.

Watchouts

The same breadth can complicate pricing. Feature flag requests, analytics events, replays, surveys, data warehouse features, and other products can all become usage lines as the team adopts more of the platform.

PostHog is strongest when you are comfortable using PostHog as a product event platform. If your experimentation metrics must come from a governed data warehouse, GrowthBook may fit better.

Pricing and implementation notes

Use PostHog when flags and product analytics should share one event system. For a proof of concept, create a flag, measure user behavior around it, and inspect whether the result answers a release question clearly enough for product and engineering.

7. Statsig

Statsig is a strong free-tier feature flagging option for teams that want gates, configs, experiments, and product analytics in one managed platform.

Best for

Statsig fits teams that want feature gates and experimentation to be part of a broader product-development workflow. It is especially relevant for data-forward teams that want release measurement, product analytics, and experimentation under one vendor.

Current Statsig pricing says new accounts start on a Developer tier with free access to feature gates, dynamic configs, experimentation, and analytics, plus 2 million metered events each calendar month.

Key strengths

Statsig's free tier is credible for evaluating a real workflow. It is not just a flag toggle. Teams can test gates, configs, experiments, analytics, and the event model together.

Statsig is also strong for teams that want managed platform depth without running infrastructure. If you want a SaaS product with sophisticated product development workflows, it belongs on the shortlist.

Watchouts

Statsig is not an open-source or self-host-first flag platform. If the requirement is code transparency, self-hosting, or warehouse-native analysis under your direct control, GrowthBook, Unleash, Flagsmith, Flipt, or GO Feature Flag may be more natural.

Events are also the scale meter. The free tier's 2 million events can be generous for a pilot, but teams should model event volume before rolling out widely.

Pricing and implementation notes

Use Statsig when a managed feature flag and experimentation suite is the goal. For a proof of concept, include one flag, one dynamic config, one experiment, and a cost model based on expected event volume.

8. Flipt

Flipt is a good free open-source option for teams that want Git-native feature flag management and self-hosted control.

Best for

Flipt fits teams that want feature flag changes to behave like code changes. Developers can manage flags through Git workflows, and every change can become reviewable infrastructure rather than a separate click path in a vendor UI.

The Flipt website positions the open-source edition as free forever, with unlimited feature flags, Git-native workflows, an intuitive UI with Git sync, real-time updates, REST and gRPC APIs, and community support. Paid Pro and Enterprise tiers add managed workflow and support capabilities.

Key strengths

Git-native feature management is Flipt's main differentiator. For teams that already use GitOps practices, keeping flag definitions close to version control can reduce confusion and improve auditability.

Flipt is also lightweight compared with broader product suites. That is useful when the team wants a flag control plane, not analytics, session replay, experimentation, and release governance all bundled together.

Watchouts

Flipt requires more operational ownership than hosted free tiers. You need to deploy it, integrate it, and decide how flag changes fit your Git workflow.

It is also not the best choice if the main need is built-in experiment analysis. Flipt can be part of a controlled rollout workflow, but you should plan how outcomes will be measured.

Pricing and implementation notes

Use Flipt when Git-native control is the point. For a proof of concept, create one flag through the UI, verify the Git commit path, test SDK behavior, and walk through a rollback.

If product managers or support teams need frequent non-code flag changes, make sure the Git workflow helps rather than slows them down.

9. GO Feature Flag

GO Feature Flag is a lightweight open-source feature flag system built around OpenFeature, with no database requirement and a configuration-driven operating model.

Best for

GO Feature Flag fits developers who want a small, infrastructure-light way to add feature flags without adopting a large management platform. It is especially useful for teams that like OpenFeature and want file-based configuration.

The GO Feature Flag homepage describes it as 100% open source, MIT licensed, OpenFeature-native, and designed to run on existing infrastructure with no database to operate and no per-seat bill. It supports targeting, rollout functionality, usage data, and multiple languages and frameworks.

Key strengths

The main strength is simplicity. A Docker container, a YAML file, and OpenFeature SDKs can be enough for small services or teams that do not need a large UI-driven platform.

OpenFeature alignment is another strength. OpenFeature providers create an abstraction between application code and the underlying flag system, making it easier to change the backend later without rewriting every flag call.

Watchouts

GO Feature Flag is best for developer-controlled workflows. If your organization needs rich collaboration, advanced permissions, audit workflows, product-manager-friendly UI, and built-in experimentation analysis, it may feel too lightweight.

File-based configuration can be powerful, but it requires discipline around review, deployment, and rollback.

Pricing and implementation notes

Use GO Feature Flag when you want a free, lightweight, OpenFeature-native setup. For a proof of concept, add it to one service and measure the operational path: define a flag, change rollout percentage, target users, observe usage, and remove the flag.

Choosing by deployment model

Free feature flagging tools usually fall into three deployment patterns.

Deployment modelBest whenTools to evaluate
Hosted free tierYou want fast setup and low operational burdenGrowthBook Cloud, Flagsmith Cloud, ConfigCat, DevCycle, PostHog, Statsig
Self-hosted open sourceYou need infrastructure control, auditability, or data residencyGrowthBook, Unleash, Flagsmith, Flipt, GO Feature Flag
Standards layerYou want to reduce code-level lock-in across toolsOpenFeature with supported providers

Hosted free tiers are the fastest way to learn. Self-hosted tools are best when control matters more than setup speed. Standards layers are useful in either model because they reduce the amount of product-specific code in your application.

Do not choose a tool only because it is open source. Choose it because your team can operate it well. Do not choose a hosted free tier only because it is fast. Choose it because the pricing model and governance path still make sense after the pilot succeeds.

Hidden costs to check before the free plan becomes infrastructure

Free feature flagging tools often pass the first demo and fail the second phase. The first demo is usually one service, one developer, one boolean flag, and no governance. The second phase is where the real operating model appears: many services, many environments, multiple teams, stale flags, support incidents, and product managers asking whether rollout changed the metric.

Flag sprawl is a workflow problem, not a storage problem

Most tools can store more flags than a team should keep. The danger is not database capacity. The danger is that old flags become permanent conditional logic across the codebase.

A healthy flag workflow needs ownership, intent, and cleanup. Every flag should have a type: release flag, experiment flag, permission flag, operational kill switch, migration flag, or long-lived configuration. Those types should not have the same lifecycle. A release flag should expire after rollout. An experiment flag should expire after the decision. A permission flag may live longer, but it should be documented as product behavior rather than temporary release machinery.

Free tools often hide this cost because early usage is small. By the time a team has 100 flags, cleanup is no longer a nice-to-have. It is release hygiene. When you evaluate a free tool, look for owners, tags, descriptions, audit history, stale flag indicators, code references, or at least API support so your team can build cleanup checks.

Client-side, server-side, and edge flags have different risks

A feature flag evaluated in a React component has different constraints from a flag evaluated in a payments service. Client-side flags may expose flag keys, targeting shape, or variation names to the browser. Server-side flags can protect more sensitive logic but may require different caching and update patterns. Edge flags can support fast personalization or routing, but the operational path for debugging is different.

This is why SDK coverage alone is not enough. Ask where evaluation happens, how often flag definitions update, what happens when the flag service is unreachable, and whether defaults are safe. A good SDK makes failure behavior boring. A bad integration turns a vendor outage, network issue, or bad default into product behavior.

Also ask whether the tool supports local evaluation, remote evaluation, streaming updates, polling, or proxy patterns. Each model has tradeoffs. Local evaluation can reduce runtime dependency on a remote decision call, but you still need a reliable way to distribute flag definitions. Remote evaluation can centralize logic, but it may add a service dependency in the request path. A free plan may support one model but not the model your production services need.

Experimentation requires exposure discipline

A flag that assigns users to variants is not automatically an experiment. Experimentation requires clean exposure logging. Exposure should be recorded when the user can actually experience the variant, not merely when the server checks a flag during a request the user never sees.

This distinction matters because noisy exposure data can dilute or bias results. If a backend service evaluates a flag for users who are never shown the feature, the experiment may include people who had no chance to respond. If a client logs exposure before the UI renders, errors or navigation can create similar problems.

Dedicated experimentation platforms tend to make exposure tracking more explicit. Flag-first platforms may leave more of the responsibility to your analytics layer. Neither model is automatically wrong. The important part is that your proof of concept includes exposure logging, not just targeting.

The free meter should match production behavior

Feature flag cost is often tied to behavior that engineers do not estimate during the first test. A frontend app may initialize the SDK for every anonymous visitor. A mobile app may fetch config at startup and again after login. A backend service may evaluate flags for internal jobs as well as user requests. A microservice architecture may multiply config requests across many services.

Before choosing a free tier, map the meter to real usage. ConfigCat counts config JSON downloads. DevCycle exposes client-side MAUs, cloud config requests, server config requests, and events. Flagsmith counts API requests on its free cloud plan. PostHog includes a monthly feature flag request allowance. Statsig uses metered events for its free Developer tier. GrowthBook's pricing page separates CDN requests, bandwidth, managed warehouse events, seats, and traffic.

The right question is not whether the pilot is free. The right question is whether the successful version of the pilot still fits the model.

Proof-of-concept checklist

Run the same proof of concept across every finalist.

  • Create one boolean flag and one string, number, or JSON flag.
  • Target internal users first.
  • Roll out to 1%, 10%, and 50% of production traffic.
  • Verify deterministic assignment across sessions.
  • Test both frontend and backend SDKs if your product uses both.
  • Confirm how often SDKs fetch or receive updated flag definitions.
  • Turn the feature off without redeploying.
  • Review the audit history for every flag change.
  • Add an owner and cleanup date to the flag.
  • Convert the flag into an experiment or connect it to an analytics readout.
  • Model cost at today's usage, 3x usage, and 10x usage.

This catches the most common failure modes before the tool becomes infrastructure. The tool should make rollback boring, targeting understandable, and cleanup visible.

The practical recommendation

For most technical SaaS teams, start with GrowthBook.

GrowthBook has the cleanest free path when feature flags are not only release toggles. You get a free hosted plan, a free self-hosted option, unlimited feature flags and experiments, and a direct path from controlled rollout to measurement. That makes it a better default for teams that want feature flags to support product decisions, not just deployment switches.

Choose Unleash if your main requirement is mature open-source feature management. Choose Flagsmith if you want open-source flags with a simple hosted start and deployment flexibility. Choose ConfigCat if you want hosted simplicity and a clear config-download pricing model. Choose DevCycle if you want a modern managed flag platform with a generous developer trial path and OpenFeature support. Choose PostHog or Statsig if you want flags inside a broader managed product suite. Choose Flipt or GO Feature Flag if you want a lighter self-hosted, developer-controlled setup.

The best free feature flagging tool is the one your team can keep using after the first rollout works. Test that by shipping one real change behind a flag, rolling it back, measuring the effect, and cleaning up the flag afterward. If the tool makes that loop clear, it is worth trying.

Table of Contents

Related Articles

See All Articles
Experiments
Feature Flags

What is mock testing? A complete guide for developers (2026)

Sep 9, 2026
x
min read

A mock can make a test fast and deterministic while letting the real integration break unnoticed.

That tension explains both the value and the reputation of mock testing. Replacing a payment API, database, clock, or feature service with a controlled double lets you force success, failure, timeout, and retry paths in milliseconds. But the substitute only behaves as accurately as the test author programmed it to behave.

Mock testing works best at a deliberate boundary. Use a mock when the interaction itself matters, a stub when you need a canned answer, and a fake when a lightweight working implementation makes the test clearer. Then pair those isolated tests with contract and integration coverage so production reality still gets a vote.

This guide uses TypeScript and Vitest examples, but the design choices apply across Jest, pytest, Mockito, Go interfaces, and other testing stacks.

Mock testing controls a collaborator and verifies the conversation

A test double is any non-production object used in place of a real dependency. Martin Fowler's test-double taxonomy distinguishes dummies, fakes, stubs, spies, and mocks. Teams often call all of them “mocks,” but the distinctions clarify what each test proves.

Mocks test observable interactions

A mock is preprogrammed with behavior and records or enforces expectations about calls. It answers questions such as:

  • Did the service publish an event after committing the order?
  • Was the payment gateway called once with the correct idempotency key?
  • Did the retry loop stop after the first successful response?
  • Was no email sent when validation failed?

This is behavior verification. The assertion concerns the messages exchanged with a collaborator, not only the final state of the system under test.

The Vitest mock-function documentation exposes both sides: a vi.fn() can return configured values and retain its call history. Jest provides the same core pattern through 1.

Stubs supply answers; spies observe calls

A stub returns a canned response needed to exercise the unit. It may return an account, throw a timeout, or report that inventory is empty. The test normally asserts the state or return value produced by the system under test.

A spy wraps or replaces behavior while recording how it was called. Framework APIs blur these terms because a single function object can act as stub, spy, or mock depending on the assertion. Name the role in the test: paymentGatewayStub, sendEmailSpy, or clockFake communicates more than mockService.

Fakes implement a simplified working system

A fake has real behavior but takes a shortcut unsuitable for production. An in-memory repository can support insert, query, and uniqueness rules without running Postgres. A fake queue can preserve ordering and retries without a broker.

Fakes often reduce test setup and implementation coupling. The tradeoff is maintenance: the fake must stay behaviorally compatible with production. Android's official test-double guidance recommends checking whether a library supplies supported fakes before inventing one.

DoubleWhat it doesTypical assertionGood use
DummyFills an unused parameterNoneRequired context object
StubReturns configured answersResulting state or valueError and edge cases
SpyRecords calls, often keeping behaviorCall historyTelemetry or callback checks
MockSimulates behavior and verifies interactionsExpected message or callCoordination with side effects
FakeImplements a lightweight working substituteState and behaviorIn-memory repository or clock

Test releases behind flags

Learn how to structure feature flag ownership, observability, and cleanup so testable release controls do not become permanent debt.

Read the Feature Flag Guide

Start with a seam, not a mocking framework

A seam is a place where code can receive another implementation. Constructor parameters, function arguments, interfaces, adapters, and dependency-injection containers all create seams. A clean seam keeps tests focused and makes production dependencies replaceable for reasons beyond testing.

Inject the dependency your unit actually needs

Consider checkout coordination. The use case needs a gateway that can charge a payment. It does not need to know which HTTP client, authentication library, or vendor SDK implements the call.

exporttype Charge = {
  orderId: string;
  amountCents: number;
  idempotencyKey: string;
};

exportinterface PaymentGateway {
  charge(input: Charge): Promise<{ transactionId: string }>;
}

exportasyncfunction completeCheckout(
  gateway: PaymentGateway,
  input: Charge,
) {
  if (input.amountCents <= 0) thrownew Error("invalid amount");
  const result = await gateway.charge(input);
  return { orderId: input.orderId, paid: true, ...result };
}

The interface is small because it describes the capability the use case consumes. It prevents a unit test from mocking an entire vendor SDK, including methods the code never calls.

Configure the smallest behavior needed by the case

Now test the observable result and the critical side-effect contract:

import { expect, it, vi } from"vitest";
import { completeCheckout, type PaymentGateway } from"./checkout";

it("charges once with a stable idempotency key", async () => {
  const charge = vi.fn().mockResolvedValue({ transactionId: "tx_test_42" });
  const gateway: PaymentGateway = { charge };

  const result = await completeCheckout(gateway, {
    orderId: "order_42",
    amountCents: 2500,
    idempotencyKey: "checkout:order_42",
  });

  expect(result).toEqual({
    orderId: "order_42",
    paid: true,
    transactionId: "tx_test_42",
  });
  expect(charge).toHaveBeenCalledOnce();
  expect(charge).toHaveBeenCalledWith({
    orderId: "order_42",
    amountCents: 2500,
    idempotencyKey: "checkout:order_42",
  });
});

The return-value assertion protects the public behavior. The interaction assertion protects a meaningful external contract: a charge must happen once with an idempotency key. Avoid asserting incidental steps, such as which helper formatted the key, unless that detail is itself part of the boundary contract.

Force failures that are unsafe or slow to reproduce

Mocks are particularly useful for rare branches:

it("does not report a paid order when the gateway rejects", async () => {
  const gateway: PaymentGateway = {
    charge: vi.fn().mockRejectedValue(new Error("gateway unavailable")),
  };

  await expect(
    completeCheckout(gateway, {
      orderId: "order_43",
      amountCents: 2500,
      idempotencyKey: "checkout:order_43",
    }),
  ).rejects.toThrow("gateway unavailable");
});

This test needs no real outage and cannot charge a card. Add separate cases for timeouts, duplicate responses, invalid payloads, and retry exhaustion when your production policy distinguishes them.

Mock boundaries, not your own business rules

The best candidates are dependencies whose real behavior makes a focused test slow, flaky, destructive, expensive, or impossible to control.

Good mock targets have operational side effects

Common boundaries include:

  • Payment, email, SMS, and push providers.
  • System clocks, random-number generators, and schedulers.
  • Cloud APIs, object stores, queues, and search services.
  • Network failures, rate limits, timeouts, and malformed responses.
  • Analytics and exposure callbacks whose payload contract matters.
  • Feature evaluation at the edge of application logic.

For HTTP behavior, prefer a network-level tool when the request itself matters. Mock Service Worker intercepts REST and GraphQL requests independently of the application's request client. Playwright API mocking can intercept browser traffic, replay HAR data, and verify UI behavior. These tests exercise serialization and routing that a mocked fetch() wrapper might bypass.

Keep deterministic domain objects real

Value objects, parsers, pricing rules, eligibility policies, and other deterministic domain code are usually cheap to construct. Mocking them replaces the behavior you most need to test. Use real objects and assert meaningful outcomes.

A suite with 8 mocks for one method often signals one of 3 design problems:

  1. The unit coordinates too many responsibilities.
  2. The test boundary is smaller than the behavior anyone cares about.
  3. Global imports or singletons make dependencies hard to substitute.

Vitest's current module-mocking guide explicitly calls out limitations around mocking methods used inside the same module and recommends dependency injection or refactoring. Treat that friction as architecture feedback, not as a puzzle to defeat with more tooling.

Test state when the outcome matters more than the conversation

Interaction assertions couple a test to how work happens. A refactor that preserves behavior but combines 2 repository calls into 1 can break dozens of mock expectations. Prefer state verification when callers care about the result rather than the sequence.

Fowler's classic “Mocks Aren't Stubs” essay frames this as behavior versus state verification and explains the broader mockist and classical testing styles. You do not need to choose a camp. Make the choice per boundary.

Test feature-flagged code at three layers

Feature flags add a decision boundary: the same code path can produce multiple experiences based on attributes, configuration, and environment. Tests need to cover local branch behavior, SDK wiring, and the assembled product experience.

Unit-test branch behavior through a narrow reader

Do not make domain code depend on a global SDK object. Inject the capability it needs:

exportinterface FlagReader {
  enabled(key: string): boolean;
}

exportfunction priceSummary(flags: FlagReader, totalCents: number) {
  if (flags.enabled("compact-checkout")) {
    return`$${(totalCents / 100).toFixed(2)}`;
  }
  return`Order total: $${(totalCents / 100).toFixed(2)}`;
}

A tiny fake is clearer than a framework mock:

import { expect, it } from"vitest";

const flagsOn = { enabled: () => true };
const flagsOff = { enabled: () => false };

it("renders both checkout variants", () => {
  expect(priceSummary(flagsOff, 2500)).toBe("Order total: $25.00");
  expect(priceSummary(flagsOn, 2500)).toBe("$25.00");
});

These tests prove the application's branch logic. They do not prove that production attributes, flag rules, and SDK initialization select the branch correctly.

Integration-test the real evaluation contract

Add tests around your adapter using the real SDK with deterministic local configuration. Cover default values, missing attributes, targeting rules, percentage assignment, and the event or callback that records experiment exposure. The GrowthBook SDK documentation is the source of truth for supported language behavior, while feature flag experiments explain how evaluation becomes measured assignment.

Keep SDK-specific test helpers in the adapter package. When a library changes configuration or evaluation semantics, a small contract suite should fail before dozens of business tests do.

Exercise complete variants before release

Use end-to-end tests for the critical user paths in both states. GrowthBook's DevTools Extension can inspect evaluations, override feature values and attributes, and help developers reproduce specific experiences. This complements automated tests; it does not replace assertions in continuous integration.

The feature flags product supports targeted and gradual releases, while the experimentation workflow measures impact. Test that control exists before relying on either: default behavior, rollback path, exposure logging, and cleanup ownership all need coverage.

Prevent mocks from becoming a second production system

Mock-heavy suites tend to fail in predictable ways. The solution is not banning mocks. It is making their contract and scope explicit.

Reset state and avoid global leakage

Mocks retain implementations and call histories unless the runner restores them. Use lifecycle hooks or runner configuration consistently. Vitest warns developers to clear or restore mock state between tests in its mocking guide, and Jest distinguishes mockClear, mockReset, and mockRestore because they remove different things.

Run tests in random order periodically. A test that only passes after another test configured a global mock is not isolated. Prefer locally constructed dependencies over process-wide replacements.

Keep mock contracts honest

Every mock contains an assumption about production. Protect important assumptions with:

  • Consumer-driven contract tests for service boundaries.
  • Schema validation for recorded fixtures.
  • Integration tests against a disposable database or sandbox.
  • Scheduled refreshes for HAR files and response fixtures.
  • A small smoke suite against real third-party test environments.

If production adds a required field and your mock continues returning the old shape, isolated tests remain green. A contract test should expose the drift.

Assert outcomes before incidental calls

Start each test with the behavior a caller cares about. Add interaction expectations only for externally meaningful effects, ordering, idempotency, security, or compliance. Avoid assertions such as “helper A was called before helper B” when the order has no user-visible or contractual meaning.

Use mutation testing or a deliberate fault to check whether the assertion can fail for the right reason. A mock that returns exactly the value later asserted, without exercising transformation or policy, may test the fixture more than the code.

Escalate to a broader test when setup tells a story

If a unit test needs a page of mock configuration, try an in-memory fake or component test. Fowler's microservice testing guidance notes that too many doubles can signal a concept that should be extracted or a component boundary that would provide more value.

The target is not a particular ratio. It is fast local feedback plus enough real integration coverage to detect false assumptions.

Use mocks where control is valuable and realism is replaceable

Before replacing a dependency, ask 5 questions:

  1. Is the real collaborator slow, nondeterministic, destructive, costly, or hard to force into the needed state?
  2. Does this test care about the collaborator's answer, the interaction, or a larger outcome?
  3. Would a stub or fake express the case with less coupling?
  4. Which contract or integration test will detect drift from production?
  5. Will the test survive an internal refactor that preserves behavior?

Mock testing is successful when it buys control without hiding the system. Keep the seam small, configure only the behavior the case needs, assert externally meaningful outcomes, and verify important assumptions against reality elsewhere in the suite.

For feature-flagged delivery, that means unit-testing both application branches, contract-testing the SDK adapter, and exercising the assembled experiences before expanding traffic. GrowthBook can support the release and measurement layer, but the reliability begins with code that remains testable when every external service is unavailable.

Ship testable changes safely

Start with feature flags and experimentation in one workflow, then expand exposure only after your automated and runtime checks agree.

Start for Free
Experiments

The SQL behind an A/B test: Writing experiment queries in Snowflake

Sep 9, 2026
x
min read

A Snowflake A/B test query is only trustworthy when its rows preserve the experiment's random assignment.

Calculating the average outcome for control and treatment is easy. Building the correct denominator is harder. A plausible result can still include outcomes before exposure, count events instead of randomized users, mix staging with production, drop non-converters, or compare a mature control window with an immature treatment window.

This guide builds the SQL in layers: first exposure, exposure-quality checks, post-exposure outcomes, one value per randomization unit, variation summaries, and operational QA. It also explains which work belongs in Snowflake and which work is safer in a tested statistical engine.

The examples assume user-level randomization and completed-order revenue. Replace database, schema, table, timestamp, environment, and business-status values before running them. Use a development role and bounded dates first.

Define the analytical contract

Assume these tables.

ANALYTICS.EXPERIMENT_EXPOSURES contains:

  • EXPERIMENT_ID VARCHAR
  • USER_ID VARCHAR
  • VARIATION_ID VARCHAR
  • EXPOSED_AT TIMESTAMP_TZ
  • ENVIRONMENT VARCHAR

ANALYTICS.ORDERS contains:

  • ORDER_ID VARCHAR
  • USER_ID VARCHAR
  • ORDER_AT TIMESTAMP_TZ
  • NET_REVENUE NUMBER(18,2)
  • ORDER_STATUS VARCHAR

An exposure means the user had a real opportunity to experience the assigned variation. A background flag refresh or an eligibility lookup is not necessarily exposure. Write this semantic rule beside the schema.

The analysis unit must match assignment. If accounts are randomized, use ACCOUNT_ID and aggregate all user events to one account value. Foreign-key joins do not make user rows statistically independent inside an assigned account.

Use half-open intervals: >= start and < end. They compose without overlap when a scheduled job advances from one analysis window to the next.

Select the first exposure and identify crossovers

This query keeps repeated exposure rows for diagnostics, counts distinct variations per user, selects the earliest qualifying exposure, and excludes users observed in both groups.

WITH raw_exposures AS (
  SELECT
    experiment_id,
    user_id,
    variation_id,
    exposed_at
  FROM YOUR_DATABASE.ANALYTICS.EXPERIMENT_EXPOSURES
  WHERE exposed_at >= '2026-08-01 00:00:00 +00:00'::TIMESTAMP_TZ
    AND exposed_at <  '2026-08-22 00:00:00 +00:00'::TIMESTAMP_TZ
    AND experiment_id = 'checkout-copy-v3'
    AND environment = 'production'
    AND user_id IS NOT NULL
    AND variation_id IN ('control', 'treatment')
),
exposure_quality AS (
  SELECT
    user_id,
    COUNT(DISTINCT variation_id) AS variations_seen,
    COUNT(*) AS exposure_rows
  FROM raw_exposures
  GROUP BY user_id
),
first_exposure AS (
  SELECT
    experiment_id,
    user_id,
    variation_id,
    exposed_at AS first_exposed_at
  FROM raw_exposures
  QUALIFY ROW_NUMBER() OVER (
    PARTITION BY experiment_id, user_id
    ORDER BY exposed_at, variation_id
  ) = 1
),
eligible_exposures AS (
  SELECT f.*
  FROM first_exposure AS f
  JOIN exposure_quality AS q USING (user_id)
  WHERE q.variations_seen = 1
)
SELECT *
FROM eligible_exposures;

Snowflake evaluates QUALIFY after window functions, so the query can filter ROW_NUMBER() without another nested select. The variation key breaks identical-timestamp ties deterministically; identical cross-variation timestamps should still trigger investigation.

Do not discard the crossover measure after filtering. It is an operational signal for unstable identity, non-sticky assignment, delayed configuration, environment overlap, or duplicated pipelines.

Create one post-exposure value per user

Extend the same CTEs with the following unit-value and variation-summary steps. The broad order bounds improve pruning; user-specific predicates enforce the fourteen-day conversion window.

WITH raw_exposures AS (
  SELECT experiment_id, user_id, variation_id, exposed_at
  FROM YOUR_DATABASE.ANALYTICS.EXPERIMENT_EXPOSURES
  WHERE exposed_at >= '2026-08-01 00:00:00 +00:00'::TIMESTAMP_TZ
    AND exposed_at <  '2026-08-22 00:00:00 +00:00'::TIMESTAMP_TZ
    AND experiment_id = 'checkout-copy-v3'
    AND environment = 'production'
    AND user_id IS NOT NULL
    AND variation_id IN ('control', 'treatment')
),
exposure_quality AS (
  SELECT user_id, COUNT(DISTINCT variation_id) AS variations_seen
  FROM raw_exposures
  GROUP BY user_id
),
first_exposure AS (
  SELECT
    experiment_id,
    user_id,
    variation_id,
    exposed_at AS first_exposed_at
  FROM raw_exposures
  QUALIFY ROW_NUMBER() OVER (
    PARTITION BY experiment_id, user_id
    ORDER BY exposed_at, variation_id
  ) = 1
),
eligible_exposures AS (
  SELECT f.*
  FROM first_exposure AS f
  JOIN exposure_quality AS q USING (user_id)
  WHERE q.variations_seen = 1
),
unit_values AS (
  SELECT
    e.variation_id,
    e.user_id,
    COUNT(DISTINCT o.order_id) > 0 AS converted,
    COALESCE(SUM(o.net_revenue), 0) AS revenue
  FROM eligible_exposures AS e
  LEFT JOIN YOUR_DATABASE.ANALYTICS.ORDERS AS o
    ON o.user_id = e.user_id
   AND o.order_at >= '2026-08-01 00:00:00 +00:00'::TIMESTAMP_TZ
   AND o.order_at <  '2026-09-05 00:00:00 +00:00'::TIMESTAMP_TZ
   AND o.order_at >= e.first_exposed_at
   AND o.order_at < DATEADD('day', 14, e.first_exposed_at)
   AND o.order_status = 'completed'
  GROUP BY e.variation_id, e.user_id
)
SELECT
  variation_id,
  COUNT(*) AS units,
  COUNT_IF(converted) AS converted_units,
  COUNT_IF(converted) / NULLIF(COUNT(*), 0) AS conversion_rate,
  AVG(revenue) AS mean_revenue_per_unit,
  VAR_SAMP(revenue) AS sample_variance_revenue,
  SUM(revenue) AS total_revenue
FROM unit_values
GROUP BY variation_id
ORDER BY variation_id;

The LEFT JOIN retains users with zero completed orders. Keep order filters inside the join. A final WHERE o.order_status = 'completed' would remove null matches, turn the analysis into a converter-only comparison, and inflate the metric.

Aggregating to unit_values before the variation summary protects the experimental sample size. Revenue events are not independently randomized; users are. VAR_SAMP returns the dispersion of user-level revenue that a statistical engine needs.

The query uses Snowflake's 0 to express the outcome window relative to each user's first exposure. Keep that per-user rule even when a broad literal predicate is added for pruning.

The summary is not a complete significance test. SQL is well suited to population construction and sufficient statistics. A tested statistical layer should handle confidence intervals or Bayesian posteriors, sequential monitoring, variance reduction, and multiple comparisons. A public discussion about warehouse-native A/B test analysis illustrates both the transparency of this approach and the platform work needed around the SQL.

Put Snowflake metrics to work

Connect governed exposures and outcomes to transparent experiment analysis without rebuilding the statistical workflow for every test.

Start Building Free

Calculate descriptive lift for reconciliation

Use a pivot only after the variation summaries are correct. This helps compare an experimentation UI with analyst-owned SQL.

WITH variation_summary AS (
  -- Replace this comment with the complete query above through unit_values.
  SELECT
    variation_id,
    COUNT(*) AS units,
    COUNT_IF(converted) / NULLIF(COUNT(*), 0) AS conversion_rate,
    AVG(revenue) AS revenue_per_unit
  FROM unit_values
  GROUP BY variation_id
),
pivoted AS (
  SELECT
    MAX(IFF(variation_id = 'control', conversion_rate, NULL)) AS control_cvr,
    MAX(IFF(variation_id = 'treatment', conversion_rate, NULL)) AS treatment_cvr,
    MAX(IFF(variation_id = 'control', revenue_per_unit, NULL)) AS control_rpu,
    MAX(IFF(variation_id = 'treatment', revenue_per_unit, NULL)) AS treatment_rpu
  FROM variation_summary
)
SELECT
  control_cvr,
  treatment_cvr,
  treatment_cvr - control_cvr AS cvr_absolute_change,
  (treatment_cvr - control_cvr) / NULLIF(control_cvr, 0) AS cvr_relative_lift,
  control_rpu,
  treatment_rpu,
  treatment_rpu - control_rpu AS rpu_absolute_change,
  (treatment_rpu - control_rpu) / NULLIF(control_rpu, 0) AS rpu_relative_lift
FROM pivoted;

Return NULL when the control mean is zero instead of manufacturing a relative percentage. Always preserve absolute differences in the original unit: percentage points for conversion and currency per randomized unit for revenue.

Observed lift alone does not answer whether to ship. Define the smallest practically useful effect before launch, then interpret uncertainty and guardrails against that threshold.

Run quality checks before interpreting effects

Sample ratio mismatch

For a nominal 50/50 allocation, calculate the Pearson chi-square statistic from eligible counts. Use a statistics library or experimentation platform for the p-value and alert policy.

WITH counts AS (
  SELECT variation_id, COUNT(*) AS observed
  FROM eligible_exposures
  GROUP BY variation_id
),
totals AS (
  SELECT SUM(observed) AS total_units FROM counts
)
SELECT
  SUM(
    POWER(observed - total_units * 0.5, 2)
    / NULLIF(total_units * 0.5, 0)
  ) AS chi_square_statistic
FROM counts
CROSS JOIN totals;

A failed sample ratio mismatch check means the observed variation counts do not match allocation closely enough for the configured threshold. It does not identify the cause. Check targeting, assignment, exposure emission, warehouse ingestion, filters, joins, and missing IDs.

Crossover rate

SELECT
  COUNT(*) AS exposed_units,
  COUNT_IF(variations_seen > 1) AS crossover_units,
  COUNT_IF(variations_seen > 1) / NULLIF(COUNT(*), 0) AS crossover_rate,
  MAX(exposure_rows) AS max_exposure_rows_for_one_unit
FROM exposure_quality;

Repeated evaluation in one variation can be normal. A unit seen in two variations has ambiguous treatment. Report and investigate it even when the main query excludes it.

Fact-table grain

If the order fact promises one row per order, test the promise.

SELECT order_id, COUNT(*) AS rows_per_order
FROM YOUR_DATABASE.ANALYTICS.ORDERS
WHERE order_at >= '2026-08-01 00:00:00 +00:00'::TIMESTAMP_TZ
  AND order_at <  '2026-09-05 00:00:00 +00:00'::TIMESTAMP_TZ
GROUP BY order_id
HAVING COUNT(*) > 1;

An empty result passes. If the source stores order versions, create a model that selects the current valid row using explicit effective-time logic. Do not add DISTINCT to the experiment query and hide uncertainty about grain.

Pre-exposure outcome leakage

SELECT
  COUNT(*) AS pre_exposure_order_rows,
  COUNT(DISTINCT e.user_id) AS affected_users
FROM eligible_exposures AS e
JOIN YOUR_DATABASE.ANALYTICS.ORDERS AS o
  ON o.user_id = e.user_id
WHERE o.order_at >= '2026-07-18 00:00:00 +00:00'::TIMESTAMP_TZ
  AND o.order_at < e.first_exposed_at;

Prior orders are valid inputs for pre-experiment covariates or eligibility. They are not post-treatment revenue. Separating these windows is essential when applying CUPED.

Handle metric maturity and late-arriving facts

A user exposed yesterday has not completed a fourteen-day outcome window. Either include only mature users or use a cumulative method that compares equal follow-up across variations.

For a mature-cohort analysis, add:

WHERE DATEADD('day', 14, first_exposed_at)
      <= '2026-09-05 00:00:00 +00:00'::TIMESTAMP_TZ

Use an as_of time that reflects source completeness, not merely CURRENT_TIMESTAMP(). Subscription renewals, refunds, offline events, and batch ingestion can update old periods. Publish a metric-lag policy and re-run historical windows when late data is expected.

Time zones need equal care. Store instant timestamps consistently, then derive business dates in an explicit zone. A revenue day based on an account locale may not align with an exposure day in UTC. Implicit session time zones make results difficult to reproduce.

Identity models must be effective-dated. Joining historical exposures to the current anonymous-to-authenticated identity map can rewrite past unit membership. Freeze or reconstruct the mapping as it was known for the analysis contract.

Make Snowflake experiment queries efficient

Snowflake automatically stores table data in micro-partitions and can prune them when predicates align with useful metadata. The micro-partition and clustering documentation explains why bounded time filters and natural clustering matter on large event tables.

Apply these practices:

  • select only necessary columns;
  • use literal or clearly bound time ranges around every large fact;
  • aggregate raw events to reusable unit-level facts;
  • avoid repeatedly scanning the same exposure and identity transformations;
  • use a dedicated, auto-suspending analysis warehouse;
  • size up only when reduced runtime offsets higher credit consumption;
  • schedule broad refreshes away from interactive workloads;
  • set a query tag for attribution.

Set the tag before an analysis session or in the service connection:

ALTER SESSION SET QUERY_TAG =
  '{"application":"experimentation","metric":"net_revenue"}';

Snowflake Query History can filter by user, warehouse, query tag, duration, and query hash. SNOWFLAKE.ACCOUNT_USAGE.QUERY_HISTORY provides longer-lived metadata such as bytes scanned, queue time, errors, warehouse size, and query tag.

Use a dedicated warehouse and attach a resource monitor with notifications and suspension thresholds. Resource monitors cover user-managed warehouses, not every serverless service, so pair them with broader budgets where necessary.

Connect the query model to GrowthBook

SQL alone can produce an audit result. An experimentation program also needs reusable metrics, diagnostics, permissions, statistical methods, result history, and decision workflows.

GrowthBook's warehouse-native architecture queries Snowflake data and exposes generated SQL. Configure:

  1. a dedicated Snowflake user, role, and analysis warehouse;
  2. an experiment-assignment query equivalent to the first-exposure population;
  3. a reusable fact table with unit, timestamp, and value columns;
  4. metric definitions for conversion and revenue;
  5. conversion windows, caps, covariates, guardrails, and statistical settings;
  6. an A/A test and a completed A/B reconciliation.

Preview the generated SQL. Compare eligible units, crossovers, mature units, sums, means, and variances with the reference. If they differ, resolve the data contract before comparing p-values or credible intervals.

GrowthBook can then reuse those governed metrics across experiment analysis and warehouse-native product analytics, reducing drift between dashboards and decisions.

Production checklist

Before a Snowflake result informs a release decision, confirm:

  • exposure represents an opportunity to receive treatment;
  • the randomization unit matches the metric grain;
  • first exposure is deterministic;
  • crossovers are measured and handled consistently;
  • environment and eligibility filters are explicit;
  • primary outcomes occur after exposure;
  • non-converters remain in the denominator;
  • follow-up windows are mature or comparable;
  • joins cannot multiply units;
  • allocation, duplicates, null IDs, and data lag are monitored;
  • every large table has a bounded predicate;
  • query tags, warehouse usage, and credits are visible;
  • statistical inference uses a tested implementation;
  • metric changes are owned, reviewed, and versioned.

Snowflake SQL is the executable expression of an experiment's population and metric rules. Treat it like production code: make assumptions explicit, test the grain, preserve zeroes, bound time, inspect cost, and reconcile against a known result. Then use a shared analysis layer to apply consistent statistics and retain the decision.

Scale beyond Snowflake SQL

Reuse governed warehouse metrics, inspect every generated query, and give teams a consistent path from exposure to decision.

Build with GrowthBook
Experiments
Analytics

A/B testing with Mixpanel data: A practical guide

Sep 8, 2026
x
min read

Mixpanel can hold both sides of an experiment—the exposure and what users did next—but only if identity and timing connect them without selection bias.

The basic workflow is simple. Randomly assign eligible units to control or treatment. Send one exposure event when the experience can first affect them. Track outcomes through the product events already used for funnels and retention. Then analyze those outcomes by variation with a method that matches the experiment plan.

Most implementation failures happen between those sentences. A user changes from anonymous to authenticated identity. Treatment logs only after rendering. A conversion event is renamed mid-test. Analysts filter to users who performed a treatment-dependent step. The dashboard still produces numbers, but the groups no longer represent the randomized comparison.

Choose the analysis topology

There are three practical paths.

Use Mixpanel Experiments

Mixpanel's current Experiments report can analyze experiments run through Mixpanel Feature Flags or detected from exposure events. It supports primary, secondary, and guardrail metrics and multiple statistical model types.

This route fits teams that want experiment review next to product analytics and whose required outcomes are modeled in Mixpanel.

Connect Mixpanel to GrowthBook

The Mixpanel and GrowthBook integration uses GrowthBook for assignment and experiment analysis while Mixpanel remains the analytics data source. An SDK tracking callback sends an experiment-start event into Mixpanel, and analysis uses the resulting data for metrics and dimensions.

This route fits teams that want GrowthBook's feature flag and experimentation workflow while keeping existing Mixpanel instrumentation.

Export or sync Mixpanel data to a warehouse

If primary outcomes combine Mixpanel behavior with billing, CRM, support, or offline facts, move the analysis to governed warehouse models. Mixpanel documents warehouse connectors and export methods for raw events, reports, and pipeline destinations.

This route adds data engineering and freshness responsibilities but gives the experiment access to broader canonical business metrics. GrowthBook's warehouse-native architecture can analyze connected warehouse data.

The choice is not permanent. Start with Mixpanel when it contains the decision metrics; move selected analysis to a warehouse when joins, governance, or scale require it.

Plan the experiment before tracking it

Write the hypothesis, eligible population, randomization unit, variations, primary metric, guardrails, minimum meaningful effect, sample and duration plan, and decision rule.

GrowthBook's A/B test design guide explains how those pieces create one causal question. A funnel report assembled after launch cannot substitute for the plan.

Choose the randomization unit

Randomize users when users can receive treatment independently. Use accounts when members share the changed experience. Use devices only when that is the intended causal unit and cross-device switching is acceptable.

The experimental-unit guide covers why outcomes must be aggregated at the same independent level. Thousands of events from one user do not become thousands of statistical observations.

Define metrics before exposure

Use a practical KPI framework to choose one primary outcome and the guardrails that protect the customer experience.

Read the KPI Playbook

Instrument one symmetric exposure event

Send exposure when the assigned variation can first affect behavior. The event should be identical in name and schema across arms.

const growthbook = new GrowthBook({
  attributes: { id: stableExperimentId },
  trackingCallback: (experiment, result) => {
    mixpanel.track("Experiment Viewed", {
      experiment_id: experiment.key,
      variation_id: result.key,
      assignment_id: stableExperimentId,
      environment: "production",
      assignment_revision: currentFeatureRevision,
    });
  },
});

The exact SDK setup varies, but the contract should remain stable. Use placeholders rather than secrets, and never send sensitive traits merely because they might be useful later.

Avoid overcounting evaluations

A component may evaluate a flag on every render. Deduplicate the exposure logically by experiment, phase, and randomization unit. Repeated raw events can remain available for debugging, but enrollment should count each unit once.

Do not log too late

If treatment logs after an asynchronous bundle loads while control logs immediately, slow or failed treatment sessions disappear. Put the event before variation-specific failure can select the sample.

GrowthBook's tracking callback documentation describes the application hook. Test its behavior in development, then verify one real event per intended unit in Mixpanel's event inspection workflow.

Align Mixpanel identity with assignment

Mixpanel's Simplified ID Merge documentation describes $device_id, $user_id, identity clusters, identify(), and reset(). That behavior matters directly to experiment analysis.

Use a stable assignment attribute and answer these questions before launch:

  1. What ID exists for anonymous visitors?
  2. Does login link that ID to the authenticated user?
  3. Can assignment change at login or across devices?
  4. Does logout call reset() on a shared device?
  5. Which canonical ID is used in analysis and exports?
  6. Is the experiment randomized by user while product behavior spreads across an account?

Run scripted journeys: anonymous exposure then signup, returning login on a new device, logout then a second user, and cross-platform use. Confirm each journey produces the intended identity cluster and one experiment assignment.

Define outcomes as metric contracts

For every metric, document event name, filters, unit, counting rule, attribution window, missing behavior, and event-schema version.

A binary 7-day activation metric might mean: among exposed users with a complete 7-day window, did at least one Activated Project event occur after exposure and before day 7? A revenue metric must specify currency, refunds, multiple purchases, outlier treatment, and whether revenue is summed per user before comparison.

Use saved metrics or a governed semantic layer where possible. GrowthBook's metric documentation covers conversion, count, duration, revenue, ratio, and guardrail definitions across analysis sources.

Keep exploration separate from the primary decision

Mixpanel funnels and breakdowns are useful for understanding mechanism: where users drop off, which platform saw errors, and which steps changed. Treat unplanned slices as exploratory. They generate hypotheses for follow-up tests rather than automatic evidence for shipping.

Community discussion about A/B testing and Mixpanel instrumentation repeatedly returns to concurrent groups and a metric chosen in advance. That principle matters more than the report UI.

Validate allocation and event quality

Before reading lift, compare observed variation counts with the planned split. GrowthBook's sample ratio mismatch documentation explains why an unlikely allocation can indicate a routing, exposure, or filtering problem.

Also check:

  • units exposed to multiple variations;
  • exposure properties missing by arm;
  • time from assignment to exposure;
  • outcome events dated before exposure;
  • platform and app-version balance;
  • identity merges and duplicate profiles;
  • event volume and conversion-rate discontinuities;
  • pre-experiment outcomes and invariant attributes.

Run an A/A test when the assignment-to-Mixpanel-to-analysis path is new. Identical experiences should produce centered effect estimates over repeated checks, while still allowing ordinary sampling variation in a single run.

Mixpanel's guidance for third-party integrations recommends a sandbox, source identification, schema synchronization, and event QA. Apply the same discipline to your internal experiment integration.

Handle time, maturity, and late events

Project time zone, event time, analysis time, and API export dates must be understood together. Mixpanel's export documentation notes that date interpretation can depend on project creation date and time-zone configuration.

For a 7-day metric, exclude units that have not had 7 days to convert or mark results preliminary. Define how late mobile events, offline sessions, and backfills change historical results. Record the data cutoff with the decision.

Avoid before-after testing. Both arms should run concurrently so seasonality, campaigns, outages, and product changes affect them together.

Compare direct and warehouse results before migrating

When moving analysis from Mixpanel to a warehouse, run both paths on completed experiments. Differences often come from:

  • canonical identity after merges;
  • time-zone boundaries;
  • event deduplication;
  • bot or internal-user filters;
  • attribution windows;
  • missing values;
  • revenue refunds and currency;
  • metric maturity;
  • unit-level aggregation.

Use Mixpanel's raw event export options or a supported pipeline rather than a UI CSV for production-scale reconciliation. Store transformation versions and automated data tests.

Do not cut over until material differences are explained. “Both dashboards are close” is not a metric contract.

Read results and close the loop

Evaluate effect size and uncertainty against the minimum useful improvement. Review guardrails, sample health, experiment duration, planned segments, and external events. Use the statistical method you declared; changing models or thresholds after seeing results increases false discovery risk.

Document the hypothesis, unit, identity behavior, event and property schema, metric versions, dates, analysis settings, cutoff, and decision. If assignment or exposure is biased, repair it and restart rather than rescuing the result with filters.

When the winner is rolled out, monitor it, remove the losing code path, and archive the experiment flag. Product analytics can then track long-term behavior without keeping temporary experiment machinery alive.

Mixpanel data becomes trustworthy experiment evidence when it retains the randomization contract: stable identity, symmetric exposure, outcomes after exposure, one independent row per unit, and a decision plan that exists before the result.

Reconcile Mixpanel with the assignment system

For each experiment, compare the flag service's assigned population with Mixpanel's first exposure population. Break discrepancies down by platform, app version, anonymous versus authenticated state, consent status, and time. A missing exposure is not random merely because overall event volume looks healthy.

Inspect sample units from both sides. Confirm that the variation property is stable, exposure precedes outcomes, and identity merges do not move a user between arms. If Mixpanel and the flag provider use different identifiers, define an effective-dated mapping instead of joining through today's profile state.

Keep an explicit control population. A user with no conversion event must remain in the denominator after exposure. Building the analysis from outcome events and then attaching variations selects only converters and cannot estimate a conversion rate.

Choose direct or warehouse analysis by metric ownership

Direct Mixpanel analysis is convenient when the required events, properties, identity behavior, and metric semantics already live there. Product teams can explore funnels and segments without waiting for another pipeline. The cost is tighter dependence on the event taxonomy and platform calculation rules.

Warehouse analysis is stronger when decisions rely on revenue adjustments, subscriptions, account hierarchies, support outcomes, or other facts governed outside Mixpanel. It also gives analysts more control over identity, attribution, late data, and unit-level aggregation. The cost is operating the export, models, compute, and statistical workflow.

A hybrid can work: use Mixpanel for exploratory product behavior and a warehouse-native platform for the declared primary and guardrail metrics. Label exploratory cuts honestly and reconcile shared metrics on completed experiments so teams understand why two interfaces may differ.

Test failure and late-data behavior

Delay an exposure event in a test project, send a duplicate, alias an anonymous user after signup, and change a property type. Observe ingestion, identity merge, deduplication, saved reports, exports, and experiment results. Document which corrections update history and on what schedule.

If data is exported to a warehouse, publish source and destination watermarks. A current Mixpanel dashboard and a delayed warehouse table should not be presented as two views of the same cutoff. Preserve transformation versions and the export job that produced the analytical fact.

Finally, rehearse cleanup. After rollout, stop temporary exposure instrumentation only when the permanent path and long-term product analytics remain intact. Archive the experiment context, decision, and metric versions so a later team can distinguish a past test from an active flag.

Use stable naming from the start. Give the experiment and variation properties machine-readable keys that do not change when a dashboard label is edited. Keep development and production values distinct, and publish accepted event and property types. A string-to-number change can fragment saved reports and downstream exports without an obvious error.

Assign an owner to every event used in a decision. The owner is responsible for trigger semantics, identity, freshness, and deprecation. This lightweight contract prevents an exploratory tracking event from becoming a permanent primary metric merely because it is convenient to query.

Review that contract when the application, SDK, consent flow, or identity logic changes; an unchanged event name does not guarantee unchanged measurement.

Add rigorous tests to Mixpanel

Use your existing product events with feature flags and experiment analysis built for transparent decisions.

Start Building Free

Ready to ship faster?

No credit card required. Start with feature flags, experimentation, and product analytics—free.

Simplified white illustration of a right angle ruler or carpenter's square tool.White checkmark symbol with a scattered pixelated effect around its edges on a transparent background.