A/B Test Development Services
A/B test development is the work of turning a hypothesis and a design into a working experiment: the HTML, CSS and JavaScript that render the variation, the targeting that decides who sees it, and the tracking that makes the result trustworthy. We build that layer for CRO agencies, consultants and in-house experimentation teams, on whichever platform you already run, at the velocity your roadmap was sold at.
CRO strategy, design and development across Shopify, WordPress, Convert, Optimizely, VWO and Adobe Target.
Does this sound familiar?
These are the situations teams actually describe when they get in touch about this, not a generic list of pain points.
The backlog is full of designs nobody can build
Strategy and design capacity almost always outruns development capacity. Hypotheses sit in a spreadsheet for weeks because the only person who can code a variation is also shipping the product roadmap.
Variations break outside the developer’s browser
A variation built against one viewport and one theme state falls apart on iOS Safari, on a logged-in session, on a product with five variants, or the second the DOM re-renders. A broken variation does not just lose the test. It costs real revenue while it runs.
Flicker quietly invalidates the result
When the control paints before the variation is applied, part of your treatment group has seen both experiences. The sample is contaminated and the reported lift no longer means what the report says it means.
The numbers do not reconcile
The platform reports a winner and analytics disagrees. Usually the goal was bound to the wrong event, or the event fires on one arm and not the other, and nobody checked before launch.
Test velocity is set by whoever is free
Some months four experiments ship, some months none, and the difference has nothing to do with the quality of the ideas. A programme that cannot hold a cadence cannot compound.
What changes when this works
Stated as the position you end up in. No percentages: we do not publish outcome figures without the measurement conditions and the client’s agreement to it.
- A predictable number of experiments live per month, decided by the roadmap rather than by who is available
- Results you can act on, because the comparison was clean and the measurement was verified before launch
- Variations that survive real traffic: re-renders, logged-in sessions, iOS Safari, mid-cart updates
- No revenue lost to a broken variation running unnoticed for three weeks
- Winners rolled into production properly, with the experiment layer removed behind them
- Your strategists back on strategy instead of writing build tickets and re-checking QA
What is included
Scope stated as specific outputs rather than as adjectives.
- Variation code in hand-written HTML, CSS and JavaScript, with no WYSIWYG editor output and no page-builder runtime
- Responsive implementation verified across mobile, tablet and desktop breakpoints, not scaled down from one design
- Support for dynamic sites and single-page applications, using each platform’s own re-application hooks rather than timers
- Variation isolation, so a change scoped to one template cannot leak into another
- Flicker prevention appropriate to the platform and its loading mode
- Audience and page targeting built and confirmed against the real URL and cookie conditions
- Goal, event and analytics validation on every arm of the experiment before launch
- Pre-launch QA evidence and a post-launch check on live traffic
- Commented, readable code handed to you, plus a hardcode-ready version of a winner
- Experiment removal when the run ends, so nothing from a finished test is still executing
How the work runs
- 1
Scope the experiment
We read the hypothesis and the design against the live site and come back with what is straightforward, what is technically risky, and what needs a decision before code starts. Ambiguity found here is far cheaper than ambiguity found in QA.
- 2
Build the variation
Code is written by hand against your real DOM, with selectors chosen for stability rather than convenience, and structured so the change reverts cleanly when the experiment ends.
- 3
Wire targeting and measurement
Audience conditions, URL targeting and goals are configured in the platform and then verified from the visitor’s side. Does the right person actually qualify, and does the event actually fire?
- 4
QA and hand over
The build goes through the full pre-launch pass described in our CRO and experiment QA service, and you receive the code, the QA evidence and anything we think the report should caveat.
- 5
Watch it run, then close it out
A re-check after any site deploy during the run, because that is when variations silently stop applying. At the end, the winner is implemented permanently or the code is handed over, and the experiment layer is deleted either way.
What you receive
The artefacts that exist at the end, all of them yours to keep and to hand to somebody else.
- Variation source code, commented, in the account you own
- The platform configuration: audiences, URL targeting, goals and allocation, documented
- A completed pre-launch QA pass with an explicit result per line
- A post-launch validation note once the experiment is on live traffic
- A hardcode-ready version of any winner, ready for your own developers if you prefer
- Anything we think the read-out should be caveated with, in writing, before you present it
Who this is for
CRO agencies
You own strategy, research and design for this engagement. You get a development bench behind it, under your own brand if you want, via white-label CRO development.
Independent CRO consultants
You sell a testing programme but do not want to write JavaScript for it, or turn down work when two clients need a build in the same week.
In-house experimentation teams
Your programme is limited by engineering priority, not by ideas. We take the variation build off the product team’s sprint.
Ecommerce teams running tests
You have traffic and a testing tool, and you need experiments built on a real storefront without breaking the theme.
Who this is not for
Saying this up front saves both of us a call. If you are on this list, the right next step is usually elsewhere on this site rather than nowhere.
- Teams who do not yet know what to test. Start with a CRO audit or full-service CRO — building the wrong experiment quickly is not progress
- Sites without enough traffic for a split test to conclude. We will tell you before you buy, and point you at the evidence-led route instead
- Anyone wanting variations built in a visual editor to save money. Editor output binds to whatever markup exists today, and that is precisely the failure mode this service exists to avoid
- Work where nobody will grant access to the analytics. If we cannot verify the goals, we cannot vouch for the result
An engagement of this kind
The problem, the hypothesis and what was actually engineered. We do not publish outcome percentages until the measurement conditions can be published alongside them and the client has agreed.
Removing product-page friction and eliminating test flicker
- Problem:
- High product-page bounce on paid traffic, and a previous agency’s WYSIWYG experiments flickered for most of a second, so the results were not trustworthy either.
- Hypothesis:
- A sticky add-to-cart with above-fold trust signals, delivered as native DOM scripts, will lift add-to-cart without contaminating the sample.
- Execution:
- Rebuilt the test layer in lightweight vanilla JS on Convert Experiences and Shopify, with every GA4 event validated on both arms before launch.
Why us rather than the alternatives
Every one of these is a legitimate way to get the work done. Here is where each of them tends to break, so you can decide honestly.
Instead of a general front-end contractor
A capable front-end developer will build the variation. They are not hired to care whether the goal fired on the control, whether the change leaked into another template, or whether the sample is contaminated. Those are the things that decide whether the result is worth acting on.
Instead of your own product engineers
Your engineers can obviously do this. The question is whether an experiment should outrank the roadmap for a week, every week. Usually it should not, which is why the backlog exists.
Instead of the platform’s visual editor
The editor is genuinely faster for a copy change on a stable page. It is also how experiments silently stop applying after a content edit, and it produces code nobody can hand to a developer at the end.
Instead of hiring in-house
Worth it when your volume is steady enough to keep someone busy. Before that, a salary is a fixed cost against a workload that swings between two experiments and eleven.
Platforms and technologies
Named specifically, because “modern tooling” tells a buyer nothing.
- Convert Experiences
- Optimizely Web Experimentation
- VWO
- Adobe Target (at.js 2.x)
- Vanilla JavaScript, HTML, CSS
- Google Tag Manager
- GA4
- Shopify / Liquid
- WordPress
- React and other client-rendered front ends
How this gets checked
Every engagement carries a QA pass. It is written down so that “QA done” means something specific.
- Functional pass on every interactive element the variation touches
- Visual pass against the design at each breakpoint
- Cross-browser and cross-device matrix, including iOS Safari
- Targeting validation: qualifying and non-qualifying sessions both checked
- Goal and analytics validation on control and every variation
- Console clean of errors introduced by the experiment
- Flicker and layout-shift check on a throttled connection
- Isolation check: pages outside the target set are unchanged
The full pass, including the post-launch checks, is described on our CRO and experiment QA service. It is also available on its own, on experiments somebody else built.
Ways to work together
The commercial shape is chosen after we know the scope, not before.
Dedicated partnership
Reserved capacity for a steady pipeline of work.
Flexible hours
Draw down specialist time as the work arrives.
Project engagement
A defined scope, owned end to end.
All three are described in full under engagement models.
Questions people ask before buying this
Related services
Experimentation platforms we build on
Further reading
Clear My Test Backlog
Send us the hypothesis, the design or just the URL. We will tell you what it takes to build, what the risks are, and when it can ship.