For teams whose test backlog is bigger than their build capacity

A/B Test Development Services

A/B test development is the work of turning a hypothesis and a design into a working experiment: the HTML, CSS and JavaScript that render the variation, the targeting that decides who sees it, and the tracking that makes the result trustworthy. We build that layer for CRO agencies, consultants and in-house experimentation teams, on whichever platform you already run, at the velocity your roadmap was sold at.

CRO strategy, design and development across Shopify, WordPress, Convert, Optimizely, VWO and Adobe Target.

What it solves

Does this sound familiar?

These are the situations teams actually describe when they get in touch about this, not a generic list of pain points.

The backlog is full of designs nobody can build

Strategy and design capacity almost always outruns development capacity. Hypotheses sit in a spreadsheet for weeks because the only person who can code a variation is also shipping the product roadmap.

Variations break outside the developer’s browser

A variation built against one viewport and one theme state falls apart on iOS Safari, on a logged-in session, on a product with five variants, or the second the DOM re-renders. A broken variation does not just lose the test. It costs real revenue while it runs.

Flicker quietly invalidates the result

When the control paints before the variation is applied, part of your treatment group has seen both experiences. The sample is contaminated and the reported lift no longer means what the report says it means.

The numbers do not reconcile

The platform reports a winner and analytics disagrees. Usually the goal was bound to the wrong event, or the event fires on one arm and not the other, and nobody checked before launch.

Test velocity is set by whoever is free

Some months four experiments ship, some months none, and the difference has nothing to do with the quality of the ideas. A programme that cannot hold a cadence cannot compound.

Outcomes

What changes when this works

Stated as the position you end up in. No percentages: we do not publish outcome figures without the measurement conditions and the client’s agreement to it.

  • A predictable number of experiments live per month, decided by the roadmap rather than by who is available
  • Results you can act on, because the comparison was clean and the measurement was verified before launch
  • Variations that survive real traffic: re-renders, logged-in sessions, iOS Safari, mid-cart updates
  • No revenue lost to a broken variation running unnoticed for three weeks
  • Winners rolled into production properly, with the experiment layer removed behind them
  • Your strategists back on strategy instead of writing build tickets and re-checking QA
Deliverables

What is included

Scope stated as specific outputs rather than as adjectives.

  • Variation code in hand-written HTML, CSS and JavaScript, with no WYSIWYG editor output and no page-builder runtime
  • Responsive implementation verified across mobile, tablet and desktop breakpoints, not scaled down from one design
  • Support for dynamic sites and single-page applications, using each platform’s own re-application hooks rather than timers
  • Variation isolation, so a change scoped to one template cannot leak into another
  • Flicker prevention appropriate to the platform and its loading mode
  • Audience and page targeting built and confirmed against the real URL and cookie conditions
  • Goal, event and analytics validation on every arm of the experiment before launch
  • Pre-launch QA evidence and a post-launch check on live traffic
  • Commented, readable code handed to you, plus a hardcode-ready version of a winner
  • Experiment removal when the run ends, so nothing from a finished test is still executing
Process

How the work runs

  1. 1

    Scope the experiment

    We read the hypothesis and the design against the live site and come back with what is straightforward, what is technically risky, and what needs a decision before code starts. Ambiguity found here is far cheaper than ambiguity found in QA.

  2. 2

    Build the variation

    Code is written by hand against your real DOM, with selectors chosen for stability rather than convenience, and structured so the change reverts cleanly when the experiment ends.

  3. 3

    Wire targeting and measurement

    Audience conditions, URL targeting and goals are configured in the platform and then verified from the visitor’s side. Does the right person actually qualify, and does the event actually fire?

  4. 4

    QA and hand over

    The build goes through the full pre-launch pass described in our CRO and experiment QA service, and you receive the code, the QA evidence and anything we think the report should caveat.

  5. 5

    Watch it run, then close it out

    A re-check after any site deploy during the run, because that is when variations silently stop applying. At the end, the winner is implemented permanently or the code is handed over, and the experiment layer is deleted either way.

Handover

What you receive

The artefacts that exist at the end, all of them yours to keep and to hand to somebody else.

  • Variation source code, commented, in the account you own
  • The platform configuration: audiences, URL targeting, goals and allocation, documented
  • A completed pre-launch QA pass with an explicit result per line
  • A post-launch validation note once the experiment is on live traffic
  • A hardcode-ready version of any winner, ready for your own developers if you prefer
  • Anything we think the read-out should be caveated with, in writing, before you present it
Fit

Who this is for

CRO agencies

You own strategy, research and design for this engagement. You get a development bench behind it, under your own brand if you want, via white-label CRO development.

Independent CRO consultants

You sell a testing programme but do not want to write JavaScript for it, or turn down work when two clients need a build in the same week.

In-house experimentation teams

Your programme is limited by engineering priority, not by ideas. We take the variation build off the product team’s sprint.

Ecommerce teams running tests

You have traffic and a testing tool, and you need experiments built on a real storefront without breaking the theme.

Who this is not for

Saying this up front saves both of us a call. If you are on this list, the right next step is usually elsewhere on this site rather than nowhere.

  • Teams who do not yet know what to test. Start with a CRO audit or full-service CRO — building the wrong experiment quickly is not progress
  • Sites without enough traffic for a split test to conclude. We will tell you before you buy, and point you at the evidence-led route instead
  • Anyone wanting variations built in a visual editor to save money. Editor output binds to whatever markup exists today, and that is precisely the failure mode this service exists to avoid
  • Work where nobody will grant access to the analytics. If we cannot verify the goals, we cannot vouch for the result
Worked examples

An engagement of this kind

The problem, the hypothesis and what was actually engineered. We do not publish outcome percentages until the measurement conditions can be published alongside them and the client has agreed.

Supplement brand · Shopify

Removing product-page friction and eliminating test flicker

Problem:
High product-page bounce on paid traffic, and a previous agency’s WYSIWYG experiments flickered for most of a second, so the results were not trustworthy either.
Hypothesis:
A sticky add-to-cart with above-fold trust signals, delivered as native DOM scripts, will lift add-to-cart without contaminating the sample.
Execution:
Rebuilt the test layer in lightweight vanilla JS on Convert Experiences and Shopify, with every GA4 event validated on both arms before launch.
Comparison

Why us rather than the alternatives

Every one of these is a legitimate way to get the work done. Here is where each of them tends to break, so you can decide honestly.

Instead of a general front-end contractor

A capable front-end developer will build the variation. They are not hired to care whether the goal fired on the control, whether the change leaked into another template, or whether the sample is contaminated. Those are the things that decide whether the result is worth acting on.

Instead of your own product engineers

Your engineers can obviously do this. The question is whether an experiment should outrank the roadmap for a week, every week. Usually it should not, which is why the backlog exists.

Instead of the platform’s visual editor

The editor is genuinely faster for a copy change on a stable page. It is also how experiments silently stop applying after a content edit, and it produces code nobody can hand to a developer at the end.

Instead of hiring in-house

Worth it when your volume is steady enough to keep someone busy. Before that, a salary is a fixed cost against a workload that swings between two experiments and eleven.

Capability

Platforms and technologies

Named specifically, because “modern tooling” tells a buyer nothing.

  • Convert Experiences
  • Optimizely Web Experimentation
  • VWO
  • Adobe Target (at.js 2.x)
  • Vanilla JavaScript, HTML, CSS
  • Google Tag Manager
  • GA4
  • Shopify / Liquid
  • WordPress
  • React and other client-rendered front ends
Quality assurance

How this gets checked

Every engagement carries a QA pass. It is written down so that “QA done” means something specific.

  • Functional pass on every interactive element the variation touches
  • Visual pass against the design at each breakpoint
  • Cross-browser and cross-device matrix, including iOS Safari
  • Targeting validation: qualifying and non-qualifying sessions both checked
  • Goal and analytics validation on control and every variation
  • Console clean of errors introduced by the experiment
  • Flicker and layout-shift check on a throttled connection
  • Isolation check: pages outside the target set are unchanged

The full pass, including the post-launch checks, is described on our CRO and experiment QA service. It is also available on its own, on experiments somebody else built.

Engagement

Ways to work together

The commercial shape is chosen after we know the scope, not before.

Dedicated partnership

Reserved capacity for a steady pipeline of work.

Flexible hours

Draw down specialist time as the work arrives.

Project engagement

A defined scope, owned end to end.

All three are described in full under engagement models.

FAQ

Questions people ask before buying this

Clear My Test Backlog

Send us the hypothesis, the design or just the URL. We will tell you what it takes to build, what the risks are, and when it can ship.

How would you like to start?

No spam. No obligation. We reply within one business day.
We use the details you submit only to respond to this enquiry. They are stored in our own database and email is sent via Resend. Email backofficeomtechservice@gmail.com to access or delete your data.