eCommerce Optimisation

Conversion rate optimisation built on evidence, not opinions

Most of what gets sold as conversion rate optimisation is a list of best practices applied without measurement. A real program starts by working out whether you have enough traffic to learn anything at all.

What is conversion rate optimisation?

Conversion Rate Optimisation is a structured program of research, hypotheses and controlled experiments that increases the share of visitors completing a valuable action. It suits Australian businesses with steady traffic and a measurable goal, where the objective is compounding verified gains rather than a redesign based on whoever spoke loudest in the meeting.

Get a fixed written quote
Typical timeline
First results in 6 to 10 weeks, then ongoing
What drives cost
How much research each cycle needs, how many templates are being tested, and whether your platform lets the checkout be changed.
Best for
Sites with enough conversions to test honestly and a goal worth money
You own
The research, the full test archive and every tool account
Built with
Experimentation tools, session replay, form analytics, server-side measurement
Where it worksSite visitorsEngaged sessionsStarted checkoutCompleted orders
Each stage is measured, then tested one hypothesis at a time against a control.

Your handover

Where good hypotheses come from

A hypothesis is only as good as the evidence behind it, and the evidence rarely lives in one tool. Analytics tells you where people leave but never why. Session replay shows the why but only for the handful of sessions someone actually watches. Talking to five customers reveals confusions that no dashboard would ever surface.

  1. 01Sample size and testing feasibility assessment
  2. 02Research findings from analytics, replay and user sessions
  3. 03Prioritised hypothesis backlog with evidence attached
  4. 04Experiment specifications with metrics and guardrails
  5. 05Built and quality assured variants
  6. 06Statistically sound result analysis for every test
  • A searchable archive of what worked and what did not
  • Monthly reporting tied to revenue or qualified leads
  • Handover documentation so your team can continue the program
Second hand summaries lose the detail that makes a test worth running

So we run a research stack rather than a single method, and the person who writes the hypothesis is the person who watched the sessions. Second hand summaries lose the detail that makes a test worth running. The output is a backlog of specific, falsifiable statements with the evidence attached, prioritised by expected impact, confidence and effort, so the conversation with your team is about sequencing rather than about who has the better instinct.

  • Analytics segmented by device, source and new versus returning, because averages hide the problem
  • Session replay watched directly by whoever will write the hypothesis
  • Form analytics showing which field people abandon and where they hesitate
  • On page polls asking what nearly stopped someone, answered in their own words
  • Moderated usability sessions with people who resemble your buyers
  • Support tickets and sales call notes, the cheapest research nobody reads

A test that wins on the primary metric and damages a guardrail has not won.

Tests that mislead, and how we avoid them

Experimentation has failure modes that look exactly like success. Stopping a test the moment it crosses significance produces winners that vanish on repetition, because a metric wandering above and below a threshold will eventually cross it by chance. Running twelve variants and reporting the best one is the same problem wearing a different hat.

A test that wins on the primary metric and damages a guardrail has not won

Then there are the effects that are real but temporary. Regular visitors notice a change and behave differently for a fortnight, which is novelty rather than improvement. And there is the trap of optimising a number that does not pay: a shorter form lifts submissions and fills the sales team's day with unqualified leads, or a discount banner raises conversion rate while reducing revenue per visitor. That is why guardrail metrics are agreed before launch, typically revenue per visitor, average order value, refund rate, lead quality and support contacts. A test that wins on the primary metric and damages a guardrail has not won.

How the engagement runs

How the program runs each month

Experimentation works as a cadence rather than a project. Research feeds a backlog, the backlog feeds a build queue, results feed back into research, and the archive of what you learned becomes an asset that outlives any individual test.

  1. 01Collect evidenceAnalytics, replay, form drop off, support tickets and whatever sales heard this month
  2. 02Stage 2Prioritise the backlog by expected impact, strength of evidence and build effort
  3. 03Stage 3Write the hypothesis as a falsifiable statement with a primary metric and guardrails agreed up front
  4. 04Stage 4Build the variant properly, including mobile, accessibility and the states nobody remembers
  5. 05Quality assure the test itselfTraffic allocation, tracking, flicker and behaviour for returning visitors
  6. 06Stage 6Run whole week cycles until the pre agreed sample is reached, without peeking at partial results
  7. 07Stage 7Report honestly, ship the winners, document the losers and feed both back into the backlog
DiscoverDesignBuildTestHandover
Two decisions on your side that keep the project moving

The discipline that matters most is agreeing the primary metric, the guardrails and the required sample before the test goes live. Deciding what counts as success after seeing the data is how organisations convince themselves of things that are not true. We also insist on full week cycles, because a Tuesday audience and a Sunday audience behave differently and a test stopped mid week is measuring the day as much as the change.

Choose the right level

The maths that decides whether you can A/B test at all

An experiment needs enough events to distinguish a real effect from noise. The smaller the improvement you hope to detect, the more conversions you need, and the relationship is unkind: halving the effect size roughly quadruples the sample required. A site with forty enquiries a month cannot detect a ten percent lift, and any tool that declares a winner in that situation is showing you randomness with a confident interface.

Monthly conversions on the page in question

01

Under 100

What a test can realistically detect

Almost nothing short of a doubling

What we do instead

Usability testing, defect fixing and offer work, measured before and after

02

100 to 400

What a test can realistically detect

Large effects, roughly a quarter improvement or more

What we do instead

Bold structural tests, one at a time, run over whole week cycles

03

400 to 1000

What a test can realistically detect

Moderate effects across a few tests each month

What we do instead

A steady program with a prioritised backlog and clear guardrails

04

Over 1000

What a test can realistically detect

Smaller refinements as well as structural change

What we do instead

Concurrent tests, segment analysis and personalisation where it pays

How we work this out during scoping

This is the first thing we calculate, before quoting any conversion rate optimisation program, because it determines whether you should be paying for experimentation at all. Plenty of Australian businesses sit below the threshold, and for them the right work is qualitative research, fixing defects that are obviously defects, and improving the offer itself. That work is cheaper and often produces larger gains. We would rather say so than sell a testing retainer that generates inconclusive results for a year.

Where the gains actually come from

Almost never from button colour. In the programs we run, the changes that move numbers are about clarity and friction. Making it obvious what is being sold, to whom and at what price. Showing delivery cost and timeframe before the checkout rather than at the last step. Cutting a form to the fields you will genuinely use. Explaining what happens after someone submits, since the fear of a pushy sales call stops more enquiries than any layout decision.

Some of the biggest wins are not conversion work at all in a strict sense

Some of the biggest wins are not conversion work at all in a strict sense. Page speed changes behaviour on mobile connections, and fixing it is engineering. Weak copy is a copywriting problem. A confusing journey across several steps is user experience design. We are happy for the answer to sit outside a testing tool, because the objective is more revenue from the traffic you already have, not a full backlog of experiments. Enrolment funnels in education are a good example, where the winning change is usually removing a step rather than restyling one.

When conversion rate optimisation is the wrong spend

Three situations come up repeatedly. The first is not enough traffic, which we covered above. If your site sees a few hundred visits a month, the money belongs in demand generation, whether that is organic search or paid search, and conversion work becomes worthwhile once the volume exists.

The second is untrustworthy measurement

The second is untrustworthy measurement. If your analytics double counts, your goals fire on page load or your eCommerce data disagrees with your accounting system, every test result is fiction. Fixing that through proper analytics implementation is a prerequisite, not an upsell, and we will refuse to run experiments on broken tracking. The third is a site that is fundamentally not fit for purpose. Optimising individual pages on a structure that confuses everyone is polishing something that needs replacing, and the honest recommendation is a redesign with research done properly at the front. Testing after that rebuild will then be worth every dollar, because you will be refining something sound rather than patching something broken.

How we scope it

Four ways to scope your Conversion Rate Optimisation project

We do not publish package prices, because the same brief can be a short build or a long one. These are the shapes the work usually takes. Tell us which one sounds like you and you will get a fixed written quote that spells out exactly what it covers.

Store launch

A first real store, set up for Australian selling

Fixed written quote, agreed before work starts

  • Sample size and testing feasibility assessment
  • Research findings from analytics, replay and user sessions
  • Prioritised hypothesis backlog with evidence attached
Request a quote
Most common

Store growth

A store with catalogue, integration or margin problems to solve

Fixed written quote, agreed before work starts

  • Everything in Store launch
  • Experiment specifications with metrics and guardrails
  • Built and quality assured variants
  • Statistically sound result analysis for every test
Request a quote

Commerce platform

Multi store, multi channel or a custom commerce build

Fixed written quote, agreed before work starts

  • Everything in Store growth
  • A searchable archive of what worked and what did not
  • Monthly reporting tied to revenue or qualified leads
  • Handover documentation so your team can continue the program
Request a quote

Store care

Merchandising, speed and conversion, month to month

Rolling monthly, quoted in writing

  • Hosting, patching, backups and uptime monitoring
  • Merchandising, campaign and catalogue changes
  • Checkout and payment paths tested after every update
  • Rolling, cancel with 30 days notice
Request a quote

These are shapes, not menus. Most quotes end up somewhere between two of them, and we will say so when the honest answer is the smallest one. Describe the problem and we will tell you which it is.

Questions buyers usually ask

Frequently asked questions

Ownership and handover

Who owns the test data and the tool accounts?

You do. Experimentation, replay and analytics accounts are opened under your organisation with your billing, and we work with permissions you can revoke. The research documents, hypothesis backlog and full result archive are yours to keep. That archive is genuinely valuable, because knowing what has already failed prevents the next agency repeating it.

Do we need to give you access to change our website?

For most tests we work through the experimentation tool, which needs a script on the site and no-code access. For anything structural, or where flicker would spoil the result, we prefer to build the variant in your codebase and use feature flags. We will tell you which approach a given test needs and why, since the tool based route has real limits on complex pages.

Detail and edge cases

How soon will we see a result worth reporting?

The first test usually launches in week three, after research and instrumentation, and needs two to four weeks to reach a valid sample. So the first defensible result lands around week six to ten. Anyone promising a lift in the first fortnight is either not running a valid test or is counting a change they made without measuring the alternative.

What determines the cost of a CRO program?

Mostly how much research and build work each cycle requires. Testing on a simple landing page is light. Running experiments across a large store, with variants that need real front end engineering and server-side rendering, is heavier. Tool subscriptions are separate and paid by you directly so you keep the data. We quote the program scope in writing before starting.

What happens when a test loses?

Roughly half of well designed tests do not win, and that is a functioning program rather than a failing one. A losing test removes a wrong assumption for the price of one build, which is far cheaper than shipping the same idea site wide and living with it. We report losses with the same detail as wins, since the reasoning behind them usually points to the next hypothesis.

Can you do this on a lead generation site rather than a store?

Yes, with one adjustment. Form submissions are easy to count and easy to inflate, so we measure qualified leads rather than raw enquiries, which means connecting outcomes back from your CRM. Otherwise you optimise your way to more junk. For professional services and trades that connection is the difference between a program that helps and one that annoys the sales team.

Find out whether your traffic can support testing

Send us your monthly conversions and the page that matters most. We reply within one business day with a feasibility read and, if it makes sense, a fixed written quote.