Publish the new theme to 10% of visitors, watch it for an hour, then decide. That is the honest use of Shopify rollouts for almost every Australian store, and it has nothing to do with A/B testing. You are not measuring a lift. You are checking that the Afterpay widget still renders on product pages, that the cart drawer opens on an iPhone, and that nobody has quietly broken GST-inclusive price display two weeks before Click Frenzy.
Shopify redesigned the Rollouts setup flow on 22 September, adding finer control over timing and over what share of traffic sees the new version. That lands about six weeks before the November Click Frenzy event and the run into a summer Christmas, which is exactly when most merchants are shipping their riskiest theme work. Good timing, if you use the tool for what it is actually good at.
Publish to a tenth of your traffic, not all of it
The normal way a theme goes live is a single click: unpublished yesterday, serving 100% of sessions today. If something breaks, every visitor sees it until someone notices. On a Tuesday in March that costs you a few orders. On the first night of Click Frenzy, when you have paid to put people on the site, it costs considerably more.
A staged release changes the blast radius. Ten percent of sessions get the new theme, ninety percent stay on the known-good one. You get real traffic, real devices, real payment methods, real Australian networks, with a tenth of the downside. If the new build is fine, you move to 50%, then 100%. If it is not, you set it back to zero and the problem affected one visitor in ten for forty minutes.
That is a deployment safety mechanism. Dress it up as experimentation and you will draw conclusions from noise.
Minimum traffic for a Shopify A/B test: the arithmetic nobody shows you
Here is the rough sample-size calculation most people skip. For a two-sided test at 95% confidence and 80% power, the usual approximation is n ≈ 16 × p(1−p) ÷ δ² per variant, where p is your baseline conversion rate and δ is the absolute difference you want to detect.
Say you convert at 2.5% and you want to detect a 10% relative improvement. That is an absolute difference of 0.25 percentage points, so δ = 0.0025.
- p(1−p) = 0.025 × 0.975 = 0.024375
- 16 × 0.024375 = 0.39
- δ² = 0.0025² = 0.00000625
- n = 0.39 ÷ 0.00000625 = 62,400 sessions per variant
So roughly 125,000 sessions to answer one question. A store doing 60,000 sessions a month splits that into 30,000 per arm and needs a bit over four months to finish. By then you have changed your ad mix, run a sale, and shipped nine other theme edits.
Loosen the standard and ask only for a 20% relative lift and the number drops to 15,600 per variant, about 31,200 total. Better. But a 20% conversion lift from a theme tweak is rare, and if your change really moves the needle that hard you will see it without a calculator. Most of what merchants want to test, button colour, badge wording, a reordered PDP, moves conversion by 2 to 5% if it moves it at all. Detecting a 5% relative lift on that same 2.5% baseline needs about 250,000 sessions per variant. Half a million sessions, for a copy change.
Australia has 26 million people and your addressable slice of it is small. That is the structural reason retention and average order value matter more here than acquisition volume, and it is the same reason split testing is a poor fit for most local stores. The traffic simply is not there.
Shopify Rollouts vs an A/B testing app
They solve different problems, and the overlap is smaller than the marketing suggests.
A dedicated testing app gives you variant assignment that sticks to a visitor, a statistics engine, segment-level reporting and usually a visual editor so a marketer can change a headline without touching Liquid. Worth paying for if you genuinely have six-figure monthly sessions and a roadmap of hypotheses. Most of those apps also inject a script that runs before paint, which is how you end up with a flash of the original content and a slower LCP on mobile. You are trading speed for measurement.
A staged theme release gives you none of the statistics and all of the safety. No extra script, no flicker, no third-party tag in the critical path. It answers "does this break anything" rather than "which one wins".
Our rule: if the sample-size maths above says you cannot finish a test inside three weeks, do not run one. Ship the change you believe in, stage it, and measure the business over a month instead. Nobody has ever regretted not running an underpowered test. Plenty of people have made a bad decision off one.
How to test a theme change before Click Frenzy
The sequence we use on client stores, compressed for peak season:
- Preview and device sweep. Theme preview link, checked on a real iPhone and a real mid-range Android, not just a desktop browser resized. Safari on iOS breaks things Chrome does not.
- Checkout a live order. Place a real order through the preview using Afterpay or Zip, not a test gateway. Instalment widgets are the single most common casualty of a PDP redesign, and they are invisible in a desktop Chrome check because the merchant already knows what the page is supposed to look like.
- Publish at 10% for a defined window. Pick a window with normal traffic. Not 2am, not during an EDM send.
- Hold at 50% overnight. Overnight traffic catches the long tail of devices and the Perth time offset, which is where regional and WA visitors show up in volume.
- Go to 100% at least 10 days before the sale. Then freeze. The point of a staged launch is to buy observation time, and you cannot observe anything if the sale starts the next morning.
The step people skip is the freeze. A theme that has been live and stable for ten days is worth more on Click Frenzy night than a theme that is 15% faster and went live on the Monday.
What to watch during the first hour
Aggregate conversion rate tells you nothing at 10% of traffic over an hour. Watch signals that break loudly instead:
- JavaScript errors in the browser console, on a product page, a collection page and the cart. One uncaught error can kill the add-to-cart handler on a single device class.
- Checkout starts as a proportion of add-to-carts. A broken cart shows up here fastest.
- Instalment messaging rendering on PDP and in cart. Afterpay and Zip copy sits in app blocks that a theme section rewrite can orphan.
- Price display including GST, and any was/now comparison pricing. If a template change makes the ex-GST figure the prominent one, you have a compliance problem as well as a conversion problem.
- Shipping threshold logic in the cart. Free shipping over $99 that quietly applies to a WA address you cannot actually ship to for $9 is a margin leak that nobody catches until the Australia Post invoice arrives.
- Mobile LCP on the template you changed. A hero image swap or a new app block can add a second without anyone noticing in the editor. If you want this tracked continuously rather than spot-checked, SwiftStore scans and monitors PageSpeed over time, which is more useful than a one-off number from a test run on a desktop connection.
One honest caveat on analytics. Splitting traffic between two themes without tagging sessions by theme version makes your reporting ambiguous for the whole window, and we get that tagging wrong on the first pass more often than we would like. Add the dimension before you start the rollout, not after someone asks why mobile conversion looks odd.
Rolling back a theme change, and what doesn't roll back
Rolling back is the easy part: set the new theme's traffic share to zero, or republish the previous theme. The old theme has been sitting there untouched, so visitors go back to a known state within a cache cycle.
What does not come back with it:
- Theme settings edited in the live editor during the window. If someone changed a homepage banner on the live theme while the rollout was running, that edit belongs to one theme only.
- App data. Metafields, review widgets, subscription configs and anything an app wrote to the store are global. Reverting the theme does not revert them.
- Orders already placed. Obvious, but worth saying. A pricing bug that ran for forty minutes still produced orders you have to honour or cancel, and cancelling at volume during a sale does more damage than the bug did.
- Customer memory. Someone who saw a $59 price they cannot get again will email you about it.
Duplicate the live theme before you start and name it with the date. Takes thirty seconds and gives you a clean restore point that is not dependent on the rollout UI behaving.
A Click Frenzy store preparation checklist
Short version, in the order we work through it:
- Code freeze date set and shared with every agency and freelancer touching the store. Written down, not assumed.
- App audit. Uninstall what you stopped using in March. Leftover scripts from dead apps are the most common cause of a store that was fast in July and is not in November.
- Mobile LCP on homepage, top collection and your three best-selling PDPs. Under 2.5s is the bar, and getting a 2.4s LCP under 1.5s usually means image format and app-script work rather than a new theme.
- Collection page filtering tested against sale-time behaviour. Shoppers filter by price and size far more aggressively during a sale, and a filter query that times out on a 3,000-SKU catalogue takes the whole collection page with it.
- Discount logic tested against stacking: automatic discount plus code plus Afterpay minimum. Find the combination that produces a $0 order before a customer does.
- Inventory buffers on anything you cannot restock, and a plan for the oversell you will get anyway.
- Shipping zones and cutoffs reviewed for WA, NT and regional postcodes, with Sendle and Australia Post rates checked against what you are promising on the cart page. Publish your own cutoff dates once the carriers publish theirs.
- Email and SMS sending reputation warmed in the weeks before, not the week of.
- Sustainability and provenance claims on landing pages checked for substantiation. The ACCC has been active on misleading environmental claims and on was/now pricing, and a sale page is exactly where a vague claim gets written in a hurry.
- A named person on call for the first four hours of the event, with admin access and the confidence to roll back without asking.
What a staged theme launch won't protect you from
Traffic splitting covers the theme. It does not cover checkout extensions, Shopify Functions, Flow automations, discount configuration, app settings or anything in the admin. Those apply to every order regardless of which theme the visitor was served. A broken shipping rate is broken at 10% and at 100%.
It also will not tell you whether a change is good. Click Frenzy traffic skews heavily to deal-seekers, so even if you had the volume for a valid test that week, the result would not transfer to February. Measure design decisions in ordinary weeks. Use peak weeks to avoid breaking things.
And if your store is doing a few thousand sessions a month, skip the ceremony entirely. Preview, check on two phones, publish, watch the console. Staging 10% of 2,000 sessions gives you six visitors an hour and no information.
Where to start this week
Pick your freeze date first, then work backwards: 100% live ten days before, 50% the night before that, 10% the morning before that. Everything that is not ready by then ships in January.
If you want a second pair of eyes on the build before you freeze, our team works with Australian merchants on exactly this kind of pre-peak hardening. A free audit will tell you where the speed and checkout risk actually sits, and the speed optimisation and Australian store pages explain how we usually sequence the work.


