Feedoptimise
Feedoptimise
menu
Start free trial Book live demo

Statistically Significant A/B Testing for Product Feeds

Measuring whether the optimisation efforts were worth it has always been one of the biggest challenges of product feed optimisation. Tweaking a product title, rephrasing a product description with the help of AI, changing an image, or reorganising product attributes is quite easy. Knowing whether the change actually improved commercial performance, or only appeared to, is harder.

This is the question Feedoptimise A/B Testing Suite was built to answer.

Built right inside the Feedoptimise product feed management system, the A/B Testing Suite integrates the process of feed transformation, running experiments, collecting item-level performance data, making decisions on their basis, and implementing winners into one seamless workflow.

Unlike simply dividing products into two segments and finding out which of the two produced the better results in terms of CTR, conversion rate, or ROAS, Feedoptimise decides whether the difference is statistically significant enough to make the call about the winning variation.

The decision engine of Feedoptimise takes into account such factors as statistical confidence, confidence intervals, minimum lift, sample size, experiment duration, conversion lag, and performance guardrails when selecting the winner.

Once the evidence is available, the winning variation is then brought back to the feed through a controlled and previewable implementation process.

It's a completely new approach to product feed optimisation since the merchants don't have to make a change and hope that it'll pay off but rather test the hypothesis and implement the changes backed up by real results.

What the latest release includes

  • Testing on any feed field. Titles, descriptions, product types, categories, images, any output attribute in your feed mapping.
  • Two experiment designs. Duplicate testing with both sides running simultaneously with different item IDs. Rotational testing with only one item ID and alternation of the entire population during phases.
  • Your own performance data. Connect Google Ads, Google Analytics, Google Merchant Center, Shopify, WooCommerce, Magento, Centra, Facebook, or a custom report.
  • A decision engine with default settings. 95% confidence, 5% minimum lift, 14 minimum days, 5,000 impressions, 300 clicks and 30 conversions per side, a 7-day conversion lag, and a 60-day maximum duration.
  • Per-item verdicts. A table showing which individual products earned enough evidence to act on, separate from the overall result.
  • A reviewed implementation path. Winners go into the Shared Override via a pre-reviewed snapshot. No changes will be made to your live feed until your approval.

Why statistical significance matters in feed testing

Feeds run within the auction. Traffic volumes fluctuate based on day of week, competitor bids, pacing of budget, seasonality, and Google algorithm updates. In this context, a CTR difference of 4% between two title formats does not represent a noteworthy signal for most catalogues.

People still make decisions off of those differences. They rollout a title format to 40,000 SKUs, wait a quarter while performance drifts, and don't know if it even caused any impact. This decision comes at the cost of both wasted rollout effort but more importantly, the wrong assumption being carried forward through the next tests.

There are three main causes of error here.

  1. Reading a result before the information is there. Guidance for running supplemental feed testing title optimization is commonly around 100 clicks being enough of a data point to draw conclusions about the trend. This is good enough for CTR. It is nowhere near enough to draw any conclusions about a conversion rate or ROAS, and this is how most feed teams lose money.
  2. Optimizing one metric into another metric's ground. Append "Black Friday" and "Free Gift" to a title and CTR skyrockets. However, conversion rate and ROAS may plummet, because the title attracted non-purchasing visitors to the site. Title optimization focused on just one metric will declare fake victory here.
  3. Stopping the test when the number looks good. Daily checks on the dashboard followed by stopping the test as soon as it hits an upward spike creates many false positives. The early test lift results from a random walk. Stop at the top of the lift and you measure the top of the lift.

Feedoptimise takes care of each of these as a built-in structural element.

Two Experimental Methods: Duplicated and Rotated Testing

Since different commerce channels imply different restrictions in experiments, Feedoptimise employs two different methodologies for testing.

1. Duplicate A/B testing

When performing a duplicated experiment, Feedoptimise releases:

A - the old product, with the same old ID and old values.
B - a duplicated product, with a new predictable ID and new tested values.

As a result, both variants can gather performance data at the same time.

The great simultaneous comparison of the experiment will be available when the destination accepts both product IDs and can direct traffic to both versions.

For example:

Old ID: 34130964
Old Title: Running Shoes
New ID: M-34130964
New Title: Men's Lightweight Running Shoes

Both products compete within the same period of time, with Feedoptimise attributing the results of the experiment.

2. Rotated A/B testing

There are destinations where duplicate products cannot exist.

In such cases, Feedoptimise will be able to retain the same ID for the experiment but rotate it between the old and new values.

The first phase may use the old version of the experiment. Then comes the new version.

Feedoptimise offers to perform rotations either in time or based on accumulated metric volume (for instance, impressions).

Rotations based on accumulated metric volume are especially valuable since the experiment can rotate depending on approximately equal amount of exposure collected.

Connect Experiments to Real Performance Data

Feedoptimise links experiments to item-level reporting sources, which include:

Advertising platforms, analytics platforms, commerce platforms, and custom reporting files.

Depending on account setup, reporting sources may include:

Google Ads, Google Analytics, Google Merchant Center, Facebook/Meta reporting, Shopify, WooCommerce, Magento,, and URL-based custom reports.

Decision-making process

There are six factors lying between “the tested line is higher” and “tested wins”.

1. Firstly, there is a primary metric selected in advance

Choose a metric that corresponds to the hypothesis: ROAS, value per click, value per conversion, conversion rate or CTR. The auto-selection algorithm prefers ROAS-based metrics at the report level. Post-result selection of a metric that favours the variant is how wins are artificially generated by the teams.

Feedoptimise calculates those based on the underlying additive roles. The mapping of your source columns is done in terms of five semantic roles:

Semantic role:
Typical source columns:
Exposure / impressions
impressions, views
Clicks / visits
clicks, sessions
Conversions / outcomes
conversions, orders
Cost / spend
cost, ad_spend
Conversion value / revenue
conv_value, revenue

Of these, the engine extracts the CTR, conversion rate, CPC, ROAS, value per click, and value per conversion. It is the mapping of quantities into numbers that allows valid resampling to be performed.

2. A confidence interval, not just one number

Perform a test for two weeks, and you get a single figure: tested performed 8% better. Perform the test for another two weeks, and it gives a different answer, since clicks and conversions are not equally distributed.

Feedoptimise analyses this uncertainty. The engine repeats your test for thousands of times on your real daily data in all possible ways and gives you an interval where the true lift can possibly be.

Observed lift:   +8%
95% CI:          +2% to +14%

All figures in this interval are positive, therefore tested wins.

Observed lift:    +8%
95% CI:           -3% to +18%

Same headline. This interval contains zero, hence the difference could be +18% and -3% and your data does not allow you to distinguish between them. The conclusion remains inconclusive.

The confidence level means how reliable the data is in confirming a certain direction of difference. It is not the probability of the profit of rollout.

3. Your minimum lift preset in advance

Statistical significance and practical relevance are different concepts. There could be a lift of 0.4% in ROAS, which will be real but not valuable enough to implement catalog-wide. The default minimum lift is 5%. Standard strictness asks to exceed zero and minimum lift. Strict setting needs the whole interval to be over the minimum lift threshold.

4. Sample gates on both sides

Default values are 5,000 impressions, 300 clicks, and 30 conversions per each side and 14 minimum days.

Per each side makes a difference. If a report says you have 10,000 impressions, it looks good till you see that 9,200 impressions belong to the original ad set. The gate will be properly kept in the Waiting status. Aggregate numbers obscure precisely that kind of information.

Per item gates work separately with 1,000 impressions, 100 clicks, and 10 conversions per each side. It's possible to have a solid report level conclusion while most of individual SKUs would be in Insufficient Data stage. It's the standard state of a long tail catalog.

5. Conversion lag

The orders come after clicks. Conversion lag keeps away the recent days from the decision sample so that late arrivals would be accounted. The default value is seven days, and you should specify it depending on the actual attribution delay of your platform. A test can reach the end date and still be Waiting for conversion lag, which means that the system refuses to evaluate an incomplete data set.

6. Guardrails

Guardrails control ROAS, conversion rate, CPC, and value per click during running the primary metric. Failure of a guardrail will prohibit the winning of a tested version. A typical example is changing a title which increases CTR by 12% but decreases conversion rate by 20%. Primary metric is fine, but the guardrail detects the problem and gives you a Guardrail failed verdict instead of Tested is winning.

Guardrails don't prevent original version from winning. Keep what you already have poses no extra risks for you.

Possible verdicts

  • Tested is winning
  • Original is winning
  • No decision yet
  • Collecting data
  • Waiting for minimum duration
  • Waiting for exposure
  • Waiting for conversion lag
  • Guardrail failed
  • Insufficient data

Seven out of nine verdicts are the refusal to make a decision. This ratio is the goal. "There is no winner" is a valid experimental result, and a testing tool which always produces a winner is useless.

Product decisions: which SKUs actually improved

Product wins aren’t always universal across a test, just positive enough on balance.

The Product decisions tab analyzes individual products separately from the aggregate analysis, either at the parent level with variant metrics rolled up or as individual parent-original to parent-tested variant pairs. Each row includes a winner, a confidence score, a lift percentage, the actual metrics, a status, a guardrail assessment, and whether or not it qualifies as a snapshot. Drill down on a row to see the base metrics for both original and tested, their ratio, the specific gate used for the status, and the decision history.

It’s right here that selective rollouts happen. Just apply the winning test title to the 340 products that won it and leave the other products as they are.

From verdict to live feed

A winning test does not touch your feed. Publishing runs through three reviewed steps.

  1. Snapshot. Decide on which tests should be rolled out with the help of Implement tested winners for a selective rollout, Save original winner decisions to record what has worked, or Apply whole-report decision, if an experiment was planned as a single decision for all people.
  2. Preview. Every snapshot dialog performs a dry run first. It informs about number of units evaluated, qualified, winners and their status, unattributable units and a sample of rows it writes. Units that cannot attribute do not contain captured field values for the selected side, therefore, they will be skipped during saving. Consider those rows prior to saving them.
  3. Override. Qualified rows are saved in a Shared Override. Applying an override to a feed is a separate step of Feed Mapping, having its own preview stage. Items that are not present in an override will retain their feed value.

Two options are worth mentioning. Values to save is independent of the winner filter: the former determines which rows are qualified, and the latter which side's values are written.

Every rollout is always auditable and can be undone. In case of important feed, use the feed preview before deploying it in live feed.

Feedoptimise versus other feed testing approaches

It should be noted that not all A/B testing features offer the same testing abilities.

The process of feed-testing, which is well-publicized and known, includes four identical steps: product division into segments, content changing, collecting performance data, comparing metrics. Feedoptimise does all of the steps mentioned above and goes further - it makes the final decision.

Capability
Google Product Data Experiments
Typical feed-tool A/B testing
Feedoptimise A/B Testing
Fields testable
Titles and images
Mainly titles
Any output field
Channels measured
Google Shopping and PMax
Varies
Any channel your feed serves
Test design
Simultaneous traffic split
Usually simultaneous
Duplicated or rotated
Same-ID rotated testing
No
Varies
Yes, scheduled
Rotation by performance volume
No
Rare
Yes, on any additive metric
Data sources
Google
Platform-connected
Google Ads, GA, GMC, Shopify, WooCommerce, Magento, Centra, Facebook, Custom reports
Primary objective fixed in advance
NoBasic
Yes, six metrics available
Confidence interval on lift
Not exposed
Rare
Yes, resampled from daily data
Conversion-lag handling
InternalRare
Yes, 7 days default
Secondary-metric guardrails
No
Rare
Yes, on ROAS, CVR, CPC, value per click
Strictness modes
No
Rare
Standard and Strict
Explicit no-winner verdict
Not documented
Rare
Yes, seven distinct refusal states
Per-SKU verdicts
No
Rare
Yes, for CTR and conversion rate
Decision timeline across the test
No
NoYes, lift and interval bounds per day
Dry-run preview before anything writes
No
Rare
Yes, on every snapshot
Selective per-item rollout
Apply variant to all
Rare
Yes, via Shared Override


Stop Guessing. Start Testing.

Feed optimisation shouldn’t just stop once a title, an image, or an attribute is created - this is where the testing begins.

With Feedoptimise A/B Testing, retailers and agencies can test their product data against reality, verify whether the changes have enough merit behind them, and safely transform qualified winners into real-life feed optimisation wins.

Start your free trial or Book a demo

Frequently Asked Questions

  • What is statistically significant A/B testing for product feeds?

    Statistically significant A/B testing for product feeds is an experiment workflow where feed variants are tested on real traffic and a decision engine checks confidence intervals, minimum lift, sample size, duration, conversion lag and guardrails before declaring a winning version or refusing to decide.

  • How does Feedoptimise A/B Testing improve product feed optimisation?

    Feedoptimise A/B Testing lets merchants test any feed field, connect real performance data, run duplicate or rotated experiments, and use a decision engine to roll out only statistically proven winners instead of relying on gut feeling or noisy CTR differences.

  • What is the difference between duplicate and rotated A/B testing in product feeds?

    Duplicate A/B testing creates a second product with a new ID so both variants run simultaneously, while rotated A/B testing keeps the same ID and alternates old and new values over time or by impression volume when a channel does not allow duplicate products.

  • Which metrics and data sources can Feedoptimise use for feed A/B tests?

    Feedoptimise can use item-level data from Google Ads, Google Analytics, Google Merchant Center, Facebook, Shopify, WooCommerce, Magento, Centra and custom reports, and optimise for ROAS, value per click, value per conversion, conversion rate or CTR with guardrails on ROAS, CVR, CPC and value per click.

  • Why is statistical significance important in product feed testing?

    Statistical significance prevents feed teams from acting on random fluctuations in CTR, conversion rate or ROAS caused by auctions, seasonality and budget pacing, reducing false positives from early spikes, underpowered samples and single-metric optimisations that hurt profitability.

  • How does the Feedoptimise decision engine determine a winning feed variant?

    The decision engine pre-selects a primary metric, builds confidence intervals on lift, enforces minimum lift thresholds, sample gates, conversion lag and guardrails, then issues verdicts such as Tested is winning, Original is winning, Guardrail failed or several no-decision states when evidence is insufficient.

  • How can I roll out winning A/B test results to my live product feed?

    Feedoptimise uses a three-step rollout: create a snapshot of winners, run a dry-run preview showing qualified items and sample rows, then save changes into a Shared Override that can be selectively applied in feed mapping, ensuring every rollout is reviewable and reversible.