Measuring whether the optimisation efforts were worth it has always been one of the biggest challenges of product feed optimisation. Tweaking a product title, rephrasing a product description with the help of AI, changing an image, or reorganising product attributes is quite easy. Knowing whether the change actually improved commercial performance, or only appeared to, is harder.
This is the question Feedoptimise A/B Testing Suite was built to answer.
Built right inside the Feedoptimise product feed management system, the A/B Testing Suite integrates the process of feed transformation, running experiments, collecting item-level performance data, making decisions on their basis, and implementing winners into one seamless workflow.
Unlike simply dividing products into two segments and finding out which of the two produced the better results in terms of CTR, conversion rate, or ROAS, Feedoptimise decides whether the difference is statistically significant enough to make the call about the winning variation.
The decision engine of Feedoptimise takes into account such factors as statistical confidence, confidence intervals, minimum lift, sample size, experiment duration, conversion lag, and performance guardrails when selecting the winner.
Once the evidence is available, the winning variation is then brought back to the feed through a controlled and previewable implementation process.
It's a completely new approach to product feed optimisation since the merchants don't have to make a change and hope that it'll pay off but rather test the hypothesis and implement the changes backed up by real results.
What the latest release includes
- Testing on any feed field. Titles, descriptions, product types, categories, images, any output attribute in your feed mapping.
- Two experiment designs. Duplicate testing with both sides running simultaneously with different item IDs. Rotational testing with only one item ID and alternation of the entire population during phases.
- Your own performance data. Connect Google Ads, Google Analytics, Google Merchant Center, Shopify, WooCommerce, Magento, Centra, Facebook, or a custom report.
- A decision engine with default settings. 95% confidence, 5% minimum lift, 14 minimum days, 5,000 impressions, 300 clicks and 30 conversions per side, a 7-day conversion lag, and a 60-day maximum duration.
- Per-item verdicts. A table showing which individual products earned enough evidence to act on, separate from the overall result.
- A reviewed implementation path. Winners go into the Shared Override via a pre-reviewed snapshot. No changes will be made to your live feed until your approval.
Why statistical significance matters in feed testing
Feeds run within the auction. Traffic volumes fluctuate based on day of week, competitor bids, pacing of budget, seasonality, and Google algorithm updates. In this context, a CTR difference of 4% between two title formats does not represent a noteworthy signal for most catalogues.
People still make decisions off of those differences. They rollout a title format to 40,000 SKUs, wait a quarter while performance drifts, and don't know if it even caused any impact. This decision comes at the cost of both wasted rollout effort but more importantly, the wrong assumption being carried forward through the next tests.
There are three main causes of error here.
- Reading a result before the information is there. Guidance for running supplemental feed testing title optimization is commonly around 100 clicks being enough of a data point to draw conclusions about the trend. This is good enough for CTR. It is nowhere near enough to draw any conclusions about a conversion rate or ROAS, and this is how most feed teams lose money.
- Optimizing one metric into another metric's ground. Append "Black Friday" and "Free Gift" to a title and CTR skyrockets. However, conversion rate and ROAS may plummet, because the title attracted non-purchasing visitors to the site. Title optimization focused on just one metric will declare fake victory here.
- Stopping the test when the number looks good. Daily checks on the dashboard followed by stopping the test as soon as it hits an upward spike creates many false positives. The early test lift results from a random walk. Stop at the top of the lift and you measure the top of the lift.
Feedoptimise takes care of each of these as a built-in structural element.
Two Experimental Methods: Duplicated and Rotated Testing
Since different commerce channels imply different restrictions in experiments, Feedoptimise employs two different methodologies for testing.
1. Duplicate A/B testing
When performing a duplicated experiment, Feedoptimise releases:
A - the old product, with the same old ID and old values.
B - a duplicated product, with a new predictable ID and new tested values.
As a result, both variants can gather performance data at the same time.
The great simultaneous comparison of the experiment will be available when the destination accepts both product IDs and can direct traffic to both versions.
For example:
Old ID: 34130964
Old Title: Running Shoes
New ID: M-34130964
New Title: Men's Lightweight Running Shoes
Both products compete within the same period of time, with Feedoptimise attributing the results of the experiment.
2. Rotated A/B testing
There are destinations where duplicate products cannot exist.
In such cases, Feedoptimise will be able to retain the same ID for the experiment but rotate it between the old and new values.
The first phase may use the old version of the experiment. Then comes the new version.
Feedoptimise offers to perform rotations either in time or based on accumulated metric volume (for instance, impressions).
Rotations based on accumulated metric volume are especially valuable since the experiment can rotate depending on approximately equal amount of exposure collected.
Connect Experiments to Real Performance Data
Feedoptimise links experiments to item-level reporting sources, which include:
Advertising platforms, analytics platforms, commerce platforms, and custom reporting files.
Depending on account setup, reporting sources may include:
Google Ads, Google Analytics, Google Merchant Center, Facebook/Meta reporting, Shopify, WooCommerce, Magento,, and URL-based custom reports.
Decision-making process
There are six factors lying between “the tested line is higher” and “tested wins”.
1. Firstly, there is a primary metric selected in advance
Choose a metric that corresponds to the hypothesis: ROAS, value per click, value per conversion, conversion rate or CTR. The auto-selection algorithm prefers ROAS-based metrics at the report level. Post-result selection of a metric that favours the variant is how wins are artificially generated by the teams.
Feedoptimise calculates those based on the underlying additive roles. The mapping of your source columns is done in terms of five semantic roles:
| Semantic role: | Typical source columns: |
| Exposure / impressions | impressions, views |
| Clicks / visits | clicks, sessions |
| Conversions / outcomes | conversions, orders |
| Cost / spend | cost, ad_spend |
| Conversion value / revenue | conv_value, revenue |
Of these, the engine extracts the CTR, conversion rate, CPC, ROAS, value per click, and value per conversion. It is the mapping of quantities into numbers that allows valid resampling to be performed.
2. A confidence interval, not just one number
Perform a test for two weeks, and you get a single figure: tested performed 8% better. Perform the test for another two weeks, and it gives a different answer, since clicks and conversions are not equally distributed.
Feedoptimise analyses this uncertainty. The engine repeats your test for thousands of times on your real daily data in all possible ways and gives you an interval where the true lift can possibly be.
Observed lift: +8%
95% CI: +2% to +14%
All figures in this interval are positive, therefore tested wins.
Observed lift: +8%
95% CI: -3% to +18%
Same headline. This interval contains zero, hence the difference could be +18% and -3% and your data does not allow you to distinguish between them. The conclusion remains inconclusive.
The confidence level means how reliable the data is in confirming a certain direction of difference. It is not the probability of the profit of rollout.
3. Your minimum lift preset in advance
Statistical significance and practical relevance are different concepts. There could be a lift of 0.4% in ROAS, which will be real but not valuable enough to implement catalog-wide. The default minimum lift is 5%. Standard strictness asks to exceed zero and minimum lift. Strict setting needs the whole interval to be over the minimum lift threshold.
4. Sample gates on both sides
Default values are 5,000 impressions, 300 clicks, and 30 conversions per each side and 14 minimum days.
Per each side makes a difference. If a report says you have 10,000 impressions, it looks good till you see that 9,200 impressions belong to the original ad set. The gate will be properly kept in the Waiting status. Aggregate numbers obscure precisely that kind of information.
Per item gates work separately with 1,000 impressions, 100 clicks, and 10 conversions per each side. It's possible to have a solid report level conclusion while most of individual SKUs would be in Insufficient Data stage. It's the standard state of a long tail catalog.
5. Conversion lag
The orders come after clicks. Conversion lag keeps away the recent days from the decision sample so that late arrivals would be accounted. The default value is seven days, and you should specify it depending on the actual attribution delay of your platform. A test can reach the end date and still be Waiting for conversion lag, which means that the system refuses to evaluate an incomplete data set.
6. Guardrails
Guardrails control ROAS, conversion rate, CPC, and value per click during running the primary metric. Failure of a guardrail will prohibit the winning of a tested version. A typical example is changing a title which increases CTR by 12% but decreases conversion rate by 20%. Primary metric is fine, but the guardrail detects the problem and gives you a Guardrail failed verdict instead of Tested is winning.
Guardrails don't prevent original version from winning. Keep what you already have poses no extra risks for you.
Possible verdicts
- Tested is winning
- Original is winning
- No decision yet
- Collecting data
- Waiting for minimum duration
- Waiting for exposure
- Waiting for conversion lag
- Guardrail failed
- Insufficient data
Seven out of nine verdicts are the refusal to make a decision. This ratio is the goal. "There is no winner" is a valid experimental result, and a testing tool which always produces a winner is useless.
Product decisions: which SKUs actually improved
Product wins aren’t always universal across a test, just positive enough on balance.
The Product decisions tab analyzes individual products separately from the aggregate analysis, either at the parent level with variant metrics rolled up or as individual parent-original to parent-tested variant pairs. Each row includes a winner, a confidence score, a lift percentage, the actual metrics, a status, a guardrail assessment, and whether or not it qualifies as a snapshot. Drill down on a row to see the base metrics for both original and tested, their ratio, the specific gate used for the status, and the decision history.
It’s right here that selective rollouts happen. Just apply the winning test title to the 340 products that won it and leave the other products as they are.
From verdict to live feed
A winning test does not touch your feed. Publishing runs through three reviewed steps.
- Snapshot. Decide on which tests should be rolled out with the help of Implement tested winners for a selective rollout, Save original winner decisions to record what has worked, or Apply whole-report decision, if an experiment was planned as a single decision for all people.
- Preview. Every snapshot dialog performs a dry run first. It informs about number of units evaluated, qualified, winners and their status, unattributable units and a sample of rows it writes. Units that cannot attribute do not contain captured field values for the selected side, therefore, they will be skipped during saving. Consider those rows prior to saving them.
- Override. Qualified rows are saved in a Shared Override. Applying an override to a feed is a separate step of Feed Mapping, having its own preview stage. Items that are not present in an override will retain their feed value.
Two options are worth mentioning. Values to save is independent of the winner filter: the former determines which rows are qualified, and the latter which side's values are written.
Every rollout is always auditable and can be undone. In case of important feed, use the feed preview before deploying it in live feed.
Feedoptimise versus other feed testing approaches
It should be noted that not all A/B testing features offer the same testing abilities.
The process of feed-testing, which is well-publicized and known, includes four identical steps: product division into segments, content changing, collecting performance data, comparing metrics. Feedoptimise does all of the steps mentioned above and goes further - it makes the final decision.
| Capability | Google Product Data Experiments | Typical feed-tool A/B testing | Feedoptimise A/B Testing |
| Fields testable | Titles and images | Mainly titles | Any output field |
| Channels measured | Google Shopping and PMax | Varies | Any channel your feed serves |
| Test design | Simultaneous traffic split | Usually simultaneous | Duplicated or rotated |
| Same-ID rotated testing | No | Varies | Yes, scheduled |
| Rotation by performance volume | No | Rare | Yes, on any additive metric |
| Data sources | Google | Platform-connected | Google Ads, GA, GMC, Shopify,
WooCommerce, Magento, Centra,
Facebook, Custom reports |
| Primary objective fixed in advance | No | Basic | Yes, six metrics available |
| Confidence interval on lift | Not exposed | Rare | Yes, resampled from daily data |
| Conversion-lag handling | Internal | Rare | Yes, 7 days default |
| Secondary-metric guardrails | No | Rare | Yes, on ROAS, CVR, CPC, value per click |
| Strictness modes | No | Rare | Standard and Strict |
| Explicit no-winner verdict | Not documented | Rare | Yes, seven distinct refusal states |
| Per-SKU verdicts | No | Rare | Yes, for CTR and conversion rate |
| Decision timeline across the test | No | No | Yes, lift and interval bounds per day |
| Dry-run preview before anything writes | No | Rare | Yes, on every snapshot |
| Selective per-item rollout | Apply variant to all | Rare | Yes, via Shared Override |
Stop Guessing. Start Testing.
Feed optimisation shouldn’t just stop once a title, an image, or an attribute is created - this is where the testing begins.
With Feedoptimise A/B Testing, retailers and agencies can test their product data against reality, verify whether the changes have enough merit behind them, and safely transform qualified winners into real-life feed optimisation wins.