Evergreen
How to Audit an Online Pricing Study
A practical method for testing claims about different online prices, including controls, screenshots, samples, calculations, causation, and company response.
How to Audit an Online Pricing Study
To audit an online pricing study, first identify the exact claim. Then check whether the comparison holds the product, seller, store, place, time, account state, promotion, fulfillment method, and total cost constant. After that, separate the observed difference from the proposed cause.
A study can prove that two prices differed without proving that personal data, demand, loyalty, or discrimination caused the difference. The audit should make that boundary visible.
1. Rewrite the headline as a testable claim
"The platform charges everyone differently" is not a testable claim. "Two logged-in shoppers using the same store location at 8:30 p.m. saw different displayed prices for the same 12-ounce product" is.
Write the smallest claim that the evidence could establish. Identify the outcome, comparison group, time, place, and proposed cause. If the article uses several labels, split them.
| Claim type | Question the study must answer |
|---|---|
| Price variation | Did equivalent sessions show different prices? |
| Dynamic pricing | Did price respond to time, supply, demand, or another changing condition? |
| Personalized pricing | Did identity, account, or segment affect the offer? |
| Surveillance pricing | Did personal data or an inference from personal data affect an individualized price or offer? |
| Loyalty penalty | Did customer tenure or expected switching behavior produce a worse comparable offer? |
| Illegal discrimination | Did the conduct meet the elements of an applicable law? |
These conclusions require different evidence. A study should not use the easiest observation to support the strongest label.
2. Define the comparison unit
Two products with the same common name can differ in size, seller, fulfillment source, package count, membership eligibility, or inventory status. Record the exact product identifier or URL, brand, variant, size, quantity, seller, and store location.
Then record the channel. A marketplace, retailer storefront, mobile app, desktop site, guest session, and logged-in account can produce different catalog and promotion rules.
The final comparison unit should read like a transaction record, not a memory: exact item, exact seller, exact store, exact fulfillment method, exact time window, and exact account conditions.
3. Hold nuisance factors still
NIST describes experimental design as planning a study in advance so the resulting data can support valid and objective conclusions. Its guidance on randomized block designs treats time, operator, and environment as nuisance factors that should be held constant or accounted for.
Online pricing has its own nuisance factors. Store location, delivery address, time, inventory, membership, coupon state, prior cart activity, tax, delivery window, and device channel can all change what a shopper sees.
Use simultaneous sessions where possible. Confirm both sessions are pointed at the same store location and fulfillment method. Open the same product through the same route. Check whether a coupon is clipped, a loyalty account is linked, or a membership benefit is active.
If a factor cannot be held constant, record it and limit the conclusion.
4. Preserve the full offer
A cropped price is rarely enough. Capture the product page, size, seller, store, promotion language, quantity, timestamp, and URL. Add the item to the cart and capture the cart total, fees, taxes, delivery terms, and discount application.
Keep the original files. An edited image can help a reader, but the underlying evidence should remain intact with its original timestamp and dimensions.
If the study involves multiple testers, use a shared naming rule and a structured data sheet. The screenshot and row should point to each other through a unique observation ID.
[[How to Document Different Prices Online]] provides a smaller evidence packet for consumers conducting a live comparison.
5. Check for ordinary false positives
Before attributing a difference to an algorithm, check whether the sessions used different stores, package sizes, third-party sellers, delivery zones, memberships, clipped coupons, loyalty accounts, inventory states, or fulfillment speeds.
Weighted grocery items deserve special treatment because the displayed estimate can change after the in-store weight is known. Instacart's current item-pricing guide also says some retailers use different marketplace and in-store prices and that their pricing policies are displayed on the storefront.
A false positive does not mean the study was careless. It means the interface contains more variables than the headline implies.
6. Inspect the sample and exclusions
Report how people were recruited, how many started, how many completed the procedure, and why any records were removed. A study with 437 volunteers and 200 usable screenshot sets has both numbers, and they answer different questions.
The sample affects what can be generalized. A convenience sample can document a platform behavior without representing the national population. A city-specific session can establish what happened there without estimating how often it happens everywhere.
Look for selection effects. Were volunteers more technical than typical shoppers? Did the procedure exclude mobile users, guests, rural locations, or people with accessibility needs? Did researchers choose products after seeing variation?
Those limits should narrow the conclusion rather than bury the finding.
7. Recalculate every headline number
Price studies often move between cents, percentages, baskets, averages, ranges, and annual estimates. Reproduce each transformation.
Check the denominator for a percentage. A 23 percent difference can mean the higher price relative to the lower price, the spread relative to an average, or another calculation. Check whether "three-quarters of products" refers to unique products, product-session observations, or analyzed baskets.
For an annual estimate, write the formula. Identify the observed difference, assumed shopping frequency, assumed annual spend, persistence of the treatment, and population to which the estimate is applied.
An extrapolation can be informative. It should be labeled as an extrapolation.
8. Separate output, assignment, and cause
The visible price is an output. The experiment group is an assignment. The data and rule that placed someone into that group are the cause under investigation.
flowchart LR
A["Screenshot or cart record"] --> B["Different outcome observed"]
B --> C["Assignment record"]
C --> D["Input data and rule"]
D --> E["Supported causal explanation"]
Researchers can often document the first two stages from outside a platform. The later stages usually require internal logs, specifications, data dictionaries, contracts, code review, discovery, or regulator access.
If the study cannot reach the assignment record, it should say so. "The reason remains unverified" is a result, not a failure.
9. Grade the evidence
Use an evidence grade before writing the conclusion.
| Grade | Available evidence | Defensible statement |
|---|---|---|
| A | Memory, report, or isolated cropped image | A possible difference deserves follow-up |
| B | Full contextual screenshots | Two offers differed under the recorded conditions |
| C | Controlled and repeated comparisons | The variation persisted across the stated controls |
| D | Assignment or system records | A named rule or input contributed to the outcome |
| E | Population analysis and independent review | The system produced a measured pattern across a defined population |
| F | Legal finding or final order | An authority reached the stated legal conclusion |
The letters are editorial grades, not legal or academic standards. Their purpose is to stop a Grade B observation from being written as a Grade D causal claim or Grade F legal conclusion.
10. Put the company response beside the claim
Send a methodology summary, representative evidence, and specific questions to the platform, retailer, and technology provider. Ask which party set the price, which rule selected the treatment, what data entered the decision, how long the test ran, and whether the practice is current.
Do not put the response in a distant paragraph after the article has already converted the disputed claim into fact. Pair the answer with the issue it addresses.
A company statement is not an independent audit. It can still resolve a factual misunderstanding, identify another responsible party, or propose a mechanism that researchers can test.
The 2025 Instacart dispute illustrates this structure. Consumer Reports documented simultaneous variation. Instacart described randomized A/B assignment and denied personal-data use for item prices. The published evidence did not include a complete independent inspection of the assignment system. [[What Consumer Reports Found in Its Instacart Pricing Test]] keeps those columns separate.
11. Check what changed after publication
A pricing article can become stale quickly. Recheck the platform policy, current interface, regulator record, and company response before publication.
Instacart ended the item-price tests discussed in the 2025 investigation on December 22, 2025. A current article that says the tests are still running would misstate the record. A current article that says every differentiated offer ended would also go too far because promotions, loyalty discounts, markups, and store-level variation remain.
Date the verification and define a refresh trigger. Corrections, litigation, enforcement, product changes, or a new independent audit should reopen the page.
12. Publish the narrowest strong conclusion
A rigorous conclusion should state what was observed, how it was tested, what the data can support, what remains unknown, what the company said, and what later changed.
The strongest sentence is often less dramatic and more useful: "Coordinated shoppers saw different item prices, while the platform said assignment was randomized and denied personal-data use."
That sentence preserves the finding and the unresolved mechanism. It gives the next researcher a clear place to start.
This guide was developed from NIST experimental-design principles, FTC consumer guidance, the Consumer Reports investigation, Instacart's response, and the E095 evidence record. It is a research and editorial method, not legal advice. AI assistance was used for research organization, drafting, and validation. Publication remains unauthorized.
Sources
Follow the evidence.
- ag.ny.gov: attorney general james demands answers instacart about algorithmic pricingag.ny.gov
- itl.nist.gov: pri332itl.nist.gov
- instacart.com: promotionsinstacart.com
- ftc.gov: ftc surveillance pricing study indicates wide range personal data used set individualized consumer pricesftc.gov
- consumer.ftc.gov: online shoppingconsumer.ftc.gov
- company.instacart.com: instacartpricingcompany.instacart.com
- gov.uk: tackling the loyalty penaltygov.uk
- company.instacart.com: instacart makes it easier for customers to save on groceries with acquisition of eversightcompany.instacart.com
- instacart.com: 1586544648instacart.com
- usa.gov: online purchase complaintsusa.gov
- company.instacart.com: ending item price tests on instacartcompany.instacart.com
- company.instacart.com: the truth about pricing tests on instacartcompany.instacart.com
- itl.nist.gov: pri11itl.nist.gov
- fca.org.uk: fca confirms measures protect customers loyalty penalty home motor insurance marketsfca.org.uk
- consumerreports.org: instacart ai pricing experiment inflating grocery bills a1142182490consumerreports.org
- investors.instacart.com: 9e9aff2c 95db 4f75 bdf1 0f4025e1468cinvestors.instacart.com