Back to the episode map

Evergreen

How to Build a Repeatable Product Testing Protocol

Build a product testing protocol that defines the decision, specimen, conditions, instruments, repetitions, deviations, scoring, and change control.

Aug 4, 20269 min readBy Dalton Anderson

How to Build a Repeatable Product Testing Protocol

A repeatable product testing protocol begins with the decision it must support. It then defines the exact specimen, conditions, instruments, sequence, observations, repetitions, deviations, and scoring handoff before the first result is interpreted.

The test is ready when a second trained person can run the written method, produce a complete record, and explain any difference without relying on the inventor's memory.

The result is a comparable record, not a persuasive score

A successful protocol produces more than a number. It creates a record that identifies what was tested, which method version was used, what happened during the run, what varied, and whether the result is comparable with earlier work.

That record lets a team separate three questions. Did the procedure run as written? Did the instrument produce stable observations? Does the observation actually support the decision being made?

Those questions are related but not interchangeable. A perfectly repeated test can measure the wrong property. A relevant test can be executed inconsistently. A stable measurement can still become a misleading score when arbitrary weights hide the part that matters.

In episode 100 of Venture Step, NapLab founder Derek Hales describes turning understandable tests with medicine balls, images, and measured movement into a broader product-testing operation. NapLab's current methodology page shows the value of making procedures, calculations, scoring versions, limitations, and corrections visible. It is a worked example, not a universal product-testing standard.

flowchart TD
    A["Decision to support"] --> B["Property to observe"]
    B --> C["Defined specimen and conditions"]
    C --> D["Instrument and procedure"]
    D --> E["Repeated observations"]
    E --> F["Deviation and uncertainty review"]
    F --> G["Interpretation or score"]
    G --> H["Decision with stated limits"]

The order matters. Starting with a preferred score invites a team to design evidence backward from the answer.

Before you write the steps, define the decision

Write one sentence that names the user, choice, and consequence. "Measure edge deflection" is a test activity. "Determine whether this chair remains within our stability threshold after normal assembly" is a decision.

Next, name the property the procedure can observe. Avoid broad labels such as quality, comfort, trust, or durability until they are decomposed. A test may measure force, displacement, temperature change, response time, breakage, or task completion. The broader conclusion requires an argument connecting that observation to the user's decision.

Define the reporting unit before data collection. If one unit is tested three times, the result describes repeated measurements of one unit. It does not establish the variation across all manufactured units. A sample of one may still provide useful product evidence, but the limitation belongs beside the result.

1. Identify the specimen so it can be recognized later

Record the manufacturer, product name, model, size, configuration, revision, serial or lot number when available, date acquired, source, price paid or consideration received, and visible condition before testing.

Photograph labels and the unopened or initial state when condition matters. Record assembly, conditioning, charging, software version, foundation, accessories, or other preparation. A product name alone is rarely enough because a review can outlive the version that was tested.

The verification check is simple. A future editor should be able to determine whether the product now sold under the same name is materially the same specimen or a replacement that requires new testing.

2. Fix the conditions that could change the observation

Name the room, surface, orientation, temperature range, humidity range when relevant, warm-up or conditioning time, power state, network conditions, and operator preparation. Do not add environmental controls merely to look technical. Include conditions that could plausibly move the result.

If the procedure uses a position or force, define how it is located and applied. "Near the edge" leaves room for drift. A distance from a documented reference point is repeatable. "Press firmly" is an instruction to improvise. A specified load, tool, or controlled action is testable.

Run a short pilot to discover the hidden choices. The inventor will make adjustments without noticing them. A second operator exposes where the written procedure depends on tacit knowledge.

3. Characterize the measurement system

Record the instrument make, model, identifier, range, resolution, units, calibration or check procedure, software, and settings. If a photograph is the instrument, define camera position, lens, lighting, scale reference, and image-analysis method.

The NIST gauge R&R guidance provides a useful vocabulary. A measurement system can vary across artifacts, operators, instruments, configurations, and time. Relevant properties include repeatability, reproducibility, stability, bias, resolution, linearity, hysteresis, and drift.

A small team does not need to run every formal study. It should still use a check object or reference condition when possible. Measure it at the start of a session, repeat it later, and record whether the system moved. The NIST variability guidance explains why short-term precision may look strong even when day-to-day handling or environmental variation dominates.

4. Write the procedure as observable actions

Use numbered steps because sequence is part of the method. Each step should name the action, the expected state, and the observation to record.

  1. Confirm the specimen, method version, operator, instruments, and environmental conditions.

  2. Complete the preparation and document any condition that falls outside the allowed range.

  3. Place the specimen and instrument using the defined reference points.

  4. Apply the stimulus or perform the task using the specified force, duration, sequence, and timing.

  5. Capture the raw observation before converting it into a category or score.

  6. Repeat the measurement according to the planned count and reset procedure.

  7. Record deviations, anomalies, damage, interruptions, and excluded runs without deleting the original record.

  8. Complete the scoring or interpretation only after the run record is closed.

Raw observations should survive a scoring change. If a team records only "good" or "8.7," it cannot recalculate the result when thresholds, weights, or the comparison set change.

5. Decide how many repeats answer the question

Repeating a measurement on one product estimates short-term variation in the method. Testing multiple units estimates product-to-product variation. Repeating on different days or with different operators begins to expose reproducibility.

These designs answer different questions. Do not present five repeated measurements on the same unit as a five-product sample.

The NIST Technical Note 1297 appendix on method-defined measurements explains that uncertainty can depend on repeatability, reproducibility, and how well the specified method was implemented. It also recommends naming the method when the measured quantity is defined by that method.

For a public review, that means "response time under the Venture Step protocol, version 1.0" is more honest than "the true response time" when the result depends on the setup.

6. Preserve deviations instead of repairing them silently

A deviation is any departure from the written method. It may be harmless, disqualifying, or evidence that the protocol needs revision.

Record what changed, why it changed, who approved it, which observations were affected, and whether the run remains comparable. If the team reruns a test, preserve both records and state which one supports the published result.

This is where many polished reviews become fragile. An operator nudges a product into place, changes a timing window, removes an outlier, or substitutes an instrument. The adjustment may be reasonable. Concealing it makes the final precision stronger than the evidence.

7. Separate measurement from scoring

Store the raw value, derived value, category, sub-score, total score, and recommendation as separate fields. Version the formula, thresholds, weights, and comparison set.

NapLab's current version 1.3 methodology provides a concrete example. The company publishes non-scoring measurements such as sinkage because it does not treat more or less as universally better. It also says the current overall score is a weighted average of eight factors and that durability is scored but not yet included in that total.

The separation makes change possible without rewriting history. A team can improve a score model while retaining the observation produced by the older test.

Verify the protocol with a blind handoff

Give the method and prepared data form to a trained person who did not write it. Do not coach during the first run. Compare their setup, observations, deviations, and questions with the inventor's run.

The protocol passes its first gate when both records are complete, material differences can be explained, and no critical step depends on an unwritten choice. It has not yet proven long-term stability, broad product validity, or fitness for every decision.

Failure signalLikely causeCorrection
Operators place the product differentlyReference point is vagueAdd a physical datum, photograph, or fixture
Repeats drift within one sessionInstrument, reset, or specimen is unstableAdd a check standard and reset condition
Results change across daysEnvironment or handling is uncontrolledRecord day-level conditions and test the suspected factor
Score changes but raw evidence cannot be recoveredData model stores only interpretationPreserve raw, derived, and scored fields separately
A new method silently replaces the old oneChange control is missingAssign versions, effective dates, and retest rules

Lock a version, then define when it expires

Once the handoff works, assign a version and effective date. Record the equipment, forms, formulas, and training material it controls. Define which changes require a minor clarification and which require a new version or product retest.

Retesting may be necessary when a product changes, an instrument changes, a material defect is found, a scoring model changes materially, or evidence shows the original procedure is not measuring the intended property.

The protocol is complete when another person can reproduce the record and challenge the interpretation. It becomes trustworthy only through continued checks, transparent changes, and correction when the evidence does not hold.

[[Derek Hales on Building NapLab from an Apartment to a Testing Operation]] shows how these concerns emerged from a real review operation. [[What a Product Review Score Actually Means]] handles the next step, where observations become a total. For another physical-testing perspective, continue to [[Episode Story - Josh Sprague and the Standard Beyond Good Enough|Building Better Products by Refusing to Accept Good Enough with Josh Sprague]].

Sources, test conditions, and updates

This guide synthesizes the episode 100 transcript with NapLab's current testing methodology, the NIST gauge study handbook, NIST measurement variability guidance, and NIST Technical Note 1297.

It is a general editorial and operational framework, not an accredited laboratory standard. Teams working under a regulated, contractual, safety-critical, or consensus-standard method must follow the controlling requirements. AI assisted with research organization and drafting under editorial review.

Sources

Follow the evidence.

  1. nist.gov: nist tn 1297 appendix d4 measurand defined measurement methodnist.gov
  2. pmc.ncbi.nlm.nih.gov: PMC4055748pmc.ncbi.nlm.nih.gov
  3. linkedin.com: naplabreviewslinkedin.com
  4. naplab.com: aboutnaplab.com
  5. naplab.com: how to choose a mattressnaplab.com
  6. itl.nist.gov: mpc4itl.nist.gov
  7. itl.nist.gov: mpc114itl.nist.gov
  8. FTC Endorsement Guides questions and answersftc.gov
  9. pmc.ncbi.nlm.nih.gov: PMC6348954pmc.ncbi.nlm.nih.gov
  10. ftc.gov: consumer reviews testimonials rule questions answersftc.gov
  11. naplab.comnaplab.com
  12. naplab.com: how we test mattressesnaplab.com
  13. FTC: Endorsements, Influencers, and Reviewsftc.gov
  14. doi.org: 9789264043466 endoi.org
  15. pmc.ncbi.nlm.nih.gov: PMC12071755pmc.ncbi.nlm.nih.gov
  16. naplab.com: derek halesnaplab.com
  17. naplab.com: how do we choose best mattressesnaplab.com
  18. linkedin.com: dhaleslinkedin.com
How to Build a Repeatable Product Testing Protocol