Article
Test Products at Real Duration and Intensity
A practical framework for testing products under the duration, intensity, variation, intervention, failure, recovery, and support conditions of real use.
Products Must Be Tested at Real Duration and Intensity
A product is not proven when it works once under presentation conditions. It earns trust when it keeps working at the duration, intensity, variation, and consequence level of the use being claimed.
That standard applies to software, physical products, services, AI systems, and operations. The exact test changes by domain, but the central question stays the same: does the evidence resemble the moment when a real user depends on the product?
flowchart TD
A["Visible successful demo"] --> B["Extend duration"]
B --> C["Raise real-use intensity"]
C --> D["Add users, objects, and environments"]
D --> E["Record intervention and failure"]
E --> F["Test recovery and support"]
F --> G["Make a bounded readiness decision"]
Short tests hide time-dependent failures
Many failures need time to appear. Heat accumulates. Batteries drain. materials shift. software state drifts. A user becomes tired. A model encounters a longer context. A support team receives several incidents at once.
A five-minute result can be honest and still be a poor predictor of an eight-hour shift or a year of ownership. Duration should match the actual unit of dependence. That unit might be one transaction, one commute, one clinical procedure, one work shift, one endurance event, or repeated daily use.
E119 makes this tangible through product testing for endurance athletes. A fit that feels fine during a short run can create pressure, heat, movement, and access problems over a longer distance. E008 presents the same problem through robot and agent demonstrations. A short selected sequence does not reveal drift, intervention, maintenance, or support.
Intensity changes the system
Intensity is not only physical force. It can mean transaction volume, request concurrency, inference load, object variability, environmental noise, user stress, speed, precision, or the consequence of a mistake.
The correct level comes from the use case. A robot moving an empty box in a closed test area faces a different task from one handling irregular objects near workers. A software assistant answering one question faces a different system load from one coordinating thousands of tasks with external tools.
Testing at low intensity can validate a component. It cannot justify a claim about a higher-intensity operating state unless the relationship has been measured.
Variation exposes brittle success
Products often pass because the demonstration conditions are stable. The starting pose is fixed. The lighting is favorable. The data is clean. The user follows the expected path. The object is familiar.
Real use adds different people, language, devices, objects, surfaces, network conditions, locations, abilities, and starting states. Variation should be chosen from the intended operating environment, not added as random noise for appearance.
The goal is not to make a test impossibly broad. It is to define the boundary honestly. A narrow product can be useful when the supported conditions are explicit.
Intervention belongs in the result
A system can look autonomous while depending on substantial setup, monitoring, resets, approvals, or rescue. Those actions are part of the operating cost and safety case.
Record who prepared the environment, selected the run, corrected the input, reset the system, approved the action, repaired the hardware, or handled an exception. A human intervention does not invalidate the product. Hiding it makes the evidence misleading.
NIST's AI Resource Center connects AI risk management with testing, evaluation, verification, and validation resources. That framing is useful because it treats evaluation as an operating discipline, not a single benchmark screenshot.
Failure and recovery are product behavior
A readiness test should define failure before the trial begins. Otherwise, teams can reinterpret an incomplete or unsafe outcome after seeing it.
The record should capture what happened, how the system detected the problem, which safeguard activated, whether a person could intervene, how service resumed, and whether the same condition can recur. Near misses matter when the consequence of a later failure is high.
OSHA's robotics overview notes that many robot accidents occur during non-routine conditions such as maintenance, programming, testing, setup, or adjustment. Those moments show why recovery and support cannot be separated from normal product performance.
Verification and validation answer different questions
Verification asks whether the system meets specified requirements. Validation asks whether the resulting system serves its intended purpose in the operational environment.
The NASA Systems Engineering Handbook gives formal treatment to both processes. A product can satisfy an internal specification and still fail its user's actual job. It can also delight in one trial while violating a critical requirement.
Teams need both views. Requirements make the test accountable. Operational validation keeps the requirements connected to reality.
Support load changes readiness
A product that works only with its builders present may be a strong prototype. It is not yet the same product that an ordinary operator can run, maintain, understand, and recover.
Measure setup time, training, calibration, maintenance, updates, monitoring, incident handling, replacement parts, support response, and the authority needed to stop use. Include accessibility and the ability to understand system state. A hidden support burden can erase the apparent gain.
The NIOSH Center for Occupational Robotics Research studies injury trends, risk profiles, human-robot interaction, standards, guidance, and training. Its scope reinforces that a robot's effect is shaped by the workplace system around it.
Build the real-condition test
Start with one claim. Name the user, task, environment, duration, intensity, variation, success measure, failure threshold, intervention boundary, recovery expectation, and support model. Then identify what decision the result can support.
Run the smallest test that resembles the actual point of dependence. Preserve failed runs and configuration changes. Repeat after material changes to hardware, software, models, data, tools, environment, or safety controls.
The result does not have to be perfect. It has to be honest enough to guide the next decision.
Connect the evidence across episodes
E008 supplies the demonstration problem. E010 turns it into an evidence ladder for high-impact technology. E025 adds operational design domains. E073 shows why a pilot boundary matters for autonomous vehicles. E117 applies the question to a high-force actuator. E119 brings the principle back to ordinary product fit and long-duration use.
This essay was developed with AI assistance from the preserved E008 transcript and the linked NIST, NASA, OSHA, and NIOSH records. It is a general evaluation framework, not a substitute for regulated testing, safety engineering, clinical validation, security assessment, accessibility review, certification, procurement diligence, or domain expertise. Publication is unauthorized.
Sources
Follow the evidence.
- NVIDIA Rubin announcementnvidianews.nvidia.com
- SIMA 2 technical reportstorage.googleapis.com
- NIOSH Center for Occupational Robotics Researchcdc.gov
- OSHA robotics overviewosha.gov
- NIST AI Resource Centerairc.nist.gov
- Google DeepMind SIMA 2 announcementdeepmind.google
- NVIDIA Blackwell Ultra announcementnvidianews.nvidia.com
- BMW Figure 02 trialpress.bmwgroup.com
- Spotify episode recordpodcasters.spotify.com
- Google DeepMind SIMA announcementdeepmind.google
- Figure news indexfigure.ai
- Figure 03 introductionfigure.ai
- NASA Systems Engineering Handbooknasa.gov
- Figure Helix 02figure.ai
- NVIDIA Blackwell launchinvestor.nvidia.com
- BMW Figure 03 projectpress.bmwgroup.com
- SIMA technical reportstorage.googleapis.com
- Figure company pagefigure.ai