Back to the episode map

Evergreen

How to Compare Robotaxi Safety Claims and Data

Compare robotaxi safety data by matching the event, exposure, operating domain, human role, baseline, reporting rule, severity, and uncertainty.

Aug 4, 20268 min readBy Dalton Anderson

How to Compare Robotaxi Safety Claims

Robotaxi safety claims are comparable only when they use compatible events, exposure, operating domains, human roles, baselines, reporting rules, severity definitions, and uncertainty. A raw crash count, a large mileage total, or "miles per intervention" cannot answer whether one service is safer by itself.

If the fields cannot be matched, the honest conclusion is not "equally safe." It is "not comparable from the disclosed evidence."

The short answer

Start by defining the event in the numerator. Then identify the denominator and the conditions that produced it. Match the comparison fleet to the same roads, places, times, vehicle types, operating mode, reporting threshold, and severity outcome. Examine sample size and confidence intervals before interpreting the rate.

Do not mix Level 2 driver assistance with rider-only automated driving. Do not divide public crash reports by a mileage number from a different period. Do not assume an intervention at one company means the same thing at another.

Two impressive numbers can still be incompatible

E073 compared a claim of 10,000 miles per intervention with another estimate of 444 miles per intervention. The difference sounded decisive.

The episode did not preserve the event definition, software version, fleet, exposure dates, roads, weather, traffic, operating mode, human role, or exclusions behind both numbers. One source could have counted direct safety takeovers. Another could have counted any precautionary human action. One could have measured a restricted launch domain. Another could have included consumer Level 2 use.

Dividing miles by an undefined event does not repair the definition.

The reusable lesson from E073 is not which number won. It is why the comparison failed.

Every safety claim needs a complete denominator

"Five crashes" is not a rate. Five crashes across 50,000 miles and five across 50 million miles describe different exposure. Even miles can be too broad if one fleet operates low-speed daylight routes while another carries passengers through dense nightlife districts.

Use a claim table before using the result.

FieldWhat must be disclosedCommon failure
SystemVehicle, ADS, software version, and operating modeCombining several versions or Level 2 and Level 4 use
EventExact crash, injury, contact, intervention, or disengagement definitionTreating different thresholds as the same numerator
ExposureMiles, trips, hours, passenger miles, or another denominator for the same periodPairing current incidents with lifetime mileage
DomainGeography, road class, speed, weather, lighting, traffic, and timeComparing a constrained urban fleet with a national average
Human roleSafety driver, rider-only, remote assistance, or direct remote drivingCalling all supported operation driverless
BaselineMatched human fleet or competing ADS recordChoosing the easiest human average to beat
ReportingThreshold, company knowledge, updates, duplicates, and exclusionsTreating public reports as a complete census
SeverityProperty contact, police report, airbag, alleged injury, serious injury, or fatalityCombining minor contact with severe harm
ResponsibilityAll involvement, preventable events, or legally attributed faultTreating involvement as fault
UncertaintySample size, confidence interval, and zero-event limitationPresenting a point estimate as settled truth

If the public claim cannot fill a material row, that limitation belongs beside the number.

NHTSA crash reports are evidence, not a leaderboard

NHTSA's Standing General Order page publishes incident data for named ADS and Level 2 ADAS manufacturers and operators. The third amended order became effective June 16, 2025.

The agency states its limitations plainly. ADS and Level 2 reporting thresholds differ. Companies have different telemetry and may learn about crashes at different rates. Initial reports can be incomplete. More than one entity can report the same crash. A later report version can change the record. Fleet sizes, miles, locations, and capabilities differ.

NHTSA says the summary data should not be assumed to be statistically representative of all crashes. It also warns readers to use caution when comparing entities.

The third amended order is designed to provide the agency timely notice of defined incidents. It is not designed as a complete exposure-matched company ranking.

An SGO report can help identify an event, system state, severity, location, narrative, and later update. It cannot produce a safety rate until compatible exposure and duplicate handling are added.

The human baseline has to drive in the same world

National human crash averages are attractive because they are available. They can be a poor comparison for a robotaxi fleet.

Risk changes by city, road type, time, weather, vehicle, trip purpose, and crash-reporting practice. Urban surface streets expose a service to pedestrians, cyclists, intersections, curb activity, and dense turning conflicts. Late-night ride service can operate during riskier hours than average personal driving.

A matched human baseline should represent the same locations, road classes, times, and vehicle types as the ADS exposure. It should also account for underreporting when the outcome depends on police or insurance records.

Waymo's July 2026 time-and-location research summary argues that these factors materially change the benchmark. That is a useful methodological point. The analysis is company-affiliated and should not be treated as independent proof for another operator.

Severity answers a more useful question than contact alone

Minor contact matters for operations and public confidence. It is not equivalent to a crash involving an injury, airbag deployment, vulnerable road user, or fatality.

A system could reduce severe crashes while producing more low-speed contacts. Another could avoid contact in a narrow domain without enough exposure to estimate serious-injury risk. One combined "crash rate" would hide the tradeoff.

Report separate outcomes. Any contact, police-reported crash, airbag deployment, any injury, serious injury or worse, vulnerable-road-user event, and fatality each have different reporting quality and statistical requirements.

Responsibility should also remain separate. A rider-only vehicle can be involved in a crash caused by another road user. Counting all involvement is valuable for measuring the rider's exposure to harm. It is not a finding that the ADS caused each event.

Interventions require their own dictionary

An intervention can mean direct steering takeover, braking, a safety driver's precaution, a system-requested handoff, a remote path suggestion, authorization for a maneuver, a stop request, or a simulation correction.

The measure also changes with the human role. A safety driver can intervene before a failure because the person is conservative. A rider-only system may instead reach a minimal-risk stop and request help. Counting only steering takeovers could make those operating modes look artificially different.

Before publishing miles per intervention, disclose the trigger, authority, severity, direction, duration, and whether the event was required to avoid harm. Then report the operating domain and exposure.

[[Remote Assistance Is Not Always Remote Driving]] provides the authority map. [[Operational Design Domains and Robotaxi Geofences Explained]] supplies the conditions that belong beside the event.

A better study shows the shape of a valid comparison

A 2025 Waymo-affiliated study compared 56.7 million rider-only miles with human benchmarks matched by vehicle type, road type, and location. It used crash-type groups, several severity outcomes, NHTSA incident records, company mileage, and statistical intervals.

The authors reported statistically lower rates for several outcomes and no statistically significant disadvantage in the studied crash groups.

That result is evidence about a defined Waymo operation and period. It is not proof that every robotaxi is safer, that every current Waymo route has the same result, or that the method contains no company bias. The authors were affiliated with the operator, and the exposure came from the company.

The study is valuable because the method can be inspected. Waymo's Safety Impact hub also exposes outcome definitions, current comparisons, confidence intervals, and downloadable mappings to NHTSA report IDs. A reader can examine how the claim was built instead of accepting one number.

Use a stop rule

The comparison process should end when the evidence no longer supports the next inference.

flowchart TD
    A["Define the safety question"] --> B["Match system, event, and exposure"]
    B --> C["Match domain, human role, and baseline"]
    C --> D["Check reporting, severity, and duplicates"]
    D --> E["Estimate uncertainty"]
    E --> F{"Are the fields compatible?"}
    F -->|Yes| G["Report the scoped comparison and limits"]
    F -->|No| H["Report not comparable from disclosed evidence"]

The stop rule prevents a dramatic claim from outrunning the data. It also protects a real safety improvement from being dismissed because a broader universal claim cannot yet be supported.

What a credible safety statement sounds like

A useful result names the system, period, domain, exposure, event, baseline, estimate, interval, and limitations in the same paragraph.

It does not say "robotaxis are safer than humans." It says that a defined ADS operating without an in-vehicle driver across named locations and dates had a measured rate for a defined severity outcome that was lower, higher, or statistically indistinguishable from a matched human benchmark under the study's method.

That sentence takes longer to read. It tells the reader what was actually learned.

For the service context behind the numbers, continue with [[Is It a Robotaxi Pilot or a Public Launch]]. E025 provides the longer autonomous-vehicle history, and E073 preserves the moment when two incompatible intervention claims first made this measurement problem visible.

This guide was freshly written from the preserved E073 transcript, the current NHTSA reporting record, and disclosed primary safety studies reviewed on July 28, 2026. AI assistance was used for research organization, drafting, and validation. Publication remains unauthorized.

Sources

Follow the evidence.

  1. statutes.capitol.texas.gov: IN.1954statutes.capitol.texas.gov
  2. rosap.ntl.bts.gov: 56823rosap.ntl.bts.gov
  3. nhtsa.gov: automated vehicles safetynhtsa.gov
  4. nhtsa.gov: national av safety forumnhtsa.gov
  5. statutes.capitol.texas.gov: OC.2402statutes.capitol.texas.gov
  6. rosap.ntl.bts.gov: dot 88518 DS1rosap.ntl.bts.gov
  7. open.spotify.com: 59g3xZ2WunOCIOrju8FTKSopen.spotify.com
  8. saemobilus.sae.org: j3016 202104 taxonomy definitions terms related driving automation systems road motor vehiclessaemobilus.sae.org
  9. daltonanderson.ghost.io: teslas robotaxi pilot hype vs reality in austindaltonanderson.ghost.io
  10. nhtsa.gov: standing general order crash reportingnhtsa.gov
  11. capitol.texas.gov: SB02807Fcapitol.texas.gov
  12. waymo.com: impactwaymo.com
  13. txdmv.gov: AVprogramtxdmv.gov
  14. nhtsa.gov: third amended SGO 2021 01 2025nhtsa.gov
  15. tesla.com: robotaxitesla.com
  16. tdi.texas.gov: auto insurancetdi.texas.gov
  17. arxiv.org: 2505arxiv.org
  18. rosap.ntl.bts.gov: 79800rosap.ntl.bts.gov
  19. youtu.be: K3Nj92fDh3wyoutu.be
  20. tdi.texas.gov: cb020tdi.texas.gov
  21. nhtsa.gov: automated driving systems 20 voluntary guidancenhtsa.gov
  22. waymo.com: time geo crash risk effectwaymo.com
  23. nhtsa.gov: av public meeting 2026nhtsa.gov
  24. tesla.com: TSLA Q2 2025 Updatetesla.com
  25. tesla.com: robotaxitesla.com
How to Compare Robotaxi Safety Claims and Data