Back to the episode map

Research Note

Tactile Model Evaluation Framework

Evaluation starts with the physical decision the estimate will influence. A force estimate used for offline inspection has a different failure cost from a slip estimate t

Aug 4, 20262 min readBy Dalton Anderson

Tactile Model Evaluation Framework

Decision first

Evaluation starts with the physical decision the estimate will influence. A force estimate used for offline inspection has a different failure cost from a slip estimate that changes gripper force near a person.

The test record must state the intended task, robot state, output unit, update rate, acceptable error, unacceptable outcome, containment, and decision owner before model comparison begins.

Baseline

The baseline may be a current estimator, a calibrated analytic method, a sensor-native measurement, a simple threshold, or a no-tactile control. It must use the same contacts, timing, split, and downstream decision rule as the candidate where feasible.

Variation matrix

DimensionRepresentative variation
ContactPosition, angle, normal load, shear load, speed, duration
ObjectShape, stiffness, material, texture, surface finish
SensorUnit, production lot, gel age, replacement surface
CalibrationFresh, aged, intentionally perturbed, failed
EnvironmentTemperature, contamination, vibration, ambient light
TimeSame session, later session, long run, restart
SystemFrame loss, timestamp skew, compute contention, queue delay
OperatorMounting, cleaning, object presentation, reset procedure

Split discipline

Random frame splits are often too weak for sequential tactile data. Neighboring frames share the same contact trajectory, object, sensor state, and environment. A defensible test groups by the unit that could leak information, such as object, trajectory, session, sensor unit, or maintenance cycle.

The split should be frozen before model selection. The final test set remains untouched until the comparison and thresholds are fixed.

Metrics and consequences

Report the metric that matches the output, but also translate it into the physical decision. Force error should be stratified by axis, load, contact location, and indenter. Slip classification needs precision, recall, F1, event timing, and false-alarm behavior. Pose needs continuous error or bin interpretation, not accuracy alone.

Latency must be measured from physical exposure through usable output at the consumer. Record median, tail latency, dropped frames, stale estimates, and recovery after interruption.

Decision record

The result is go, revise, or stop. A go decision names the allowed operating envelope and residual risk. A revise decision names the failed slice and next experiment. A stop decision records the unacceptable failure and prevents quiet reuse of the same evidence.

This framework supports research evaluation. It is not a safety certification and does not authorize closed-loop control around people.

Sources

Follow the evidence.

  1. arxiv.org: 2206arxiv.org
  2. ai.meta.com: sparsh self supervised touch representations for vision based tactile sensingai.meta.com
  3. NIST AI Risk Management Frameworknist.gov
  4. arxiv.org: 1803arxiv.org
  5. gelsight.com: GelSight Datasheet GSMinigelsight.com
  6. github.com: sparshgithub.com
  7. open.spotify.com: 4M1AacvVLwWqI8GrQSVopmopen.spotify.com
  8. ai.meta.com: fair robotics open sourceai.meta.com
  9. arxiv.org: 2410arxiv.org
  10. openreview.net: forumopenreview.net
  11. sparsh-ssl.github.iosparsh-ssl.github.io
  12. daltonanderson.net: metas sparsh a new era for robotic touch sensingdaltonanderson.net
  13. youtu.be: psjHxZL1j0wyoutu.be
  14. daltonanderson.ghost.io: metas sparsh a new era for robotic touch sensingdaltonanderson.ghost.io
  15. arxiv.org: 2005arxiv.org
Tactile Model Evaluation Framework