Research Note
Tactile Model Evaluation Framework
Evaluation starts with the physical decision the estimate will influence. A force estimate used for offline inspection has a different failure cost from a slip estimate t
Tactile Model Evaluation Framework
Decision first
Evaluation starts with the physical decision the estimate will influence. A force estimate used for offline inspection has a different failure cost from a slip estimate that changes gripper force near a person.
The test record must state the intended task, robot state, output unit, update rate, acceptable error, unacceptable outcome, containment, and decision owner before model comparison begins.
Baseline
The baseline may be a current estimator, a calibrated analytic method, a sensor-native measurement, a simple threshold, or a no-tactile control. It must use the same contacts, timing, split, and downstream decision rule as the candidate where feasible.
Variation matrix
| Dimension | Representative variation |
|---|---|
| Contact | Position, angle, normal load, shear load, speed, duration |
| Object | Shape, stiffness, material, texture, surface finish |
| Sensor | Unit, production lot, gel age, replacement surface |
| Calibration | Fresh, aged, intentionally perturbed, failed |
| Environment | Temperature, contamination, vibration, ambient light |
| Time | Same session, later session, long run, restart |
| System | Frame loss, timestamp skew, compute contention, queue delay |
| Operator | Mounting, cleaning, object presentation, reset procedure |
Split discipline
Random frame splits are often too weak for sequential tactile data. Neighboring frames share the same contact trajectory, object, sensor state, and environment. A defensible test groups by the unit that could leak information, such as object, trajectory, session, sensor unit, or maintenance cycle.
The split should be frozen before model selection. The final test set remains untouched until the comparison and thresholds are fixed.
Metrics and consequences
Report the metric that matches the output, but also translate it into the physical decision. Force error should be stratified by axis, load, contact location, and indenter. Slip classification needs precision, recall, F1, event timing, and false-alarm behavior. Pose needs continuous error or bin interpretation, not accuracy alone.
Latency must be measured from physical exposure through usable output at the consumer. Record median, tail latency, dropped frames, stale estimates, and recovery after interruption.
Decision record
The result is go, revise, or stop. A go decision names the allowed operating envelope and residual risk. A revise decision names the failed slice and next experiment. A stop decision records the unacceptable failure and prevents quiet reuse of the same evidence.
This framework supports research evaluation. It is not a safety certification and does not authorize closed-loop control around people.
Sources
Follow the evidence.
- arxiv.org: 2206arxiv.org
- ai.meta.com: sparsh self supervised touch representations for vision based tactile sensingai.meta.com
- NIST AI Risk Management Frameworknist.gov
- arxiv.org: 1803arxiv.org
- gelsight.com: GelSight Datasheet GSMinigelsight.com
- github.com: sparshgithub.com
- open.spotify.com: 4M1AacvVLwWqI8GrQSVopmopen.spotify.com
- ai.meta.com: fair robotics open sourceai.meta.com
- arxiv.org: 2410arxiv.org
- openreview.net: forumopenreview.net
- sparsh-ssl.github.iosparsh-ssl.github.io
- daltonanderson.net: metas sparsh a new era for robotic touch sensingdaltonanderson.net
- youtu.be: psjHxZL1j0wyoutu.be
- daltonanderson.ghost.io: metas sparsh a new era for robotic touch sensingdaltonanderson.ghost.io
- arxiv.org: 2005arxiv.org