Research Note
TacBench Task and Result Record
TacBench tests whether a pretrained tactile representation can support several downstream tasks across DIGIT, GelSight 2017, and GelSight Mini data. It uses task-specific
TacBench Task and Result Record
Benchmark purpose
TacBench tests whether a pretrained tactile representation can support several downstream tasks across DIGIT, GelSight 2017, and GelSight Mini data. It uses task-specific labels and metrics, so its aggregate is a summary across unlike problems rather than one universal score.
Task map
| Task | Sensor and data | Output | Reported metric |
|---|---|---|---|
| T1 force estimation | DIGIT and GelSight Mini, 75,000 samples for each sensor | Three-axis normal and shear force | Average RMSE |
| T1A force-field visualization | DIGIT and GelSight Mini images | Dense normal and shear fields | Qualitative visualization |
| T2 slip detection | DIGIT, 125,000 samples with 13 percent slip | Slip class and force change | F1 score |
| T3 pose estimation | DIGIT, 49,000 samples | Relative SE(2) pose bins | Multiclass accuracy |
| T4 grasp stability | GelSight 2017, about 9,300 grasp trials | Success or failure | Accuracy |
| T5 textile recognition | GelSight 2017, 4,467 clips and 20 classes | Textile class | Accuracy |
| T6 bead maze | DIGIT, 50 demonstrations and about 34,000 pairs | Joint-angle trajectory | Position error and distance before failure |
The paper uses "six tasks" for T1 through T6 and treats T1A as an additional force-field visualization. A public explanation should retain that naming.
Frozen evaluation
Most tasks freeze the Sparsh encoder and train a small attentive decoder with labels. The end-to-end baseline uses an encoder and decoder with comparable capacity initialized from random weights and trained for the task.
The standard headline uses 33 percent of labeled data for T1 through T5 and 50 percent of demonstrations for T6. The point of the comparison is data efficiency and representation reuse, not elimination of supervision.
Metric differences
Lower is better for force RMSE and bead-maze position error. Higher is better for slip F1 and the accuracy measures. T1A is qualitative and does not enter the six-row aggregate.
Slip uses F1 because only 13 percent of the samples are labeled as slip. Pose accuracy comes from discretized relative translation and rotation bins, not continuous pose error. Grasp stability uses a randomized split because the source dataset did not provide an official split. These choices affect interpretation.
Aggregate discrepancy
The abstract, introduction, and Meta announcement state an average improvement of 95.1 percent over task and sensor-specific end-to-end training. Appendix D Table 13 lists per-row improvements of 28.31, 59.74, 242.70, 235.89, 5.14, and 19.72 percent and displays a 98.75 percent average.
These two published figures do not match. Venture Step should repeat the 95.1 percent headline only with attribution and a note that the appendix reports 98.75 percent for its six-row summary. No new composite should be calculated without a declared method and technical review.
Failure and transfer boundaries
The benchmark does not establish performance after sensor replacement, gel wear, contamination, temperature change, different contact geometry, new object families, different collection timing, or a new robot.
The paper reports that none of the models completed the full bead maze in real rollouts. It also documents a slip-label failure case caused by an imperfect friction boundary. These are useful reminders that ground truth and physical execution can fail independently of representation quality.
Sources
Follow the evidence.
- arxiv.org: 2206arxiv.org
- ai.meta.com: sparsh self supervised touch representations for vision based tactile sensingai.meta.com
- NIST AI Risk Management Frameworknist.gov
- arxiv.org: 1803arxiv.org
- gelsight.com: GelSight Datasheet GSMinigelsight.com
- github.com: sparshgithub.com
- open.spotify.com: 4M1AacvVLwWqI8GrQSVopmopen.spotify.com
- ai.meta.com: fair robotics open sourceai.meta.com
- arxiv.org: 2410arxiv.org
- openreview.net: forumopenreview.net
- sparsh-ssl.github.iosparsh-ssl.github.io
- daltonanderson.net: metas sparsh a new era for robotic touch sensingdaltonanderson.net
- youtu.be: psjHxZL1j0wyoutu.be
- daltonanderson.ghost.io: metas sparsh a new era for robotic touch sensingdaltonanderson.ghost.io
- arxiv.org: 2005arxiv.org