Research Note
Point Tracking Evaluation Framework
An evaluation must state the tracker, checkpoint hash, code revision, model mode, query policy, frame resolution, preprocessing, window length, hardware, metric implement
In this article
Point Tracking Evaluation Framework
Evaluation contract
An evaluation must state the tracker, checkpoint hash, code revision, model mode, query policy, frame resolution, preprocessing, window length, hardware, metric implementation, dataset version, and split.
Scores produced with different query modes, resolutions, support points, visibility rules, or multi-point grouping are not automatically comparable.
Evidence layers
| Layer | Required evidence |
|---|---|
| Benchmark reproduction | Official dataset, metric code, declared query mode, expected range |
| Representative set | Workflow videos, hard conditions, held-out sequences, rights |
| Localization | Distance thresholds and visible-point error |
| Visibility | Occlusion classification and false visibility costs |
| Continuity | Re-entry, long occlusion, drift, fragmentation |
| Runtime | End-to-end latency, throughput, memory, initialization, queue behavior |
| Workflow | Downstream error, human correction, abstention, and harm |
TAP-Vid context
TAP-Vid includes real and synthetic videos, human and perfect synthetic annotations, and metrics for tracking arbitrary points. It is a strong common reference, not a substitute for workflow evidence.
Average Jaccard combines localization and occlusion behavior across thresholds. Position accuracy measures visible-point localization at declared thresholds. Occlusion accuracy measures visibility classification. The official implementation should define exact calculations.
Failure reel
Every aggregate score should link to sampled failure clips. The reel should cover fast and abrupt motion, long occlusion, re-entry, camera motion, blur, deformation, lighting change, repeated texture, featureless regions, reflective surfaces, crowded scenes, and long sequences.
Each clip needs query points, ground truth, prediction, visibility state, error category, severity, and workflow consequence.
Decision rule
The model should pass declared benchmark-reproduction tolerance, representative-set thresholds, runtime budgets, and high-severity error limits. A high mean score cannot compensate for an unacceptable failure in the region that drives the workflow.
Boundary
High-stakes domains require independent domain evidence and safety review. This framework does not validate a medical, autonomous, surveillance, labor, or other consequential use.
Sources
Follow the evidence.
- youtu.be: BNTcjZ0Ym38youtu.be
- ai.meta.com: sam2ai.meta.com
- proceedings.neurips.cc: 58168e8a92994655d6da3939e7cc0918 Abstract Datasets and Benchmarksproceedings.neurips.cc
- arxiv.org: 2410arxiv.org
- open.spotify.com: 26JgnnwjvK5vYIdRofV8ntopen.spotify.com
- github.com: co trackergithub.com
- cotracker3.github.iocotracker3.github.io
- NIST AI Risk Management Frameworknist.gov
- vggsfm.github.iovggsfm.github.io
- daltonanderson.ghost.io: metas cotracker 3 a leap in ai object trackingdaltonanderson.ghost.io
- arxiv.org: 1803arxiv.org
- ecva.net: 3526 ECCV 2020 paperecva.net
- github.com: tapnetgithub.com
- raw.githubusercontent.com: LICENSEraw.githubusercontent.com
- tapvid.github.iotapvid.github.io
- NIST Privacy Frameworknist.gov
- arxiv.org: 1504arxiv.org
- openaccess.thecvf.com: Karaev CoTracker3 Simpler and Better Point Tracking by Pseudo Labelling Real Videos ICCV 2025 paperopenaccess.thecvf.com