Behind the Pooled Table
The slide shows the pooled Table 4. Tables 2 and 3 of the paper give the same within-context Spearman ρ for every architecture and encoder state, in appearance space and in motion space; the supplement repeats them with the from-scratch configurations.
positive ρ: the distance predicts transfer (darker = stronger)negative ρ: anti-correlatedbold = best per cell, as marked in the papergreen box = coverage d(T→S), our claim
Read Table 3 row by row: coverage d(T→S) is the best or tied-best motion predictor in nearly every configuration, on real-motion and semantic targets alike, and the whole appearance table (Table 2) sits near zero or below it with DINOv3 representing appearance and the same estimator used in both spaces. Every column of these tables is one context of the 684 measurements you can explore below.
The 684 Measurements
Every trained model on every benchmark. Pick an architecture and encoder state: the grid is peak PCK; the plot beside it is transfer against a distance for one target, a stratum, or all nine, with the Spearman ρ.
hover a point for its source
Rows are training sources, columns are target benchmarks, cells are peak PCK@5% (the best checkpoint). Use the target buttons, or click a column header, to choose what the plot shows; "All 9 targets" overlays every context with ρ computed per target and averaged, which is how the paper pools it. Coverage ρ is computed within a (target, architecture, encoder-state) context, so no benchmark or model effect can inflate it; the pooled values reproduce the slide's table. Dashes mark models that were not trained (RAFT was not trained on SPair; FlowFormer-frozen skips it).