Skip to content
Home ↗
Slide 17 the transfer study

The transfer study: 684 measurements

Train everything on everything; score each source ranking within a fixed (target, architecture, configuration) context.

4
architectures: CATs++, GLU-Net, FlowFormer, RAFT, each with a frozen and a fine-tuned backbone → 8 configurations, 76 trained models
11
training sources: the SDF-Fractals3D family (base, +2-D warp, +large-zoom, +small-zoom, +flip), FlyingThings, PointOdyssey, Sintel, Kubric MOVi-F, ImageNet-2D, SPair
9
targets: KITTI-2012, KITTI-2015, FlyingThings, PointOdyssey, SDF-Fractals3D; SPair-71k, PF-PASCAL, PF-Willow, TSS

Transfer = peak PCK@5%. Rankings scored by Spearman ρ within each (target, architecture, encoder-state) context, so no benchmark or model effect can inflate the correlation. Every distance (Chamfer, sliced W₂, FID, and the two directed halves) is computed in both representations, motion space and DINOv3 appearance space, with the same estimator.

Interactive · try it hereOpen standalone ↗
Behind the Pooled Table
The slide shows the pooled Table 4. Tables 2 and 3 of the paper give the same within-context Spearman ρ for every architecture and encoder state, in appearance space and in motion space; the supplement repeats them with the from-scratch configurations.
positive ρ: the distance predicts transfer (darker = stronger)negative ρ: anti-correlatedbold = best per cell, as marked in the papergreen box = coverage d(T→S), our claim

Read Table 3 row by row: coverage d(T→S) is the best or tied-best motion predictor in nearly every configuration, on real-motion and semantic targets alike, and the whole appearance table (Table 2) sits near zero or below it with DINOv3 representing appearance and the same estimator used in both spaces. Every column of these tables is one context of the 684 measurements you can explore below.
Interactive · try it hereOpen standalone ↗
The 684 Measurements
Every trained model on every benchmark. Pick an architecture and encoder state: the grid is peak PCK; the plot beside it is transfer against a distance for one target, a stratum, or all nine, with the Spearman ρ.
hover a point for its source
Rows are training sources, columns are target benchmarks, cells are peak PCK@5% (the best checkpoint). Use the target buttons, or click a column header, to choose what the plot shows; "All 9 targets" overlays every context with ρ computed per target and averaged, which is how the paper pools it. Coverage ρ is computed within a (target, architecture, encoder-state) context, so no benchmark or model effect can inflate it; the pooled values reproduce the slide's table. Dashes mark models that were not trained (RAFT was not trained on SPair; FlowFormer-frozen skips it).