Skip to content
01 · Introduction/Evidence · 4
Home ↗
Slide 4 the supervision problem

Supervised matching needs pairs that barely exist

Dense ground truth is close to impossible to annotate by hand, so the field renders its training data, ever more realistically.

Every generation chases appearance. The motion is sampled, never designed. Is realism actually what makes the data transfer?

Each step improves how the data looks: exact but flat 2-D warps, then rendered 3-D objects with real parallax and occlusion, then physically simulated scenes with scanned assets and path-traced light, then procedural worlds with no asset library at all. What no step chose is how the data moves. The displacement statistics of every one of these sources are an accident of its scene sampler, and that accident is what this dissertation measures and then manufactures.

Interactive · try it hereOpen standalone ↗
Four Generations of Rendered Supervision
Step along the lineage. Each generation chose its appearance more carefully than the last; none of them chose its motion. The fingerprint column is the motion that fell out.
What the designer chose
    What was left to chance
      Why render at all: how scarce real correspondence labels arenumbers
      200
      KITTI-2015 training pairs with flow ground truth, derived from LiDAR and a fitted 3-D model, and sparse even then
      1,351
      PF-PASCAL image pairs, each with a handful of hand-clicked keypoints
      400
      TSS pairs with dense semantic flow, the largest dense semantic set
      40,302
      FlyingThings pairs in this study's pool, every pixel labelled, rendered for free
      1.28 M
      ImageNet 2-D-warp pairs in the pool: a photograph plus a random warp, exact by construction

      Dense ground truth cannot be clicked by hand: a 512 × 512 pair has 262,144 correspondences. Real benchmarks get theirs from sensors (LiDAR, hidden textures, tracked markers) or from a few keypoints per image, which is why every training-scale correspondence dataset is rendered. Pool sizes are the training sources of Chapter 3 (paper Table, supplement); benchmark sizes are the published dataset counts.

      Read the lineage left to right and the appearance column improves at every step: cut-outs, then rendered 3-D objects, then physically simulated photoreal scenes, then endless procedural worlds. The motion column never improves because nothing in any of these pipelines ever chose it; the displacement statistics are whatever the sampled camera and objects happened to produce. This dissertation asks whether that unchosen quantity, not the realism, is what decides transfer (Chapter 3) and then makes it the designed quantity (Chapter 4). Infinigen appears as the current ceiling of synthetic rendering; it is not a correspondence dataset and is not trained on here.