Four Generations of Rendered Supervision
Step along the lineage. Each generation chose its appearance more carefully than the last; none of them chose its motion. The fingerprint column is the motion that fell out.
What the designer chose
What was left to chance
Why render at all: how scarce real correspondence labels arenumbers
200
KITTI-2015 training pairs with flow ground truth, derived from LiDAR and a fitted 3-D model, and sparse even then
1,351
PF-PASCAL image pairs, each with a handful of hand-clicked keypoints
400
TSS pairs with dense semantic flow, the largest dense semantic set
40,302
FlyingThings pairs in this study's pool, every pixel labelled, rendered for free
1.28 M
ImageNet 2-D-warp pairs in the pool: a photograph plus a random warp, exact by construction
Dense ground truth cannot be clicked by hand: a 512 × 512 pair has 262,144 correspondences. Real benchmarks get theirs from sensors (LiDAR, hidden textures, tracked markers) or from a few keypoints per image, which is why every training-scale correspondence dataset is rendered. Pool sizes are the training sources of Chapter 3 (paper Table, supplement); benchmark sizes are the published dataset counts.
Read the lineage left to right and the appearance column improves at every step: cut-outs, then rendered 3-D objects, then physically simulated photoreal scenes, then endless procedural worlds. The motion column never improves because nothing in any of these pipelines ever chose it; the displacement statistics are whatever the sampled camera and objects happened to produce. This dissertation asks whether that unchosen quantity, not the realism, is what decides transfer (Chapter 3) and then makes it the designed quantity (Chapter 4). Infinigen appears as the current ceiling of synthetic rendering; it is not a correspondence dataset and is not trained on here.



