Skip to content
Home ↗
Slide 11 results

Cross-collection translation leaves a larger gap

Every submission rescored from its own outputs under one frozen suite: LPIPS ↓ · KID ↓ · scene-retrieval rank ↓ (0.5 = chance).

0.02
RGB→IR retrieval: scene solved, every method alike
0.22
SAR→IR retrieval: best of any method
SAR→IR (hard)LPIPS ↓KID ↓Retr ↓
Ours, token student0.5710.1380.22
USTC-IAT (winner)0.6130.1370.30
NJUST-KMG (winner)0.6400.4220.48
pix2pixHD baseline0.6470.2760.41

Green = ours, bold = best in column, lower is better.

Plentiful pairs: nearly solved, every method alike. Scarce pairs: all degrade, differing only in how.

Interactive · try it hereOpen standalone ↗
Every Method, Every Sample
Pick a direction and step through the test set. Tap a tile to enlarge it; the row of numbers is the frozen suite that rescored every submission from its own outputs.

Where aligned pairs are plentiful (RGB→IR, same collection) the task is essentially solved and every method looks alike. Where they are scarce (SAR→IR, SAR→RGB, cross collection) every method degrades, differing only in how: the GAN family averages texture away into mush, our discrete-token student commits to one, and neither escapes the band. "Oracle" is the target tokenizer reconstructing the ground truth itself: the ceiling any token method can reach, far below every translation error, so representation is not the bottleneck.