Skip to content
Home ↗
Slide 10 our approach

Our approach: translate in a discrete token space

Three stages. The design point: the tokenizer's reconstruction ceiling is 0.214 LPIPS, far better than any method's translation error, so representation is not the bottleneck. Prediction is.

Interactive · try it hereOpen standalone ↗
Token Translation, Stage by Stage
Step through the three stages. Blue blocks are trained at that stage, grey ones are frozen. Drag the dot in the lattice to feel what FSQ does.
Stage diagram
FSQ lattice: continuous z → nearest lattice point
Finite scalar quantization rounds each latent channel to one of a few fixed levels. There is no learned codebook. The frozen target decoder maps the predicted tokens into the target modality. The tokenizer's own reconstruction ceiling is 0.214 LPIPS, far below any method's translation error: representation is not the bottleneck, prediction is.