GW6 5-Modality Scaling

ImageNet-only view of the three directions we care about: DINOv2 to SONAR BLEU, BGE to RepTok image metrics, and DINOv2 to symbolic classification accuracy. BigGAN rows are excluded.

Missing in current eval files: clip, fid.

DINOv2 -> SONAR BLEU, higher is better 9.940 10.57 11.20 11.82 12.45 10.53 32M 11.02 51M 11.86 100M
BGE -> RepTok LPIPS, lower is better 0.6117 0.6271 0.6425 0.6578 0.6732 0.6501 32M 0.6401 51M 0.6348 100M
BGE -> RepTok CLIP similarity, higher is better not run in these eval outputs
BGE -> RepTok FID, lower is better not run in these eval outputs
DINOv2 -> Symbolic classification accuracy, higher is better 0.7703 0.7977 0.8250 0.8523 0.8797 0.8450 32M 0.8500 51M 0.8000 100M