RepTok Prior vs GW Direct BGE On COCO

COCO val2017 captions, n=1000. The grounded GW rows use the ImageNet-only GW checkpoints plus their short_T5_keep 5-step grounding modules. The 12h row is the activated-start 51M GW run.

GW direct BGE

37.8M
LPIPS ↓
0.7015
CLIP ↑
0.2864
FID ↓
77.60

GW short_T5 grounded

38.3M
LPIPS ↓
0.6959
CLIP ↑
0.2888
FID ↓
67.09

GW 12h direct BGE

51.0M
LPIPS ↓
0.6972
CLIP ↑
0.2864
FID ↓
75.61

GW 12h activated short_T5

51.5M
LPIPS ↓
0.6953
CLIP ↑
0.2916
FID ↓
68.93

RepTok prior 796M

796.4M
LPIPS ↓
0.7215
CLIP ↑
0.3034
FID ↓
65.89

RepTok prior 37M

36.8M
LPIPS ↓
0.7332
CLIP ↑
0.2774
FID ↓
81.88

RepTok prior BGE memory

850.1M
LPIPS ↓
0.7245
CLIP ↑
0.2903
FID ↓
70.26

Metric Barplots

COCO metric barplots including 12h GW direct and grounded rows

Parameter Bubble Plot

Parameter count versus CLIP similarity bubble plot including 12h GW rows

Side-by-Side Generations

COCO reference, GW direct BGE, grounded GW, RepTok priors, and captions

12h Activated-Start GW Samples

COCO reference, 12h GW direct BGE, 12h grounded GW, and captions

RepTok Prior Breakdown

image/time/text projections14.8M
latent transformer blocks780.6M
learned helper/null/position tokens0.1M
output head1.0M
RepTok prior 37M: image/time/text projections2.2M
RepTok prior 37M: latent transformer blocks34.2M
RepTok prior 37M: learned helper/null/position tokens0.0M
RepTok prior 37M: output head0.3M
RepTok prior BGE memory: image/time/text projections14.1M
RepTok prior BGE memory: latent transformer blocks780.6M
RepTok prior BGE memory: learned helper/null/position tokens0.0M
RepTok prior BGE memory: output head1.0M

metrics.json · summary.json · metrics_with_12h.csv · metrics.csv · metrics_grounded_short_t5.csv · metrics_12h_activated_grounded.csv · metrics_37m.csv · metrics_bge_memory.csv · model_size_metrics_with_12h.csv · model_size_metrics.csv · per_sample_metrics.csv · per_sample_metrics_12h_activated_grounded.csv · manifest_public.jsonl · manifest_public_12h_activated_grounded.jsonl