How do scalar grounding losses shape RepTok decoding?

Question 1: how do the validation metrics agree with each other?

Metric correlation heatmap LPIPS and CLIP against latent MSE

Question 2: how does each objective correlate to each metric?

Objective coefficient effects on metrics

Question 3: how does each learned coefficient's level correlate to each metric?

Learned coefficient correlations with metrics

Question 4: what do pure-objective coefficient runs learn?

Pure objective learned scalar coefficients Pure objective qualitative RepTok decodes

Question 5: which loss recipes rank best by latent MSE?

Best latent MSE ranked loss coefficients

Question 6: which loss recipes rank best by LPIPS?

Best LPIPS ranked loss coefficients

Question 7: which loss recipes rank best by CLIP?

Best CLIP ranked loss coefficients

Question 8: how well does learned fusion actually work?

Best learned fusion against equal fusion

Question 9: does Bayesian optimization directly on LPIPS beat equal fusion?

1000-image Bayesian optimization held-out metrics 1000-image Bayesian optimization coefficients bo1000_bge_bo_trace.jpgbo1000_sonar_bo_trace.jpg