back to home

Unified Evaluation of Text-to-Image Models

Text-to-Image Evaluation Metrics Online Learning

My final-year project asked a practical question: given a text prompt, which text-to-image model should you use? I compared FLUX.1, OmniGen and Stable Diffusion v1.4, and sorted prompts into simple, abstract and technical ones. The plan was an online-learning recommender: generate images, score them, label each prompt with the model that did best, train a text classifier on the prompt alone, and keep fine-tuning as new samples arrive.

The recommender did not reach practical performance. On 7,000 samples, validation accuracy stayed around 40 percent while training accuracy went to nearly 100, only a little better than picking one of the three models at random. All three are strong general-purpose models, and for a simple prompt such as a bike parked under a tree they produce nearly the same picture, so there is little for a classifier to learn.

Training and validation loss and accuracy over epochs
Training accuracy climbs while validation stays flat.

The more useful result was about the labels. Each prompt was first labelled with a weighted mix of CLIP, BLIP and ImageReward scores. These metrics agree with each other far less than I expected, with correlations between 0.35 and 0.62, and BLIP and ImageReward carry subjective judgments of their own. Mixing them made the supervision inconsistent, and labelling with a single metric worked better. Next time I would check how far the metrics agree before combining them, and test on harder prompts and on models with more distinct styles.

Correlation matrix of CLIP, BLIP, ImageReward and the mixed label
How far the three scoring metrics agree with each other and with the mixed label.