Unified Evaluation of Text-to-Image Models
My final-year project asked a practical question: given a text prompt, which text-to-image model should you use? I compared FLUX.1, OmniGen and Stable Diffusion v1.4, and sorted prompts into simple, abstract and technical ones. The plan was an online-learning recommender: generate images, score them, label each prompt with the model that did best, train a text classifier on the prompt alone, and keep fine-tuning as new samples arrive.
The recommender did not reach practical performance. On 7,000 samples, validation accuracy stayed around 40 percent while training accuracy went to nearly 100, only a little better than picking one of the three models at random. All three are strong general-purpose models, and for a simple prompt such as a bike parked under a tree they produce nearly the same picture, so there is little for a classifier to learn.
The more useful result was about the labels. Each prompt was first labelled with a weighted mix of CLIP, BLIP and ImageReward scores. These metrics agree with each other far less than I expected, with correlations between 0.35 and 0.62, and BLIP and ImageReward carry subjective judgments of their own. Mixing them made the supervision inconsistent, and labelling with a single metric worked better. Next time I would check how far the metrics agree before combining them, and test on harder prompts and on models with more distinct styles.