Who got close?
Ten drawings per model.
Higher scores mean closer to Bufo.
| Rank / model | USD / image | Bufo Score |
|---|
About these prices
“Measured run” is what the API charged us, including the reference image when billed. Prices marked ≈ come from three fresh samples because the old runs didn’t save their bills. Microsoft returned no images, so we show its published token rates.
“Good value” marks a useful tradeoff: no other listed model offers a higher score for the same price, or the same score or better for less. These are September 2026 prices, and some are estimates. They aren’t a promise of how many good Bufos your dollar will buy.
The score measures likeness, not the percentage of good frogs. How it works
Pick your frog.
Both got the same prompt. Pick the more convincing Bufo, then see who made it.
Your picks stay in this browser. They don’t change the leaderboard.
The family resemblance varies.
How do you score a Bufo?
Bufo can hold a coffee. Bufo can be the coffee. The good ones feel like someone had a joke and opened an image editor. The familiar face helps, but a rough redraw can work too.
Each ranked model got the same ten prompts and the original image. We average its ten scores to get a Bufo Score out of 100. Missing images mean no rank. The scorer learned from human labels; your picks don’t change it.
The scorer can miss the joke. It can also miss whether Bufo actually brought the coffee we asked for. Small score gaps don’t settle much, so look at the drawings too.
Download scoring method