My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."(frogs.vaguespac.es) |
My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw."(frogs.vaguespac.es) |
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
edit: nevermind, definitely "monkey-puppet side-eye" vibe.
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
That looks like something from Machinarium or Robots :)
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
That's a pretty good benchmark
Would've wanted to see also DS4 flash.
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.