GPT-6 Astra on OpenRouter(openrouter.ai) |
GPT-6 Astra on OpenRouter(openrouter.ai) |
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
Treat it like a bit as is
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
Coz who knows if astra low will produce max like output if tried once more.
Reminded me of https://clocks.brianmoore.com/
I'm wondering if this is being trained on by the models today.
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
Isnt Netherlands the leader in bike riders and they dont wear helmets.
Apparently the crowd agrees because they keep upvoting these.
Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...
And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.
Compared to what?
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
Edit: nevermind it JUST gave me a notification to use it!
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
Edit:
GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens
GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens
So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive.
Why is the burden of proof on me tho!?
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
I'm a Business plan user with Cyber verification enabled, FWIW.
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
but again, seems like there's no word from Azure if this applies to them.
very confusing.
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
Sure seems like Astra is expensive AF.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
>Playing ping pong from the side of the table
I think marketing might be getting a bit absurd at this point
I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
Luna — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 7.83 | 1.57 |
| xhigh | 4.24 | 0.85 |
| high | 2.46 | 0.49 |
| medium | 1.26 | 0.25 |
| low | 0.76 | 0.15 |
| none | 0.71 | 0.14 |
+--------+--------+---------+
Per million tokens:
Before: $1 input / $6 output
After: $0.20 input / $1.20 output
Sol — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 48.55 | 32.37 |
| xhigh | 24.11 | 16.08 |
| high | 10.38 | 6.92 |
| medium | 10.55 | 7.03 |
| low | 8.33 | 5.55 |
| none | 5.90 | 3.93 |
+--------+--------+---------+
Per million tokens:
Before: $5 input / $30 output
After: $4 input / $20 output
Terra — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 32.09 | 25.67 |
| xhigh | 14.67 | 11.74 |
| high | 3.74 | 2.99 |
| medium | 3.46 | 2.77 |
| low | 3.47 | 2.78 |
| none | 2.60 | 2.08 |
+--------+--------+---------+
Per million tokens:
Before: $2.50 input / $15 output
After: $2 input / $12 outputastra high is also 3x cheaper than opus max at basically the same intelligence.
astra high is also about as expensive as sol max while being more intelligent.
astra medium is cheaper than sol max while also being cheaper and roughly same intelligence.
im going to replace my sol usage with astra high/medium i think
caveat: benchmarks are really fuzzy with llms
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
Yes