I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
You would not expect the developers of the model to optimize for a well known benchmark?
y1n0 [3 hidden]5 mins ago
I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
mudkipdev [3 hidden]5 mins ago
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
awakeasleep [3 hidden]5 mins ago
You gotta look up how RLHF works before you ask a demanding question like this.
mi_lk [3 hidden]5 mins ago
It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it
Treat it like a bit as is
vb-8448 [3 hidden]5 mins ago
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
thimabi [3 hidden]5 mins ago
This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.
pizza234 [3 hidden]5 mins ago
Astra is the only model that correctly depicts occlusion of crank and leg, although interestingly, at max and medium levels (not in between).
vessenes [3 hidden]5 mins ago
Looked to me like it missed the chain though.
jsdalton [3 hidden]5 mins ago
Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
Luna — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 7.83 | 1.57 |
| xhigh | 4.24 | 0.85 |
| high | 2.46 | 0.49 |
| medium | 1.26 | 0.25 |
| low | 0.76 | 0.15 |
| none | 0.71 | 0.14 |
+--------+--------+---------+
Per million tokens:
Before: $1 input / $6 output
After: $0.20 input / $1.20 output
Sol — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 48.55 | 32.37 |
| xhigh | 24.11 | 16.08 |
| high | 10.38 | 6.92 |
| medium | 10.55 | 7.03 |
| low | 8.33 | 5.55 |
| none | 5.90 | 3.93 |
+--------+--------+---------+
Per million tokens:
Before: $5 input / $30 output
After: $4 input / $20 output
Terra — costs in cents
+--------+--------+---------+
| Effort | Before | After |
+--------+--------+---------+
| max | 32.09 | 25.67 |
| xhigh | 14.67 | 11.74 |
| high | 3.74 | 2.99 |
| medium | 3.46 | 2.77 |
| low | 3.47 | 2.78 |
| none | 2.60 | 2.08 |
+--------+--------+---------+
Per million tokens:
Before: $2.50 input / $15 output
After: $2 input / $12 output
fHr [3 hidden]5 mins ago
Luna is my go to daily model it's great value
tuo-lei [3 hidden]5 mins ago
Me toooo, for all my personal projects I need to pay for the tokens~
At workplace I use sol because I don't need to pay
ghthor [3 hidden]5 mins ago
Mine as well, it’s fast and keeps me in flow; and cheap!
leoqa [3 hidden]5 mins ago
[flagged]
samuelknight [3 hidden]5 mins ago
How are we supposed to know if Astra is frontier without the pelican?
satvikpendem [3 hidden]5 mins ago
It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
satvikpendem [3 hidden]5 mins ago
That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.
dprkh [3 hidden]5 mins ago
I thought all the designers have been using vector graphics for a long time now.
mceachen [3 hidden]5 mins ago
Aldus Freehand 1.0 was released in 1988. Adobe Illustrator was first released in 1987.
kulahan [3 hidden]5 mins ago
That photo looks pretty hilariously stupid, so this appears to be more of a first toe dip rather than some indication we can one-shot a previously difficult process.
appplication [3 hidden]5 mins ago
Sure, it is a bit cartoonish but it’s relatively impressive. I do wonder what you would get if you asked for photorealism
embedding-shape [3 hidden]5 mins ago
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.
XCSme [3 hidden]5 mins ago
Do you prefer the fable one?
It's more "correct" but looks a lot worse in my opinion:
I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!
readams [3 hidden]5 mins ago
The Astra one looks pretty good except it's standing on the wrong side of the table
kingstnap [3 hidden]5 mins ago
Its also available finally to Pro users! Just took 24 hours.
InsideOutSanta [3 hidden]5 mins ago
They gave out bankable resets for every day people on pro plans didn't get Astra. Given that, I wish they'd waited a few more days before activating it on my account :-D
wincy [3 hidden]5 mins ago
They haven’t activated Astra for me yet, I have two resets now. I’ve been using the opportunity to test out how good 5.6 Sol is at computer use asking it to generate stuff in Blender which has been… interesting
Edit: nevermind it JUST gave me a notification to use it!
paxys [3 hidden]5 mins ago
That's a pretty genius internal incentive to move fast.
embedding-shape [3 hidden]5 mins ago
> They gave out bankable resets for every day people on pro plans didn't get Astra.
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
sumedh [3 hidden]5 mins ago
Just got access to it on Plus plan in Australia. 2 Banked resets as well.
vb-8448 [3 hidden]5 mins ago
Played in codex app a couple of hours today: it feels much faster than SOL, even if the TPS is half of it.
algoth1 [3 hidden]5 mins ago
Just got it. European plus user here. Only codex, no chatgpt
jaesonaras [3 hidden]5 mins ago
Anyone had success using Astra as a Foundry model via Github Copilot? The error I get is that tooling is not available if reasoning has a value.
gavinray [3 hidden]5 mins ago
I have GPT-6 access in Codex and OpenAI API now
I'm a Business plan user with Cyber verification enabled, FWIW.
embedding-shape [3 hidden]5 mins ago
Same just got access literally this minute, Pro user here, no Cyber verification but have passed my ID over to them back in 2024 or something, maybe at the ChatGPT 3 API launch or something?
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
r_lee [3 hidden]5 mins ago
is Azure for this actually ZDR?
starik36 [3 hidden]5 mins ago
What is the actual utility of using this model on Azure? It's twice as expensive, according to the link.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?
hhh [3 hidden]5 mins ago
It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.
claiir [3 hidden]5 mins ago
They’re ZDR and the OAI ones aren’t
olalonde [3 hidden]5 mins ago
WTHIT?
spdustin [3 hidden]5 mins ago
ZDR = Zero Data Retention — they don't store your inputs/outputs.
itsjustkev [3 hidden]5 mins ago
Compared to OpenAI flex? I'm pretty sure that is their batch processing endpoint, which is naturally cheaper.
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Treat it like a bit as is
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
Reminded me of https://clocks.brianmoore.com/
I'm wondering if this is being trained on by the models today.
Isnt Netherlands the leader in bike riders and they dont wear helmets.
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
Edit: nevermind it JUST gave me a notification to use it!
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
I'm a Business plan user with Cyber verification enabled, FWIW.
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?