I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Btw, they had 3D pelican-on-the-bicycle easter egg in one of the promo videos: https://youtu.be/bOC3DisEOfg?t=117 so I'm pretty sure that they spent some small amount of resources to train the model to produce good svg version as well. :D
When you see how good the output of the Luna model without reasoning is compared to SotA just a year and a half ago, it's pretty clear that it's been trained on explicitly.
I don't see how this is proof, and not just that the model got the better. You're comparing models a year and a half apart; this is a lifetime in LLM development.
It seems to start understanding how bicycles work. Look at the front fork. There are reasons why it's shaped the way it is in real bicycles. If you do physical simulations in 3D with reinforcement learning you will start understanding that most of the arrangements that the other agents use just won't work.
Yes, the curve of the front fork on the max pelican was what I noticed also. It's impressive (if not an accident), and a remarkably accurate bike overall. The pelican just needs to raise their seat a little.
I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
Do you really think there aren't data annotator services specifically training for SVG drawing? And that those annotators don't have a rich set of frequently requested icons/graphics/etc. that they review and train on?
I don't think it's a great rebuttal, quite the opposite actually: it's the same prompt with the most obvioy tweak you can imagine and the results are all the same. That definitely looks like it's been something the models have been trained on, and not some emergent capability (and if you take into account how nonthinking Luna compares to the SotA models from a year an half ago: https://simonwillison.net/tags/pelican-riding-a-bicycle/?pag... it becomes even more obvious)
It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.
Simon do you have a page somewhere that shows _all_ of the penguins created by various models across time?
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
Would be great to spend some money (other people's money) on paying some illustrators and artists to do the same for comparison. It could be straight images converted to SVG or direct vector graphics.
I'm not clear what you are saying. I am curious about what humans can do in the same context. It's becoming a standard I think for some tasks. E.g. how does a Gen AI stack up against a human in quality and "cost". Perhaps I'm missing what you are asking.
I find curios that Astra's pelicans are basically the same (yellow sun top right corner, green bike, same bike shape, same legs style, very similar background) while in other there is more randomness.
This tracks with what OpenAI has been saying about Astra: that it tends to do things in a certain way and it’s up to you to prompt it to change its style.
Shouldnt you give the AI a different task each time? otherwise the model companies just optimise for this benchmark because its in their best interest to...
Most images of bicycles on the web are arranged that way because the drivetrain is always on the right of the bicycle and you want that displayed prominently.
It's simonw. It's interesting to see their pelican benchmark, another comment by a different author elsewhere here shows some very good SVG generation too.
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
Wdym no room for error or experimation? Changes are even cheaper, and in my experience they are faster and way cheaper with models than with designers.
I would say opus was in some ways more accurate, but missed the higher level curvature feel of the site that astra picked up on.
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
That's quite shocking, at a sufficiently advanced level we can make all non-realistic graphics purely out of SVGs, as they'd have good scaling for things like logos and app icons. I know it was technically and theoretically possible before AI but most people weren't spending hours tweaking SVG HTML. I remember making an SVG dark mode toggle icon and it took days to get it right, I assume it's one shottable now.
"Designer" is a very broad term and can mean very different things to different people, not everyone is doing stuff that can be represented nicely by vectors.
That photo looks pretty hilariously stupid, so this appears to be more of a first toe dip rather than some indication we can one-shot a previously difficult process.
The problem is not photorealism. The SVG is outright dumb, its on the wrong side of the table, the table has fucked up geometry (its tilted) and many more minor flaws.
At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.
Incredible, I've played table-tennis since I was like 14, and I didn't notice that egregious error yet I noticed all the other small ones! Thanks for pointing that out, should have been very obvious.
Yeah, ranking is for entire model, not only SVG generation. Maybe I remove it entirely and add generation date instead.
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
$10/$50 is incredibly expensive compared to Chinese models which are cents.
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
An official motto of a president of one of the biggest software companies in Poland (Comarch) was "you can replace any experienced engineer with a finite amount of sutdents... for half of the cost".
Not really comparable IMO. Astra and Fable are not the every day workhorse you reach for to do basic tasks (unless your company has fuck you-money), they’re the tool you break out when you need the absolute strongest performance. There are plenty of tasks where finding and fixing one or two extra edge cases saves the business a lot of money, even if the cost is high. The best example would be scanning for vulnerabilities, if these models weren’t kneecapped in that area.
We have ChatGPT Pro at work and I usually use Terra medium/high and only bring out Sol High when the big or feature actually requires "thinking"/complex behavior. This has worked pretty well for me and it's very token efficient
in the middel of a coding session with 5.6 sol. astra popped up, swapped, asked review, it fixed a few really critical bugs immediately. pretty nice. one around some resource lifetimes in a rendering pipeline which would have been a nightmare to find manually.
almost had the feelin it was watching its little brother fail and had to 'step in' for a moment :').
I tried it with some humanities questions and with the default OR system prompt (no prompt?) it seems to be rather mild - lacking usual AI mannerisms. Kinda cool.
For those on plus, are you seeing astra limited to medium? Considering the rate limits that probably makes sense, but I'm wondering what different groups have access to.
They gave out bankable resets for every day people on pro plans didn't get Astra. Given that, I wish they'd waited a few more days before activating it on my account :-D
I must have been one of the last ones to get it because I managed to snag two bankable resets. Already reset once after having it clear up geometry in Blender, but it did a fantastic job!
They haven’t activated Astra for me yet, I have two resets now. I’ve been using the opportunity to test out how good 5.6 Sol is at computer use asking it to generate stuff in Blender which has been… interesting
Edit: nevermind it JUST gave me a notification to use it!
The higher price seems less important if it actually gets the job done with fewer tokens. I'm still very worried that this will end up coming back to bite us, by becoming more expensive once they inevitably nerf it. Every major model provider does that now after all
This is the second comment I see where you write "no it isn't", without providing a source for your statement. Then you follow it by asking for a source. So is there a source you can provide to back up your statement?
(This applies to Sol but probably works for Astra too)
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
Same just got access literally this minute, Pro user here, no Cyber verification but have passed my ID over to them back in 2024 or something, maybe at the ChatGPT 3 API launch or something?
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
It cracks me up that you have to ask in a public forum because the vendors purposefully obscure this critical information.
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
Threw $10 at this to help me prepare for my league’s fantasy auction this weekend. It spend $3.50 and then said “this action would cause you to go above your spending limit”.
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
It's the same price as regular processing. You get guarantees microsoft give you, which are ones OpenAI won't (or require dedicated spend,) and you can use azure identities for access.
Pelicans from Astra, plus 5.6 Sol, Terra, Luna for comparison: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...
I think this is a genuinely interesting comparison grid. Astra may be more expensive, but if you have a budget of 10 cents for a Pelican Astra low gives you something SO much better than the other models.
Astra uses less tokens overall too, for better results.
Astra transcript here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
https://www.svgviewer.dev/s/i8t1VXzQ
It feels silly to say that about making a pelican image but it really shows the difference in output and cost in an easy to understand way.
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
Yes
If that was the case, the models would have been producing near perfect outputs for it a year ago.
Instead they are just training on general SVG generation, which in no way should be viewed as "benchmaxxing".
I'd be shocked if they didn't myself.
Treat it like a bit as is
You’ve been doing this public service for so long (well, for so long in “AI hype” years anyway) that it’d be fascinating to see the evolution of this artifact across time.
I'm running out of excuses not to build a proper comparison site though. Maybe I'll have Astra do rhat.
I used Light mode and it used up all my limits for the day and had to continue the following day.
Reminded me of https://clocks.brianmoore.com/
Coz who knows if astra low will produce max like output if tried once more.
I'm wondering if this is being trained on by the models today.
https://openai.com/index/advancing-the-price-performance-fro...
Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog
The improved front fork design mentioned by Threatripper is about the only thing Astra is doing better, IMO.
Isnt Netherlands the leader in bike riders and they dont wear helmets.
Apparently the crowd agrees because they keep upvoting these.
Here's an image design source of truth: https://image.non.io/78f4cd8b-2560-4643-9a51-96a89171f994.we...
And here's the page it build from it: https://image.non.io/e7d3a9e5-f9df-4fd8-b79f-1f90280f978f.we...
Note the flowing svg lines, and how accurately it recreated them. Here's Opus 5 for comparison - you can really see how while Astra really recreated the flow that was in the original design, opus only got the general vibe: https://image.non.io/dfe13de0-4487-431f-8b69-544ff3030dac.we...
One thing I will say is you are paying for quality. That site build cost $24 - extremely non-trivial for a simple frontend.
I would say that $24 is trivial IF that's the final design. The truth is that the cost doesn't leave much room for error or experimentation.
Yeah if you are solo developer without budget.
For any business this is nothing, the ROI is massive.
Compared to what?
It sounds completely trivial and likely I'm wrong here, but could it be that opus saw the reference image squished? That might explain the sharper horizontal curvature
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
It took a while to test it, initially OpenRouter was giving Not Found errors for this model ID.
It's more "correct" but looks a lot worse in my opinion:
https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...
https://aibenchy.com/showcase/?page=2#showcase=67fc6d6c8e4c3...
https://aibenchy.com/showcase/?page=3#showcase=c215b5c915da6...
Good point about the mouths, I just noticed, lol
Imo, it's still better than most models, I personally like the stylized perspective.
You can view here all generations for all models: https://aibenchy.com/showcase/
What panel of judges are you using for scoring/ranking this? Seems subjective enough to not be able to be ranked/scored at all
I was thinking to manually grade/rank the SVGs, but I decided against it, as it is indeed subjective.
I was thinking it could have at least a simple objective check (hamster doesn't have extra or missing parts, table has 2 sides, and net is in the middle, etc.).
>Playing ping pong from the side of the table
I think marketing might be getting a bit absurd at this point
I even keep seeing obvious stealth marketing like this: "<topic> and how do I use it with <product> in <product>"
The entire mainstream media and political establishment, and every normie I meet is convinced Skynet is already upon us
I think they’re really going to struggle selling these models long-term. My company is already massively cutting down on access because they’ve realised most people don’t actually produce any value using it. All the tokenmaxers have ruined it for the rest of us now that accounting have seen the costs.
It must be super interesting working there rn :D
almost had the feelin it was watching its little brother fail and had to 'step in' for a moment :').
time to go play outside...
Yeah, when I saw that Tweet I knew the person was saying it because they knew it'll be available within 24h.
Edit: nevermind it JUST gave me a notification to use it!
Edit:
GPT-6 Astra (low): 57 Intelligence Index, $7.70/M tokens
GPT-5.6 Sol (high): 57 Intelligence Index, $3.08/M tokens
So for the same measured intelligence, Sol costs only 40% as much — i.e. ~60% cheaper, while Astra is ~2.5× more expensive.
Why is the burden of proof on me tho!?
astra high is also 3x cheaper than opus max at basically the same intelligence.
astra high is also about as expensive as sol max while being more intelligent.
astra medium is cheaper than sol max while also being cheaper and roughly same intelligence.
im going to replace my sol usage with astra high/medium i think
caveat: benchmarks are really fuzzy with llms
https://artificialanalysis.ai/models/releases/gpt-6-astra
I've created with Sol a skill called Low Quota Mode that intends to reduce the use of tokens usages by the frontier (intelligent model) and delegate the use of bulk reading of docs/code and implementation to a sub-agent running Luna Max. Sol is asked to supervise, read the diffs and approves the commit/pr.
The skill might need some iterations while you use it, for example at the end of a rough session you can ask Sol how did it went, which were the points of conflict with Luna and try to iron them little by little by editing the skill.
Also in difficult tasks, ask to babysit the sub-agent model, I've seen it makes more effort into communication between frontier and sub-agent to guide the task with more care.
So far it has reduced my tokens usage a lot (have not quantified but the quota lasts more).
I'm a Business plan user with Cyber verification enabled, FWIW.
Has there been anything published about if Astra uses different amount of usage from your subscription plan compared to Sol? Don't recall coming across that in the press releases.
The only reason most of my customers would use Azure Foundry instead of OpenAI directly is the ZDR assurance but it is so incredibly difficult to extract out of their model menu.
There is no trivial way to block non-ZDR models either so every customer has to “vet” and individually approve models.
If anyone from Microsoft is reading this: get your act together! You’re failing at the one thing people might want to pay you to do!
but again, seems like there's no word from Azure if this applies to them.
very confusing.
Then I threw $100 for a Codex Max sub and it included Astra and it did it for me.
Sure seems like Astra is expensive AF.
Do Azure offer something that simply hitting the OpenAI endpoint doesn't provide?