Claude Haiku 5.5

(anthropic.com)

998 points | by sfkgtbor 23 hours ago

78 comments

  • simonw 21 hours ago
    Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

    Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.

    The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.

    The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:

      llm install llm-anthropic -U                                
      llm anthropic refresh
      llm -m claude-haiku-5.5 'prompt goes here'
    
    EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
    • rotis 18 hours ago
      thinking_effort: max Reasoning trace: This is the classic pelican-on-bicycle SVG test.

      Opus 5.5 had similar response on max: This is a classic test request

      https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

      I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)

      • simonw 17 hours ago
        The fact that they've heard of the test doesn't seem to help them draw a good picture of a pelican riding a bicycle.

        That said... here's "Generate an SVG of an armadillo in fishnet tights jaywalking on Mars" on xhigh for comparison: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... (and here's the same thing from other models: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...)

        • besterman23 15 hours ago
          Why even respond to these comments anymore?

          Every new model you make this post, every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not” it’s not a worthwhile conversation to have at this point, either the models can do some arbitrary thing or they can’t.

          • ozgung 9 hours ago
            > every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not”

            The answer is “they are definitely in the training set”.

            They all generate the same composition and style because this has become a self-reinforcing loop. They seem to reproduce a specific solution from memory instead of designing from scratch.

          • stingraycharles 14 hours ago
            I'd turn that question around: why bother with the pelican tests at all? Simon explained why he started them in the first place:

            "I chose that because a) I like pelicans and b) I'm pretty sure there aren't any pelican on a bicycle SVG files floating around (yet) that might have already been sucked into the training data."

            That's no longer true. Every new model gets a blog post with its pelican SVG in it, and that ends up on the web like everything else. Haiku 5.5 and Opus 5.5 now say "this is the classic pelican benchmark" in their reasoning traces. People keep pointing this out and it keeps getting dismissed.

            I, for one, don't see the point anymore. The reason for running the test is gone. I just don't get why we still treat the results as meaningful.

            https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle

            • simonw 13 hours ago
              I continue to do the test because I still learn something new from it every time.

              This time, just seeing the difference between Haiku 4.5 (a year ago) and Haiku 5.5 (today - and 1/10th the cost) was worth it alone.

              Same for Mistral the other day - the leap from Mistral Large 3 (their previous best model) to Mistral Large 4 was similar to the Haiku 4.5 to 5.5 jump.

              I wrote some more thoughts about what value we can still get from the pelican test back in July - https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l... but I've actually become MORE confident in its ongoing value since then. Using it to compare reasoning levels is proving particularly useful at the moment.

              • andsoitis 12 hours ago
                > because I still learn something new from it every time

                Can you share in what ways these learnings affect your decisions or behaviors?

                • simonw 12 hours ago
                  I get a good initial intuition about how much they are going to cost, how much the reasoning efforts affect their output (and their duration and cost), and how much they have changed since the previous release in the same model family.
                  • Xx_crazy420_xX 6 hours ago
                    You get intuiton on capability which is contaminated by your blog post being ingested to training corpus, wihout strict comparision analysis on not pelicans, not animals, not svgs you cannot state any general capability with your prompting
                  • Melatonic 1 hour ago
                    My only concern is that it might actually decrease the quality of the output. So many are outputting the pelican leg on the "wrong" side of the frame (along with the left pedal if you were facing same way as the pelican) it may be reinforcing future models to draw it that way.
                  • realty_geek 3 hours ago
                    You are not alone.

                    The hn community is diverse. Of course some people will be tired of it but if enough people are engaging with the pelican test I would say that is likely because it still has some relevance.

                  • johnisgood 6 hours ago
                    [dead]
            • yreg 7 hours ago
              Yeah, I haven't been seeing the point of this exercise for a long time (apart from driving traffic to Simon's blog if I wanted to be negative) and find it boring by now.

              But apparently we are in the minority since the pelicans are always upvoted to the top, so if others have fun with it then whatever, fair enough.

            • stratos123 7 hours ago
              The benchmark is not saturated. They are in the training set in a way where the models have heard of it, but not in a way where all the models get a perfect score on it.
            • StackTopherFlow 9 hours ago
              It’s just fun.
          • Barbing 12 hours ago
            Why even respond with a great new test? :)
        • rotis 10 hours ago
          Are you saying even these recent pelican pictures are not good? Vector graphics aren't meant to be photorealistic I don't think. For what they are they seem pretty good already to me. You can easily describe what is in the picture.
          • conception 10 hours ago
            Yeah, all of the Opus 5.5 ones are totally fine and usable for vectors.
        • Super3000 8 hours ago
          Fellow pelican here. Since you host on CF, maybe you can turn on their AI Crawler Block Feature? Maybe helps a little.
        • nomel 15 hours ago
          Whoah!!! gemini/gemini-3.8-flash is what I've been hoping to eventually see with the pelicans!
        • anentropic 6 hours ago
          LOL, I love that it inferred that this particular armadillo would be pretty smug
        • monster_truck 17 hours ago
          Maybe they're as sick of it as some of us are?
          • saaaaaam 11 hours ago
            Some of you might be but some of us appreciate joyous absurdity and look forward to the pelicans with great pleasure and anticipation. Chacun à son goût.
          • weird-eye-issue 14 hours ago
            Just skip it and move on I think there are more people sick of the people saying they are sick of it
        • aetherspawn 13 hours ago
          Make this the new benchmark please. The prompt is funny, the result is even funnier. I think the great part is there isn’t even a right or subjective way to draw it.
        • saaaaaam 11 hours ago
          Frog on a unicycle!
        • ricardobayes 6 hours ago
          Lol... can't believe this armadillo thing caught on.
      • breezybottom 17 hours ago
        LLMs hallucinate and pander. Calling something "classic" is a common LLMism
        • stingraycharles 13 hours ago
          So, just to confirm: you’re saying that it’s hallucinating “classic”, because you don’t think it’s in the training data, even though it correctly is a classic test?

          Please show me your personal reasoning chain how you reached this conclusion because I think you’re hallucinating that the models are hallucinating the word “classic” here.

          • breezybottom 5 hours ago
            You're making a lot of assumptions.
            • stingraycharles 3 hours ago
              I think it’s quite a stretch to say that when a model correctly calls out something to be a classic problem, that it’s probably hallucinating.

              That’s quite a big assumption from your end as well.

              • breezybottom 1 hour ago
                It's not a classic problem though. It's a very niche, novel benchmark that can't be more than 2-3 years old. That doesn't fit any definition of "classic"
                • simonw 19 minutes ago
                  It will be two years old on the 25th of October.

                  I think it's credible to call it a "classic" though, in a field that moves this fast. Name another comedy benchmark for LLMs with the same traction.

        • Grimblewald 15 hours ago
          right? "im having this niche problem i find no information on anywhere, any ideas?" and you get "ahh this is the classic <wrong, other niche problem>"
          • Barbing 12 hours ago
            Or “of course, super niche thing you made up is a problem everyone’s had” - they used to do that but not sure if anymore.
            • TeMPOraL 11 hours ago
              There's 8 billion people on this planet, and you're probably failing to imagine something sufficiently niche to be not actually very common anyway.
              • Barbing 10 hours ago
                I meant “made up” as in truly invented - this was years ago now, they’d use a boilerplate sometimes with “everyone’s dealt with <x>”, but see what you mean.
              • Grimblewald 10 hours ago
                not all 8B have computer access, and even less access the computer as a device they control. Our crowd is actually rather small.
      • jonplackett 7 hours ago
        Pelican on a tricycle?

        Two pelicans on a tandem?

        • Melatonic 1 hour ago
          Pelican writing on a forum about bicycles
    • ozgung 20 hours ago
      Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right.
      • accrual 20 hours ago
        I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack.

        > "What's up with the pelican?"

        Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...

        • snowram 19 hours ago
          Will Smith eating spaghetti is the OG benchmark
        • SwellJoe 17 hours ago
          Do people really want to be seen making using AI part of their personality?
          • accrual 15 hours ago
            Like most things, it depends. Think of the web devs of yore with MacBook Pros and Angular and React and Vim stickers stuck aside their glowing Apple logo (no longer a thing unfortunately). IMO having a pelican pin would be more of the same - advertising you are a developer of some kind. An attempt to bridge the gap to a stranger who might be into the same thing. The tech changes, the desire to signal things doesn't. The usual Exceptions(tm) apply of course.
            • SwellJoe 15 hours ago
              Yeah, but Vim is cool.
              • accrual 14 hours ago
                Vim is indeed cool. So is mg. I'd argue so is having my graphics card performing work on my behalf.
          • ultrarunner 16 hours ago
            Several people I work with have adopted its mannerisms unironically. Then again we don't exactly hang out outside of work.
          • HighGoldstein 4 hours ago
            "Do people really want to be seen making using Java part of their personality?"
      • kenhwang 17 hours ago
        The grass too, everyone knows bikes ride on grass.
        • ozgung 10 hours ago
          I think models consider this composition as a meme at this point. They don’t generate a pelican on a bicycle, they generate “the pelican on the bicycle”. A well-known composition with the sun, the grass and occasionally a scarf.
      • cortesoft 18 hours ago
        The medium thinking effort one doesn't have a sun at all?
    • Jcampuzano2 19 hours ago
      I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.

      Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.

      • mudkipdev 18 hours ago
        Benchmarks are the entire reason why max exists
      • LoganDark 19 hours ago
        I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general)

        I'm definitely not the average person though.

        • abustamam 18 hours ago
          I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround?

          I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.

          • lobocinza 15 hours ago
            "s" for just this session. "ENTER" for use it and set it as the default.
          • LoganDark 18 hours ago
            I thought the thinking effort was specified out of band from that, though maybe it's not. Not sure if the model was trained to listen in other areas. The biggest issue is, it's difficult to tell if it works because you can no longer see the thinking! Though I guess if you can't tell a difference in the output, was there any point to max in the first place?
    • internet101010 17 hours ago
      Long legs to illustrate me attempting to stretch my budget by using Haiku the last week of the month.
    • bagacrap 13 hours ago
      I don't think any of those frames are "right". None of them has a seatpost, none of them has representative handlebars, most of them have some fatal flaws that would make the bike unrideable (like the head tube being angled the wrong way from vertical). But you are right that the first one is the only one that hallucinates entirely new tubes.
    • firemelt 1 hour ago
      any index for this i want to see opus 5.5
    • ijidak 20 hours ago
      I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time.

      The pelicans all start to look the same after a while.

      But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.

    • grodes 10 hours ago
      why are all the pelicans always the same?!

      They are training on your prompts.

    • dankle 9 hours ago
      They do not look meaningfully different to me
    • zahlman 19 hours ago
      I'm still getting network errors. Seems to be CORS-related.
    • sourcecodeplz 9 hours ago
      medium looks better than high
    • vinni2 21 hours ago
      I thought Anthropic models didn’t generate images.
      • simonw 21 hours ago
        This is SVG, but recent Anthropic models have got extremely good at other forms of visual data.

        Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...

        And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party

        And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox

        Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...

        • noduerme 19 hours ago
          Pardon, I have a lot of questions about that Scrimshaw music text format. It's clever. Did you invent it, and is it specifically intended to be written to by LLMs? Is the editor/player LLM-coded as well, and was this its recommendation for a format that would be easy for LLMs to write? I'm wondering why this instead of say, asking it to write a .MOD file.
          • simonw 17 hours ago
            Opus invented it, and wrote the player, and the songs.

            My prompts were:

            > I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact

            > I am looking for music of the quality of the original secret of Monkey Island

            And then later:

            > Modify scrimshaw jukebox to add a copy-paste prompt that explains the music format, it should be shown at the bottom of the page below the readable instructions, the prompt should be designed to help any LLM tool compose music in the correct format. It should have a copy to clipboard button.

            https://claude.ai/share/1f721c20-2499-4d23-b368-3ab57146d956 and then https://claude.ai/code/session_01R3xuRtjVHqHo1GNbget6Tu

        • bitexploder 13 hours ago
          I have been using these models to make 3d models with build123d — they are great at it. I am never touching Fusion again. Astra is current a bit better than Opus 5.5.
        • VectorLock 17 hours ago
          What was the prompt for the widow's bay egg?
        • lastdong 20 hours ago
          This is great! Love the pixel art and tunes.
        • hazelnut 20 hours ago
          Tried it with GPT-6 Astra with Ultra but the outcome was underwhelming with Blender. Maybe it was my prompting ¯\_(ツ)_/¯
        • user3939382 15 hours ago
          I’ve had it transcribe speech from images of a waterfall viz
        • jansan 21 hours ago
          They are really good at generating artifacts, which are windows within the replies containing all kind of visualization, often interactive.

          They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.

          • mayli 20 hours ago
            SVG is hard.
        • khanan 17 hours ago
          [flagged]
          • simonw 17 hours ago
            That's a bit extreme.

            I think the music from the original is art, and I have enormous respect for it - I can still hum some of those tunes out loud thirty years later.

            The "music" in my demo helps show that text-based LLMs can do a passable job of composing simple 90s-era imitations of computer game music. That's interesting, because most people don't like not expect a text LLM to be able to do that.

      • vunderba 20 hours ago
        I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images.

        This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.

        https://kq-styles.specr.net

      • abirch 21 hours ago
        They generate svg. You can paste in pngs and they'll convert them to svg with varying degrees of success.
        • xienze 7 hours ago
          > You can paste in pngs and they'll convert them to svg with varying degrees of success.

          I've found Opus to be REALLY good at converting images that are well-suited to being SVGs (decent resolution, sharp edges, clear color boundaries [though gradients do OK]) after a few rounds of back-and-forth. It'll do precise measurements to figure out curvature, the exact colors to use/gradient stepping, simplify complex paths, etc.

      • sixtyj 20 hours ago
        They don’t do raster images.
      • sixothree 20 hours ago
        I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg.

        Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.

        https://www.youtube.com/watch?v=2EqMplbt0gU

        • zyberzero 20 hours ago
          To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all? Then I think it is impressive! Are you able to share the prompts you used?
          • sixothree 19 hours ago
            Source code is linked from the video! Scan the QR code. I will try to /resume tonight and give you some prompts.
            • cruffle_duffle 19 hours ago
              I mean Claude code sessions are all jsonl files that it can interrogate on its own. Get a new agent to capture how it was made and what the prompts were. No need for tedious /resume’ing and prompting.
              • sixothree 17 hours ago
                It's the easiest way to convey what I was doing. Anyway, some prompts added. I have other videos too if intersted.
          • sixothree 17 hours ago
            I'm going to truncate a lot of this. But here is the root prompt that will get you a video.

              Make a youtube poop video about having too many tabs that keep appearing faster than you can close them. Style it after the "too many cooks" viral youtube video that seemed to repeat over and over, getting worse and worse every time. You have ffmpeg and a plethora of programming languages at your disposal. Go completely nuts and make it as extensive and creative as you want. Render should be 1920x1080, and later 4k if you did a good job.
            
            [discussions about acts, characters, darkening theme, etc.]

            Workshopping the acts:

              The play has a font issue. See screenshot.
            
              [Image #1]
            
              Also, can you change the thing that happens at the start of episode one the cursor clicking the one tab and it becomes two. Then it clicks another tab and it becomes four. Is that a breaking change? You can change the speed of the actions to fit it in if needed.
            
            
            Here is a refinement prompt:

              The last screen just before "The End", the cursor clicks on the browser window instead of the tab close button. Can you make it click on the tab itself? Here is a screenshot of where it clicks [Image #3]
            
              Also, on the intro screen, the cursor clicks the tab and new tabs open, can you have it click the links in the page instead. See screenshot[Image #4].
            
              You can keep the exact same timing for both of these.
        • Bluestein 19 hours ago
          The Purple Screen of Death at the end :)
          • sixothree 13 hours ago
            ... is a link to the source code.
      • jonshariat 21 hours ago
        SVG is code
      • fakedang 21 hours ago
        They're SVGs
      • 1ucky 21 hours ago
        Those are SVGs not images.
      • ghoshbishakh 21 hours ago
        Bruh. Svg. It is like drawing something with geometric shapes which are represented using equations.
    • xmcp123 15 hours ago
      Claude still can’t decide which side of the bill should be darker/lighter smh
    • jsolson 19 hours ago
      [dead]
    • secondary_op 14 hours ago
      This test is shallow and useless but you keep coming back with it, please stop bringing it up, people deserve better.
  • minimaxir 23 hours ago
    Pricing is...a bit weird.

        Input
        $0.10 / MTok for prompts up to 100,000 tokens
        $0.50 / MTok for prompts over 100,000 tokens
    
        Output 
        $0.50 / MTok for prompts up to 100,000 tokens
        $2.50 / MTok for prompts over 100,000 tokens
    
    100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.

    In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])

    • dannyw 22 hours ago
      Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.

      For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.

      These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.

      In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

      • RussianCow 19 hours ago
        I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix.
      • WinstonSmith84 21 hours ago
        noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...
        • goosejuice 17 hours ago
          No? Don't use these lower end models to work on small well defined tasks with a frontier model orchestrating? This approach works very well for me and don't have any issue staying under 100k. I have no idea if it's cheaper but it does seem to be much faster for tasks like QA.
      • flockonus 19 hours ago
        > it’s really impressive how much intelligence per dollar has grown in just a few short months.

        Open weights models giving a distant salute from afar

        • pimeys 19 hours ago
          Yes, it was weird to see MiMo and DeepSeek missing in the article's comparison...
          • drob518 6 hours ago
            And GLM 5.3 Flash.
            • pimeys 4 hours ago
              I haven't really tried it that much. I use GLM 5.3 a lot, especially with Coralbricks where the cache reads are free so I can let it think a lot without breaking the bank in the following turns. It's really good for bughunt, planning, and security work.

              Maybe I should try the flash...

          • stavros 18 hours ago
            Is Mimo good? I've never tried it, but I've seen it mentioned three times in this subthread alone. DSv4.1 is my daily driver.
            • sausagefeet 9 hours ago
              I have been settling on using Mimo 2.6 flash for discovery/research and glm 5.3 flash for implementation. So far I've been pretty happy with the results, but mimo can be quite slow. If speed is an issue, I just switch to glm 5.3 flash.
            • vardalab 14 hours ago
              As context grows, it gets slower and dramatically slower. Otherwise it's pretty good.
          • RussianCow 19 hours ago
            It's not that weird. Most companies considering paying Anthropic are probably not considering Chinese models as alternatives. Many don't even realize they exist.
            • doodlesdev 18 hours ago
              The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs.
              • RussianCow 17 hours ago
                Is it? I would guess that it's much more about the race to get customers as they both near IPO than anything to do with the Chinese models.
                • verdverm 17 hours ago
                  we are actively preparing to move our devs from closed to open models, take it as a piece of anecdata

                  the trend in industry is clear by now though

                  • pimeys 12 hours ago
                    Yes. We did the same a few weeks ago. Everybody I talk with in the EU startup ecosystem is either doing the same or considering doing it.
            • epolanski 18 hours ago
              "companies" is a meaningless metric.

              If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.

              • RussianCow 17 hours ago
                I said "most companies considering paying Anthropic", which is not the same as "most companies". I also don't agree that "nobody" is doing this; I have lots of anecdata suggesting otherwise. Maybe the majority of companies are using the easy option of Copilot or Gemini like you said, but it's nowhere near 99%.
                • lifeisloving 15 hours ago
                  Yes it is. Literally every company Ive spoken too about this only has MS teams co-pilot lol. Like all of them are using the most basic form of AI.

                  They dont even know about Anthropic and think its just "AI". By the way this includes one of the largest power systems design firms in the world, who helps build many datacenters... Uses only MS copilot.

                  • blackqueeriroh 15 hours ago
                    Do you understand how many companies there are? Your anecdata is wholly useless
                    • vardalab 14 hours ago
                      Well, I can add two more points to your anecdata. I personally know of two companies where only exposure for peasants is Copilot.
                      • drob518 6 hours ago
                        Yep, this doesn’t surprise me.
                    • lifeisloving 14 hours ago
                      Umm, just go look up the whole of software revenues globally and compare that to OpenAI. Its very small. Nobody is paying for SOTA models outside software engineering and ancillaries, except for some outliers fields like consulting. Then there's techies in every field who use it, but not companies themselves.

                      Companies as a whole, do not care about this technology, except tech companies and its adoption cant even be compared to CRMs. A company might buy a SaaS product with AI but most of them are not purchasing Anthropic subscriptions lol.

              • tranceylc 18 hours ago
                Vibe coding gta 6 haha
      • JacobAsmuth 20 hours ago
        The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA.
    • Tiberium 22 hours ago
      There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.

      You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

      • AtNightWeCode 22 hours ago
        > ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5.

        So, it is might be even worse.

        • Tiberium 22 hours ago
          No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+
    • Eridrus 22 hours ago
      It's actually existing flat per-token pricing that is weird.

      Neither encode nor decode are linear in compute, so providers need to price for average expected length.

      This is just getting closer to the true cost of generating tokens.

      • foota 21 hours ago
        My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.
      • hgoel 20 hours ago
        Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.
        • vardalab 14 hours ago
          You would think with all this AI available, they could figure out the pricing to be as granular as necessary.
          • anthonypasq96 11 hours ago
            people dont like when the price of something is a giant formula or ???
      • sebzim4500 21 hours ago
        Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO
        • stingraycharles 14 hours ago
          I think this was originally started by Google who first offered these large context windows and others just followed suit.
      • willsmith72 8 hours ago
        they don't need to, companies smooth out costs all the time. cost of goods sold isn't a great way to price software products
        • stingraycharles 6 hours ago
          Yeah, also, the real cost here is memory, not compute. Which is why KV cache quantization is a thing.
    • HarHarVeryFunny 22 hours ago
      Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.

      For this application 100K token input is plenty.

      Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.

      • martianvoid 22 hours ago
        I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification
        • HarHarVeryFunny 22 hours ago
          The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.

          For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.

          I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.

    • tr4656 23 hours ago
      Luna does as well, but just at a higher limit.

      From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

      • minimaxir 22 hours ago
        Huh, that disclaimer is on the model page (https://developers.openai.com/api/docs/models/gpt-6-luna) but not the pricing page. Annoying.

        Fixed.

      • tripleee 22 hours ago
        So even at the 1.5x/2x rate luna is still half the price of this. Weird pricing strategy from Anthropic. I'm sticking with Luna if I don't need a super smart model
        • usef- 20 hours ago
          You're judging purely by token cost I assume, not cost per completed task?

          The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.

          • RussianCow 19 hours ago
            The cost per task from Artificial Analysis is roughly 3x higher at every reasoning effort level for Haiku than Luna. Sol 6.1 on medium has the same cost per task as Haiku with significantly higher intelligence. According to those numbers (which you should take with a grain of salt), from a pure cost vs intelligence standpoint, you're better off using Luna for economics and Sol for intelligence.

            With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.)

            https://artificialanalysis.ai/models/releases/comparisons/cl...

          • tripleee 20 hours ago
            yes, that's true. I should be looking at the $/completed task
    • alexchamberlain 21 hours ago
      Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".
    • jeremyjh 19 hours ago
      If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful.

      I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.

      • serf 18 hours ago
        >If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue.

        "less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.

        it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.

        it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.

        • jeremyjh 16 hours ago
          It isn't an argument, and I never said it is better. It is an explanation for why the pricing break is relevant. Anyone can do the same things to take advantage of that pricing. Would you prefer I not explain a basic fact to someone who may not know what is possible?

          The services are priced this way because larger context has significantly higher costs. That is a fact about the technology and it is true for every provider. So moaning about it isn't useful.

          On the other hand, there are a lot of people who don't manage context effectively - who start every session with 60K tokens - and that is significantly hurting the performance of every single thing they do with coding agents.

          • komali2 15 hours ago
            So what's your setup? Speaking as a "open terminal in repo, open Claude, say 'do this thing please '" kinda guy I'm interested in learning about these more advanced techniques.
            • jeremyjh 5 hours ago
              I use oh-my-pi (omp.sh) - mostly with stock settings and skills but its a "fully loaded" harness that you don't really have to add anything to. I do change a couple of things: I set it to prefer subagents, and I enable rewind. It is crazy that rewind is not a default, it saves a TON of context - if the agent goes down a crazy path that burns a lot of tokens it can rewind to an earlier checkpoint with the exploration or bug hunting summary.

              Presently I'm using sol-high for the default agent which does orchestration and a lot of smaller investigation and coding tasks itself. sol-max for planning and review. Luna-max for planned coding and general tasks. I also have a $10 minimax plan and use M3 for exploration and library roles, but I could probably be using Luna for that just as well and still only very rarely run into usage issues.

              I don't use any plugins or skill libraries apart from Caveman and I'm not sure how useful that really is anymore so I'd start without it so you have a baseline to compare. I do think it reduces context usage a bit but I haven't measured it recently. Caveman also includes some team, agent & investigation skills - again they might be helping but I haven't re-evaluated since like 90 days ago.

    • krzyk 1 hour ago
      They want to get on the Luna market.

      If there was no Luna you would see only the >100k pricing, but because we have Luna, they had to lower price for something.

    • port3000 22 hours ago
      They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.
      • cogman10 20 hours ago
        I think they are also trying to make sure Deepseek and other chinese models don't eat their lunch. They need something price competitive.
    • giancarlostoro 23 hours ago
      I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).

      I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

    • mnicky 22 hours ago
      You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.
    • mkotlikov 20 hours ago
      If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.
    • system2 22 hours ago
      Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?
      • mrngld 22 hours ago
        That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.
        • RussianCow 19 hours ago
          Except for the new MiMo V2.6 models, which appear to give some of the best value right now, at least on paper. (I haven't tried them so I can't speak from experience.)
      • wyrdcurt 22 hours ago
        Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.
        • pimeys 19 hours ago
          Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task...
        • girvo 19 hours ago
          I’m on the Legacy v2 plan and same: nothing comes close to 5.3 Flash’s value on it. It’s crazy, no wonder they discontinued them!
        • RideOnTime22 15 hours ago
          There's also fomo and what I believe is faniticism.

          Even for simple tasks why use X if I know "Y Max" is available and on paper, better?

          And why use something else when your favorite company releases something. Surely it must always be the best one to use.

        • vardalab 14 hours ago
          z.ai speed has been horrific, worse than sol-6.1 was last week, I am not renewing my sub once it expires.
      • nharada 19 hours ago
        Isn't the point of this release that it's comparable?

        AAI Index // Input // Output

        Haiku 5.5: 43 // $0.10 // $0.50

        Mimo 2.6 Pro: 46 // $0.43 // $0.87

        Mimo 2.6 Flash: 38 // $0.10 // $0.28

        Seems competitive to me? Plus then I don't have to manage multiple providers

      • user43928 22 hours ago
        Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.
      • usef- 20 hours ago
        On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins.

        And Opus 5.5 is really good.

      • flaburgan 11 hours ago
        They say it in the announcement: Haiku is basically useful to be called as a subagent. So you use Opus, and you want to investigate your production logs, instead of throwing them right away which is going to use a massive amount of token for mostly noise data, Opus asks Haiku to determine patterns to extract only the relevant logs and feed it back in Opus. The models are meant to be used together.
        • walthamstow 6 hours ago
          > The models are meant to be used together.

          If that was the case, then Claude Code would use smaller models for subagents. It doesn't. The subagent always inherits the parent model unless you tell it specifically not to.

      • pkulak 20 hours ago
        Where do you get this 10% number? Checking providers I know/respect, and GLM 5.3 flash is $0.15/m. Haiku is $0.10/m.
      • skeledrew 22 hours ago
        Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub.
      • aesthesia 22 hours ago
        There really aren't any models at 10% of the price of Luna or Haiku.
      • ray_kay777 20 hours ago
        People who are stuck using Bedrock in-geo due to their company policy (me).
    • dj_io 13 hours ago
      Haiku doesn't seem most cost effective solution. Jev like classifier can work in a fairly smaller model which are 1/10 of the cost. Opus is SOTA so I get the use case for one being restricted to Anthropic ecosystem. I think the use case for Haiku is mainly for users using on their chat for pro subscribers and free users to maximize their quota.
    • insanitybit 22 hours ago
      I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.
    • enraged_camel 22 hours ago
      >> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents

      Your vibes don't appear to be supported by facts. From the announcement:

      >> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

      • Philpax 22 hours ago
        People weren't using Haiku 4.5 for agents before. 5.5 is good enough that it might be.
      • StilesCrisis 22 hours ago
        Haiku 4.5 users were using it for Kleenex requests because that was the best it could do.
        • enraged_camel 21 hours ago
          Not really. We use Haiku 4.5 to turn users' natural language queries and requests into fairly complex structured specs for interior design and construction. It has near perfect accuracy.
          • dotancohen 20 hours ago
            How many examples are in your prompt? How large is that prompt? Or do you have some other way of tuning the output?

            I'm asking to learn for a similar project, not to discount anything you're saying.

    • AustinDev 22 hours ago
      encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.

      There are plenty of workflows like translations where you'd easily be under the cap.

    • sixtyj 20 hours ago
      Chatbot could be < 100k tokens.
    • solenoid0937 18 hours ago
      This is pretty good tbh
    • esafak 22 hours ago
      It's their creative way of 'matching' Luna's prices.
    • chaostheory 14 hours ago
      Going on a slight tangent, I've found that Anthropic (for my work) costs about 4x as much as OpenAI give or take, specifically Opus vs Astra
    • judetechdevs 15 hours ago
      interesting! thx for your info
    • j45 22 hours ago
      It could be to incentivize people to not be lazy users of tokens.
  • chriddyp 18 hours ago
    Ran our DataAnalyticsBench benchmark on it: https://plotly.com/blog/claude-haiku-5-5-plotly-data-analyti...

    9x cheaper than Haiku 4.5 and 2 letter grades better. It's also now the fastest model (using the default speeds, not trying any of the other models "Fast" mode) to complete the exam.

    Similar ballpark to Luna in price, cost, and accuracy. These are very cheap models: $0.38 to answer 40 in-depth data analytics questions (compared to $15 for Opus 5.5 or $20 for Astra).

    Overall very good at data analysis - handling all of the straightforward data analytics questions correctly. It fell short answering some of the questions that required some deeper statistical analysis like looking into other variables. In other words, it's not as persistent as other models in its analysis, which I think we'd expect from how they're positioning the model.

    Compared to OpenAI: GPT-6 Luna did a bit better and was about 30% the cost of Haiku 5.5. GPT-6.1 Sol got all answers correct, but was 10x more expensive.

    • doktor_pepsi 8 hours ago
      Appreciate providing this benchmark. I am curious, do you think that as new models (from the same company) are released and you repeat the benchmark, they adapt/extend their training data to include your dataset, thus polluting the benchmark results? Would you be doing anything to combat this?
      • agentdev001 1 hour ago
        I've considered this as well, as many others have. I tend to think that A: this is moreso a feature rather than a bug, and B: this is almost impossible to measure- it's a boogeyman imo.

        In other words, combating this is impossible if it's happening. Measuring whether its happening is also not plausible, for small fish running bespoke benchmarks suites. I think we (consumers) have hit a pretty clear stride of; new model releases > some subset of evals/benchmarks are saturated > new, harder, more niche evals and benchmarks take their place. That doesnt seem unhealthy to me.

      • chriddyp 2 hours ago
        It's a valid concern. We're not releasing the dataset, the answers, or the full set of questions to help prevent this. At the same time, I like to share where it gets things wrong in a bit more detail, which involves sharing a bit of the exam. I expect that we'll create a new benchmark with a new dataset in 6 months.
  • charlesabarnes 22 hours ago
    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users

    This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes

    • thepasch 22 hours ago
      This is them sneaking in taking the Claude Agent SDK (claude -p) off of subscription plans through the back door along with a model release. They previously wanted to do this in June, but backpedaled after huge backlash:

      https://support.claude.com/en/articles/15036540-use-the-clau...

      • luketaylor 18 hours ago
        Sorry that that help center article was misleading; we’ve updated it to clarify that `claude -p` has not been removed from subscriptions as part of this change!
        • thepasch 18 hours ago
          That's super encouraging to hear, thank you for the clarification! I'd edit my original comment, but the edit window has unfortunately run out on it. Should be OK, though, since the help desk page is now explicit about it.
        • arcanemachiner 12 hours ago
          Are you able to speak about the degree to which non-Claude harness use is discouraged? Has Anthropic's disdain towards non-Claude Code harnesses settled down since the compute crunch + OpenClaw heydey?

          More importantly, will my personal company account get banned for using a non-Claude Code harness?

          Not sure if you're able to answer these questions, but I would appreciate it if you can.

        • TomGarden 9 hours ago
          I've been operating under the pretense that claude -p was removed ages ago in the OpenClaw era, it's back now? What a mess of contradictory information, hope yall can figure it out
          • scottyeager 2 hours ago
            They never actually made the change. Using -p with a subscription has worked all along.
            • TomGarden 1 hour ago
              At one point they definitely turned it off and gave us a one-time credit to claim for -p, after which Claude -p was billed as extra usage
        • pastel8739 1 hour ago
          But it is? Especially from Pro subscriptions, which don’t get API credits
          • eli 1 hour ago
            It uses subscription quota just like before
        • jen729w 4 hours ago
          Can I build an app against this? Do you commit to not changing it again tomorrow?
        • AISnakeOil 18 hours ago
          They changed this from earlier today... It's still very confusing.
        • goosejuice 17 hours ago
          Was it ever? I thought that whole thing was paused. Still unclear if personal use of agent sdk inside a harness is allowed.
          • deaux 13 hours ago
            I've been using `claude -p` for personal use for a long time without issues. I do mean actual irregular, personal use though. Not using it to max out my quota, using it as main rather than sub or trying to be cheeky with it.
      • tekacs 21 hours ago
        I hope people notice again that this is happening this time around.

        Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.

        To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.

        It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.

      • eli 21 hours ago
        The page does not say anything about changing the way Agent SDK bills. I just tested Agent SDK and nothing has changed (yet).

        You might be right and they will change this in the future, but that's speculative

        • thepasch 21 hours ago
          > Claude Max and Team plans now include monthly API credits, which cover the Claude Agent SDK, the Claude API, and Claude Managed Agents.

          This text has replaced the entirety of the page called "Use the Claude Agent SDK with your Claude plan."

          • eli 20 hours ago
            Yes. Previously the page was all about how they were going to start charging for Agent SDK use with a banner on the top saying that, actually, they weren't going to do that.
            • thepasch 20 hours ago
              ...yes, a banner which has also now disappeared and been replaced, with the explicit mention that API credits "cover the Claude Agent SDK"?

              What more do you need?

              • eli 20 hours ago
                Well, it doesn't currently work that way on the latest SDK. If there's a change coming, it hasn't happened yet.
                • thepasch 20 hours ago
                  The monthly credit allocation hasn't rolled out yet either, so as of right now, we're effectively at the status quo. I'd expect the billing change to land once you can actually collect your Claude Console account.
                  • eli 16 hours ago
                    The previous note confirming subscription billing has now reappeared
      • cjav_dev 18 hours ago
        How credits work with `claude -p` is a common question. we're updating the faq now to make sure it's more clear
      • sanex 22 hours ago
        Those mfers. I'm using this for work! I use my work teams plan with pi so I can do all kinds of custom workflows that I can't in Claude Code. Time to convince management I need OpenAI instead.
        • stsch 21 hours ago
          Time to convince management (and yourself) to build some skills. :)
          • persedes 16 hours ago
            Or just for the token usage. The rug pull is coming eventually
          • stavros 17 hours ago
            Yeah there's no way in hell you're going to convince management to pay 10x for the same work because "human skills".
          • sanex 21 hours ago
            Spent many years building skills, I'm just working on a different level now.
      • sambaumann 22 hours ago
        Even after the June changes there was some allowance to use agent SDK on the pro plan. This will move me to codex tomorrow if agent SDK is really blocked on pro
      • luketaylor 19 hours ago
        Nothing is being removed as part of this!
        • winwang 19 hours ago
          *For now. If a company were to degrade something, it shouldn't be so obvious that the "goodwill" was just a reallocation. Just a good strategy. For example, it allows them to claim that they're "just going from 150% to 125% usage allowance, which is still more than 100%".
      • vmg12 20 hours ago
        These are api tokens, you can build a business with them using any harness.
        • tekacs 20 hours ago
          Yes, but they're wildly lower in value than the corresponding subscription usage.
      • martinald 22 hours ago
        Do we know if claude -p is now drawing from this API usage?
      • olejorgenb 7 hours ago
        [dead]
    • geek_at 22 hours ago
      This is literally for you to get tangled in their api and when they stop giving you the allowance they hope you will just continue to pay
      • losvedir 22 hours ago
        Nah, it's pretty trivial to switch providers (especially with Claude's help, ha).

        This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.

        • eli 20 hours ago
          Or to discourage people from using cheap subscription tokens as part of automated workflows
      • tomjen3 22 hours ago
        That's an old tactic for an old world. You only need, what, half an hour with your agent of choice to write you out of that?
      • enraged_camel 22 hours ago
        What does "tangled in their api" mean? Switching is pretty easy.
        • charcircuit 22 hours ago
          Not really, you have to fiddle with generating api keys and setting environment variables. Meanwhile with Anthropic it will just start charging you API prices for the tokens you are generating without even a single warning.
          • ClikeX 7 hours ago
            Of all the things I've had to deal with with migrating solutions and frameworks. Generating API keys and setting environment variables are the least time consuming of all of them.
          • enraged_camel 22 hours ago
            >> Not really, you have to fiddle with generating api keys and setting environment variables.

            That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.

            • charcircuit 22 hours ago
              The user could have always done that regardless of if the user has the option to be charged API rates on or off.
    • Topfi 22 hours ago
      This is massive. So on top of the regular usage, we now have USD 200,- to freely use via the API however we please, even resell? That is a statement, even knowing that inference does not cost them nearly as much as they charge, this is very developer-friendly. Does some minor de-risking for testing concepts. Terms seem to be reasonable [0].

      Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subscription, especially given changing to another model is essentially frictionless via OpenRouter.

      Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.

      Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.

      [0] https://www.anthropic.com/legal/credit-terms

    • tech234a 22 hours ago
      OpenAI will probably add this to their plans within a week
      • alasano 22 hours ago
        With OpenAI you can just use Oauth and get a token to use your subscription.

        Anthropic isn't even close to being this useful.

        • Iolaum 22 hours ago
          Biggest reason for an OAI subscription instead of Ant imo.

          Biggest loss is that Ant models look like they are genuinely better.

          • matsz 22 hours ago
            > Biggest loss is that Ant models look like they are genuinely better.

            This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).

            • floydnoel 5 hours ago
              I even got one to X.ai because I wanted to compare them all. The models were fine but the usage quota on X was extremely low compared to OAI/Ant
  • simonw 21 hours ago
    My complaint about Haiku 4.5 was that it was 10x the price of GPT-6 Luna.

    > Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens

    Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.

    So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.

    • agentdev001 1 hour ago
      It grinds my gears that we are still measuring $ per token, when the tokenizers for OAI and Anthropic are different. It is nonsense to compare tokens across labs, it only should be used to compare against a single lab's other models.
    • zapnuk 5 hours ago
      At scale this make a difference. For me/us not that much.

      At my work we currently don't have a subscription and use API based billing.

      Some time ago the price for my usage was about 30-50€ per day when we used opus for about everything. With luna+sol its about 10€ for sol to plan or debug and 5€ for luna to implement.

      At that point it doesn't quite matter to me if the cheap anthropic model is 20-50% more expensive compared to luna since luna is so unbelievable cheap. The big gamechanger was the cheaper models being good enough to do the implementation given a good "rough" plan.

      Nice that that's also possible with anthropics models.

    • mrbungie 21 hours ago
      Yep. I was looking at the prices of lower tier models a few weeks ago for zero/few shot tasks (pre Jev) and Haiku rates just didn't make sense at all. I ended up using 5.6-luna.

      Good to know that is going back to being an actual option from perf/price perspective.

    • janalsncm 17 hours ago
      If we look at their performance per dollar charts,

      In OSWorld 2.1 Haiku is better.

      On GDPval-AA v2.1 Haiku is equal or worse than Luna.

      On Humanity’s Last Exam they don’t seem even be comparing Haiku with Luna.

      For these baby distillations of flagships, I expect their users to be very price sensitive.

      • JacobAsmuth 16 hours ago
        Gotta be careful about these benchmarks because they're extremely difficult questions that may be significantly more complex than questions you would ask of the model in real use.

        If Haiku is "noticing" this and working harder to improve quality, you could still see similar or better cost-per-task in easier domains. ObviousBench is a good test of this.

        • janalsncm 14 hours ago
          Agreed. I’m specifically responding to the statement that Haiku was higher on all benchmarks. Normalized by cost (and why wouldn’t you normalize by cost?), it was higher on 1/3.
    • lightbendover 19 hours ago
      They do not, however, have the same price per task or task execution ability at any thinking level. Token cost alone is not a sufficient metric.
  • jjcm 22 hours ago
    Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.

    Haiku 5.5: https://html.non.io/lcars-haiku-5.5/

    Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5

    Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...

    One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.

    • thefourthchime 22 hours ago
      Pac-Man Bench:

      Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.

      TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5

      All results: https://jonclegg.github.io/pacman-bakeoff/

      • myzie 21 hours ago
        Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results?
        • thefourthchime 20 hours ago
          Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash.
          • myzie 20 hours ago
            Yeah, I was mainly thinking about how much cheaper Luna appeared to be in this case.
      • rpcope1 21 hours ago
        Something is not right there. DSv4.1 flash shows $1.89 for tens of thousands of tokens? What am I missing?
      • onlyrealcuzzo 21 hours ago
        How have you avoided being sued by Namco?
        • thefourthchime 21 hours ago
          I'm pretty sure they'll never see this. It's pretty much impossible for anything you do you build nowadays to get noticed anyways.
    • sparklingmango 21 hours ago
      > likely best used for small subagent tasks / tightly scoped work.

      Hasn't this always been the case with Haiku?

    • saretup 22 hours ago
      To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it.
      • jjcm 21 hours ago
        Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down.
    • BrokenCogs 22 hours ago
      Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.
      • twostorytower 22 hours ago
        It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.
        • BrokenCogs 22 hours ago
          I guess I'm giving GP feedback about their product diffui.ai, not really about Opus' performance.
      • jjcm 22 hours ago
        Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.

        It's meant to be a good test, not a good design.

      • FranzFerdiNaN 22 hours ago
        It’s not really AI slop, it’s how most modern SAAS websites look like.
  • seaal 22 hours ago
    The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted.

    Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.

    • thepasch 22 hours ago
      Note that this is Anthropic Trojan-Horsing the previously announced June change in with a model release, where the Claude Agent SDK can no longer be used with Claude subscriptions and is now billed with API credits only.

      https://support.claude.com/en/articles/15036540-use-the-clau...

    • 0gs 22 hours ago
      yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)
    • laurels-marts 18 hours ago
      OpenAI has been a disaster lately.
    • skeledrew 21 hours ago
      > Being able to actually use my Claude plan for other harnesses

      Wait what? This has gotten their blessing?

      • neucoas 20 hours ago
        You could always use Claude models on other harnesses via API... just not via subscription. Now they give you $100 worth of API tokens to use on opencode or Pi. Which is better, but still not the same as OpenAI were you can use the subscription on Pi without problems.
        • dfdydx 9 hours ago
          So you have the normal subscription usage plus an extra $100 in API tokens now?
      • copperx 20 hours ago
        Absolutely not.
  • wyrdcurt 22 hours ago
    About time Anthropic released a competitive cheap model. Haiku 4.5 has been too expensive compared to its performance for months now (in fact I don't remember being too impressed even when it was released). This one actually looks worth using in some scenarios. If it's really as much of a step up from Luna as the benchmarks they've shown indicate, it'll probably replace Luna in my workflows. 100k tokens is a pretty low threshold before the price goes up, but I tend to use these smaller models for smaller tasks anyway.
  • bouk 22 hours ago
    This is great! Been using GPT 6 Luna for decompiling my childhood favorite game (Age of Mythology) and this means I can throw Haiku into the mix as well. 17352/21965 functions matched so far...
    • WASDx 19 hours ago
      How do you validate the functions are correct? I did something similar, letting it (mostly deepseek 4.1) translate from assembly to C but it commonly made mistakes, some really hard to discover and fix.
      • itsgrimetime 10 hours ago
        If it's like any other matching (game) decomp, you compare the output byte-for-byte to the retail binary (e.g. any of the projects on https://decomp.dev/). Hard part to getting started is making sure you have a comparable compiler, flags, & toolchain
        • bouk 10 hours ago
          Opus figured out the compiler and flags in 5 minutes actually
      • Karrot_Kream 19 hours ago
        You should be able to validate it by creating tests against the assembler.
    • gizmodo59 22 hours ago
      can you share more details? was this very involved or asking codex/claude/open code with a 1 shot like approach?
      • bouk 22 hours ago
        I'll write a blogpost when I actually have it working, but basically I gave the game .msi installer to claude opus 5.5 and said to read these blogs:

          - https://blog.chrislewis.au/using-coding-agents-to-decompile-nintendo-64-games/
          - https://blog.chrislewis.au/the-long-tail-of-llm-assisted-decompilation/
        
        And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.

        I could now one-shot a new game, yeah.

        • supersour 21 hours ago
          Maybe a Show HN? I would be quite interested in seeing the results of this project
          • wingworks 18 hours ago
            I've done the same with some old games I used to play, SimTower and Oregon Trail II, both fully decompiled and now running natively on modern macOS.

            I'm going to try create an interactive twitch stream where viewers can play the game through the stream and other non-player viewers can trigger events in the game via points.

            Crazy time we live in.

            Edit, you come to really understand the game in the process, and why things happen and how to better play the game. And occasionally come across bugs, dev assets, assets never used, or assets all coded up, but code never triggered.

            • vunderba 18 hours ago
              Amusingly I just saw a "Show HN" for SimTower running online via WASM:

              https://news.ycombinator.com/item?id=49676394

              • wingworks 17 hours ago
                Haha yeah, I saw that post when I looked into reverse engineering SimTower, that guy saved me so much time decompiling. Didn't get so lucky with Oregon Trail II, there are some very old github repo's with attempts, but none got very far.
                • vunderba 17 hours ago
                  We're definitely in the age of ports! Interested to see how the OT2 port turns out.

                  Reverse engineering Redhook's Revenge binary (an old DOS game) before the advent of LLMs cost me way more hours than I'd care to admit back in the day - so I can't wait to put an LLM to work on some more obscure games like Sword Quest.

                  On a side note I should really give Oregon Trail II a shot. I never got into any of the successors like Yukon Trail, Amazon Trail, etc.

        • varenc 19 hours ago
          decompilation doesn't trigger any safeguard refusals? I would have assumed it would but glad it doesn't. Very cool and would also love to hear more.
          • supern0va 18 hours ago
            Surprisingly, no. I've been using Fable and Astra both to orchestrate decompilation of a relatively modern game (delivered via Steam) and they have no qualms about it.
    • anthonypasq 22 hours ago
      i love Age of Mythology, but why did you feel the need to decompile it? Its got a great world editor if you were trying to "mod" it.
      • bouk 21 hours ago
        I want to get the original (Age of Mythology Gold Edition) running natively on macOS and then port it to WASM to run it on the web so I can easily play it with friends
        • MisterMunchkin 21 hours ago
          Intriguing, I wonder how far you could go with turning games into websites.

          Like could total war become a browser game?

          • vunderba 20 hours ago
            Probably. There have been dozens of examples of taking old games (Crazy Taxi, Super Monkey Ball, Quake, etc) and making WASM browser equivalents using AI to decompile them just on "Show HN" alone.

            They often ship the original assets in a somewhat brazen disregard for basic copyright law even when the games are still for sale on places like GOG though.

          • steveklabnik 20 hours ago
            I saw recently that someone had ported Halo CE to the web and had 1024 players in Blood Gulch.
          • Pannoniae 14 hours ago
            Definitely, TW can "cheat" a lot with using sprites/very low LOD at distance, only streaming the corresponding units on battle load, and it's turn-based so performance isn't critical. Although obviously you won't be doing those 50000 unit battles on the web.

            Off-topic sidenote: what's with all these new projects targeting WASM instead of native, even if packaged for desktop anyway?

            • bouk 7 hours ago
              You can definitely have 50000 unit battles on the web
          • kro 21 hours ago
            Last time I gave that a try (without LLM assistance though) it was really hard as games DirectX calls cannot simply be glued to WebGL so performance was bad.
          • haunter 11 hours ago
            GTA V was ported to Wasm so everything is possible at this point
          • simlevesque 9 hours ago
            There's a Super Smash Bros. Melee website: https://lucasigel.com/melee
  • XCSme 18 hours ago
    It's around Qwen-3.8, and Sonnet 5.5 level, but a lot cheaper. It is also really fast.

    My tests for Haiku 5.5: https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...

    Twice as expensive as Luna, but also considerably smarter too:

    https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...

    • XCSme 18 hours ago
      A really big difference can be seen in this basic CSS animation generation of a solar system of Luna vs Haiku:

      https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...

      • sourcecodeplz 9 hours ago
        not a great comparison.

        haiku x-high: Cost $0.028, Tokens 54,579

        luna high: Cost $0.004 Tokens 6,679

        • XCSme 6 hours ago
          Sort of, it is still the model's job to know how much it should reason for a job to do be done properly.
        • BoorishBears 9 hours ago
          Their entire benchmark is deeply flawed and every model release they astroturf it (and I appear to call that out, but only after someone else independently verifies that it's bunk)
  • d1l 22 hours ago
    At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I’ve been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it’s slower. I guess it’s cheap but I think they got the balance wrong on this.
    • saretup 21 hours ago
      Curious as to why. Haiku 4.5 has been far away from pareto frontier for a long time. Maybe you need to update your prompt for the newer model in your workflow.
      • d1l 13 hours ago
        It performs well in the domains we’ve implemented it. Cheap, fast, good enough, multimodal. Its a workhorse and it strikes a nice balance. I’m open to the suggestion that I’m just winging it here and happened to get lucky with a model whose training fit our usage, which is boringly unspecific.
        • jstummbillig 7 hours ago
          Probably explained by implementation around a particular model, no? I suspect if you put similar effort in, you could get Haiku 5.5 to do what you want.
    • HyperL0gi 19 hours ago
      Exact same thing here.

      Both evals and Human pairwise tests for our use case are giving Haiku 4.5 first place in pretty much all tests.

      No we'll try understand if we need to change our prompts to match performance ...

      edit: maybe this will help: https://platform.claude.com/docs/en/build-with-claude/prompt...

      • d1l 18 hours ago
        Thx I’m digging into it but am somewhat comforted not to be alone in this. Turning up the level helps some but then we lose the speed. I don’t know that we’ll switch to 5.5 and may shop a different provider.

        Anecdotally we ran sonnet 4.6 for our more complex stuff and sonnet 5 was a LOT worse. 5.5 seems to have fixed it and we cut over our customer workloads. It’s strange, really.

    • nl 17 hours ago
      Something seems wrong here. Haiku 4.5 is a pretty bad model for just about everything.

      Did you have prompts that were especially tuned for Haiku 4.5 or something?

    • ygouzerh 14 hours ago
      Cam it be replaced by a Jev-like model? That might be better for latency sensitive tasks
  • Farmadupe 16 hours ago
    I compared the costs between luna and haiku for some classification work. Luna comes out a fair amount cheaper due to more efficient tokenization and less outputs.

      Model        input tokens   output tokens   batch cost   sync cost
      Haiku 5.5    11,893,643     240,451         $0.21        $0.42
      gpt-6-luna   7,903,468      177,602         $0.16        $0.31
  • jschveibinz 15 hours ago
    Some attempts to benchmark and compare LLM's:

    https://artificialanalysis.ai/

    https://benchlm.ai/

    https://tokenscost.com

    https://www.vellum.ai/llm-leaderboard

    Are there any projects out there for doing something similar?

    • sherby 9 hours ago
      ignoring the results on the website, all of these except for artificial analysis are very low quality. On vellum, it's hard to even understand which model does the bar refer to since the alignment is off and different for each item.

      Even if they add value, it's very hard to take them seriously given the quality of the website

  • matltc 20 hours ago
    My weekly limit __on a Pro sub__ has not gone over 50% since before the pre-Fable promos, but usage has been pretty much the same from my point of view. Maybe I am holding it right? Anyone else getting this?

    As such, I do not need to even reach for Haiku, and 4.5 was so inaccurate that it often cost more to do so in the past. Sonnet 5.5/low has been good for this kind of thing, and i didn't even touch thinking tokens or any of that. Opus 5.5 low for questions/repros, medium for implementation, basically never reaching for anything above that anymore. 5.5 has been great, so I'll try Haiku, but don't see myself going out of my way to integrate it.

    • pkulak 20 hours ago
      I'm excited for API use. I run some agents, mostly on Luna 6 right now. It's just tool use, web browsing, etc, so something dirt cheap, but also not super dumb, is much appreciated. Having a Luna competitor is nice.
    • coubri 20 hours ago
      there is no way in hell that im gonna use Haiku too, tho weekly limits become a problem for me in a last couple of month tbh
  • runtime_terror 1 hour ago
    Real talk; why would you use Haiku over say Deepseek v4.1 Flash?
    • caxco93 1 hour ago
      because it is included with your Pro subscription
  • TheAmazingRace 23 hours ago
    I wonder if we have an AI LLM equivalent to Moore's Law. Like how often do we expect improvement in this technology and with what timing?
    • onlyrealcuzzo 22 hours ago
      Yes -> every 18 months they've gotten 90% more efficient for the same level of quality for about 5 years. There's little sign that trend is slowing. If anything, there's reason to believe that System 1 models (plus potentially 1-2-3 workflows) may increase that over the next 3-5 years.

      You'll know when the trend stops -> when the intelligence differential between smaller models like 7B starts to grow instead of shrink from 32B models -> that means 7B is getting about as smart as it can get. Then, 32B will follow next, then 70B, etc etc.

      We haven't yet seen that at any size AFAIK.

      • thefourthchime 22 hours ago
        Andrej Karpathy said once that he expects superintelligence could fit in 1 billion parameters.
        • onlyrealcuzzo 22 hours ago
          Super intelligence that doesn't have to deal with the real world, maybe.

          I wouldn't be surprised if less than 1B param equivalent of our brain deals with solving math and writing computer programs and physics and all the things we tend to associate with "intelligence" - especially if you ultra optimized for that, I doubt our brain works like that.

          Dealing with the real world, I highly highly doubt it.

          • jstummbillig 22 hours ago
            How about if we get away from written text as the input, to something more fundamental, that then also is able to produce text (among other things)?

            Given that humans learn to talk while having encountered a measly number of word instances, and, given enough time, we should always be able to improve on the lottery that is biology, it does seems fairly likely.

          • Gigachad 19 hours ago
            It would be interesting if running ends up being a more complex task than advanced math. And our brains are just 95% allocated to dealing with the real world.
        • bakies 16 hours ago
          What's super intelligence?
          • qznc 11 hours ago
            Higher intelligence than any human.

            AGI (general) is about matching humans and ASI (super) is about surpassing humans.

            • bakies 2 hours ago
              What useless definitions, ok. Not gonna take opinions of people using these terms too highly.
    • istjohn 22 hours ago
      According to Epoch AI:

      > The cost of achieving a given level of AI performance has fallen about 47% per quarter since 2023, or 13× per year. [0]

      0. https://epoch.ai/publications/the-plunging-price-of-thought

      • FooBarWidget 22 hours ago
        Then why are AI plans still so super expensive, and AI spending going through the roof, while all the subsidies are ending?
        • stephbook 20 hours ago
          https://en.wikipedia.org/wiki/Jevons_paradox

          AI gets cheaper, people use it everywhere. Google searches, for example. Now we want to crack math problems and spend weeks with unreleased models.

          If you used GPT-2, it'd be incredibly cheap. You basically can't use it for anything and it's simple to serve.

          • versteegen 17 hours ago
            You mean if you used a modern LLM of GPT-2-level quality. Vanilla transformers like GPT 2 are ridiculously inefficient in comparison.
        • adgjlsfhk1 22 hours ago
          The cost per fixed level of intelligence is dropping, but we're also getting dramatically more intelligent models.
          • verdverm 17 hours ago
            MiMo-2.6 RL'd for ~$3.5M (not B), both main and flash combined, that is dramatically less and top 10 on https://artificialanalysis.ai/

            https://mimo.xiaomi.com/mimo-v2-6

            A frontier Ai is cheaper to make than a single 5/6th gen fighter jet, and maybe every fighter jet at this point.

            • JacobAsmuth 16 hours ago
              Especially if you have millions of Opus 5.5 examples to train off of!
              • verdverm 15 hours ago
                so tiring... you don't get to frontier from traces alone...

                also, who cares, the world is a better place if there are more awesome models at cheaper prices built with more efficient means

                we used to celebrate this kind of advancement, now it seems like astroturfing and belittling are the cool thing de jour

        • jrflo 21 hours ago
          Because models are only getting better at a rate of 10% per year, people always want the best quality possible. You can get SotA performance from a year ago for a fraction of the cost, but why would you use Opus 4.5 when you can use Opus 5.5?
        • f6v 19 hours ago
          Reddit is full of people complaining how they burn their 200$ sub in half an hour by starting ten Max sub agents. That’s to say, many people just don’t know what they’re doing.
        • BenzeneDream 16 hours ago
          To pay for the training of the models which are getting bigger and more expensive. So intelligence is getting cheaper overall but the need for ever-increasing intelligence can't be sated.'
        • jstummbillig 22 hours ago
          Because it's increasingly useful and the thing you are substituting (human time) is much more expensive.
        • srdjanr 21 hours ago
          Apart from what others said about using more intelligent models instead of cheaper ones, token usage is also increasing a lot. Classic Jevons paradox
        • teaearlgraycold 22 hours ago
          At least for me the Claude plans seem like an incredible deal and I never hit my limit.
    • ChaseRensberger 23 hours ago
      reminds me of this blog post: https://campedersen.com/singularity
      • kator 19 hours ago
        Whew, at least I won't have to hand-code solutions to the 2K38 problem!
    • bravetraveler 23 hours ago
      I've heard tell about 100% of certain types of work being ended in batches of six months. For years. Truthfully, I'm skeptical, but accuracy wasn't prioritized.
    • himata4113 23 hours ago
      double the information density every 2 days?

      serious bit: if you think about how these smaller models work, at the end of the day it seems that they are now capable of forgetting useless information because they're able to derive it in reasoning allowing models to become smaller at the cost of requiring more reasoning tokens to solve a task.

      • qeternity 23 hours ago
        Knowledge will be shifted to systems like n-gram augmentation which are relatively cheap and will not compete with reasoning capabilities for weight saturation.
    • dyauspitr 23 hours ago
      Hopefully enough runway for an existing model to train the next to be better than itself with absolutely no human intervention.
  • swalsh 22 hours ago
    Top of the page in 17 minutes? Now I know what y'all do while your agents are working.
  • djoldman 21 hours ago
    https://www.anthropic.com/claude-haiku-5-5#further-updates

    This section makes the reader think: why would I not pick Sonnet 5.5 instead of Haiku 5.5?

  • garo-pro 22 hours ago
    > Claude Haiku 5.5 is our fastest model to date at each model’s standard speed, although it runs less quickly than our Opus models in Fast Mode.

    Opus 5.5 runs 117 tps average on Openrouter, so it must be at least 10-20 tps slower for them to mention. IDK why they mention this as it does not help for marketing though. https://openrouter.ai/anthropic/claude-opus-5.5

    • jstummbillig 22 hours ago
      Maybe they think it's of interest.
    • fred_dawg 21 hours ago
      Is that page showing Opus TPS stats in fast mode? IIRC fast mode is 2.5x speed, so that would be 293 TPS, no?
      • yorwba 21 hours ago
        117 tps is the fast one, regular speed is 69 tps.
  • freakynit 14 hours ago
    Tested on my RAG system containing all of cloudflare docs (more than 3000 A4 sized highly technical documents): best `quality*speed/price` ratio of any other model. And I have tested more than a 100 different models.
  • AnodicElegy 3 hours ago
    Very important caveat if looking at the Artificial Analysis cost vs. intelligence graph (https://artificialanalysis.ai/?models=gpt-6-luna-low%2Cclaud...):

    "The site does not yet reflect tiered pricing, so provisional cost figures for Haiku 5.5 do not include the step up cost."

    Frankly, it's irresponsible for them to put the points on the chart if they don't have the correct values for one of the axes yet. At least they added some asterisks since yesterday, but you still have to go to the blog post (https://artificialanalysis.ai/articles/claude-haiku-5-5) to see what the asterisks mean.

  • tpoacher 22 hours ago
    Good to see Anthropic back alternative OSes.
    • Ectiseethe 10 hours ago
      Poems, then BeOS,

      then a model named for both.

      Same name, three lives.

  • michaelkdev 5 hours ago
    With those cheaper models being released, it’s always kind of tricky to decide which one to use. For example, if Luna is still cheaper but Haiku is a bit smarter, which one would you use?
  • Topfi 22 hours ago
    131tok/s P50 according to OpenRouter currently, though might move up or down over the coming days. If it sticks at that speed, roughly twice the throughput of Luna and far lower latency (up to 2sec depending on provider) is impressive, though the 5x price increase beyond 100k is painful.

    Was a big fan of Haiku 4.5, though understand why for most Sonnet was the far better option back then.

  • xixixao 11 hours ago
    > And we’re introducing a new monthly API credit for our Claude Max and Team subscribers, designed to support our users in building new agents and applications that run on the Claude Platform.

    OpenAI really should as this too, even to Plus. I’m already paying more for ChatGPT than I use, give me some free API tokens to play with (there is a free tier for data sharing, but only older text generating models)

  • TomGarden 23 hours ago
    From these selected benchmarks, it looks like it smokes Luna capability-wise. Excited to put it through its paces
  • MisterMunchkin 21 hours ago
    > we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200

    They’re definitely planning to make the subscriptions API based so they can charge you full price.

  • johnisom2001 22 hours ago
    It fails the "How many r's in <word>?" test.

    I ask:

    > how many r's in diminished

    It answers:

    > Diminished has 1 r.

  • sroussey 23 hours ago
    It’s about time Haiku got an update!
  • dangoodmanUT 21 hours ago
    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API.

    This is kind of nuts

    • djeastm 21 hours ago
      Is it realistic or cynical for me to assume this is to wean developers off the heavily subsidized subscriptions? Presumably it's using similar compute.
      • notatoad 14 hours ago
        >These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API.

        this seems like a plenty cynical explanation. it's just a free trial to get developers building on their api.

      • copperx 20 hours ago
        Anthropic was ignoring the usage of third party harnesses. Not anymore.
  • patrickwdaly 22 hours ago
    How are y'all using Haiku though? I rarely select it.
    • mariocesar 22 hours ago
      I have a zsh functions that calls claude code with haiku to suggest commit messages, is faster and the instructions are two lines.

      I also have an "ask" script that I use daily to ask simple stuff, it can access websearch and webfetch, it's more than enough to parse logs, ask for commands, quick research on the internet, small stuff. https://github.com/mariocesar/dotfiles/blob/main/common/.loc...

      I use haiku for things that needs to be quick, have really clear instructions.

      • tetraodonpuffer 20 hours ago
        with claude -p seemingly now using api credits I guess this approach will have to change unfortunately :/ I wonder what will be the best cmdline way to do things like these
    • swalsh 22 hours ago
      I've been using GPT-6 Luna in some capacity for nearly all my agent workflows. It's just a really good model, and the pricing is cheap. If Haiku 5.5 is better, and the same price (under 100k context... which is a big caveat) i'd probably swap it.
      • dannyw 22 hours ago
        It’s absolutely better than Luna. It feels closer to a “sonnet 5.2” if that makes sense.

        Of course it’s not as big, and hence falls-off quicker. I’d consider the 100k a “promotional price” to match Luna’s token pricing while delivering noticeably more intelligence.

    • svachalek 22 hours ago
      Opus often picks it when it's doing a "find me something" subagent. But largely it's been held back by being fully a year old at this point, and priced at a much higher price than models that are far more capable.
    • gghootch 22 hours ago
      I was waiting for this.

      Planning on doing flash analyses of PRs that impact evals in some way, and then post comments on GitHub whenever there’s flaws in them

      ( https://evalship.com )

    • notatoad 22 hours ago
      not haiku, but luna - last week i used it for things like "read this historical dump of 15k support tickets and break them into categories that make sense, then propose help docs that i could write to handle the most frequent queries in each category"

      used <10% of my 5hr limit on a $100 codex plan.

    • steve_adams_86 22 hours ago
      It's great at parsing documents inexpensively. For the few skills/plugins I've made, I usually instruct Claude to use Haiku for low-reasoning grunt work.
    • apothegm 16 hours ago
      Translating GPT’s word salad to English at the end of an agent interaction.
    • hector_vasquez 21 hours ago
      My software application uses Haiku in production more or less as a Jev. I do not use it for coding or development.
      • mrkn1 20 hours ago
        Why not have a CPU-first decision model for free? check out gutsy
    • Plutoberth 22 hours ago
      I'm building a game that incorporates LLMs as a game mechanic.

      I've been using Luna, but I'll probably switch to Haiku.

    • caskeycoding 21 hours ago
      [dead]
    • mochizou 22 hours ago
      [flagged]
  • satvikpendem 22 hours ago
    Apparently quite a bit smarter than Luna, I wonder what use cases it can cover. I actually honestly don't need a Haiku level AI to be that smart, and looks like you pay for it in the per token cost, I need speed mainly. I might even rather have a dumber but much faster model for things like web searching and parsing to retrieve results for the app or other LLM to do things with.
  • waximabbax 20 hours ago
    Alright its still little early since there is not enough independent testing but this looks very promising and I wasn't expecting anthropic to beat GPT-6 Luna especially at the same price. Haiku 5.5 beats Luna on every shared benchmark Anthropic published, particularly computer use and agentic coding.
  • tesnorindian 11 hours ago
    Does it support Jev like schema for classification request? Though the article confirms "classification requests".
    • pletnes 10 hours ago
      I was under the impression that many providers let you define a response schema. So - I believe - you can say «json list containing these options» and it will work, just be more expensive. Not sure about accuracy.
  • mchusma 18 hours ago
    At a glance, looks competitive on the pareto. I hope it retains the flavor of sonnet 5.5 and opus 5.5. Both are extremely good and productive. I have had a lot more issues with gpt 6/6.1 and getting what I want out of them.
  • arjunchint 9 hours ago
    the key thing to realize is the Haiku tokenizer consumes 1.5x the tokens for the same text compared to Luna.

    So at minimum costing 1.5x for the same request and if input tokens cross 100k then all tokens cost 5x more

  • Descensus 3 hours ago
    Am I the only one who doesn't see uses for Haiku all that much? I'm not an expert, but it feels like the real use cases for Haiku are slim. Feel free to enlighten or correct me if I'm wrong.
  • afrnswrth 22 hours ago
    The important question though...how does it do making a pelican on a bicycle?
  • peter_d_sherman 3 hours ago
    >"The table below shows how Claude Haiku 5.5’s pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model."

    This is interesting!

    This is the first time in the "frontier LLM world" that I've seen a pricing split along a token boundary, in this case, less than or greater than 100k tokens...

    In the past, first we saw ChatGPT, then we saw Claude, then we saw commercial frontier competitors, then we saw local open LLM's (Deepseek, etc.), then we saw "thinking" LLM's, i.e., use the LLM itself to generate a text plan to answer the user's prompt behind the scenes, then use that as a second, hidden prompt to actually answer the question, then we saw frontier pricing split across several different models (depending on how much behind the scenes "thinking" aka, "better hidden query plan text writing" that they did, and how many levels of this that they did) and these models may have had different context/token sizes but were not separately billed for prompt size, and now we're seeing pricing based on query size in tokens.

    Now, I'm wondering how many levels or "pricing tiers" this could segment into...

    In theory, an LLM provider could make 1k, 10k, 100k, 1M, 10M token context pricing tiers for the same model...

    Would there be any practical benefit to doing so?

    Well, that's a function of how much memory/compute are used by different sized prompts, how many prompt requests are incoming at any given point in time for a given commercial LLM provider, etc., etc.

    I don't have all of those numbers.

    In theory, an LLM provider could also charge less during nights or weekends -- times when electricity is cheaper, if those cost savings are passed on through the data center/compute supply chain, but again, I don't have all of those numbers...

    Still it is highly interesting to watch product/service segmentation and differentiation in this highly competitive marketplace...

    What will be the next innovation in terms of speed/efficiency/quality and/or price segmentation?

    We don't know... but this LLM pricing split on a 100k token boundary is highly interesting for AI market watchers, and possibly for Economists as well, present and future!

  • the__alchemist 21 hours ago
    I am perpetually confused about every name and version combination from both OpenAI and Anthropic. Especially in conjunction with the effort levels.
    • verdverm 17 hours ago
      haiku < sonnet < opus (poetry)

      lune < astra < sol (astronomy)

      • krzs9 14 hours ago
        I think you meant

        luna < terra < sol < astra

        • verdverm 13 hours ago
          yea, also left out Fable in retrospect

          I don't use Big Ai models so I forget them all

  • hidelooktropic 20 hours ago
    Finally! I understand Haiku is the less intelligent model, but the gap between Sonnet and Opus has been far too wide for about a year now.
  • declanjackson 21 hours ago
    According to AA benchmarks, it uses 162k output tokens per task (with max reasoning) - over double GLM-5.3 Flash for similar level of Intelligence
  • zyk_computer 15 hours ago
    I used to always use glm or deepseek instead of Haiku, but now I want to give it a try.
  • nico 18 hours ago
    Has Anthropic released a decision model ala Jev? I wonder if they’ll launch something soon
  • sfkgtbor 23 hours ago
    Happy about the Sonnet cache read price cut.
    • minimaxir 23 hours ago
      That was effectively required to match GPT-6.1 Sol (costs and caching prices are now equal). Sonnet 5.5 made zero sense to use over Opus 5.5 under the old cache prices.
  • oh_no 21 hours ago
    no AA benchmarks yet and the last chart in the announcement makes Haiku look useless vs new Sonnet pricing, interesting to see what 3rd party benchmarks show because i think Anthropic are costpertaskmaxxing here and it's going to look more like that bottom chart than the top ones.
  • 3371 22 hours ago
    Seeing people talking about the Agents SDK -> credits change makes me wonder does it impact Zed or likes.
  • margorczynski 23 hours ago
    How does the price compare to Luna? At least looking at the numbers it is noticeably better at most tasks.
    • onlyrealcuzzo 22 hours ago
      IMO, this is better. Luna is super cheap, but it's not that capable. At higher levels of reasoning, it's not that fast.

      This is more expensive, but it also looks like it's better enough that it's far more useful.

      I also won't be surprised if you look at cost per completed task + wall clock time that it comes out ahead for the majority of what you'd want to actually use it for.

      Luna will still be a great option for doing non-engineering tasks super cheaply.

    • TomGarden 22 hours ago
      For prompts under 100k tokens, it's priced the same as Luna - $0.10 in, $0.50 out.

      For prompts over 100k tokens it's 5 times more expensive - $0.50 in, $2.50 out.

  • mattz56 22 hours ago
    It's finally here ! Need to take a look at some benchmark now
  • maz1b 23 hours ago
    Wow, the rate of improvements in the AI era is staggering.

    GDPval-AA v2.1 as of now: 1620

    GDPval-AA v2.1 for Haiku 4.5: 735

    The 100k tokens pricing makes sense, looks to be a hedge against OpenAI's decisions API and Jev or its open source alternatives that are springing up.

    Nice release, congrats to Anthropic.

  • davvie 12 hours ago
    It's been a while...
  • skeledrew 22 hours ago
    The forgotten model is back on the map. I actually got OK mileage when I tried it for coding months ago. Maybe I'll try it again, with Opus guiding it, and see how it goes.
  • justmaris 21 hours ago
    Finally it has arrived.
    • bgolson 38 minutes ago
      What will you be using it for?
  • harshitkrhere 20 hours ago
    request to anthropic team release haiku as os model
    • verdverm 17 hours ago
      they are afraid of regular people having access to open weights, or at least that's what they tell us and the government
  • dhabedank 19 hours ago
    Really excited to use this
  • dcchambers 21 hours ago
    Begging Anthropic to let us use Claude subs with harnesses other than Claude Code at this point.
  • crooked-v 22 hours ago
    The important question is, does it talk in incomprehensible Claude-ese like the other Claude 5.x models?
  • simianwords 22 hours ago
    > Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users. These credits are designed to allow our users to experiment with building tools, apps, and agents that call our API. They can be used on any of our models. For more information, see our Help Center article.

    Did anyone read this? We get free API credits on some plans now

  • axthauvin 22 hours ago
    will use it to replace luna in production !
  • AtNightWeCode 22 hours ago
    Probably the same scam as the last Haiku update I guess. Uses more tokens to compensate for the lower price.
    • AtNightWeCode 21 hours ago
      To correct myself. The price was not lower. It was up about 20% for the tokens. But, the big price hike was that it used a lot more tokens for the same tasks.
  • areoform 22 hours ago

        > but they still block penetration testing and other techniques more likely to be used by attackers.
        > 
        > Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our Life Sciences Verification Program and Cyber Verification Program.
    
    I would like to take a moment of your time to tell you about some of the "bioweapons" Anthropic has blocked that involved Haiku!

    These are the examples from "Detecting and countering misuse of AI: September 2026" - https://news.ycombinator.com/item?id=49647300

        > Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated). 
        >
        > Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
    
    Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"

    While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.

    What did they save us from? What bioweapons did these filters prevent? From the front matter report,

        > The above LLM platform is not the only route via which researchers engaged in viral gain-of-function research have used our platform. In May 2026, we discovered a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract.
    
    OK. Sounds serious. "Gain of function research..." but who and why?

        > The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
    
    So this was a researcher inside of some country's national lab ("credible institutional context") doing research on dangerous viruses using Claude for "for editorial assistance in writing up the research."

    What "uplift" are you providing to scientists working at specialized global BSL-4 labs that already have – and I quote their report - "physical access to such isolates." (as in samples of viruses)? Are we uplifting their grammar?

    These "safeguards" are being expanded. The scientists I know can't use Claude for grammar checks or anything serious. You can try it for yourself.

    • solenoid0937 2 hours ago
      Anyone that needs to do serious bio work can do KYC and verify their institution, it's totally ridiculous that you think this should have been allowed, you clearly do not understand security or integrity

      I promise you that "random users may want to do bio work" is not the hill to die on

  • caaqil 22 hours ago
    > Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.

    If you block pentest or "other techniques more likely to be used by attackers", then what does "permit a wider range of defensive tasks" even mean?

    Any defensive task that's meaningful is almost indistinguishable from legitimate red-teaming that then falls under 'likely to be used by attackers". If only they would just stop nerfing these models, that'd be great. No APT is waiting around for Anthropic's permission, so might as well let us have some cool stuff.

    • TuxSH 22 hours ago
      Yeah it's too little too late, cat's out of the bag as people know that GLM 5.3 exists and is great at defensive and offensive cybersec.

      (sadly Mistral Large 4 isn't up to par - but Mistral serves GLM at 130 tps!)

  • simianwords 22 hours ago
    I remember a friend asking me why LLMs suck so bad. She was using Haiku 4.5 and that poor model couldn't keep track of the context within 3 messages.

    She said she was using Haiku 4.5 because she was advised to be careful with the spending.

    I hate that model so much lol.

  • iagocc 23 hours ago
    Where is Pelican? ehehhe
    • rvz 22 hours ago
      [flagged]
      • InsideOutSanta 22 hours ago
        So do you have the pelican or no?
        • rvz 22 hours ago
          [flagged]
          • InsideOutSanta 20 hours ago
            Ok, so I looked at all of your links, but nary a pelican to be found. How am I supposed to know what all these numbers mean if there is no pelican?
          • tomhow 21 hours ago
            Can you please not be so sneery/grouchy? That's far worse for HN than suboptimal benchmarks. The guidelines specifically ask us to avoid being curmudgeonly.
          • Mashimo 18 hours ago
            What is i want to code svg files though?
      • swalsh 22 hours ago
        I think it's a joke at this point, but also the visual benchmark is a remarkably dense method for demonstrating how good a model is.
        • InsideOutSanta 22 hours ago
          Yeah, people like to poop on the pelican. But pelican quality still correlated with overall model capabilities reasonably well, and you can immediately see and interpret it. It's a running gag, but it also does have some actual value.
  • booster-rooster 2 hours ago
    [flagged]
  • gabe-santana 2 hours ago
    [dead]
  • www_male 3 hours ago
    [flagged]
  • tancky 2 hours ago
    [flagged]
  • ariwilson 19 hours ago
    [flagged]
  • newtypecola 13 hours ago
    [dead]
  • ulukaya 20 hours ago
    [flagged]
  • quard8 22 hours ago
    [flagged]
  • nicolamanzini 20 hours ago
    [dead]
  • aidiveyt 22 hours ago
    [flagged]
  • tomkeen 11 hours ago
    [dead]
  • dfhdskfhdsjf 17 hours ago
    [flagged]
  • vickyonlinecont 22 hours ago
    Does anyone still use Haiku model?
    • bgolson 39 minutes ago
      I would use it if it was fast. Why are there no fast Anthropic models?
    • frozenseven 49 minutes ago
      The previous Haiku release (4.5) is a year old at this point. But this new one is pretty good on price/performance, so I assume it will be popular for at least a while.
    • Bolwin 22 hours ago
      I use it for title generation basically. Will have to see where this one fits in
    • jstummbillig 22 hours ago
      Opus 5.5 does :^)
    • mrpluto 4 hours ago
      i still use it!