8 comments

  • WithinReason 9 minutes ago
    Mixed signals, here it's performing below even GPT-5.4 Nano:

    https://livebench.ai/

    while here it outperforms Fable by a significant margin:

    https://oxalpha.com/

    but if the latter is true, will people still say it was "distilled" from Fable?

    • sunbum 5 minutes ago
      the 2nd website is not official, just something someone slopped together for some reason.
  • esskay 12 minutes ago
    I'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.
    • rfoo 8 minutes ago
      lol don't shout out the obvious
    • daveyoung 7 minutes ago
      [dead]
  • tosh 4 minutes ago
    my guess is this is a small model punching way above its weight

    on toy benches it made quite a few mistakes but was able to fix all of them on its own

    (meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)

  • garo-pro 57 minutes ago
    Unfortunately I can't find sources other than this for now but this seems to be legit.
  • j_maffe 21 minutes ago
    Anyone has a link to a report of its capabilities? I can't find a reliable source.
    • vblanco 12 minutes ago
      completely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range.
    • daveyoung 5 minutes ago
      [dead]
  • dgellow 25 minutes ago
    Do we know the size of the model?
  • hncsiocp9x 10 minutes ago
    [dead]