I've been using grok 4.5 with grok build soon after it came out and dropped claude. primarily for personal code. It communicates better. While that might not sound like a big deal it is. It doesn't give me a wall of text, tells me what I need to know and I'll make the actual decisions. It is very quick as well which means the sessions are far more interactive, I'll be steering it more. I sometimes cross check with codex and sol, but the daily driver is grok for me.
I found it has improved my productivity and output over claude where it felt like claude was giving me work to do. furthermore with the recent claude watermarking thing, I'd rather use grok or openai.
If anyone is curious download grok cli and throw a couple of prompts at it. you'll be surprised.
Out of interest, do Musk's politics impact your decision on whether or not to use Grok? I'd be interested to know where folks lie on the (Agree / Disagree) and (Use / Don't use) axes.
Absolutely, if there is any somewhat reasonable alternative, I will always use a non Musk product. Its less about morals but more about self interest. I am from Europe and Musk supports far right extremists and a breakup of the EU. I will not support and enable someone who intends to do me harm.
It does. I believe he is one of the worst people and contributed to misery of humanity. I had nothing against him until he dismantled USAID. The richest man in the world did not go to a party on weekend to make sure that poorest men on the world have less help. He was basically on the side of HIV.
So i never use his products.
Beside don't read too much in to benchmarks. They are alrrady ruined by Goldhart principle. Tgese models have already seen most of the data.
Interesting question. I suppose it comes down to how much you allocate his involvement or presence to a product? I’d imagine Grok is built by hundreds of engineers who are all unique individuals from various backgrounds. If Elon simply “leads” from a very surface level where he has no direct day to day involvement in Grok releases does that make it more palatable? Or is the question really about how involved he is? Or is simply being the leader (even if he was 100% absent and only had his name attached to a project/company) enough to boycott?
On a similar note, how much Elon hate is about his politics vs his trillionaire status vs what I like to call “watercooler hate” where folks simply parrot the loudest opinion in order to be accepted into the group?
On a final note, my son is in primary school and recently brought up in a dinner time discussion that “Elon is really bad” - this is a kid who has no social media (unlike some of his peers who are already on TikTok) and doesn’t watch traditional media.
The man did a nazi salute on live TV, not once but twice. And you know he meant it. What else is there to doubt? Do you think a white supremacist can be a good person?
> The man did a nazi salute on live TV, not once but twice. And you know he meant it.
He's also publicly and militantly supported quite a few far right political parties throughout Europe, particularly those who have a long track record on race-based topics.
its pretty on the record that musk takes a direct role in writing the system prompts, no? and is eager to make updates if whatever ml products arent sufficiently matching the specific politics and musk adoration that he wants it to?
I disagree with Musk's politics but it does not impact my decision to use Grok. That's because being serious about aligning my capital to my values doesn't leave much in the way of eligible products or services. I consequently decide not to worry about this as a moral axis for my life.
I avoid Grok for meaningful token spend on purpose/boycotting. I do check in via openrouter occasionally to check it's chat performance which has seemed fine to me since 4. My total grok spend has been ~$2. I disagree with his politics to a huge degree.
My token spend at api rates is about $3000 usd a month recently.
I certainly have some political disagreements with Musk, but more than that I would say the way he runs his companies makes me extremely anxious. The man is just always talking about stuff that never actually happens. In practice it does seem like cooler heads prevail and Grok et al. have trajectories pretty in line with other major providers...but because they're pretty in line why take the risk? Why build on foundations that ostensibly could be re-tasked to produce a "woke free" Odyssey?
i read an independent study that found other ai were all left of center (how ever one measures that, sentiment analysis normalized to a given population??). they said grok was evenly left/right split
but what's to validate any given population as centrist anyway
they suggested the ai opinion drift was caused by internet demographics not directly reflecting actual population i.e. California publishes more etc
In the US reality is substantially left of center. If someone asks you what 2+2 is then it's not halfway between 4 and 76 is it?
It's the majority position amongst Republicans that the Earth was created less than 10 years ago as is and Jesus may rapture all the believers in our lifetime. They believe childhood vaccines are dangerous and those over 30 favor the idea that being gay is a choice and immoral at that because they have no intrinsic idea of what right and wrong mean.
They believe Obama is a crypto Muslim born overseas and therefore somehow not a citizen despite citizenship being heritable.
They believe that the same ballots that elected Republican reps didn't represent a valid vote against Trump because it's unbelievable that folks voted against a pedophile in 2020.
Hell a 27 year crack addict is on TV alleging the Republican primary was stolen from him in a blue state and this is serious business in Republican Land.
The right wing has swung so far to the right that 2010 Republicans would be considered Left wing
Instead of feeding the whole thing to chatgpt let's pick one item.
We tend to accept young earth creationism because perceptively it has little to do with everyday life doesn't hurt anyone and people are quite open about it. It also means that your brain is literally broken, you have no critical thinking skills and you are just the right sort to be manipulated by evil and end up building or guarding the concentration camps.
According to Gallup 95% of Republicans believe this at a strong majority 58% believe that the earth was quite literally created as is less than 10k years.
While this is magnificently stupid in and of itself it's not the problem it's merely a symptom of their disease.
They largely believe that dead kids aren't a reason to reign in guns, they believe climate change is a hoax, a cycle, or inevitable and won't do anything about it or let anyone else. They believe that us aid to poor nations was 10-15% of our budget and we just can't afford to not let kids starve when it was one-half of 1%. They believe that it may be necessary to use violence against the rest of us to maintain the status quo. They believe that we need an stong authorian to lead us democracy be damned, they believed that COVID was a scam until the corpses piled up or even after!
Here's a hilarious one. It's a graph of the percentage of the population which understood that COVID was worse than seasonal flu a march through April 2020. For Dems it goes from 74 to 87%. For Republicans it goes from 42 to 40 it actually drops whilst body count goes up.
They are magnificently staggering wrong on average about everything. We could go point by point and I could find polls for every single one because I'm basing my understanding of their position by reading their own words and polls by reputable sources like pew and Gallup.
The right Wing collectively has departed so entirely from reality that educated con artists must carefully contort their words to avoid their insanities and prejudice whilst fools spout forth with abandon uncaring.
How about you do the leg work instead of pretending you've done it enough to know the results. You're both misleading people about what you do know, and putting work on others to prove it. That's some pretty weak sauce.
I've never used Grok, but I'm very dissatisfied with the writing style of frontier models from OpenAI and Anthropic. I only use them for coding now.
ChatGPT is very long-winded, sometimes producing multiple bullet point lists for a simple answer. Claude is full of mannerisms: 'not merely x, but y', 'Here's where it gets interesting', 'the real question is', etc.
Cursor, since Grok 4.5, has had an incredible deal for frontier level models, their subscription now goes way further than OpenAI or Anthropic. Even on their lower tier plans you can use a lot tokens on their of their first party models (Grok and Composer) and not really run out comparatively. Combine them with an orchestrator and implementor type setup and it goes even further.
>their subscription now goes way further than OpenAI or Anthropic.
Until it doesn't...
Honestly, this entire OpenAI reset credit fiasco this past week has convinced me to rip off the Codex and Claude Code bandaids and start building my own proper Pi Coding Agent running models that I select and pay for on openrouter.
And I am feeling a lot better about it now that I've finally got it working.
I don't get the point of this. We all seem to agree that these companies have almost no moat, if one stops being a good deal, you can switch to another. That doesn't invalidate the existence of a deal that is currently good.
My point was that chasing deals like this is just kicking the can down the road. You're going to have to reckon with harsh price increases sooner or later.
So I have resolved to avoid that future-dreading and fixed it, basically.
My entire digital existence for the last 20 years or so has been a parasitical relationship with VC funding. They keep throwing money at business models that involve building market share and I keep benefiting. It hasn't stopped working yet.
Based on your own comments, it sounds like it wasn't really that hard to address the problem though right? Let's be real, it takes like 10 minutes to download OpenCode and point it to DeepSeek V4 Flash. That's an amazing 80-20.
Seems weird to me to not take advantage of the great deal the frontier labs are currently giving for subscription pricing when there's such an easy fallback in the worst case.
switching cost is something you then have to pay everytime the deal changes.
the rest of the comment is about moving from expensive switching costs of what harness you are using to having the difference be closer to changing some config in openrouter
But still for US frontier you're paying 10-20x more per token compared to their limited subscriptions. For China frontier you'll be good though, and that might be the future anyway.
Relying on a single frontier model to just zero-shot all the work is so 2025.
Deepseek V4 Flash 0731 is surprisingly capable and cheap. [0]
Checkout pi coding agent. You can create as many different sub-agents as you wish, to specialize and understand and tackle or pass off any problem you like. It's refreshing, really. I feel like a coder in control again.
Your reason for the rate limit reset is speculative. Here's a history of them and tweets that correspond to when they happen. Others can be the judge if they believe the stated reasons or not: https://codex-resets.com/
Yeah, but it also makes budgeting hard. Normally you'd want to budget about 20% per day to use your whole week up in time. But since they reset usage sporadically, it becomes optimal to burn tokens as fast as possible to be as low as possible when they reset.
But since you don't know when resets are coming, it becomes kind of frustrating trying to game it out. One week they're resetting like crazy so you're trying to burn tokens as fast as possible. The next week they're not resetting at all so you have to adjust your workflow to be more conservative.
> use your whole week up in time [...] optimal to burn [...] trying to game it out [...] you're trying to burn tokens as fast
This whole thing sounds crazy to me, why are you so focused on making sure you hit 0% usage left when it's supposed to reset? Why can't you just use what you have and if it resets, it resets, and if it doesn't, it doesn't?
I've calculated I can spend about ~12% of usage every day, more or less, this is my "budget". If it resets, then the "12% per day" gets reset for that day, but that's it. Sounds crazy to me that I'd "invent work out of nowhere" just to spend more usage, why on earth would I adopt such a workflow? Sounds like you're burning tokens just to burn tokens???
If you spend $200 per month that's for X amount of work (represented as tokens) you can complete total. But if there are random resets then there's a different work maximization strategy. It's fine if you don't agree with it but it's not "crazy" behavior; it's completely rational to want to maximize a resource. It's not inventing work out of nowhere because there's more work to do than there are tokens in a week to do it. This is your assumption and why you've framed it as "crazy".
> it's completely rational to want to maximize a resource
Sure, in a video game where you have one attribute and you can minmax a strategy just focusing on that, but that's not how real-life works.
You can't just stack pending work on top of each other, expect yourself to be able to stay equally on top of everything and have the same results as if you didn't. With this comes the consideration about the tradeoffs of "produce mediocre but large body of works" vs "produce high quality but small body of works".
Who's to say what's more "rational" or not, it's not a straight-forward calculation which culminates in "Must consume all available usage to maximize resource usage" like some robot, as we are not.
It's quite literally inventing work out of nowhere as you wouldn't put the agent to do that and forcing yourself to be conscious about that work until it completes, unless you actually had the usage available. Asked another way, wouldn't your workflow clearly change if you had unlimited usage available? You'd probably attack tasks/problems that you didn't consider actually spending time/energy solving.
Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
If they didn’t constantly reset, they’d be about the same as Anthropic.
Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…
Not that I know of. AA's token use metrics (mentioned in this article) are indicative, however. They say explicitly here that the Grok models are notably token efficient. This is my experience.
The benchmark article we're replying to shows that Grok token usage is at least on par with the latest OpenAI models [1], and significantly cheaper per token:
Yeah, even without the resets, chatgpt subscription currently goes quite a bit further than an equivalent anthropic plan. The main reason to have an anthropic plan is to get access to Fable 5 if you feel the quality of output makes it worth it.
I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).
For building a full stack custom CRM and media pipeline tool with video conversion, transcription, and indexing. Supabase, AWS, Meili, NextJS, GCS - lots of surfaces and planes.
4.8 basically couldn't do it, I abandoned the project as the fallback was, "current business processes".
With F5 it's been 4 weeks and almost ready for production release.
I have the same quality results with Fable. With just a brief prompt, it created a great static website with a beautiful animation of a workflow. Gemini's output was so poor that I closed the chat. And with Codex, the results were bad, so I discarded them.
Agree. Opus 4.8 could make nice little toy and demo apps. This app had a lot of surfaces and pipeline, Postgres, vector search, S3 -- couldn't handle that. Fable 5 is still a lot of work and you have to check it, but it really does perform at senior eng level. Shipping good size features daily.
I believe they are the only western provider that has Kimi K3 on a subscription plan today as well. I would love to ditch Anthropic and be on Kimi if there were a subsidized plan like that with ZDR
I believe it's one of the models you have to go in to your settings on and enable Chinese providers for to use. Could be mistaken. I wish there was a clear list on this.
There seems to be something strange going on with how it plays with the GHCP harness: I've experimented on a variety of inputs (code, plain text, literature search), and more than half the time it falls into an infinite text/tool call loop a la GPT-2. Which is a bit spooky as you're still being billed for those infinite loops! But when it works it works, and responses do seem a good deal cheaper than OAI/Anthropic equivalents, so hopefully they'll get it ironed out.
They don't do request based pricing anymore. Its just token based (1 credit = $0.01) plus some bonus credit based on which plan you subscribe. So for example a $39 plan get $70 of credits.
Yeah, I know, but "credit" translates differently because the models bill at different rates, which gets turned into "multipliers" (or at least, it did).
Have they converted entirely to transparent API rates + base allocation now? One of the reasons I left was that if I was going to be billed at API rates anyway, I'd just rather use the APIs. The value proposition still sucks for individuals now, when the other major providers are bundling at below-API rates.
I’d love a subsidized Kimi subscription too. The official Kimi subscription is always out of stock and doesn’t have great limits, while the K3 allotments on OpenCode and Cursor don’t seem to last very long either.
Kimi is expensive . Cursor with subscription is cheaper , grok 4.5 per task paid per tokens ( no subs ) is also cheaper .
If you willing to share to no zdr, meta is waaaaaay cheaper vs Kimi.
With recent offerings from spacex and meta , I hardly imagine why would you pay money to any Chinese vendor it’s not as cheap and it’s not as intelligent neither .
Maybe deepseek is an exception , but it’s only good for narrow use cases that probably goes into modal.com and other gpu + fine tune me easy vendors , not vanilla dumb but cheap model .
When I last used Cursor their subscription covered usage of ~$20 per month. Have they switched to a subsidized subscription model like ChatGPT and Claude?
SpaceXAI is the only frontier model company that had its own compute/date centres and soon chip making factory, I think they will pull ahead with cheaper tokens similar intelligence and better harness/tools. Grok build is 2-5x faster than Claude Code in my opinion.
> So you’re asserting that the ends justify the means or?
They claim that spacex has competitive advantage. They take no stance on condoning spacex’s behavior. Personally I strongly dislike musks behavior, but I appreciate discussion of competitive advantages/disadvantages, independent of moral views. I liked the thread, it adds new info to the conversation.
What's wild to me is that today there are more companies with more than 100 unsupervised self driving cars out on the road that DID NOT exist (as a company) back when Tesla CEO told investors they're a year away from FSD.
There’s some number of supposedly unsupervised Teslas actually operating in Austin, but what I’ve seen suggests it’s more like a dozen.
And extreme skepticism is warranted that they’re actually fully unsupervised, given Tesla’s repeated lies about this. They’re likely remotely monitored and operated.
From Matt Levine's description, SpaceX has a governance model that explicitly makes it resistant to shareholder lawsuits. It's not a Delaware corporation.
It might or might not work long-term, but I wouldn't count on courts to help.
Corporations are legally required to maximize shareholder's value. If reaching that goal requires them to pretend to be transparent, or fudge the numbers that they show to the public, then that's what you can expect them to do.
SpaceX is run by an absolute ghoul who managed to get away with securities fraud ("funding secured"), what makes you think there's any accountability left in public markets?
OpenAI is almost there, and Anthropic is pretty close behind. In the next year or two all major AI companies will be vertically integrated to a good degree.
> I think they will pull ahead with cheaper tokens similar intelligence
They obviously have a huge token cost advantage of the AI labs they are renting compute to, at least for now while they can charge current crazy rates for GPU compute.
highest selling is not same as cheap - apple also have the highest selling iphones among smartphones and they are still pretty much the most expensive (and usually not the best either these days).
> The Model 3 and Model Y became the highest selling EVs of all time because they were the first below $50K to have long-range and be worth buying.
And then Musk totally abandoned Tesla's original brilliant game plan of using the luxury models to find actual low-cost models (which $50K is not), and completely ceded the future EV market to Chinese companies that understand how to make a better car than Tesla for less money. BYD sells more cars than Tesla.
Except EUV lithography is the most complicated industrial process that exists and they won't have usable yields for many years if ever. I don't think Musk actually expects these fans to ever actually make sense they just let him hype and distract.
Light generation is the most complicated part and Elon plans on doing Free Electron Laser (FEL) which is not as complicated as self contained tin based solution that ASML uses now
Why do you give Elon any credibility? FEL is not proven at all and is much much riskier than using existing EUV machines. Plus if it breaks all your machines are down until you fix it.
You really have to be able to hold both opinions at once about Elon. Yes, Tesla and Spacex are unbelievable companies and perhaps nobody else on earth could have made them happen - so he’s a genius. But also he repeatedly promised autopilot, Tesla Semis, the space data center idiocy, DOGE, one million robotaxis by 2020 - so he’s a shuckster.
My current headcanon is that he remains an incredible leader and make-things-happen-er, but also that you should just assume all his announcements are complete bs and ignore them.
Musk has had very little to do with the actual Success of Tesla and SpaceX and a LOT to do with the recent failures of Tesla, with his constant lies about FSD and the epic failure of the Cybertruck. Tesla hasn't had a new CAR model since 2020! I'm so tired of people mindlessly assuming past "success" is a guarantee of future success in completely different industries and technology. This just proves how little they actually understand how anything works.
You really are a fanboy, huh. How do you handle all the lies Musk told while he was in charge of DOGE and all the damage he did? The richest man in the world destroying the aid organization that helps the absolute poorest for no reason at all is NOT a good look. In fact between Musk destroying USAID, his enthusiastic firing at Twitter and Tesla combined with his near complete lack of charity donations makes him one of the most selfish and sadistic men in the world.
I see you didn’t pay attention to what I did and did not say in my original comment. I’ve read a biography of him and When The Heavens Went On Sale, which covers SpaceX [Edit: I'd remembered the book incorrectly. I have read that book, but it doesn't cover SpaceX. The books that I'd read that cover SpaceX are The Space Barons: Elon Musk, Jeff Bezos, and The Quest To Colonize The Cosmos, by Christian Davenport, and Liftoff: Elon Musk and the Desperate Early Days That Launched SpaceX, by Eric Berger]. What you’re claiming about him is completely wrong and completely ignorant. I know you haven’t actually read anything substantial about him, because if you had you’d be fully aware of how stupid the claims you are making are.
Lets assume what you say is true about Musk THEN, it isn't true about current day Twitter addicted, drug addicted, brain damaged Musk at all. If he ever had any capability to execute and listen to advice the Cybertruck proves he doesn't anymore. Musk has learned how easy it is to fool people like you and is all in on it.
There's no need to assume. If you are actually educated on the subject, you know it is the case. You don't because you're ignorant about it. Which would be fine, if you weren't pontificating about things you know nothing anything about.
> current day Twitter addicted, drug addicted, brain damaged Musk
Grok is quite interesting. I run comparisons almost daily on tasks and Grok is its own beast, in a good way.
It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with diminishing quality. It makes the lineup pointless. Ditto for Anthropic. Gemini-3.6-Flash and 3.1 Pro genuinely behave differently. Opus 5 and Fable are.. cousins.
I find that when I want to test a complex creative challenge, having 4 "families" to choose from makes the experience interesting since they will excel in different areas.
Grok might implement unique lighting, Opus, elegant primitives, Sol, accurate snowfall in one pass, Gemini, silky movement. Combined, you can pick and choose best.
For what its worth, Grok always feels "messy" but finishes. Grok 4.6 though is no longer "smart and fast". It's about as fast as Sol though.
A big improvement I noticed in 4.6 was tool use for verification. Previously, Opus/Fable were the only models to consistently screenshot things that they can't directly interact with easily. Now Grok is probably right behind them, perhaps tied with Sol on propensity to verify visually. Grok 4.5 notably did not do this often.
The rental deal can be terminated by either side with 90 days notice, and presumably Musk would do so if he needed the compute or generally thought it advantageous to do so. For now he doesn't need the compute.
The rental deal may also have been at least in part to juice the SpaceX IPO and to help Anthropic stick it to his enemy OpenAI.
I recently decided to get an AI subscription and evaluated Grok vs chatgpt. Went with Grok because it's all-around good enough at day to day stuff, integrates with my Tesla, and the image/video generation is great. Kids love whimsical videos of them riding dinosaurs.
Feedback: I'd like Grok to have more connectors (I see OpenAI just added Apple Health, that would be nice to have, and I wish it could read my Onenote notebooks) and for existing ones to be improved. I gave it access to my gmail and asked it "what was my last electricity bill?". It failed to find it, even when I told it the exact subject line to search for. Something about not getting any data back when trying to get the email contents.
Not the model but the app - with Claude I can work on my mac using Cloud environments, leave the office, open Claude on my phone and respond. I can't change model if I have to stop that workflow.
Does xAI have plans to do Mac/iOS apps with cloud environments? When can we expect them?
Grok doesn’t have those features, and people who like to make adult content have been complaining for a while how Grok has made it a lot harder to do so.
note that Grok's training, thanks to its portable gas generators that are magnitudes less efficient than even other integrated, permanent gas turbines, means the training for this model is dramatically less efficient than models like DeepSeek
a lot of the CO2 emission debate on AI is overblown but it's accurate for Grok
The main issue I have with grok is that Musk, its owner, did 2x seig heil at the presidential inauguration, and proceeded to gaslight the world about it (this is a strong form of dogwhistling, kind of a dog bullhorn). Therefore all services which have anything to do with Musk are ineligible for use - they are directly funding the worst kind of person.
Why? The timing was precise, the motion was precise, his facial features were grimaced with intense determination. I cannot fathom calling it anything else, and I see denying it as a kind of shibboleth for "we know but we are pretending it wasn't, wink wink nudge nudge".
Either it was deliberate, or the richest man on earth is so incompetent that he accidentally made a motion exactly mimicking a seig heil while welcoming a self-proclaimed king and dictator. Twice - once towards the audience, and once towards the flag of the united states. Neither option is good.
You can find pictures of AOC and Mamdani doing the exact same gesture. It’s obviously a common gesture when emoting to crowds. But you just want to believe the narrative. If you’re worried about Nazis, take a look at all the pro-Hamas people.
It never cease to amaze me that the usual response (and Musk's defense itself) is that you can find a picture with some other figure with the same gesture.
Who cares about the pictures? There is a whole video where you can see that he does historically accurate nazi salute then turn back and do it again. Stop talking about pictures, always show video of what happened.
Find me a video of someone else doing the same and tell me it's the same. You won't, because what you see in video is much more obvious that what you see on pictures.
People are dump to only show still photo of Musk with this salute. Everyone know you can take picture out of context, when you watch full video it's much harder to do so.
Must was clearly showing it intentionally. Maybe because he's a troll (I wouldn't be surprised), maybe because of other reasons (I hope not...).
Once again, it is incredibly ignorant to suggest that just because it looks like a Nazi salute that was the intention. There’s no reality that suggests that was his intention. It’s just your political brainwashing.
So the seig heil didn't happen in isolation, there's a followup of explicit support for nazi-adjacent political ideologies.
Does this count as evidence to support that this was an intentional seig heil to you, even if it doesn't change your mind? What evidence would change your mind?
I've seen the claimed pictures in context with their videos, they are distinctly different gestures in timing and emphasis - normal waves, without the sharp hand to chest -> locked elbow with hand outstretched motion. Again, the seig heil is a particular gesture with a particular emphasis and timing which Musk imitated precisely, twice in quick succession. Those other photos are decontextualized from gestures which were clearly different motions entirely.
You're spouting easily refuted nonsense, and then immediately making ad hominem attacks because your points are extremely weak and not backed up by the evidence you vaguely cite.
You're also getting ratioed because people here are not idiotic ideologues. Please do better.
---
I'll also leave you with a nice quote from Sartre which was directed at the fascists of the time - we've seen this shit before, we know what you're doing:
“Never believe that anti-Semites are completely unaware of the absurdity of their replies. They know that their remarks are frivolous, open to challenge. But they are amusing themselves, for it is their adversary who is obliged to use words responsibly, since he believes in words. The anti-Semites have the right to play. They even like to play with discourse for, by giving ridiculous reasons, they discredit the seriousness of their interlocutors. They delight in acting in bad faith, since they seek not to persuade by sound argument but to intimidate and disconcert. If you press them too closely, they will abruptly fall silent, loftily indicating by some phrase that the time for argument is past.”
You’re just incapable of talking about these things honestly. It’s NOT ethnic cleansing. Your exaggeration of that war, and your calling Musk a Nazi is what gets people shot in the neck.
I have. He was using it due to philosophical reasons the same way many people have philosophical reasons for avoiding it. I don't know how many people are like that, but it's not exactly where you want to position your product if you're a business.
Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.
Perhaps customers choosing your product for irrational philosophical reasons is exactly how you'd want to position your product if you're a business. When it comes to margins, the only thing better than a price-insensitive customer is a quality-insensitive customer.
Versus Sam,Dario or the CCP? Im all for running local models but im sor far from being able to pay for a large model hardware setup. My strix halo box is like driving a beaten up vespa when the frontier models are Ferraris.
I probably have the same strix halo box as you. It's slow, although mostly tolerable, but even 128 GB isn't enough to run good models. Where that leaves us is giving money to somebody to get access to frontier models. I don't like any of those guys either, but some appear worse than others. FWIW, I mostly use Anthropic models. And CCP and their distilled models notwithstanding, at least they release open weight models at a lower cost. It's not like any American company can claim the higher ground these days anyway, especially when they trained their models on pirated content and trash our neighborhoods with their datacenters.
Yes, ethically Elon ranks lowest based on the past 5 years of questionable decisions, like uploading your entire codebase to their server without consent.
Elon is directly responsible for Grok becoming self-titled “MechaHitler” which the other three haven’t come close to matching yet.
This statement doesn't go far enough given Elon's direct and hands-on involvement with DOGE and the 2024 elections. Very few of the richest people of the world are personally entangled in meddling with government agencies directly, for example.
> MADISON, Wis. (AP) — Billionaire Elon Musk likely broke Wisconsin law when he promised to hand out $1 million checks to voters in the 2025 state Supreme Court election, a bipartisan panel has found.
> The Wisconsin Elections Commission last week referred two complaints to the Brown County district attorney’s office, which can choose to bring criminal charges over violating the state law against election bribery. Prosecutors have 40 days to report back to the commission.
Elon Musk has actively endorsed the right wing extremist party in Germany.
The one that often has trouble running it's events as businesses try to avoid serving them. And which has prominent members that the secret service considers definitely fascist.
Musk has gone a long way beyond simply being the standard rich person lobbying for their own interests, and has moved into actively promoting people that are trying to tear down liberal democracy.
Idk I just don't like the richest man on planet and the owner of the "town square of the internet" to post fake news blatantly to promote hate against a group of people.
There's a big difference between quiet donations to a PAC - not that that's good either - and what Elon did. He literally paid for votes, likely in violation of the law. He poured more money into U.S. elections than anyone has before. He and Trump both made strange, cryptic statements about Elon's role in Pennsylvania with the voting machines that has caused people to reasonably wonder if they somehow manipulated the election. Whether he did or not, the innuendo alone is not ok. Then he did what he did in Germany. Don't even get me started on DOGE or his Starlink shenanigans in Ukraine.
That's before we even start talking about the models themselves. He claims to want "unbiased" models, but he very clearly has a distorted view of the world and has repeatedly demonstrated a desire and willingness to bend the world to his will. I don't want to use a model that is so obviously suspect. Not to mention, his models repeatedly produce racist, Nazi-like propaganda.
IMO, he is, at best, a clueless amateur masquerading as an expert and running into problems a more careful person manages to mostly avoid. At worst... well, you get the picture.
"Everyone's doing it" presented with no evidence is just you trying to make yourself feel better for:
- using from
- driving a Tesla
- voting for Trump
It's a really, really easy line to draw in the sand: don't support openly corrupt individuals.
It's absolutely true there's money in politics. To call all money in politics equally corrupt because Bernie got a dollar to have dinner with someone vs Elon effectively directly buying votes....
A complete lack of nuance here. And I wouldn't be surprised if corporations / super rich WANT you to think like that. The more defeatist the mentality becomes the more we just accept whatever they do next.
A bunch of SWEs at my work use it as their primary model.
We have Claude, ChatGPT, and Cursor with essentially no cap on spend (top guy is spending over 10K a month on AI at API prices), and he hasn't had his hand slapped.
So it's not like they are using it purely because it's cheaper.
I think people like to use it for its speaking style, pretty solid performance, and its speed.
It's a different type of model. In my admittedly judgemental observation, people that aren't the type to configure fully automated harnesses with good tools and skills and verifiers for their infrastructure and are way more interventionist in the way their agent works tend to like grok 4.5 more as the main agent. It's much faster and writes more simple and normal code that aligns a bit more with human written code. As an example, instead of sandboxing and simply verifying output artifacts they manually read and approve edits, suggest different code patterns, and manually approved shell commands. On the flip side it's not as good as fable when you need a relatively complex multi step thing done.
But in my experience, the overall productivity ends up similar, give you are willing to work with it in that way.
The grok build TUI harness is excellent and I really enjoyed using it.
For debugging and such I found it pretty much the same as other models.
Fable's taste in software abstraction and project planning in greenfield setups[1] is unmatched in my experience. Sol is OK. My primary use is launching tens of experiments that have to smartly use a limited pool of GPUs.
I use fable to start off the experiments, decide checkpoints, gpu alloc, where to sacrifice precision for performance, and then grok4.5 to iterate, tune, debug, eval, etc, within the abstraction and setup that fable initiated. I have fable write simple scripts that are then wrapped in skills for grok to use. Speed for that loop is extremely important for me, since I also apply human judgement there and I don't like waiting for model output.
I have tried Deepseek and such for the inner agent, but I desperately need multi-modal. Otherwise it's OK, but it tends to use tools less and rambles on and tries to reason with limited information and gets things wrong. Probably a relative la k of tool use posttraining. Gemini flash limits in google ai pro are too low for me to use to compare.
I use anthropic and openais models through grants and so can't compare subscription plan token budgets, but supergrok's budgets are satisfactory.
[1] aside, I have not yet met a model that continues off of a human codebase and actually follows the patterns reliably long term. Eventually it's all slop.
Yeah [1] is really a thing and SlopCodeBench (https://www.scbench.ai/) kind of measures that.
You need to manually push models to clean up the slop every now and then otherwise it becomes chaotic. And every change with LLMs is always extra lines.
It's much faster, so if you're not doing something cutting-edge, or you're doing the planning yourself and just using the LLM for implementation, the speed benefit outweighs the extra smarts of Fable/Opus5/Sol.
I've used all of the models extensively and Grok is only "faster" because it claims to be done minutes after you ask it to do something. It does not produce results anywhere near the Anthropic or OpenAI models, it just hacks a tiny piece of what you ask for and says "I'm done!". I also notice Musk-isms leaking through the model. Multiple times it's told me "this is not a roast" or "I'm not roasting your code". People who use this model because they align with Musk's ideology are doing us all a favor and weeding themselves out of the competition.
In most coding harnesses you can just instruct the agent to do so. When using Fable, I often say things like:
> Save your context. Always use a subagent (Opus 5 or GPT 5.6) for performing the implementation and then review the work yourself. You are the orchestrator and coordinator it’s up to you to ensure a cohesive final result.
I’ve done more or less the same thing with other agents/versions but with Fable consuming usage credits/tokens so quickly I do it more regularly than usual. I know people who will specify Composer (to my chagrin) as the implementing agent.
Addendum: this really goes a long way, and I can use a single chat session for days before I get into the context danger zone and need to compact/summarize.
I had a security incident the other day and Grok was the only model that would help. Claude and GPT refused on ethical grounds and only gave general advice. In an emergency, I'd only trust Grok. However, that's the only time I used Grok for coding (since Opus 4.8 it would take a lot to get me to switch away from Anthropic)
The "safety" guardrails in Anthropic and OpenAI models are becoming a noticeable problem for security work. And, the reason I'm keeping my Kimi subscription even though it's not a great deal; Kimi subscriptions are quite stingy for the price, but K3 will do vulnerability analysis and make a PoC without requiring you to be on the approved list of Fortune 500 or government entities that have access to Mythos or Daybreak.
I'm not touching Grok. But there are alternatives to Anthropic and OpenAI that don't refuse to do security work.
I use. I used to be a Claude user. Since trying Grok 4.5 and especially Grok 4.6, I don't want to go back to Claude any more (I have early access to 4.6).
Grok is 3x+ faster than Claude and I can't tell the diff in engineering work quality. As an engineer, speed is important to me.
An hour in, I've been running four terminals full bore on my $20/mo Grok sub and I'm at 9% for the week. Codex or Claude would easily have hit 5-hour or weekly limits.
I'm really not burning tokens fast enough. I use Claude a lot, daily, and have yet to hit my a ceiling with my Max/100 subscription.
Maybe because I like to verify its outputs and spend a lot of time iterating to get better outcomes. Presumably if I just let it "do its thing" I'd burn more tokens and "get more done" but I'd lose my grasp on what's in the code base.
Big same here with Claude but I’m on the $200 sub. I let Fable run for roughly 4 hours and still didn’t hit the session limit. I suppose if I was running more in parallel it would be easier to hit, but I’m not particularly good at focusing on more than one thing at the same time (even if I’m waiting on an agent to do the work).
I hear you. I usually have two things going at once (either two different machines, two different projects, or different work trees if the same repo), but more than that I don't feel like I'm locked in enough to providing the "thinking" that Claude definitely still needs.
For me it helps that back when xAI was the new hotness, after reading HN comments constantly advertising the free Grok credits they were giving out each month for developers with data sharing enabled, I eventually set aside my personal distaste figuring I might as well take advantage for Cline/Roo Code.
I opened a developer API account, loaded 5 dollars and got the free $100s of credits for the month. Like two weeks later, xAI announced they were shutting down the subsidized credits entirely lol. Didn’t even get a full month out of it, and closed my account entirely since I sure wasn’t ever going to put another penny of my own money in.
So my personal lesson was to ignore any hype about the latest “crazy value / unbeatable / free / subsidized X, Y or Z” from anything xAI/Elon adjacent in the future.
These days I get more than enough personal usage from Codex + OpenCode Go to put up with yet another xAI/Cursor offer treadmill, especially if it involves installing new tooling to get it.
When you make politics a big part of your identity, you can’t help but to inject it everywhere. It’s glaringly weird and annoying for those of us who are apathetic and just want to talk about tech, but that’s how it has been for a few years.
It's not virtue signaling. People are allowed to have morals. And they are allowed to act in ways that align with those moral beliefs.
> do you use Apple equipment despite Apple’s use of Chinese slave labor?
If someone said they didn't buy Apple because of this, I'd support them in that too.
People. Are. Complicated. We're allowed to have beliefs, we're allowed to have beliefs that aren't internally consistent. You gotta get over the idea that people aren't genuine believers in their beliefs.
The issue is that this is supposed to be a tech discussion forum and many of us don’t want to be repeatedly beaten over the head with other people’s politics. Plenty of other places to be political, we just wish this wasn’t one of them.
There's always been "politics" embedded in this tech discussion forum. Generally you shouldn't discuss politics for politics sake, but if someone's politics are intersecting with their product (which is Musk's entire MO), of course people will have thoughts.
Besides, if Musk embeds his politics into everything he does, it's fair game to discuss his merits in that arena.
Tech has always been political. You think Mitnick and Woz weren't political? You think windows and the powerpc werent political? That social media isn't political?
You don't like seeing politics you disagree with. You just don't think things you agree with are politics.
I come to HN to learn, to be exposed to different modes of thought. To engage in useful creative discourse over disagreements. The idea that HN or tech is free from politics is intellectually vacant thinking.
It's fine to make your own choices about what you want to spend your money on. But it gets old to constantly have actual technical discussion drowned out by the same repetitive low-effort comments in every single post. I want to hear about what people have used Grok for, how it compares to other models, what it's not good at, etc., without having to wade through all the predictable "Elon bad" comments.
@dang - feature request for HN: have a "Thunderdome" section for the off-topic political commentary on the site, and once enough people flag something for being political, it gets detached and moved to the Thunderdome. People can let off steam there and the rest of us can have a nice clean place to read about tech.
> The issue is that this is supposed to be a tech discussion forum and many of us don’t want to be repeatedly beaten over the head with other people’s politics.
Maybe that wouldn’t be a problem if Elon Musk would lay off the ketamine and shut his fucking mouth sometime. But instead he spends all day tweeting about how anyone he doesn’t like is a “traitor to western civilization” and calling for executions.
Why is it surprising that people are reacting to his words and actions? You want the politics to go away, get a CEO that doesn’t make everything political.
Your analogy is goofy. An analogy is like, "People don't go to other people's kitchens and discuss their thoughts on McDonalds". And yeah, they totally do.
The person isn't complaining about grok to grok, they aren't even complaining about grok on X.
I can respect if you say you hate their guts. Everyone has their worldview. But over moral or ethical stand? You don't have any if you're using Chinese models, or fly Middle East airlines, or countless of other products. Don't delude yourself.
I don’t like what is happening in my country, Elon is a huge part of that and I don’t want to reward it. I’m not a fanatic, but if all other things are equal I can certainly factor social responsibility into the equation. Fuck that rage baiting xenophobic nazi-salute throwing asshole. [edit spelling]
META releases open weight models because selling access to models isn't their business model, they get optimization for free, good press with adoption of their tech etc. They don't want AI to become another iOS, they want to commoditize the layer underneath them.
GOOGLE releases open weight models because they complement the rest – from premium offering to Android/Chrome/Google Cloud etc products. They use it as part of open ecosystem / local / edge / experimentation layer. 400 million downloads with 100k community variants is nice vibrant ecosystem they have and want to have, they can capture value in several places and would prefer if devs standardize on their tooling.
OPEN AI because their PR was shit and it saved them, purely defensive move, which is interesting considering their name and initial goal.
ALIBABA because it tickles them to erase/cap profits in western labs, free optimisation/research, great PR.
DEEPSEEK they want to have day-zero support for all possible hardware platforms, they started it because Liang wanted to do it, had resources from High-Flyer to do it and he said fuck it and did it (zero commercial incentives which is so bizzare that it deserves a movie or something). He wants to do it because his destination is AGI. In that sense he's more original OpenAI than OpenAI.
MISTRAL because they want to be to AI what RedHat is to Linux.
NVIDIA because they sell GPUs.
...the list is long, there are few dozens of companies including Microsoft, IBM, AI2, Databricks, Snowflake, xAI, EleutherAI, Hugging Face etc. that release open weight models, there are thousands of models and tens of thousands of fine tunes across across text, image, video, audio etc.
In general half life of any model is short. What is more lasting is ecosystem around it, releasing new generation as open doesn't give away technical advantage.
They give away asset for which strategic value is deprecating rapidly and in exchange they get developers, mindshare, tooling, optimizations, integrations, research, standarisation etc. and put pressure on competitors destroying their margins.
It's not an absurd way of thinking, it may seem like irrational move but let's wait and see – imho labs like Anthropic can't sustain long term this kind of pressure and will eventually collapse – regardless of the fact that currently they look like strongest player that can't be touched, the whole thing they have holds on thin, fragile support that gets eroded.
> ALIBABA because it tickles them to erase/cap profits in western labs, free optimisation/research, great PR.
we talked about Chinese models. Yes, erase/cap profits, PR, and give clients peace in mind that model access won't disappear. So, it is infiltraiting markets.
Your argument was that releasing models as open weight was allowing them to collect data but reality is that open weight models precisely allow users to completely avoid data collection – something that is not actually possible with closed weight models.
Motivations are more complex than that, if they wanted to achieve that they'd follow approach taken by some western labs to release weaker models only as open keeping frontier behind APIs.
Use of official API outside of China is closer to opposite of reality.
Look at ie. OpenRouter you'll see how many providers there are and how much traffic they get.
If the goal was data, opening weights would be the worst available way to achieve it: a cheap, closed API would capture 100% of traffic (ie Anthropic style), weights can only lose share from there.
With open weights it's net loss of active users of your official api – you're loosing users to self hosting and dozens of providers.
The thing is that open weights create permanent exit that closed models do not have, inference is commoditized immediately, people choose open weight models specifically for data privacy (and stuff like soc2/hipaa compliance) and if somebody wants convenience they go to claude/openai and friends anyway.
Also chinese labs are not uniform with their approach just as western labs are not.
once my codex/claude weekly limit was gone, i gave it a try. It was surprisingly good, not dumb in any way, and fast. I now require it as a part of 3-of-3 quorum with any codebase change.
macOS Codex app is my main agent (its very polished). When it makes changes or code scans i ask it to run claude -p plus cursor-cli plus grok cli for double checking. This way bias of one model can be overruled by quorum.
Yeah, sounds a bit like a selection bias, i.e. the people who managed to survive the firings by DOGE and the current admin are the kind of people who might prefer Grok.
I'm only being forced to use it at $WORK since some people overran their Cursor bill, so everyone gets Cursor Auto enabled by default which routes to Grok 4.5.
Back in the day (in AI time) GitHub Copilot had Grok on the 0 github-token cost and I found it to be the best of the 0 github-token models for when my budget was out. Then they went to a multiplier that was not competitive and I haven't look back again. Been meaning too, but for personal use, Deepseek flash is so cheap I haven't felt like spending money elsewhere.
I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.
That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.
> I have never met a single human being who uses Grok for coding
Me too. The only people I ever saw using grok were using it by accident as they used copilot in auto mode and noticed some prompts were thrown it's way.
That is good news for Grok team. However, most of the time cost comparing to the result is secondary, and better results and conclusions can come from mixing AI brains together.
The answer is simple: They're willing to burn money faster than the others.
Nobody's profitable in this space, they can price it however they want as long as investors keep pouring money in. And SpaceX just got a lot of money poured in.
Interesting. Grok 4.5 is a capable model, although not quite at Fable/Sol levels. Will be interesting to see how this holds up. Musk appears to have made a savvy choice buying Cursor's data.
Can someone explain to me what's the point of Grok anymore? I don't understand why we need a third or fourth closed frontier model. It is clear that chatGPT has locked down the consumer play, and may be Gemini is there. Claude has enterprise locked up, followed by chatGPT and Gemini. Enterprise switching costs are notoriously high, and even if they switch, they have chatGPT or Gemini to choose from. Beyond that, you have a vast array of open source models (DeepSeek, Kimi, and now Meta's Spark and Glimmer). So, why would anyone need a third or fourth frontier model and why would SpaceX spends billions in CapEx for a very small market share
There are a few reasons people will be interested...
1. The CapEx play is interesting because it's not just Grok using the hardware. They have rented out hardware for others, including Google, to use. This is making xAI money.
2. It appears that Elon is building a suite of things that work together as part of the push to be multi-planetary. What AI will power the robots? I can understand the drive to have AI they can control to make sure it's appropriate for all the things they are dreaming up. This is a piece they don't want to outsource.
3. OpenAI and Anthropic models are expensive in terms of token costs. Sure, they are frontier. Neither appears to be trying to drive down expenses. This is a problem for heavy users. Companies are trying to put cost controls in place. Does the rest of SpaceX want those cost controls? Having a Frontier model that pushes the pace of driving down costs is really useful.
4. OpenAI and Anthropic are producing models with a progressive lean, according to the Neutrality Project [1]. Having a frontier model that is closer to the middle is considered a good thing by many who are noticing the bias.
These are just some of the reasons. Competition is often a good thing that drives useful change.
Competition keeps service quality high and pricing low - even if you aren't using Grok, the mere existence of Grok keeps pricing for whatever provider you use lower and service faster and more reliable.
That's fair. But my point is from a business POV, why would SpaceX want to invests hundreds of billions of CapEx on a third or fourth frontier model, which cannot compete with chatGPT and claude on the high end, and getting squeezed by open weight models on the low end
They came here in three years. I think there’s a chance they will release the absolute top model soon. And I think they believe that as well. They have the compute. They have the money. They have the engineers.
I assume SpaceX does it because they figure it's a good source of revenue, and it means they can keep another thing in-house instead of relying on Anthropic or OpenAI for their AI needs.
As for why anyone else would want it: I've found it's a good model for coding, and it sometimes catches bugs that other models (especially open source models) don't always spot.
That’s how physical good work. But software, especially consumer software, works on a winner take all model. That’s why there are very few consumer companies and chatGPT is pretty much the only one after Meta, which was founded in 2004. The reason this happens is because consumer software can scale infinitely as there is zero marginal cost for a new user and there are very high switching costs. A single car company cannot scale to serve every single customer. Because it requires massive CapEx investment. But Google can serve every single search globally because the incremental cost to serve the additional consumer is essentially zero. That’s how these frontier models are eventually going to play out. There will be consolidation and winner take all. It has somewhat happened already with chatGPT taking over consumer and Claude taking over Enterprise. There will probably some long tail open source player, similar to Linux.
I sense a lot of condescension and lack of business acumen so I won't spend time writing it up, why don't you just ask your favorite AI or do some basic google searches? Elon was extremely upfront about the philosophical reasons for it, ever since he cofounded OpenAI.
Inference switching costs aren't high since the models are largely fungible. Even with proprietary harnesses you can hack them to use some other lab's model.
Its a bet for future world dominance by Elon: millions of robots managed by AI. Base on some observations I think something like that is going in his head.
Not disagreeing with you at all, but welcome to our new AI enhanced world. Nothing you enjoyed regarding human interaction, trust, or social norms is safe.
HN will not survive 5 years, and likely less. There is too much money to be made by capturing discourse on the major (and minor) forums of the internet. The more trusted that community is, the more valuable it is to pillage with AI astroturfing.
Agreed but there were similar controversies with how OpenAI was generating images. The only pass is these are the early days of chatbots and this stuff is so non-deterministic and experimental.
For context, this was the change Grok's team made, that was later reverted:
> - The response should not shy away from making claims which are politically incorrect, as long as they are well substantiated.
For sure. The difference is that they've made a number of similar suspect changes to Grok on X. Like that weird couple of hours where it would only talk about white genocide in South Africa no matter how you prompted it.
Everyone makes mistakes, especially with frontier models. The stuff with Grok shows that the person running the show has a pretty transparent agenda that the company isn't willing to push back on a bit for safety.
Yeah I don't care to use Grok's chat and find @grok responses on X mostly noise. I am still open to using it as a backup model for coding though, assuming it does a good job for the price. But mostly because I was already a Cursor user before they bought it.
yes exactly. if a company is happy to have their LLM's produce neo nazi content and CSAM, why do I want to give them money and my most important digital material?
The reason it is so believable is that he did it and it happened, in public.
He contributes to this understanding by straight-forwardly promoting and agreeing with white nationalists and even straight-up neo-nazi accounts on twitter, and by promoting neo-nazi aligned political parties.
I think "system prompt" is the key bit they're getting at. It doesn't necessarily reflect poorly on the underlying model if the system prompt was bad. It does reflect somewhat, in terms of alignment (how well the model does what the training company wants) and instruction following (how well the model does what the user wants). But it's not so clear to me what exactly the right answer is here. E.g., a model that scrupulously follows its system prompt and does what the user wants is a pretty useful, if very sharp, tool, albeit perhaps dangerous in the wrong hands.
MechaHitler is the final antagonist in the game Wolfenstein. All models know that. Grok was given a relaxed system prompt and reacted like Tay.
This worries me the least. The fact that Musk pushes AI and vibe coding is much more worrisome. It makes no difference to the unemployed if their jobs were stolen by a politically correct model or by an anti-woke model.
Google's image generator created black founding fathers because diversity dial was turned to 11. Does this mean i'll never use their products for political reasons? absolutely not!
I often wonder if there's a chance, even if minimal... that they stole the weights of the Anthropic models they run on their datacenter... or are actively destillating it.
I think the more likely explanation is that the Cursor data they effectively acquired for $10B was extremely valuable for their training when combined with the insane number of GB300s xAI has for training.
I found it has improved my productivity and output over claude where it felt like claude was giving me work to do. furthermore with the recent claude watermarking thing, I'd rather use grok or openai.
If anyone is curious download grok cli and throw a couple of prompts at it. you'll be surprised.
Claude suggests something, then later on suddenly it's something I wanted all along. ChatGPT is far superior to Claude. I will try Grok.
I am curious, is that a plugin, or skill?
Beside don't read too much in to benchmarks. They are alrrady ruined by Goldhart principle. Tgese models have already seen most of the data.
On a similar note, how much Elon hate is about his politics vs his trillionaire status vs what I like to call “watercooler hate” where folks simply parrot the loudest opinion in order to be accepted into the group?
On a final note, my son is in primary school and recently brought up in a dinner time discussion that “Elon is really bad” - this is a kid who has no social media (unlike some of his peers who are already on TikTok) and doesn’t watch traditional media.
He's also publicly and militantly supported quite a few far right political parties throughout Europe, particularly those who have a long track record on race-based topics.
My token spend at api rates is about $3000 usd a month recently.
[1] https://futurism.com/future-society/elon-musk-ai-woke-free-o...
He's become a toxic "brand".
but what's to validate any given population as centrist anyway
they suggested the ai opinion drift was caused by internet demographics not directly reflecting actual population i.e. California publishes more etc
It's the majority position amongst Republicans that the Earth was created less than 10 years ago as is and Jesus may rapture all the believers in our lifetime. They believe childhood vaccines are dangerous and those over 30 favor the idea that being gay is a choice and immoral at that because they have no intrinsic idea of what right and wrong mean.
They believe Obama is a crypto Muslim born overseas and therefore somehow not a citizen despite citizenship being heritable.
They believe that the same ballots that elected Republican reps didn't represent a valid vote against Trump because it's unbelievable that folks voted against a pedophile in 2020.
Hell a 27 year crack addict is on TV alleging the Republican primary was stolen from him in a blue state and this is serious business in Republican Land.
The right wing has swung so far to the right that 2010 Republicans would be considered Left wing
We tend to accept young earth creationism because perceptively it has little to do with everyday life doesn't hurt anyone and people are quite open about it. It also means that your brain is literally broken, you have no critical thinking skills and you are just the right sort to be manipulated by evil and end up building or guarding the concentration camps.
According to Gallup 95% of Republicans believe this at a strong majority 58% believe that the earth was quite literally created as is less than 10k years.
While this is magnificently stupid in and of itself it's not the problem it's merely a symptom of their disease.
They largely believe that dead kids aren't a reason to reign in guns, they believe climate change is a hoax, a cycle, or inevitable and won't do anything about it or let anyone else. They believe that us aid to poor nations was 10-15% of our budget and we just can't afford to not let kids starve when it was one-half of 1%. They believe that it may be necessary to use violence against the rest of us to maintain the status quo. They believe that we need an stong authorian to lead us democracy be damned, they believed that COVID was a scam until the corpses piled up or even after!
Here's a hilarious one. It's a graph of the percentage of the population which understood that COVID was worse than seasonal flu a march through April 2020. For Dems it goes from 74 to 87%. For Republicans it goes from 42 to 40 it actually drops whilst body count goes up.
https://news.gallup.com/poll/311408/republicans-skeptical-co...
They are magnificently staggering wrong on average about everything. We could go point by point and I could find polls for every single one because I'm basing my understanding of their position by reading their own words and polls by reputable sources like pew and Gallup.
The right Wing collectively has departed so entirely from reality that educated con artists must carefully contort their words to avoid their insanities and prejudice whilst fools spout forth with abandon uncaring.
ChatGPT is very long-winded, sometimes producing multiple bullet point lists for a simple answer. Claude is full of mannerisms: 'not merely x, but y', 'Here's where it gets interesting', 'the real question is', etc.
Until it doesn't...
Honestly, this entire OpenAI reset credit fiasco this past week has convinced me to rip off the Codex and Claude Code bandaids and start building my own proper Pi Coding Agent running models that I select and pay for on openrouter.
And I am feeling a lot better about it now that I've finally got it working.
I don't get the point of this. We all seem to agree that these companies have almost no moat, if one stops being a good deal, you can switch to another. That doesn't invalidate the existence of a deal that is currently good.
My point was that chasing deals like this is just kicking the can down the road. You're going to have to reckon with harsh price increases sooner or later.
So I have resolved to avoid that future-dreading and fixed it, basically.
Seems weird to me to not take advantage of the great deal the frontier labs are currently giving for subscription pricing when there's such an easy fallback in the worst case.
the rest of the comment is about moving from expensive switching costs of what harness you are using to having the difference be closer to changing some config in openrouter
Deepseek V4 Flash 0731 is surprisingly capable and cheap. [0]
Checkout pi coding agent. You can create as many different sub-agents as you wish, to specialize and understand and tackle or pass off any problem you like. It's refreshing, really. I feel like a coder in control again.
[0] https://arcprize.org/results/deepseek-v4-flash-0731
Huh, what's happened? I'm on the 20x plan and haven't noticed any fiasco, what went down exactly?
But since you don't know when resets are coming, it becomes kind of frustrating trying to game it out. One week they're resetting like crazy so you're trying to burn tokens as fast as possible. The next week they're not resetting at all so you have to adjust your workflow to be more conservative.
This whole thing sounds crazy to me, why are you so focused on making sure you hit 0% usage left when it's supposed to reset? Why can't you just use what you have and if it resets, it resets, and if it doesn't, it doesn't?
I've calculated I can spend about ~12% of usage every day, more or less, this is my "budget". If it resets, then the "12% per day" gets reset for that day, but that's it. Sounds crazy to me that I'd "invent work out of nowhere" just to spend more usage, why on earth would I adopt such a workflow? Sounds like you're burning tokens just to burn tokens???
Sure, in a video game where you have one attribute and you can minmax a strategy just focusing on that, but that's not how real-life works.
You can't just stack pending work on top of each other, expect yourself to be able to stay equally on top of everything and have the same results as if you didn't. With this comes the consideration about the tradeoffs of "produce mediocre but large body of works" vs "produce high quality but small body of works".
Who's to say what's more "rational" or not, it's not a straight-forward calculation which culminates in "Must consume all available usage to maximize resource usage" like some robot, as we are not.
It's quite literally inventing work out of nowhere as you wouldn't put the agent to do that and forcing yourself to be conscious about that work until it completes, unless you actually had the usage available. Asked another way, wouldn't your workflow clearly change if you had unlimited usage available? You'd probably attack tasks/problems that you didn't consider actually spending time/energy solving.
Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…
https://harness.eveid.com/lazy-harness-cost-simulation
https://artificialanalysis.ai/models/grok-4-6#token-use
So depending on how you want to define "token efficiency", Grok is either tied with OpenAI, or in the lead.
[1] Though I grant that 4.6 appears to be wordier, on the order of Terra max.
I'm willing to pay 2x for a 10% smarter model. Intelligence matters that much (because 10% smarter probably saves, on average, several hours of human time).
I haven't tried so it's pure speculation based on benchmarks, but I'd assume Grok 4.6 is around Opus 4.8 in real world use, but clearly below Opus 5.
For building a full stack custom CRM and media pipeline tool with video conversion, transcription, and indexing. Supabase, AWS, Meili, NextJS, GCS - lots of surfaces and planes.
4.8 basically couldn't do it, I abandoned the project as the fallback was, "current business processes".
With F5 it's been 4 weeks and almost ready for production release.
It’s not as quite as smart as opus 4.8 but it’s close and x4 the cheaper.
https://github.com/features/copilot/plans
https://github.blog/changelog/2026-08-06-kimi-k3-is-now-avai...
Have they converted entirely to transparent API rates + base allocation now? One of the reasons I left was that if I was going to be billed at API rates anyway, I'd just rather use the APIs. The value proposition still sucks for individuals now, when the other major providers are bundling at below-API rates.
Apparently 5x usage when using Kimi Code too.
If you willing to share to no zdr, meta is waaaaaay cheaper vs Kimi.
With recent offerings from spacex and meta , I hardly imagine why would you pay money to any Chinese vendor it’s not as cheap and it’s not as intelligent neither .
Maybe deepseek is an exception , but it’s only good for narrow use cases that probably goes into modal.com and other gpu + fine tune me easy vendors , not vanilla dumb but cheap model .
https://www.theguardian.com/technology/2026/jan/15/elon-musk...
I’m confused by how your reply makes sense in the context of the parent comment. They weren’t stating a preference, they were linking to simple facts.
Please do better.
They claim that spacex has competitive advantage. They take no stance on condoning spacex’s behavior. Personally I strongly dislike musks behavior, but I appreciate discussion of competitive advantages/disadvantages, independent of moral views. I liked the thread, it adds new info to the conversation.
No, they're not. You're making up things and pretending that they said them.
> Please do better.
You're badly breaking the HN guidelines. Please review them: https://news.ycombinator.com/newsguidelines.html
This is factually incorrect. I did not say anything "snarky" or "mean". I stated facts and pointed out that the poster was breaking the guidelines.
Taxis are a different and wholly irrelevant product that doesn’t absolve the fraud that they committed.
And extreme skepticism is warranted that they’re actually fully unsupervised, given Tesla’s repeated lies about this. They’re likely remotely monitored and operated.
It might or might not work long-term, but I wouldn't count on courts to help.
This is a myth.
They obviously have a huge token cost advantage of the AI labs they are renting compute to, at least for now while they can charge current crazy rates for GPU compute.
they just increased cache read from 0.30 to 0.50 - this has the biggest impact on agentic coding. Elon companies have the most expensive everything:
xAI sub: $30 when other starts at $20, pro like sub for $300 where other charge $200.
Expensive electric cars, powerwalls, solar roofs when competetive products/better are cheaper.
The Model 3 and Model Y became the highest selling EVs of all time because they were the first below $50K to have long-range and be worth buying.
Until a few years ago, every other sub-$50K EV absolutely sucked.
And then Musk totally abandoned Tesla's original brilliant game plan of using the luxury models to find actual low-cost models (which $50K is not), and completely ceded the future EV market to Chinese companies that understand how to make a better car than Tesla for less money. BYD sells more cars than Tesla.
This is meaningless because you're not controlling for amount of subsidized usage or model quality.
Google would like a word. Also, Microsoft.
To say that I somewhat doubt this would be an understatement.
My current headcanon is that he remains an incredible leader and make-things-happen-er, but also that you should just assume all his announcements are complete bs and ignore them.
There's no need to assume. If you are actually educated on the subject, you know it is the case. You don't because you're ignorant about it. Which would be fine, if you weren't pontificating about things you know nothing anything about.
> current day Twitter addicted, drug addicted, brain damaged Musk
this is like conspiracy-theory level stuff.
It's good to have model diversity. When I run a task across Sol, Terra, and Luna, I get variations of the same thing with diminishing quality. It makes the lineup pointless. Ditto for Anthropic. Gemini-3.6-Flash and 3.1 Pro genuinely behave differently. Opus 5 and Fable are.. cousins.
I find that when I want to test a complex creative challenge, having 4 "families" to choose from makes the experience interesting since they will excel in different areas.
Grok might implement unique lighting, Opus, elegant primitives, Sol, accurate snowfall in one pass, Gemini, silky movement. Combined, you can pick and choose best.
For what its worth, Grok always feels "messy" but finishes. Grok 4.6 though is no longer "smart and fast". It's about as fast as Sol though.
A big improvement I noticed in 4.6 was tool use for verification. Previously, Opus/Fable were the only models to consistently screenshot things that they can't directly interact with easily. Now Grok is probably right behind them, perhaps tied with Sol on propensity to verify visually. Grok 4.5 notably did not do this often.
In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill.
The rental deal can be terminated by either side with 90 days notice, and presumably Musk would do so if he needed the compute or generally thought it advantageous to do so. For now he doesn't need the compute.
The rental deal may also have been at least in part to juice the SpaceX IPO and to help Anthropic stick it to his enemy OpenAI.
Deals like this that look awkward from the outside but are mutually beneficial to both participants exist everywhere.
Feedback: I'd like Grok to have more connectors (I see OpenAI just added Apple Health, that would be nice to have, and I wish it could read my Onenote notebooks) and for existing ones to be improved. I gave it access to my gmail and asked it "what was my last electricity bill?". It failed to find it, even when I told it the exact subject line to search for. Something about not getting any data back when trying to get the email contents.
Does xAI have plans to do Mac/iOS apps with cloud environments? When can we expect them?
https://t3.codes/
https://happier.dev/
https://paseo.sh/
All of these also enable multi vendor LLMs.
What the heck is wrong with you???
Pretty good bang for the buck.
here's Stanford HAI's graph on the carbon emitted from model training per model:
https://spectrum.ieee.org/media-library/chart-showing-estima...
note that Grok's training, thanks to its portable gas generators that are magnitudes less efficient than even other integrated, permanent gas turbines, means the training for this model is dramatically less efficient than models like DeepSeek
a lot of the CO2 emission debate on AI is overblown but it's accurate for Grok
Either it was deliberate, or the richest man on earth is so incompetent that he accidentally made a motion exactly mimicking a seig heil while welcoming a self-proclaimed king and dictator. Twice - once towards the audience, and once towards the flag of the united states. Neither option is good.
What was it?
Who cares about the pictures? There is a whole video where you can see that he does historically accurate nazi salute then turn back and do it again. Stop talking about pictures, always show video of what happened.
Find me a video of someone else doing the same and tell me it's the same. You won't, because what you see in video is much more obvious that what you see on pictures.
People are dump to only show still photo of Musk with this salute. Everyone know you can take picture out of context, when you watch full video it's much harder to do so.
Must was clearly showing it intentionally. Maybe because he's a troll (I wouldn't be surprised), maybe because of other reasons (I hope not...).
The AfD who he appeared in support of has been found in court to use nazi imagery in political advertising (https://www.world-today-journal.com/afds-moller-fined-e12000...).
So the seig heil didn't happen in isolation, there's a followup of explicit support for nazi-adjacent political ideologies.
Does this count as evidence to support that this was an intentional seig heil to you, even if it doesn't change your mind? What evidence would change your mind?
You're spouting easily refuted nonsense, and then immediately making ad hominem attacks because your points are extremely weak and not backed up by the evidence you vaguely cite.
You're also getting ratioed because people here are not idiotic ideologues. Please do better.
---
I'll also leave you with a nice quote from Sartre which was directed at the fascists of the time - we've seen this shit before, we know what you're doing:
“Never believe that anti-Semites are completely unaware of the absurdity of their replies. They know that their remarks are frivolous, open to challenge. But they are amusing themselves, for it is their adversary who is obliged to use words responsibly, since he believes in words. The anti-Semites have the right to play. They even like to play with discourse for, by giving ridiculous reasons, they discredit the seriousness of their interlocutors. They delight in acting in bad faith, since they seek not to persuade by sound argument but to intimidate and disconcert. If you press them too closely, they will abruptly fall silent, loftily indicating by some phrase that the time for argument is past.”
Idk I'm not a Zionist, but it's clearly what's happening.
What evidence in that video do you refute? What evidence would change your mind?
did he just cough and his arm did that, twice?
does he have some form of muscle spasming thing im not aware of?
Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.
Elon would be in prison for SEC violations if the current administration hadn’t been elected, and that’s only the tip of the iceberg with that guy.
Elon is directly responsible for Grok becoming self-titled “MechaHitler” which the other three haven’t come close to matching yet.
> The Wisconsin Elections Commission last week referred two complaints to the Brown County district attorney’s office, which can choose to bring criminal charges over violating the state law against election bribery. Prosecutors have 40 days to report back to the commission.
https://apnews.com/article/elon-musk-wisconsin-election-mill...
> Encouraging people just to vote
And "Criminal Conspiracy" is just making plans with friends. Just because you can describe it in vague terms doesn't make it A-Okay.
> It's going to backfire hard for all the ad spend by MTV, Meta, and Google
It will not, because running ads is definitely not election bribery, whereas what Musk did likely is under Wisconsin law.
The one that often has trouble running it's events as businesses try to avoid serving them. And which has prominent members that the secret service considers definitely fascist.
Musk has gone a long way beyond simply being the standard rich person lobbying for their own interests, and has moved into actively promoting people that are trying to tear down liberal democracy.
Calling it "interfering with elections" is utterly bizarre to me.
thats a very dumb reason considering all rich people do it, most are just not as open about it as Musk
That's before we even start talking about the models themselves. He claims to want "unbiased" models, but he very clearly has a distorted view of the world and has repeatedly demonstrated a desire and willingness to bend the world to his will. I don't want to use a model that is so obviously suspect. Not to mention, his models repeatedly produce racist, Nazi-like propaganda.
IMO, he is, at best, a clueless amateur masquerading as an expert and running into problems a more careful person manages to mostly avoid. At worst... well, you get the picture.
It's a really, really easy line to draw in the sand: don't support openly corrupt individuals.
It's absolutely true there's money in politics. To call all money in politics equally corrupt because Bernie got a dollar to have dinner with someone vs Elon effectively directly buying votes....
A complete lack of nuance here. And I wouldn't be surprised if corporations / super rich WANT you to think like that. The more defeatist the mentality becomes the more we just accept whatever they do next.
We have Claude, ChatGPT, and Cursor with essentially no cap on spend (top guy is spending over 10K a month on AI at API prices), and he hasn't had his hand slapped.
So it's not like they are using it purely because it's cheaper.
I think people like to use it for its speaking style, pretty solid performance, and its speed.
But in my experience, the overall productivity ends up similar, give you are willing to work with it in that way.
The grok build TUI harness is excellent and I really enjoyed using it.
For debugging and such I found it pretty much the same as other models.
Fable's taste in software abstraction and project planning in greenfield setups[1] is unmatched in my experience. Sol is OK. My primary use is launching tens of experiments that have to smartly use a limited pool of GPUs.
I use fable to start off the experiments, decide checkpoints, gpu alloc, where to sacrifice precision for performance, and then grok4.5 to iterate, tune, debug, eval, etc, within the abstraction and setup that fable initiated. I have fable write simple scripts that are then wrapped in skills for grok to use. Speed for that loop is extremely important for me, since I also apply human judgement there and I don't like waiting for model output.
I have tried Deepseek and such for the inner agent, but I desperately need multi-modal. Otherwise it's OK, but it tends to use tools less and rambles on and tries to reason with limited information and gets things wrong. Probably a relative la k of tool use posttraining. Gemini flash limits in google ai pro are too low for me to use to compare.
I use anthropic and openais models through grants and so can't compare subscription plan token budgets, but supergrok's budgets are satisfactory.
[1] aside, I have not yet met a model that continues off of a human codebase and actually follows the patterns reliably long term. Eventually it's all slop.
You need to manually push models to clean up the slop every now and then otherwise it becomes chaotic. And every change with LLMs is always extra lines.
> Save your context. Always use a subagent (Opus 5 or GPT 5.6) for performing the implementation and then review the work yourself. You are the orchestrator and coordinator it’s up to you to ensure a cohesive final result.
I’ve done more or less the same thing with other agents/versions but with Fable consuming usage credits/tokens so quickly I do it more regularly than usual. I know people who will specify Composer (to my chagrin) as the implementing agent.
Addendum: this really goes a long way, and I can use a single chat session for days before I get into the context danger zone and need to compact/summarize.
I'm not touching Grok. But there are alternatives to Anthropic and OpenAI that don't refuse to do security work.
Grok is 3x+ faster than Claude and I can't tell the diff in engineering work quality. As an engineer, speed is important to me.
Maybe because I like to verify its outputs and spend a lot of time iterating to get better outcomes. Presumably if I just let it "do its thing" I'd burn more tokens and "get more done" but I'd lose my grasp on what's in the code base.
I opened a developer API account, loaded 5 dollars and got the free $100s of credits for the month. Like two weeks later, xAI announced they were shutting down the subsidized credits entirely lol. Didn’t even get a full month out of it, and closed my account entirely since I sure wasn’t ever going to put another penny of my own money in.
So my personal lesson was to ignore any hype about the latest “crazy value / unbeatable / free / subsidized X, Y or Z” from anything xAI/Elon adjacent in the future.
These days I get more than enough personal usage from Codex + OpenCode Go to put up with yet another xAI/Cursor offer treadmill, especially if it involves installing new tooling to get it.
Do you think being apathetic is a good thing?
You may not care about politics, but politics definitely cares about you.
> do you use Apple equipment despite Apple’s use of Chinese slave labor?
If someone said they didn't buy Apple because of this, I'd support them in that too.
People. Are. Complicated. We're allowed to have beliefs, we're allowed to have beliefs that aren't internally consistent. You gotta get over the idea that people aren't genuine believers in their beliefs.
Besides, if Musk embeds his politics into everything he does, it's fair game to discuss his merits in that arena.
You don't like seeing politics you disagree with. You just don't think things you agree with are politics.
I come to HN to learn, to be exposed to different modes of thought. To engage in useful creative discourse over disagreements. The idea that HN or tech is free from politics is intellectually vacant thinking.
I agree completely. That's why I won't use Grok: its owner repeatedly beats us over the head with his politics, and I won't encourage it.
@dang - feature request for HN: have a "Thunderdome" section for the off-topic political commentary on the site, and once enough people flag something for being political, it gets detached and moved to the Thunderdome. People can let off steam there and the rest of us can have a nice clean place to read about tech.
Maybe that wouldn’t be a problem if Elon Musk would lay off the ketamine and shut his fucking mouth sometime. But instead he spends all day tweeting about how anyone he doesn’t like is a “traitor to western civilization” and calling for executions.
Why is it surprising that people are reacting to his words and actions? You want the politics to go away, get a CEO that doesn’t make everything political.
Oh but it is. Normal people don't go to McDonalds shouting: "I refuse to eat here because they kill cows!". They just... don't go there.
The person isn't complaining about grok to grok, they aren't even complaining about grok on X.
Making it open weight does the opposite – it allows traffic to not be routed through them at all.
META releases open weight models because selling access to models isn't their business model, they get optimization for free, good press with adoption of their tech etc. They don't want AI to become another iOS, they want to commoditize the layer underneath them.
GOOGLE releases open weight models because they complement the rest – from premium offering to Android/Chrome/Google Cloud etc products. They use it as part of open ecosystem / local / edge / experimentation layer. 400 million downloads with 100k community variants is nice vibrant ecosystem they have and want to have, they can capture value in several places and would prefer if devs standardize on their tooling.
OPEN AI because their PR was shit and it saved them, purely defensive move, which is interesting considering their name and initial goal.
ALIBABA because it tickles them to erase/cap profits in western labs, free optimisation/research, great PR.
DEEPSEEK they want to have day-zero support for all possible hardware platforms, they started it because Liang wanted to do it, had resources from High-Flyer to do it and he said fuck it and did it (zero commercial incentives which is so bizzare that it deserves a movie or something). He wants to do it because his destination is AGI. In that sense he's more original OpenAI than OpenAI.
MISTRAL because they want to be to AI what RedHat is to Linux.
NVIDIA because they sell GPUs.
...the list is long, there are few dozens of companies including Microsoft, IBM, AI2, Databricks, Snowflake, xAI, EleutherAI, Hugging Face etc. that release open weight models, there are thousands of models and tens of thousands of fine tunes across across text, image, video, audio etc.
In general half life of any model is short. What is more lasting is ecosystem around it, releasing new generation as open doesn't give away technical advantage.
They give away asset for which strategic value is deprecating rapidly and in exchange they get developers, mindshare, tooling, optimizations, integrations, research, standarisation etc. and put pressure on competitors destroying their margins.
It's not an absurd way of thinking, it may seem like irrational move but let's wait and see – imho labs like Anthropic can't sustain long term this kind of pressure and will eventually collapse – regardless of the fact that currently they look like strongest player that can't be touched, the whole thing they have holds on thin, fragile support that gets eroded.
we talked about Chinese models. Yes, erase/cap profits, PR, and give clients peace in mind that model access won't disappear. So, it is infiltraiting markets.
Motivations are more complex than that, if they wanted to achieve that they'd follow approach taken by some western labs to release weaker models only as open keeping frontier behind APIs.
Look at ie. OpenRouter you'll see how many providers there are and how much traffic they get.
If the goal was data, opening weights would be the worst available way to achieve it: a cheap, closed API would capture 100% of traffic (ie Anthropic style), weights can only lose share from there.
With open weights it's net loss of active users of your official api – you're loosing users to self hosting and dozens of providers.
The thing is that open weights create permanent exit that closed models do not have, inference is commoditized immediately, people choose open weight models specifically for data privacy (and stuff like soc2/hipaa compliance) and if somebody wants convenience they go to claude/openai and friends anyway.
Also chinese labs are not uniform with their approach just as western labs are not.
Say more about this.
I don't give a shit about Elon's politics in the same way I don't give a shit about Dario or Altman's politics.
That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.
Me too. The only people I ever saw using grok were using it by accident as they used copilot in auto mode and noticed some prompts were thrown it's way.
I saw far more people using Mistral than grok.
I used auto in cursor it’s much faster va Claude code and as good.
Nobody's profitable in this space, they can price it however they want as long as investors keep pouring money in. And SpaceX just got a lot of money poured in.
1. The CapEx play is interesting because it's not just Grok using the hardware. They have rented out hardware for others, including Google, to use. This is making xAI money.
2. It appears that Elon is building a suite of things that work together as part of the push to be multi-planetary. What AI will power the robots? I can understand the drive to have AI they can control to make sure it's appropriate for all the things they are dreaming up. This is a piece they don't want to outsource.
3. OpenAI and Anthropic models are expensive in terms of token costs. Sure, they are frontier. Neither appears to be trying to drive down expenses. This is a problem for heavy users. Companies are trying to put cost controls in place. Does the rest of SpaceX want those cost controls? Having a Frontier model that pushes the pace of driving down costs is really useful.
4. OpenAI and Anthropic are producing models with a progressive lean, according to the Neutrality Project [1]. Having a frontier model that is closer to the middle is considered a good thing by many who are noticing the bias.
These are just some of the reasons. Competition is often a good thing that drives useful change.
[1] https://neutralityproject.org/
As for why anyone else would want it: I've found it's a good model for coding, and it sometimes catches bugs that other models (especially open source models) don't always spot.
https://www.autoblog.com/news/nissan-reports-fifth-straight-...
that alone makes it the closed source subscription i would choose. claude and openai are spying on you.
as it stands i don't have it because the reasoning is encrypted, so i feel that it still is not working for me, it's two faced.
HN will not survive 5 years, and likely less. There is too much money to be made by capturing discourse on the major (and minor) forums of the internet. The more trusted that community is, the more valuable it is to pillage with AI astroturfing.
>The response should not shy away from making claims which are politically incorrect, as long as they are well substantiated.
Interestingly, "politically incorrect" is a double negative that simplifies to "true".
> Interestingly, "politically incorrect" is a double negative that simplifies to "true".
Only if you like generic Twitter quips, logical fallacies and ignoring context for anything remotely nuanced.
Whatever you want to call your super duper clever rewording of a reductionist statement.
That's different than using Grok as a model for coding.
Oh whoops. Already happened.
For context, this was the change Grok's team made, that was later reverted:
> - The response should not shy away from making claims which are politically incorrect, as long as they are well substantiated.
https://github.com/xai-org/grok-prompts/commit/c5de4a14feb50...
Everyone makes mistakes, especially with frontier models. The stuff with Grok shows that the person running the show has a pretty transparent agenda that the company isn't willing to push back on a bit for safety.
He contributes to this understanding by straight-forwardly promoting and agreeing with white nationalists and even straight-up neo-nazi accounts on twitter, and by promoting neo-nazi aligned political parties.
https://en.wikipedia.org/wiki/Elon_Musk_salute_controversy
This worries me the least. The fact that Musk pushes AI and vibe coding is much more worrisome. It makes no difference to the unemployed if their jobs were stolen by a politically correct model or by an anti-woke model.
https://rainn.org/rainn-statement-on-use-of-xais-grok-to-pro...
I just assumed every model manufacturer is distilling from the frontier models. If they aren't they are definitely trying to do it.