It may surprise some people here to see that Andrew is warming up to using LLMs to discover bugs (inspired by results from SQLlite) and considers it a tool on the path to getting to bug free software.
> It may surprise some people here to see that Andrew is warming up to using LLMs to discover bugs (inspired by results from SQLlite) and considers it a tool on the path to getting to bug free software.
Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered. Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?
> Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered
I don't know about this particular case, but if I saw someone report a bug and as evidence claim they had X Y and Z LLMs verify it I would be pretty upset. If you're going to use an LLM to make a replication, just do that and give me the replication, don't point to your notoriously error-prone tools as though they lend your report credence.
It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"
Human reproductions are notoriously error prone whether ai assisted or not right? I’m not sure what the analogy is between “llms are error prone” and “thoughtlessly copy-pasting something from Claude” is.
Human reproductions don't have to be bad. It feels just as justified to push back on a bad bug report whether human or AI and say "I don't have enough to go on here".
If a report is improved and becomes actionable, that's great.
I feel very torn as a maintainer on this, since the only way to respond to the increased noise from AI has been to have AI do the research for me to extract all the links and line numbers I used to have to find by hand to explain why the PR needs more effort to be completed. But I also will be highly dismissive of any submitter who just posts AI text without cleaning it up first. It feels hypocritical, but the alternative is that I just can’t respond to most people instead due to limited bandwidth. I mark if something comes directly from the LLM though and to try to express my degree of confidence in its claims.
He specifically says in that presentation that they are open to LLMs helping them get to bug free, but because the language is still in flux, they would rather prioritise bugs users actually find rather than those found by LLMs, essentially with the intent of unblocking people rather than wasting time fixing things that may need to be fixed again or be wasted work come the next update.
Shouldn't everybody already know that while using AI to find bugs for oneself is amazingly efficient, using AI to submit bug reports for others is quite the opposite? Burden of verification and all.
Exactly, the main issue isn't the LLM creating or confirming the report. Rather, the maintainer has no idea how it was prompted, and the results may be completely wrong. Some LLMs also have a bad habit of trying to please the user, confirming their biases.
A lot of bug reports aren’t valid, or handle a case that can’t realistically happen or realistically be handled (eg what do you do if you detect a crash while the prior crash is crashing and the logging pipe is throwing errors)
Sure but this is true regardless of the model. Either it found a bug or it didn't—who cares about the intermediary steps or tools used so long as the reporter can reproduce it?
I personally think using AI is of no problem for certain cases. It is absolutely pissing off when someone tries to generate slops that too verbose to review only to increase the complexity of the codebase meaninglessly.
I’ve written software for a living in JS, C, Pascal, and Go, and I’ve tried many, many more languages. After working on a project in Zig for a year, I’m convinced that Zig is the best-designed language I’ve tried so far. Haskell comes close. At least, it’s the best-designed language for humans.
It’s not for everyone yet. It’s still unstable, and its ecosystem is small. However, both are improving.
I'd like to learn more about the connection between Zig and Haskell as they seem almost opposite in design and philosophy. I'm pretty sure you just meant they're both "good languages" but if there are parallels I'm missing I'd love to know!
I'm not the person you asked, but I see them both as languages that try to get a lot of mileage out of a few features. Both of them try to have small cores, instead of taking a "maximalist" approach like C++.
Congratulations to zig team. Good to know they are trying pragmatic approach to LLM now. I left the zig eco-system due to zig core members hostile behaviour toward humans not just LLM, so good to see the change they are becoming pragmatic. For me I am slowly porting the same work to odin language [1].
Background is I created issue and one pull request to fix them in zig compiler version 0.16.1 issue numbers 36812, 36811 (you cannot access them as my account is banned can see my fork at [2]). Respecting the community’s no AI stand. For these specific issues I wrote the issue and code myself and not let AI write it. Spend a lot of time on it. Subsequently without any notice my account was banned because my projects on github using zig uses LLM. This was done without message or any information. My account was banned on zig repository.
I can now understand the other side of coin how bun team might have been treated with disdain when they used LLM.
I wrote an email and left the zig community, have many work in zig but slowly moving them to odin.
I feel personal disdain should not be spilled on to people who are pragmatic on using LLM. I was a very big evangelist of zig for their no LLM stand and promoted them among my community, but with poor treatment by community I just left. You can still see projects I wrote in zig [3].
I have worked with postgreql community since 1997 and python community since 1998. Never felt such hostile community. So all the best and I wish zig continue its progress
I don't think the zig community is necessarily hostile to LLMs, they just don't want it in their "maintained by 10-ish core people not even full time" language impl. Mitchell Hashimoto, a big zig contributor (both money and effort), for example, uses LLMs a lot and no particular shade is thrown.
They seem to be doing fine. Tigerbeetle and Ghostty are two projects that continue to do well and are written in Zig. The foundation’s funding also seems to be doing well. They are continuing to make releases.
Strange in what way? Mitchell has also stated that he reviews every line of code that goes into Ghostty so I don't think its quite the same as the Bun rewrite.
Filed a compiler bug related to dwarf tables that screws up debugging and line of code coverage that they completely ignored, just because I mentioned that I had every LLM check it to confirm it's a bug, since all I know is what kcov and every coverage tool generates incorrect coverage data for my repo, for 100% certain.
Reading their comment it sounds like they wanted to confirm the bug existed with AI, not sure they ever said they had AI write the code. Why have such a dumb policy?
The things you've quoted and your conclusion feel at odds. They just don't want AI contributions, and they, like a lot of the world, are bored of hearing about AI. Is it really too much to ask?
The second part affirms that it is centralized. There is nothing wrong with that except saying that because you can leave the community and go elsewhere, it is ipso facto decentralized.
If there's an edict that no one is allowed to bring up the topic, how can someone change this part of the code of conduct?
Asking non-rhetorically. It seems like one position "ai in any circumstance = bad" is being enforced. The commenter above didn't even understand why he was ignored.
Interesting to hear! We maintain the project Antfly (entirely zig) and have been nervous about bringing issues to the zig folks or asking questions because of our ai usage.
We’re quite knowledgeable and thoughtful folks fwiw
I think I’ve seen you are a core team member or contributor? I remember your tag?
So for instance because of the size of our codebase our project has pushed Zig to some of the edges, specifically we end up hitting a bug when using llvm and zig on arm64 (Mac and Linux) where it seems to be caused by some configuration Zig passes through to LLVM. I’ve used codex and claude to help me diagnose and find the bug (we use nix’s glibc zig to circumvent the problem now). I now understand the root cause but am not sure what the proper fix would be. But I’ve not known whether or not even raising the issue would break the terms of contributing? Would raising the issue break the implicit agreement?
I’ve not found the zig folks to ban dissent, they engage in a lot of thoughtful dialog. Just because they’ve made a different decision for how they take contributions than other people agree with doesn’t make them a cult?
It's clearly a hobby project (constant breakages, the maintainer getting into politics, rejecting some safety mechanisms, the anti-LLM crusade, a strange focus on esoteric targets with little to no commercial significance), but the maintainer does not admit that it is a hobby project.
Wow I thought surely they wouldn't object to using AI to confirm bugs, but they really do.
Tbf I guess as a popular open source project not using AI to fix bugs, they probably already have more open bugs than they can ever fix so it doesn't really help them for people to find more.
I would imagine his bug was actually ignored just because Zig has 2700 open bugs, rather than some AI policy violation.
They don’t, at least not any more. Andrew Kelley has explicitly stated that he sees the value of using LLMs to uncover bugs. Inspired by sqllite project.
They may just be taking the slow route of rejecting by default until they can be sure that the usage of LLMs provides long term value. I don’t see anything wrong with that. If you’re writing robust software, using LLMs at this stage is a bit of a gamble. We don’t fully know the long term effects on code quality yet.
Project maintainers know what level of AI use is acceptable for themselves, and quantifying it and enforcing it for random contributors is very difficult.
Encountering something that seems incorrect, and proving something is a bug and not just YOUR user error, AND reproducing it minimally - turns out to not be easy when it's a low-level compiler issue, and YOU'RE not an expert, especially when it relates to Dwarf tables...
Typically, any time you think you've found a compiler error, you're using it wrong...
If it's a real bug, you can just report what you know and leave it at that. You don't need to embellish it with random guesses about a root cause.
Okay: "I compile this input and the linker crashes."
Not okay: "I compile this input and the linker crashes. Also here's 10 paragraphs of slop about dwarf tables, which I can't even evaluate the accuracy of since I'm not an expert."
> Typically, any time you think you've found a compiler error, you're using it wrong...
Yes, if you can't figure out if it's a bug, the bug tracker is the wrong place to get help. Ask in a community forum instead.
> Not okay: "I compile this input and the linker crashes. Also here's 10 paragraphs of slop about dwarf tables, which I can't even evaluate the accuracy of since I'm not an expert."
It's a 20-line file with a few commands to run it to reproduce it. Not 10 paragraphs of slop.
Presumably that is much more helpful than - here's my gigantic repo, good luck running my tests, also good luck finding the bug.
> here's my gigantic repo, good luck running my tests
What about my comment made you think I was suggesting not to give repro steps?
> also good luck finding the bug
But yes actually, this half is true. It's better to give them no extra information, than to give too much information that you have no idea if it's true or not.
Seems like "parallel construction" where you show you actually understand and can explain the issue independent of an LLM would be the way to go. No need to mention how you found the bug as long as you can explain what the bug is, why it matters, and how to repro.
Lying to get around a project policy you disagree with is immature. They have the right to run things the way they want, either accept their rules or leave the project be.
If they really insist on no LLM involvement at all then they're going to fall behind and lose out. Your LLM use case seems really very conservative - you didn't write any code with it, you just use it to confirm the bug. There are many OSS projects that are also taking a similar hardline against LLMs - people are going to fork them and then move on. I've had LLMs fix bugs/add features to a couple of projects like this and, well, they're missing out on the fixes/added features that I'm using locally.
Can you offer any specifics? The ToC looks quite extensive, which to me indicates a lot of design decisions needed to be made (probably involving many people) and unclear how that process might be accelerated by AI.
Unless you're suggesting the language design should also be vibed together?
> Unless you're suggesting the language design should also be vibed together?
Sounds like a fun little project, have a bunch of AI pushers fork Zig and see if they can do a better job. I want to see results, not snarky HN comments. After all this progress, ChatGPT should be able to one-shot a better language since AI is so good now... right?
One-shot, no, but there are a bunch of people pretty much solo-building their personal ideal language with AI and it's going quite well. You need to know just enough about language design to be dangerous, but you don't need to be a seasoned pro.
It seems like such a strange thing to do, building a language that you aren't going to write by hand. It's guaranteed to perform worse at higher cost, fill up a lot more of the context window, and burn a ton more reasoning tokens.
If you're using AI, a language with a large training set is going to win.
I guess it depends on whether you will write everything in an AI assisted fashion or not. There's benefits to languages that are quick and easy to read by the author even if the LLM is doing the writing, because most code still benefits from human review above and beyond the review that agents provide. Language popularity certainly helps but it seems like for moderately popular languages [1] the cost you pay for a lack of popularity is quite modest.
A language you just created isn't going to be moderately popular, so it's just going to put you at a disadvantage -- and you're not even going to be writing in it, so why the self-kneecapping?
I mean what does "put you at a disadvantage" even mean concretely? To use a less popular language, it means you need to load up context related to the semantics of your language, load context on how to invoke tools to make sure the syntax with your language is correct, load up context related to each tool call you make (which will be more numerous in a niche language), and load up context on architectural decisions that might be specific to your language. All of this is simply a token cost. By forcing a model to load an initial amount of context per harness turn you also effectively shorten the max context window beyond which the model becomes stupid (which itself is much shorter than the max context length.)
Obviously it's not like people are specifically trimming each and every prompt they give a model to tokenmax their models to get the best output / input prompt, we instead live in a spectrum of how many tokens of input and context we're willing to provide to a model to make progress. If the cost of those tokens is low enough for the problem domain you're working in, then it's fine. For some the readability of a personal language may outstrip any of the token costs that one needs to pay to use it. Alternatively maybe you want something like an array language (J, K, APL, etc) which allows array programming and optimizations that conventional PLs just can't do. Maybe you want your language to compile to a target that is highly portable. There's actually a lot of stuff out there that previously wasn't feasible but with LLMs-as-force-multiplier absolutely is.
I also suspect the space is a continuum. There may be pareto optimal points, such as DSLs built atop languages, that are both highly readable but also fairly token efficient.
Been thinking about doing this myself (did some PL in grad school but it's been a long time), but I find myself wanting to reach for a Scheme (using macros to grow the language I want) and customize it or build something atop Janet.
Curious why you wanted a more ML / Rust / Scala inspired syntax. (Personal preference here is totally valid btw, just curious.)
None of these systems can. They need enormous training. They need alignment and reinforcement. They need harnesses. And most importantly they need a human that knows how to write and develop a C compiler.
The ISO specifications are not sufficient. Neither are the System V guidelines. Not even spec tests and compcert.
You're not going to one-shot it, but over the course of a couple of months of evenings you could come up with something usable/interesting if you manage the LLM well.
Yep, it's true, but the reality of what it would take to actually fork and commence with healthy use of AI alongside with the elusive soft/hard stance needed to lead technically and socially is probably a bigger challenge than anyone is willing to take on! Soo.. slow and (rather?) well it goes for Zig.
The fix could not be backported to LLVM 22 because it changed the LLVM library ABI. We also could not skip straight to LLVM 23 because that would make life harder for distro package maintainers.
Was loop vectorization by chance enabled in 0.15.2 and disabled since 0.16.0? One of my projects got 14% slower after upgrading to 0.16. I didn't dig in yet but I assumed the Io vtable was the cause (it does a lot of file io). But if the versions line up maybe it's actually loop vectorization.
I just noticed it today upgrading to 0.17.0 from 0.15.2, while doing bitwise operations a lot of large integers. I haven't looked at a debugger yet, but my guess is that would benefit a lot from loop vectorization.
Nope, but why would it matter? Nobody is stopping you from writing Zig code with an LLM, the policy you are referring to is only relevant to the compiler source code.
Back in a more adult age, the way that this was handled was "We've benefited from our past collaboration, but we believe a new direction is needed going forward."
For some reason--COVID brain rot, poorly-socialized people coming online, general increase in viciousness in the population, who knows!--people have forgotten the utility and purpose of boring polite manners and communication.
https://youtu.be/zwi5b5xSsKA?is=PTjJJjSnVMdRuZag
It may surprise some people here to see that Andrew is warming up to using LLMs to discover bugs (inspired by results from SQLlite) and considers it a tool on the path to getting to bug free software.
Someone in the thread below says their bug was closed because of mentioning that they use AI to confirm the bug they had encountered. Is he going to go back and reopen all of those now that he learned what pretty much everyone else already knew?
I don't know about this particular case, but if I saw someone report a bug and as evidence claim they had X Y and Z LLMs verify it I would be pretty upset. If you're going to use an LLM to make a replication, just do that and give me the replication, don't point to your notoriously error-prone tools as though they lend your report credence.
It's in a similar vein to people who reply to questions with "well Claude says: <chat transcript dump>"
or substantially worse: "<chat transcript dump>"
If a report is improved and becomes actionable, that's great.
In the state of the tagged video he says still not accepting AI submissions until a certain set of preconditions is met. So... No?
https://youtu.be/zwi5b5xSsKA?si=w6zZN6AtIvJS9MxP&t=2084
Actually, on second thought, it really doesn't. Indoctrination and herd mentality are quite the things.
It’s not for everyone yet. It’s still unstable, and its ecosystem is small. However, both are improving.
Background is I created issue and one pull request to fix them in zig compiler version 0.16.1 issue numbers 36812, 36811 (you cannot access them as my account is banned can see my fork at [2]). Respecting the community’s no AI stand. For these specific issues I wrote the issue and code myself and not let AI write it. Spend a lot of time on it. Subsequently without any notice my account was banned because my projects on github using zig uses LLM. This was done without message or any information. My account was banned on zig repository.
I can now understand the other side of coin how bun team might have been treated with disdain when they used LLM.
I wrote an email and left the zig community, have many work in zig but slowly moving them to odin.
I feel personal disdain should not be spilled on to people who are pragmatic on using LLM. I was a very big evangelist of zig for their no LLM stand and promoted them among my community, but with poor treatment by community I just left. You can still see projects I wrote in zig [3].
I have worked with postgreql community since 1997 and python community since 1998. Never felt such hostile community. So all the best and I wish zig continue its progress
[1] https://github.com/insanai/sqlodin
[2] https://codeberg.org/vyomtech/zig
[3] https://github.com/insanai/zenfmt
I'm looking forward to see what the new build integration can unlock on the tooling side.
What I'm looking for the most for the next release(s):
- New stackless coroutine IO implementation
- First class fuzzer tooling
Outcompetes C even? I'm especially exited for SpirV. Would be great to use Zig for both CPU and GPU programming.
Especially in WebGPU, where WGSL tooling is very early.
>No LLMs for finding bugs.
>No talking about use of chatbot/LLM services.
I've said it before and I'll say it again- it's a cult that bans dissent
Asking non-rhetorically. It seems like one position "ai in any circumstance = bad" is being enforced. The commenter above didn't even understand why he was ignored.
To clarify, do you mean someone who isn't part of the core team?
We’re quite knowledgeable and thoughtful folks fwiw
I think I’ve seen you are a core team member or contributor? I remember your tag?
Yes, I'm a core team member.
So for instance because of the size of our codebase our project has pushed Zig to some of the edges, specifically we end up hitting a bug when using llvm and zig on arm64 (Mac and Linux) where it seems to be caused by some configuration Zig passes through to LLVM. I’ve used codex and claude to help me diagnose and find the bug (we use nix’s glibc zig to circumvent the problem now). I now understand the root cause but am not sure what the proper fix would be. But I’ve not known whether or not even raising the issue would break the terms of contributing? Would raising the issue break the implicit agreement?
I've written code a long time and that's probably the dumbest rule I've seen.
It's clearly a hobby project (constant breakages, the maintainer getting into politics, rejecting some safety mechanisms, the anti-LLM crusade, a strange focus on esoteric targets with little to no commercial significance), but the maintainer does not admit that it is a hobby project.
It makes me respect the Rust community even more.
Tbf I guess as a popular open source project not using AI to fix bugs, they probably already have more open bugs than they can ever fix so it doesn't really help them for people to find more.
I would imagine his bug was actually ignored just because Zig has 2700 open bugs, rather than some AI policy violation.
Show me a popular open source project that doesn't have a large number of open issues and I'll show you one that has a triage bot auto-close them.
https://youtu.be/zwi5b5xSsKA?is=PTjJJjSnVMdRuZag
They may just be taking the slow route of rejecting by default until they can be sure that the usage of LLMs provides long term value. I don’t see anything wrong with that. If you’re writing robust software, using LLMs at this stage is a bit of a gamble. We don’t fully know the long term effects on code quality yet.
Typically, any time you think you've found a compiler error, you're using it wrong...
Okay: "I compile this input and the linker crashes."
Not okay: "I compile this input and the linker crashes. Also here's 10 paragraphs of slop about dwarf tables, which I can't even evaluate the accuracy of since I'm not an expert."
> Typically, any time you think you've found a compiler error, you're using it wrong...
Yes, if you can't figure out if it's a bug, the bug tracker is the wrong place to get help. Ask in a community forum instead.
It's a 20-line file with a few commands to run it to reproduce it. Not 10 paragraphs of slop.
Presumably that is much more helpful than - here's my gigantic repo, good luck running my tests, also good luck finding the bug.
What about my comment made you think I was suggesting not to give repro steps?
> also good luck finding the bug
But yes actually, this half is true. It's better to give them no extra information, than to give too much information that you have no idea if it's true or not.
> gets ignored
Who could have forseen this.
Unless you're suggesting the language design should also be vibed together?
Sounds like a fun little project, have a bunch of AI pushers fork Zig and see if they can do a better job. I want to see results, not snarky HN comments. After all this progress, ChatGPT should be able to one-shot a better language since AI is so good now... right?
I'm doing it myself: https://zena-lang.dev/
If you're using AI, a language with a large training set is going to win.
[1]: https://danluu.com/pl-tokens/
Obviously it's not like people are specifically trimming each and every prompt they give a model to tokenmax their models to get the best output / input prompt, we instead live in a spectrum of how many tokens of input and context we're willing to provide to a model to make progress. If the cost of those tokens is low enough for the problem domain you're working in, then it's fine. For some the readability of a personal language may outstrip any of the token costs that one needs to pay to use it. Alternatively maybe you want something like an array language (J, K, APL, etc) which allows array programming and optimizations that conventional PLs just can't do. Maybe you want your language to compile to a target that is highly portable. There's actually a lot of stuff out there that previously wasn't feasible but with LLMs-as-force-multiplier absolutely is.
I also suspect the space is a continuum. There may be pareto optimal points, such as DSLs built atop languages, that are both highly readable but also fairly token efficient.
When I design my own languages (I have written several, all terrible!) it's typically to learn about language design.
Curious why you wanted a more ML / Rust / Scala inspired syntax. (Personal preference here is totally valid btw, just curious.)
None of these systems can. They need enormous training. They need alignment and reinforcement. They need harnesses. And most importantly they need a human that knows how to write and develop a C compiler.
The ISO specifications are not sufficient. Neither are the System V guidelines. Not even spec tests and compcert.
Thanks for your work btw.
Just don't be surprised if the project BFDL talks shit about you or your company later.
(Still a good language though, credit where credit is due.)
A simple, "this entity is a sponsor, and therefore there is a conflict of interest and we will not comment on recent controversy" is enough.
If it's big enough, refuse to take further contributions.
I know it's not entertaining, but that's why we have video games.
For some reason--COVID brain rot, poorly-socialized people coming online, general increase in viciousness in the population, who knows!--people have forgotten the utility and purpose of boring polite manners and communication.