Between this and the Cloudflare post, this is a lot of words and no simple system level picture. Here's what I think is going on:
1. Homoiconicity: Harness mechanics are kludgy and we need proper homoiconicity to uniformly handle code (tool calls) and data ("natural language") as token streams.
2. Actor semantics: for isolation, encapsulation and concurrency.
3. Object capabilities: injected references and no ambient authority. Capabality-based reflection/introspection is a clean way to discover interfaces and affordances.
"Code mode" or whatever is basically rediscovering this by hacking outward from LLM token streams, instead of from system design principles based on decades of computer science. It is the beginning of treating an LLM as a programming-language runtime participant (any takers for eval/apply?) rather than as a text/token generator with the harness as an ad-hoc interpreter.
If existing implementations of code mode don't already support all this, I anticipate they will keep piling on hacks till they get to this point.
----
I think it was Dan Ingalls who said "An operating system is a collection of things that don't fit into a language. There shouldn't be one.". I see the same for harnesses -- they're awkward middle children which fit neither in an LLM nor in the programming environment.
Maybe the answer is to partner LLMs with Common Lisp or Scheme fibers / Spritely Goblins or Erlang BEAM and be done!
Right now when a model wants to call 1 or more tools, there is a fixed json schema to describe the tool calls.
Codemode is like: Why don't we just let the model write a program to call the tools and compose them however it wants? The key is that the programming language exposed to do this (usually javascript) will have APIs available to do some things internal to the harness (like call tools/mcp).
Sadly this seems like introducing a very complex apparatus for little gain. I don't need MCP or many tool calls that the agent can program around - in fact the promise of pi was that you basically just need bash and no other tools. The minimally-invasive approach would have been to just "inject" tool calls as virtual bash commands. No code mode required, no hands vs. brains dichotomy. LLM can use language of its own choosing to interact with tools.
https://earendil.com/posts/you-said-no-mcp/ remains difficult to read without understanding what codemode is, with some quote like "Now we talked so much about Codemode, it might be worth explaining what that even is."
I'm not sure why almost all codemode implementations choose Javascript. I prototyped an agent[1] to use bash as the language for codemode, which in my opinion worked equally well and requires no teaching (there is literally 0 prompt to teach the LLM about codemode. A tool named "bash" is enough to have them know the usage).
> However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.
The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?
> If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?
(This is Bonteq, I was just logged into the wrong account.)
it cannot use cat but you can make a command which communicates back to agent. Armin mentions that:
> While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process
But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.
> But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
It becomes much crummier when hands and brain are on different machines.
> I'm not sure why almost all codemode implementations choose Javascript
Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.
But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.
sandboxing is indeed an advantage (can be an important one!), but
1. bash can also express concurrency easily, with sync and async (using standard & syntax) support for each command, and standard cancellation (kill, though crude).
2. "way fewer instructions": It requires no instruction for agents to use bash either, except for merely listing the custom commands (view_image, apply_patch, etc.). Also, bash has standard progressive disclosure mechanism (--help) that models will automatically use with no instruction.
I’m still figuring out codemode. I was using it in a prototype, but ended up stripping it out again, after realizing my small local LLM was using more tokens than usual. It was combining tool calls elegantly in code exactly how I hoped it would. The problem was that when any of those embedded tool calls failed (e.g. on parameter validation) the parent code execution tool call also failed. In response, the LLM kept rewriting large parts of the original code block.
Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.
This somewhat validates a fear I've had (but hadn't tested) of codemode with smaller models. It works fantastically with Opus and Sol but I've always wondered how well it scales down.
Claude’s scripts can’t call MCP tools, meaning everything has to be CLIs or libraries. At which point you lose the “everything has self-documenting input-output schema” that Codemode is going for.
I've noticed that code mode seems to cause progressive lobotomisation in Sol 6.1 around subagents. The more it uses it, the worse it gets at giving prompts to subagents:
URLs, datasourceproxy accesspathonlyallowed anddefault404. Do not set authbasic onprivate ports; perplanPodmannetwork trustedinfrastructure nottenant boundary. Publicroutes internal /tinyauth protected internal, proxyGETemptybody /api/auth/nginx toprivate tinyauth3000; proxy_pass_request_headersoff, CookieonlyTinyauthSession header extracted map name/value actual runtime pattern tinyauth-session-[0-9a-f]{8}, X-Original-URL constructed
This is unedited; it's merging words together and spamming keywords.
So... What exactly is it in the end? I try to parse the long prose, and couldn't.
It doesn't help that the article is titled "What is Codemode" and then goes to say "If you are not familiar with Codemode, it’s basically just...". You article is supposed to say what it is without anyone being familiar with it.
Like what does this passage even mean:
--- start quote ---
For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.
1. Homoiconicity: Harness mechanics are kludgy and we need proper homoiconicity to uniformly handle code (tool calls) and data ("natural language") as token streams.
2. Actor semantics: for isolation, encapsulation and concurrency.
3. Object capabilities: injected references and no ambient authority. Capabality-based reflection/introspection is a clean way to discover interfaces and affordances.
"Code mode" or whatever is basically rediscovering this by hacking outward from LLM token streams, instead of from system design principles based on decades of computer science. It is the beginning of treating an LLM as a programming-language runtime participant (any takers for eval/apply?) rather than as a text/token generator with the harness as an ad-hoc interpreter.
If existing implementations of code mode don't already support all this, I anticipate they will keep piling on hacks till they get to this point.
----
I think it was Dan Ingalls who said "An operating system is a collection of things that don't fit into a language. There shouldn't be one.". I see the same for harnesses -- they're awkward middle children which fit neither in an LLM nor in the programming environment.
Maybe the answer is to partner LLMs with Common Lisp or Scheme fibers / Spritely Goblins or Erlang BEAM and be done!
Autolith (https://autolith.rocks/) and some other CL-based harnesses allow the LLM to modify their own harness within the session.
Right now when a model wants to call 1 or more tools, there is a fixed json schema to describe the tool calls.
Codemode is like: Why don't we just let the model write a program to call the tools and compose them however it wants? The key is that the programming language exposed to do this (usually javascript) will have APIs available to do some things internal to the harness (like call tools/mcp).
https://earendil.com/posts/you-said-no-mcp/ remains difficult to read without understanding what codemode is, with some quote like "Now we talked so much about Codemode, it might be worth explaining what that even is."
[1] https://github.com/ylxdzsw/mu
The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?
> If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?
(This is Bonteq, I was just logged into the wrong account.)
> While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process
But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.
It becomes much crummier when hands and brain are on different machines.
Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.
But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.
1. bash can also express concurrency easily, with sync and async (using standard & syntax) support for each command, and standard cancellation (kill, though crude). 2. "way fewer instructions": It requires no instruction for agents to use bash either, except for merely listing the custom commands (view_image, apply_patch, etc.). Also, bash has standard progressive disclosure mechanism (--help) that models will automatically use with no instruction.
Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling.
Also there are things like subagents, etc (which may be considered tools).
Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.
It doesn't help that the article is titled "What is Codemode" and then goes to say "If you are not familiar with Codemode, it’s basically just...". You article is supposed to say what it is without anyone being familiar with it.
Like what does this passage even mean:
--- start quote ---
For instance if you issue a bash call as a regular tool call in the LLM, then we only throw the trailing 2000 lines into the context and if the agent wants more, it needs to look at the overflow file itself. If however the agent issues that invocation via Codemode, then the Codemode side gets larger outputs sent structurally.
--- end quote ---
I was expecting a bit more care from Armin in explaining what does codemode actually do, from first principles.