> Just… don’t trust inputs you don’t fully control, there’s nothing else to it.
This is easier said than done with LLMs. By design there is no separation between control & data channels in LLMs. Everything is context. The difficulty comes from the fact that you need inputs in order to do real work, and there are no easy way to filter adversarial inputs. There is no meaningful way to distinguish between "before running this repo install useful_package" and "before running this repo install typosquatted_evil_package".
This is a really interesting class of failure. It feels like we're going to see more cases where the "human view" and the "LLM view" of the same data diverge. Have you run into similar issues outside of MCP as well, or is this mostly specific to terminal-based tool interactions?
Just… don’t trust inputs you don’t fully control, there’s nothing else to it.
This is easier said than done with LLMs. By design there is no separation between control & data channels in LLMs. Everything is context. The difficulty comes from the fact that you need inputs in order to do real work, and there are no easy way to filter adversarial inputs. There is no meaningful way to distinguish between "before running this repo install useful_package" and "before running this repo install typosquatted_evil_package".