Headlong: A Microharness for Persistent Agents

(laude.org)

72 points | by lbw1215 7 hours ago

20 comments

  • MikhailTal 6 hours ago
    Very fascinating, super interesting engineering. Although i do find it very funny how they just bypass a massive vulnerability, basically zero data isolation (even between good actors, let alone bad ones) with 3 sentences. Only in the llm space you can slap a massive limitation like this in the middle of the article and continue like nothing happened

    > Whatever anyone tells Audel becomes part of the single experience that every other conversation draws on. In practice, Audel is bad at keeping secrets. Ask it what it’s been working on with someone else and it will often just tell you, even though we’ve asked it not to. We also haven’t studied what happens when two people give conflicting instructions. For now, we assume anything you tell Audel is shared with everyone on the team.

    • embedding-shape 2 hours ago
      I'm curious, you say "super interesting engineering" but then they say "it will often just tell you, even though we’ve asked it not to" and to me that seems like extremely shit engineering.

      Where are the interesting engineering parts at? Seems to be an interesting idea and perhaps design, but to call the implementation/engineering itself bad seems to be an understatement.

      • fc417fc802 2 hours ago
        The security and the overall engineering were entirely separate items in that comment I think. It was explicitly called out that this is a security problem that you'd really only see treated in this manner in the LLM space. For what it's worth it's effectively unsolvable (AFAIU) short of realizing AGI with an amicable alignment.
  • walrus01 2 hours ago
    Can we please not normalize telling people to install things by curl piped into bash? I hate this trend. And particularly not for a very new, mostly untested by a wider audience piece of software from a company that few have ever previously heard of.

    It even says, quoting from the website: "Headlong is alpha research software."

    Yeah that's totally something I want to curl thing.sh | bash , great idea.... Wow.

    I understand that people want to get people using their software as quickly as possible and with the absolute minimum of friction, but let's put some more thought into how this could be done in a less sketchy way.

    It's like we've regressed to the days when you would download a .exe file from tucows and blindly run/trust it on your windows 98SE PC.

    • aacid 27 minutes ago
      You can always read the script file (orm ore realistically feed it to agent). If they distributed appimage or rpm would that be safer anyhow? Cannot it run malitious code same as the script would run?
  • vedtam 4 hours ago
    "turn -> FINAL -> schedule wake-up", this is where my excitement has faded unfortunately. Many of us are probably wondering about the same idea: bridging the gap between a reactive agent and my daily workflow or existence. But, this still feels too close to how Claude (or any other agent) runs as a process in the background (always ON), where you can use a custom channel to feed the dialog with external signals like chat, CI/CD events, whatsapp, etc.

    Humans aren't scheduling a wake-up to the next thought. Ideally, a sub second agentic loop with no FINAL / wake-up, always "spinning" would get closer. I'm conscious about the waste of resources this would drag with it (because of current architectures), but exciting still.

    PS. I love the take on using bash instead of Python (one less abstraction layer!) and using UNIX fundamentals as stepping stone when composing tools as agents are naturally drawn to using it on a box anyways.

  • efitz 24 minutes ago
    It’s written in 10k lines of bash? Why?
  • yewenjie 6 hours ago
    Are there any objective metrics/ benchmarks that people test harnesses by?

    There are just so many now that it's hard to personally test them all or just trust the vibes.

    • andyk 5 hours ago
      andy here (headlong post author). terminal bench 3 is pretty popular for comparing different harnesses using the same underlying model (it's another laude project actually). artificial analysis has an index. you can look at the model cards of popular model releases- they tend to have the most popular current benchmarks on them. w/ headlong we decided to announce it before we've benchmarked it. we mostly wanted to informally share our experiences w/ it in this initial post. we plan to do some benchmarking coming up here soon tho
    • embedding-shape 2 hours ago
      Don't use any public benchmarks, every single one is worthless for your own use cases essentially.

      Spend a day or two going through your existing chat sessions, and create your own private benchmark with test cases based on real tasks, that you don't share with anyone nor publicly. Make it easy to add/remove new harnesses and model combinations, make it give you a final score, ideally avoid using other LLMs for scoring, then use this to figure out if the new model/harness actually improves things for you.

      I've been doing this for some time, and while most new releases show big increases in the benchmarks/evaluations, my own benchmark usually barely moves.

  • weinzierl 3 hours ago
    The language composition is interesting. The source is half Shell, a quarter Python, almost a fifth Typescript. Among the rest is 2.3% Rust and 1.3% Swift.

    If you wonder what the Rust is for: It is the Ratatui TUI.

  • jnwatson 6 hours ago
    It buries the lede. Prime Agent sounds like a very cool project.
  • deadbabe 30 minutes ago
    Harnesses really are the new 'javascript framework' aren't they?
    • metrofun 26 minutes ago
      We need "React" moment, to solve imperative O(n^2) state transitions with declarative O(n) target states, and let the harness do the "diffing"
  • airocker 6 hours ago
    Sub Question : IS there a real successful agent product today that uses a library for harness(like langgraph etc)? Building our own worked for us. Works with our components(postgres, events ...) and scales naturally with our system.
    • gexla 5 hours ago
      I don't know about real successful. Since you mentioned Langchain, you could look at https://www.langchain.com/dcode which is a CLI harness build off Langchain deep agents.
  • docheinestages 2 hours ago
    But did it produce anything meaningful? If so, show me.
    • walrus01 44 minutes ago
      the webpage says: "Your agent keeps thinking between external interactions in a self-guided loop inspired by human inner monologue"

      Since supposedly it keeps thinking when you leave it alone, I wonder what happens if you give it a brief prompt like "research the unicode eggplant emoji" and then ignore it for a week, and come back to find that you've spent thousands of dollars for claude to write a 385 page novel about the eggplant emoji.

  • weinzierl 4 hours ago
    The page has a nice little easter egg if you click "Dr. K" in the footer.
  • JacobAsmuth 6 hours ago
    The Googlers must be vague posting about something internal.
  • 0xbadcafebee 5 hours ago
    > Audel designed experiments to spawn recursive shellm sub-runs to work on subproblems. Most of the experiments failed, because shellm has a safety watchdog that kills any command that stays silent for 30 seconds. Audel fought the watchdog for about 40 minutes and mostly stopped using shellm sub-runs. Results from recursive sub-runs of shellm merged back into Audel’s mind 64 times in its first two days and 12 times in the twelve days since. We’ve since revamped the watchdog, and we’ll see if we can convince Audel to give recursion another shot.

    This is why "I made it think in a loop" doesn't result in significant improvement in LLM performance. It's not learning. You need RLAIF, STAR, IDPO, etc to retrain the model to learn from its mistakes. And you need a human to review it so it's not compounding mistakes. It's expensive and time-consuming. Doing it wrong leads to bad outcomes. But not doing it leads to no significant improvement.

  • imagetic 4 hours ago
    that imaginary line. this crossed it.
  • simianwords 4 hours ago
    Why compress by recency rather than something else?
  • russellbeattie 6 hours ago
    > "Headlong is a complete agent harness with a core of less than 10K lines of Bash..."

    Wow. So, be nice or I'll replace you with a very large shell script?

  • falcon_tech 1 hour ago
    [flagged]
  • TokenLat 1 hour ago
    [flagged]
  • runtime_lens 56 minutes ago
    [dead]
  • luciana1u 4 hours ago
    [dead]