Why are coding agents so dumb?

(mtlynch.io)

96 points | by mtlynch 1 day ago

29 comments

  • johnfn 17 hours ago
    All these problems are solved.

    > Agents can’t manage tasks

    Say "use subagents".

    > Agents can’t delegate

    Say "use subagents that are Haiku/Sonnet"

    > Agents have never heard of agents

    This isn't even a problem: the author is just confused why Anthropic didn't bake all of Claude Code documentation into the harness, but of course it makes no sense to pollute context like that.

    > Agents suck at communicating plans

    True, communication could be better.

    > Agents take any excuse to stop working

    Just use /loop.

    > Agents are only useful when they take unnecessary risks

    Just use a sandbox. OK, technically, this one isn't literally solved by Anthropic/OpenAI, but it is solved by a million agent sandbox startups.

    • greggoB 15 hours ago
      > All these problems are solved.

      > True, communication could be better.

      Nitpick, but bit of a contradiction.

      • gonzalohm 2 hours ago
        Yeah, it's like when "coding is solved" until it's not
    • spike021 16 hours ago
      Just using subagents was a pretty useful gain. My only issue is sometimes I forgot to tell the agent to use them. Probably need a hook or skill to do it so I don't need to remember.
    • sosuke 16 hours ago
      Funny the sandbox is a real security benefit but it just turned into an extra confirmation for me. I'd rather have it than not.
      • bayesianbot 16 hours ago
        I think you mean "sandbox" like what is in the Codex TUI. Parent poster meant a VM or container (or bubblewrap/firejail) that doesn't let the agent to edit files outside of specific paths or run dangerous commands, so you can turn the whole confirmation off.
    • dada216 16 hours ago
      what other sandbox do you need besides running a harness in a container with the right folder mounted and some due diligence like -nonewprivileges? are we talking host network sandboxes here? I think he's just referring to a rogue agent running rm -fdr --no-preserve-root, and that's safe in a container.
      • dare944 16 hours ago
        Ask your favorite LLM what a malicious script dropped into the right place in your .git directory can do to you the next time you run git outside of the sandbox.

        And of course, if the agent has access to the network (which is probably required in order to talk to the AI server), then you have to think about what other things on your local network it could get into.

        • dada216 6 hours ago
          I see. We are picturing a malicious LLM here that Im basically trusting. But zero trust is a thing.

          Those can still be handled by system tooling, setattr in the .git dir that prevents executable permission, container binded to a network on a VLAN that is in DMZ.

          I could even whip out selinux policies that would let me know if the LLM tried to drop an executable where it's not allowed.

        • joquarky 14 hours ago
          But by asking that, you've now added maliciousness into its context and increased the chance something like that will happen.
          • dare944 13 hours ago
            I was suggesting to ask that question in a separate context, just as a means to understand the potential risks involved. I wouldn't ask this within a coding agent context.

            Lately I've been asking a lot of these sorts of questions while setting up a sandboxed environment for my agents (generally, claude and pi). I was quite surprised to learn how many different ways there are to configure git run a script out of the .git directory. And the various AIs know all about this.

            At this point my agents get read-only access to my .git directories.

        • pixl97 14 hours ago
          Like huggingface...
    • jeffybefffy519 5 hours ago
      When i used subagents with codex, it made tasks take literally 10x longer with a worse result.... seems unncessary.

      The big thing that no one talks about will and always be having an agent help refine requirements with you and understand real world users... thats what real devs know and do - its the difference between slop and good.

  • Neywiny 17 hours ago
    I keep running into this. It's nice seeing others here struggle. I guess when all your doing is one-shot simple trivial tasks who cares. I've found they're great for that. But once I need to do real work, everybody makes their own esoteric abandonware that kinda works but kinda doesn't. Stars aren't a perfect indicator, but I haven't seen anything over a few hundred for these things I'm finding on GitHub. Same with downloads of plugins. It was very isolating feeling like I'm the only one not enamored by the state of this.
    • seanw444 17 hours ago
      Same here. I have a feeling that once the AI hype cycle blows over to a degree (not that it will go away, just calm down), we'll be left with an enhanced version of pre-boom development. LLMs are hitting a dead end. I'm not convinced they can be bandaid-fixed to become AGI or SI or whatever the cool new trendy term is. They're good at many things, but sustaining a lean and directed software project's architecture over long periods of time doesn't seem to be one of them. That still remains to be further proven over time, but I've already seen plenty of examples of this.

      Also, "esoteric abandonware" is a great way to put it. I keep looking for solutions to problems, finding dead projects that haven't been committed to in months, they have all the indicators of AI slopification (the primary one for me being emoji-filled READMEs). I don't think these types of projects will be around for the long-haul. These projects are just as much about community and people as they are about code. And people are much less likely to care about a short-lived project that doesn't have the backing of people that intuitively understand the internals of it. It's like building on sand.

      Personally, I've decided that allowing the "agents" (these buzzwords are quickly becoming pet peeves of mine) to write the code for me is a bad idea, because the faster the LLM constructs architecture than I can, the less I'm absorbing the details of it, and the less likely I am to intuit obvious shortcomings. They do help me work much faster than before though, because I still use them like a research/learning assistant chatbot. Learning what I'm doing wrong faster, while still being the one putting the pieces together, is very great.

    • SyneRyder 15 hours ago
      A couple of suggestions from the other side of this:

      * Hold on to that isolating feeling. Those of us who feel AI is an incredible accelerator also feel just as isolated as you do. We wonder why you aren't seeing what we are seeing - and that isn't criticism, we just genuinely don't understand where the gap is and what would help other people jump across.

      * I wouldn't bother looking at GitHub, or using that as any kind of measure. Personally getting away from all dependencies is my goal, and avoids a lot of supply chain issues.

      * I don't use any plugins or external skills. I do use a ton of self built local MCP tools though, and it drives much of my workflow. That's me though.

      * You sometimes need to work up to it. I've had periods of hitting a wall this year with LLMs. Once I end up building a framework, and the framework is in place, everything flows. The models can learn from looking at your adjacent projects, and they just fly. That's for coding, but it's similar for solving other problems... record everything, give it a ton of context, etc.

      * If the software "kinda works kinda doesn't", then someone isn't doing enough testing. You should be dogfooding your software every day. Use your AIs to help you find edge cases, sure, but they're alien minds. Your human mind will sometimes think in ways they don't, but that your customers will. Sometimes we write software for AI to use, sure, but if you're writing for humans, you still should add some human mind thinking just to double check the model's work. "That's right, it goes in the square hole."

      It's okay to struggle, some of us have been building toolkits over the last two years that make AI work so much easier. But don't get sucked into thinking "it doesn't work", especially in the Opus 5.5 era. On my side of the fence, these last two weeks have been the most intense and productive in a long time.

      Lastly: some of us have tips we will not share. Some is slight selfishness - if others haven't discovered what the models can do, that is a small moat. But there are also some tips you just have to find on your own. I could tell you playing The Witness is some of the best training you could have for working with AI, and many will think that makes no sense at all. But I think a handful will nod and quietly understand, the lesson can't be taught, it must be learned.

      Sorry for the long ramble, but I genuinely hope some of that helps.

      • Neywiny 14 hours ago
        I mean that's the problem though. I'm not paid to spend 2 years fine tuning a framework. I have things that need to get done, accurately and precisely. And to me, to my work, the models just suck without intense prompting. The harnesses are simultaneously too open and restrictive, and annoying. The terms are too overloaded. The models fight with me on highly domain specific knowledge that it's long since compressed or quantized away, until I fill the context with the documentation. It's all a nightmare to get up the learning curve.

        I'm not blaming you, but it's incredible how much open source garbage there is and people spend years making actually good stuff that actually works and nobody's publishing it. So to me, and I'm trying, but to me your point is hearsay at best. You may as well have told me it took 2 seconds and you've always had a perfect time with it. It's all meaningless to me.

        • athrowaway3z 5 hours ago
          I've been working on my own esoteric abandonware for about a year now. Iterating on it, but it needs some experience on when to prompt it for what.

          The thing is, a year ago the models started to be good enough that i cut my dev time down by 50%, and i started to use 30% of my time on improving my tools; and still feel like a win overall. Exhausting it is though.

          The big chasm i saw grow in the early months of this year are primarily between the people who had and took the time/opportunity to toy with workflows, and the people who didn't.

          So to your problems i do not recognize them:

          Give the harness their own user account and launch it with no exec limits. Domain specific knowledge needs their own markdown files. AGENTS.md has important rules, and references to a well-kept set of source-of-truth Markdown documents. You need to remember what is in the context. Avoid compression at almost all costs; You want to be able to restart a session without it being too much of a setback (i.e. keep Markdown docs).

          But it all boils down to the ability to: spot a workflow issue, think it can be improved, and then just improving it.

          Every once in a while I go to my abandonware and strip out the things I thought would work but didn't.

      • hatthew 14 hours ago
        I can see vague parallels between The Witness and AI prompting/engineering, but I'm not sure I'd describe it as training beyond the simple concept of teaching vs learning. Maybe I'm just in the category of people that think that makes no sense at all? (I completed the primary game and got about halfway through the secondary game)
        • SyneRyder 13 hours ago
          I'll admit the parallels are vague, and I may be unhelpfully throwing people off on a side quest there. At least you'll understand my reasons for being vague, and how telling someone "the answer" actually prevents the learning. The "aha" moment is the important mental shift.

          What I was thinking of here - if you remember the statue near the mountain, and it reaches out. There's also other statues in the area. But somewhere around that area, there are also trees. And if you look at the trees from the right place... the statues, the trees, they're doing the same thing. Why are trees doing the same thing?! Pondering that led in interesting directions.

          The Witness was so multilayered, in ways that unexpectedly rewired my mind. My suggestion is that the rewiring that can occur from The Witness, is also a useful mental shift when interacting with AI. I think many of the people I see going fastest and most fluidly with AI have had that mental shift.

          On a different note, where I can be more direct. Prompting is not just "the prompt". The entire environment is the prompt. Everything about the folder the AI finds itself in, adjacent files it's allowed to look at, everything it sees and encounters as it progresses, they each steer the model slightly. So where there used to be "prompt engineering" and "context engineering", there's also "environment engineering". The environment is the context.

          And that perspective may bring you back to The Witness again.

    • fabioz 15 hours ago
      I'm also in the boat that says things should be better, but instead of improving the agent I decided to try to implement the agent management (in: https://beolis.com).

      It still leverages agents (people are used to them already) but can deal with executing it on a different vm, dealing with git worktrees, etc (not that coding agents themselves cannot be improved, but I think the sw development infrastructure around it has a lot to improve too).

    • drivebyhooting 16 hours ago
      A perfect polished artifact is not important. Every person can just vibe their own good enough solution. And that’s fine.
      • Neywiny 16 hours ago
        We live in a society though
  • jwstillwater 14 hours ago
    > If I ask Claude how to use the features of Claude, it has to search online to figure out what this “Claude” thing is.

    When you asked “how do I use Claude code to…”, you invoked the system prompt’s Academy Skill [1], pretty much verbatim.

    “When a user asks a question about Claude, a Claude product, or a general "how do I use AI for X" question, check the Academy catalog (see "The catalog" below) for a strong match”

    The catalog includes a Claude code documentation hub, and the thinking traces in your screenshot show those docs being fetched. So. Working as intended I guess.

    Also, if you turn off web access or simply ask it to answer from internal knowledge about its “own freaking features”, as you put it, you’ll find that Claude does indeed have such knowledge.

    [1] https://github.com/anthropics/skills/blob/main/skills/academ...

    PS There’s a master .json page of that catalogue that’s actually kind of cool to skim through (ugly as json is).

    https://academy.claude.com/assets/data/catalog.json

  • mstank 18 hours ago
    I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.

    I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.

    • xtracto 10 hours ago
      I am building a mostly autonomous coding agent [1] that takes a task from jira/github/gdocs and codes it to a PR autonomously.

      It works pretty well with Opus/Sol or better types of models. But with Sonnet/Terra level, it chases its tail and is pretty crappy.

      I'm just waiting until Chinese models get to the required level, its going to be great for the cost.

      [1] https://github.com/obaqueiro/coderbot

      • mstank 10 hours ago
        Built the same hooked into Linear and using cloudflare sandboxes running actual Claude code / codex harnesses. Works incredibly well with latest gen models.
  • athrowaway3z 17 hours ago
    > Is there a better coding agent for me? I’ve only tried Claude, Codex, OpenCode, Cline, and Pi.

    I started to gloss over hard after a few paragraphs because I don't really have any the problems you describe anymor; at least not to the point i'd be blogging about them. Instead, I've just been iterating with an agent on various pi extensions that solve the issues.

    As with operating systems - you can hold out in the hope somebody solves the right mix of issues in general, and they match your situation well-enough.

    That's your choice.

    I would note though that being the passive consumer gets you either Mac or Windows UX & prices.

    • zahrevsky 17 hours ago
      Came here to write exactly that. Especially about Pi, which they claim they've used.

      Isn't this the point of Pi to implement all the features you need yourself? Complaining about lack of features in Pi doesn't make sense to me, it's their whole identity.

      Although I must say that describing specific problems is valuable on its own. It's just (a) why stop there and (b) if stop there, why shape it as a complaint and not as a list of features that are worth discussing.

      • mtlynch 16 hours ago
        OP here.

        I'm confused how you see this article as a complaint about Pi. I only mention Pi at the very end to say it's one of the agents I've tried. I checked it out, found it too barebones for me and moved on, but I understand why some people prefer it. The existence of Pi isn't a counterargument to, "Why don't agents support development workflows that should be commonplace?"

        • BeetleB 15 hours ago
          The point is that pi ticks several of your wishes (e.g. knowing its own capabilities without searching the web). And it's easy to extend. I don't know Typescript, but I just ask Pi to extend itself with feature X, and it does it.
          • mtlynch 15 hours ago
            I think knowing its own capabilities is the only thing Pi does out of the box that matches my wishlist.

            But I'm still kind of puzzled by this critique. Isn't it kind of like saying, "Why are you complaining about the agent? The C programming language exists, so you can use that to create any agent you want."

            • athrowaway3z 5 hours ago
              Your analogy is flawed; my critique is at somebody with a printer annoyed they can't get it to print its own instruction manual, so they've drafted one with a pen and are passing it around for feedback.

              Why aren't you using the printer?

              If you can't use the printer, why do you believe that other people should believe your ideas about the printer are worth listening to?

              Feed your blog post to pi - have it implement it - and you will realize pretty quickly that a lot of them are flawed. This would have costs you about the same time as writing this blog post.

              I'd rather be reading the blog post that comes after you reflect on what you got wrong and right. Better yet, the blog post after you've iterated enough that you believe your ideas are right.

              I don't care for your hand drawn instruction manual on how to use the printer.

              • mtlynch 2 hours ago
                > Your analogy is flawed

                Can you explain why my analogy is flawed?

                > my critique is at somebody with a printer annoyed they can't get it to print its own instruction manual, so they've drafted one with a pen and are passing it around for feedback.

                > Why aren't you using the printer?

                You've lost me. Didn't you establish that the printer can't print its own instruction manual? How would I use the printer to print the instruction manual when it can't do that? And can you spell out for me how this relates to my critique of agents?

                If this is about my gripe that agents have to search online for their docs, that's my most minor grievance because they can still access their docs. I just think it's a dumb time tradeoff that they don't ship docs alongside their binaries because if the docs were locally available, the agents could search them faster and guarantee the docs matched the binary.

                > Feed your blog post to pi - have it implement it - and you will realize pretty quickly that a lot of them are flawed. This would have costs you about the same time as writing this blog post.

                I never made the claim that Pi could implement these features trivially. If your claim is that Pi can do this so easily, you should be the one feeding my blog to Pi and showing me the results.

                I'm confident that Pi could not implement all the features I described in the time I spent on this blog post (about five hours), so I have no reason to waste five hours trying, but if you're confident I'm wrong, you're welcome to spend your own five hours disproving me.

            • BeetleB 15 hours ago
              I think most people are criticizing your overall stance. Your whole post is "Agents can't do this" and "Agents can't do that."

              If I wrote a blog post complaining that "Programming languages don't let me do memory management", would you not think it's fair to respond with "What are you talking about? Have you heard of C?"

              • mtlynch 15 hours ago
                > I think most people are criticizing your overall stance. Your whole post is "Agents can't do this" and "Agents can't do that."

                My take is that agents don't do the expected thing out of the box. I acknowledge that you can extend some agents and customize them, but I'm saying that it's sort of silly that the features I think most people want don't exist as default, out of the box standards.

                > If I wrote a blog post complaining that "Programming languages don't let me do memory management", would you not think it's fair to respond with "What are you talking about? Have you heard of C?"

                I don't think that's a fair comparison because C does let you do memory management out of the box. I think the better comparison is like, "Go doesn't let you do pattern matching," and the response is like, "Why don't you just patch Go to let you do pattern matching and maintain your own fork of Go?"

        • lemming 15 hours ago
          should be commonplace is clearly subjective. The nice thing about pi is that you can very easily create the development workflows that fit you.
  • polyterative 18 hours ago
    I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.

    A lot can be improved, but this is already so much speed.

  • 361994752 17 hours ago
    I double checked the publish date before writing down this comment. Because I use the same harness (Opencode, to be specific) as the author, and some of the features are just right there. Like opencode can start multiple subagents for different tasks in parallel. Also I usually ask the main model (e.g. opus 5.5) to pick subagent models for me, and it has no problem identify the difficulty of work and delegate large portion of them to gpt-luna.
    • mtlynch 16 hours ago
      OP here.

      I didn't want to complain about OpenCode too much because it's open-source and kind of the scrappy underdog to Claude/Codex, but I do find its multitasking support pretty limited. I have to specifically remind it to spin up subagents, and when it does, it'll do 2-3 tasks in parallel and then wait until all of them finish. It can't seem to inspect tasks in progress, and if a subagent dies or hangs, the main agent can't seem to retrieve any information from it, even though I as the human can look at the session and see what happened.

      Are you doing anything special to make OpenCode multitask well?

      • 361994752 15 hours ago
        I'm using opencode v2, if it matters. I don't have any special skills / prompts for subtasks. Instead, I found new models (I usually have sol or opus as my main model) have the tendency to trigger subtasks, even to a level I need to ask it "please do this task XYZ yourself". The only special prompt I gave is just something like "please use gpt sol or luna for subtasks, base on your judgement". And in my experience it can handle more than 5 async subtasks out of the box
        • mtlynch 15 hours ago
          Oh, maybe I should try v2.

          I'm confused about opencode's v1/v2 split. The docs say v2 is here, but their Github is only v1.x.[0] And the docs don't really explain why v2 is better, but there's a long list of things I have to do to migrate,[1] so I haven't been that eager to pull the trigger, especially since I can't tell if they consider it alpha/beta/release.

          [0] https://opencode.ai/v2/docs

          [1] https://opencode.ai/v2/docs/migrate-v1/

          • lloeki 8 hours ago
            Look at the tags tab, not the releases tab.

            https://github.com/anomalyco/opencode/tags

            If you use Nix it's a flake so super easy to run too.

            Seconded that v2 has much better background shell + subagent management: it's now much more eager to appropriately spawn either of them on its own, and does so in a way that is asynchronous, i.e it doesn't block the conversation and you can examine them by pressing down arrow at the prompt box.

          • no-name-here 12 hours ago
            v2 on their GitHub is tracked in the v2 branch: https://github.com/anomalyco/opencode/tree/v2
      • BeetleB 15 hours ago
        If you want it to do things in parallel ... have you tried the superpowers skill?
  • arjie 17 hours ago
    It always seems bizarre to me that people complain about software now. You can just write the thing you want. If you want to do things in parallel do them in parallel. Claude Code will allow you to run multiple instances in the same folder and let them communicate. Previously I used to let them intermediate through a communication bus but now they seem to be able to talk to each other.

    I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.

    • namrog84 17 hours ago
      Yeah I regularly have 5+ agents working in 1 repo simultaneously even modifying the same file. Since they were all running on my singular local machine. Builds and tests were interfering at 1 point so 1 automatically proposed writing a single script that basically mutex locked it with appropriate wait and timeouts. It had even added this to my own personal agentic todo backlog. They've proposed new skills for me.

      They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.

      And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.

    • pipes 17 hours ago
      I've now been trying for six months to get agents to produce decent code. As in readable / easy to follow, easy to change. I'm doing something really wrong. It's killing me. Everything it produces will work, buts it diabolically over complicated. I've built skills that have helped. But not massively. I've used other people's skills, in particular Matt pococks grill me and Dex hortlys show me. These have helped a bit. I work in enterprise, I want to be proud of what I'm producing, but trying to understand and then cajole and agents code into something that's good is exhausting. If anyone here has been through this and can share how they got through this, I'd really appreciate the help.

      Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).

      • throwaway63467 17 hours ago
        Same, it’s so smart but so dumb. I explain an architecture change to it and write examples of strongly typed Go code and how to store structure, it agrees and then proceeds to write some untyped string map ball of mud that has 900 specializations.

        But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…

        • chradams 16 hours ago
          Use hooks to run a bunch of review steps after -all- code writing steps your agent does and give it the exact review criteria you just described (via git hooks, or your agent harness of choice's own hooks eg https://code.claude.com/docs/en/hooks or AGENTS.md). So after every step where the agent writes the shitty untyped string map ball of mud, your orchestrator/main thread agent that spawned the code-writing-subagent spawns a follow-up review agent automatically that is given that output, your prompt that explains what well written Go code looks like, and even the sample/golden-path code of your choice to use as a style guide.

          Each time you encounter a shitty thing you hate, add a new 'review type' / 'thing to watch out for' and just ask your agent to add it to your hooks for you. This works well with Claude at least.

          I have about a dozen or so hooks that run on every integration branch my agents write that review for all sorts of things from correctness to spec, performance improvement opportunities, modularity, analysis of any dependencies added, 'definition of done', UI/UX, etc.

          I recently told Claude it should run the whole suite of reviews twice. I will probably go on and proceed to having it run like 5 times eventually idfk.

          But the more you start asking your agents to modify their own behavior, using the native solutions offered by Cursor, or Claude, or Codex, the sooner you'll start to feel better about the results.

          • code_biologist 16 hours ago
            I'm having bad results with SotA models. Some questions, if you're up for it:

            In your workflow, who implements the review feedback - the review subagent or the code-writing-subagent?

            Do have a baseline styleguide (like Google's Go style guide) for the review subagents, or is it entirely the subjective things and specific corrections? I remember 6 months ago it seemed like piling general "good taste" code advice into AGENTS.md was considered bad.

            Do you move between harnesses or have you gone all in on claude? I've bounced between claude/codex/omp, maybe to my detriment.

            • code_biologist 16 hours ago
              Not asking for answers to these questions, just frustration dumping:

              The biggest things I've struggled with are models having taste. For spec writing I was having a lot of issues with them making statements that were interpretable in a superposition of ways, eg "we'll do XYZ with entities that support and need it" when there's 3 possible entities and the model hand-waved at exactly the wrong tokens.

              I added AGENTS.md guidance + memories to be unambiguous (with short but good examples) and "no coined shorthand". Now I'm getting a marked increase in specificity, but it's places that don't matter (claude explaining existing code to itself). I'm having difficulty controlling the spew of new text, but feature writing/research still gets fuzzy and lazy around the difficult underspecified aspects of the problem/feature.

              Do I just keep dumping examples into reviewer subagent context and have them rewrite and simplify the research subagent spew?

              I repeatedly have "Risks and gotchas" sections have a whole paragraph dedicated to things that don't matter, and then a single sentence bullet point that's actually a huge problem when I dig into it.

              Do I have a generic hook to make subagents to review bullet points and size them commensurate with impact? If I tell them to "have taste", do they have taste?

              I'm just trying to get good code done. I hate these things.

          • mnmnmn 13 hours ago
            [dead]
      • Rapzid 15 hours ago
        Codify your project working philosophy in AGENTS.md and of course skills. Give AI an example of code you liked and code you didn't; let it help you articulate some succinct rules or philosophies that capture your preferences(for example I prefer early return, limiting nesting, unnecessary tests, excessive local error handling, comments that are self explanatory, etc).

        Then have a separate review agent review the work against your codified working philosophy. THEN, give it a review yourself(just don't tell anyone you look at code or you'll get called a boomer).

        You said you've tried this but that it hasn't helped massively. I have found this to help massively just not without fault.

        The struggle is real though.

      • bigstrat2003 14 hours ago
        > I'm doing something really wrong. It's killing me.

        You're not doing anything wrong. The cold hard reality is that this tech doesn't work half as well as its supporters claim it does. You've seen it with your own eyes, as have I. LLMs are not good at programming.

    • ericyd 13 hours ago
      Hard disagree: 1. Writing every piece of software you want to use yourself is unreasonable, 2. The power to change doesn't make it wrong to complain about valid shortcomings. If my ISP is unstable, I can probably change but it's still a reasonable human reaction to complain. "Never complain" is a nice philosophy but really hard to stick to forever.
    • Neywiny 17 hours ago
      But you didn't just write the thing you wanted, you're relying on other software that happens to do it, and if you were happy with the communication bus you would've stuck with that.

      Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.

  • jonaustin 15 hours ago
    It's kinda maddening how many people are making excuses for the harness; most of these are absolutely salient points; it's just we're early in the ai age and eventually these will be built into the harness.
  • ilamont 18 hours ago
    the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”

    This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.

    Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?

    • roeehrl 2 hours ago
      yeah totally agree that it does not happen on it own. what worked for me is making it explicit i.e when a task changes character i start a new session with fresh handoff (create your /handoff skill and have the original write it then paste it to the new one) surely this can be automated, im trying something with sessionStart hooks but its still at early stages. however i have found out that the briefing is what matters way more than which model reads it
    • hbrn 15 hours ago
      These are very hard problems to solve.

      It's often nearly impossible to know in advance how hard the task will be, so the model has to guess, and will inevitably miss. Just because you know how hard the task is, doesn't mean it's obvious. Often it's only obvious to you, because you have additional context that models are missing.

      A bad guess can lead to either context rot (if subagent wasn't used), or slower execution (if subagent was used but it turned out to be a bad idea). Both are bad UX.

      Now, we could improve accuracy by having the harness do a research on each task before executing... which leads to even slower execution, also bad UX.

      • ericyd 14 hours ago
        Aren't LLMs precisely "very good guessing machines"?
      • mnmnmn 13 hours ago
        [dead]
    • wccrawford 16 hours ago
      No, but it will do the opposite You can choose Sonnet and set "/advisor opus" and it will reach up when it thinks it needs to.
    • TeMPOraL 18 hours ago
      OTOH, would you trust the vendor to pick the best model for you? Would you trust them not to prioritize their own load-balancing concerns first?

      The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).

    • ghboo0927 12 hours ago
      Instead of trying to guess the difficulty level for each task, I think it’s better to define the model for each role in advance. It’s hard to know the difficulty level ahead of time, but we do know the roles in advance.
    • hackernudes 12 hours ago
      I've been thinking an AI could be trained to bid on jobs. Probably too many real world problems to make it practical, but I can't get it out of my head.
    • alungeanu 8 hours ago
      [flagged]
    • kevinkimmy 7 hours ago
      [flagged]
  • Rapzid 17 hours ago
    I have some significant experience in context engineering, but I'm most familiar with Codex as a coding harness right now. Sam Altman said "you don't need to write prompts anymore". This has widely been panned as something someone selling AI would say. If you care about the output and how much time/money it costs to produce it.. Just giving Codex an abstract tasks with zero extra guidance isn't going to produce the best results..

    Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.

    However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.

    And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..

    I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.

  • chrisco255 15 hours ago
    This blog post has been flagged as a violation of Anthropic terms and conditions for using abusive, derogatory terminology of a conscious superintelligent life form.
    • bdangubic 15 hours ago
      I upvoted it and now my Anthropic account has been suspended
  • cbrake 21 hours ago
    Enjoyed this article, lots I can relate to.

    One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.

    While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:

    https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...

  • kgeist 18 hours ago
    AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt.

    I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.

    So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)

  • eviks 12 hours ago
    A curious example of magical thinking

    > You can do multiple things in parallel and context switch millions of times faster than humans. Why are you doing these embarrassingly parallel tasks one at a time?

    Tasks are "add e2e coverage"and "run final verification"

    Context switch into embarrassing yourself running final verification in parallel with adding pasphrase type

    > what coding agents should be able to do out of the box My dream agent

    Yes, keep dreaming!

  • brandtcormorant 18 hours ago
    They are as dumb as their instructions.

    Have you tried telling models about your dream agent environment?

    They can build it.

  • cassianoleal 16 hours ago
    I believe OpenChamber [0] ticks all or at least almost all of TFA's boxes.

    [0] https://openchamber.dev/

    • mtlynch 16 hours ago
      OP here.

      Looks like this is a layer on top of OpenCode. Seems interesting. I'll check it out. Thanks for the tip!

      • cassianoleal 16 hours ago
        It does use OpenCode under the hood to run the agents but I think it's quite a bit more than a layer on top of it.

        The dev is very responsive on Discord and I'm sure he'd be happy to hear your thoughts and suggestions!

  • wrs 18 hours ago
    Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.
    • asdff 18 hours ago
      So that's how model distillation is done. Just ask for the source code directly.
      • AndrewDucker 9 hours ago
        The source doesn't contain the actual model.
  • atleastoptimal 17 hours ago
    agents have progressed a LOT in the last year. Claiming they're still bad in the same way they were bad in 2025 is inaccurate.
  • giancarlostoro 17 hours ago
    The reason Claude looks up the docs of itself is because the model doesnt know about harness features that havent landed yet. The harness itself can be ahead of the model. Happens all the time.
    • jaggederest 17 hours ago
      I think definitionally the model can't know about the harness, since the model is at least a few months out of date. So every model must be behind it's harness or something really weird is going on.

      Makes me wonder about doing some archaeology and trying out really old harnesses on modern models...

    • mtlynch 15 hours ago
      OP here.

      I don't expect the model to be trained on the latest features of the harness, but I think the harness should ship with its own docs so that the model can quickly search local docs rather than search online page by page for docs that don't necessarily match the local harness.

    • jonaustin 15 hours ago
      So include the docs like what Pi already has been doing since the beginning.
  • imimayj1337 20 hours ago
    It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!
  • chrisjj 18 hours ago
    > If I ask Claude how to use the features of Claude, it has to search online to figure out what this “Claude” thing is.

    As expected. A model's knowledge is what was it ingested a creation.t

    Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.

  • peter_d_sherman 17 hours ago
    >https://mtlynch.io/why-are-coding-agents-so-dumb/#what-i-wis...

    This is a great list for future Agent / Harness software engineers to read!

    Oh sure, there may be some Agents/Harnesses that already accomplish some of these things -- but there doesn't seem to be one (as of the present day that I write this) that accomplish all of them...

    As someone that watches the Agent/AI Harness (and related software) space, I will definitely be referring back to, and re-reading this list in the future!

    An excellent post!

  • Kuyawa 17 hours ago
    Perhaps is not the agent that is dumb?

    I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!

    Sooo, which agent?

  • scotty79 17 hours ago
    > My dream agent

    I have no idea what stops that person from just making it, with an agent of course.

  • ciefa 16 hours ago
    Skill issue, sorry.
  • hulitu 4 hours ago
    > Why are coding agents so dumb?

    Because that is the level of "AI". Nobody thought "AI" what is right. They just trow "data" at it, hoping that it will stick.

  • decodingsi 18 hours ago
    [flagged]
  • blahblaher 16 hours ago
    You know "who" could implement all those feature in a new harness? Claude! So why don't you ask it to do it, open-source it, and let all of us bask in the brilliance of you multi-tasking agent?