> Overall, our Rust rewrite and following performance hillclimbing has made Prime Agent significantly faster and more resource efficient. With time to input roughly 14x faster than TypeScript and using over 80% less memory after startup
So they used Prime Agent and GLM 5.3 to swarm and rewrite the code in Rust. This also shows Prime Agent doing what it preaches by rebuilding itself. Since Prime Agent is Pi under the hood, will they push a Rust rewrite to Pi? Pi extensions use Typescript so I wonder if they will work.
I don’t see many people talk about Prime Agent, I always wondered if it could just be a Pi extension cause it seems to be a subagent orchestration agent.
I feel like I talk about it often enough people might think I have an agenda. For a while it felt like a superpower compared to other harnesses, although they've caught up with persistent/backgrounded agents.
Very happy user. Loved it with deepseek-v4-flash, although I've moved on because newer models are so compelling.
I just get good results from it; I think it's the python requirement.
It has been very good with sub-agents, and also finding old context, for a long time. And had backgrounded agents that you can run from one instance without herdr. Herdr is kinda redundant (better UI than PA though).
I see nobody else talk about it. It's so little talk that when I mention it on X, the devs comment on my posts sometimes.
I'm interested in their Planner -> Implementer -> Reviewer -> Verifier process they used for this transition to Rust. I've see similar but curious how they actually implemented this.
Curious how this could be applied to greenfield coding rather than just making a copy in a new language or performance optimizing.
Kinda wish people break down token usage into input (cache hit), input (cache miss) and output when talking about it. Giving a total 200B tokens number doesn’t help gauge costs.
I was checking openrouter a few days ago to see if they provided that information yet. Would be a really nice breakdown to have under total task price. Doesn't map to every task, but averages would still be nice.
For Anthropic's reported benchmark numbers for Sonnet 5.5, when running terminal-bench 4.0, they had a comparison between $0.10 and $0.20 cache read for total task cost. That 50% cost reduction resulted in a 20.9-26.2% total cost reduction. Pretty specific though... single benchmark, model, harness, and provider. Cache reads make up ~42-52% of the total cost at the standard $0.20 price.
I know people tend to hate on rust rewrites. But having something that compiles to a single binary that you can just copy over and get started has its advantages especially when working with sandboxes etc.
That sounds like a nightmare for comprehensibility, incremental development, and maintenance.
Combine that with the fact that, unlike other code generators, LLM output can't as a rule be reliably reproduced from the original input in the future, sounds like a recipe for unpredictable long-term costs and regular regressions.
very nice outcome of this port
Naturally then there isn't source material for "we rewrote yet another slow scripting project into Go/Rust/Zig/C/C++/..." blog posts.
I don’t see many people talk about Prime Agent, I always wondered if it could just be a Pi extension cause it seems to be a subagent orchestration agent.
Very happy user. Loved it with deepseek-v4-flash, although I've moved on because newer models are so compelling.
I just get good results from it; I think it's the python requirement.
It has been very good with sub-agents, and also finding old context, for a long time. And had backgrounded agents that you can run from one instance without herdr. Herdr is kinda redundant (better UI than PA though).
I see nobody else talk about it. It's so little talk that when I mention it on X, the devs comment on my posts sometimes.
Curious how this could be applied to greenfield coding rather than just making a copy in a new language or performance optimizing.
https://ctx.company/blog/introducing-ctx-traits/
For Anthropic's reported benchmark numbers for Sonnet 5.5, when running terminal-bench 4.0, they had a comparison between $0.10 and $0.20 cache read for total task cost. That 50% cost reduction resulted in a 20.9-26.2% total cost reduction. Pretty specific though... single benchmark, model, harness, and provider. Cache reads make up ~42-52% of the total cost at the standard $0.20 price.
Now there is no excuses to not use Rust and it just shows in raw performance alone.
Eventually AIs will generate Assembly code directly anyway, no need for intermediate 3 GL languages.
Combine that with the fact that, unlike other code generators, LLM output can't as a rule be reliably reproduced from the original input in the future, sounds like a recipe for unpredictable long-term costs and regular regressions.
JIT compilers, GC workflows, PGO, and ML optimising compilers passes are also non deterministic, and yet work gets done.
If you are curious, there is already enough work out there into this direction.