Pi.dev: You Said No MCP(earendil.com) |
Pi.dev: You Said No MCP(earendil.com) |
Just generate CLI tools, with docs, from MCP servers on demand.
Swagger was the new old name for the concept
While we are waiting on that to become stabilized, we implemented a inspired/co-evolved way to do that in our tool[0], where you mark individual fields in the request/response schema as being file payloads, so that file exchange can be properly orchestrated by the harness and doesn't pollute the context. We just do inline base64 uploads of the required payloads, which in practice we've seen to work quite will until ~100MB files (which is otherwise also the size limit we usually recommend for file processed).
It's annoying that it's not stabilized yet, but for most bigger customers we've seen, they implement 80% of the MCP servers they connect in-house, so doing adjustments to the tool surface, and metadata has been less of a pain for them than we expected.
[0]: https://erato.chat/docs/features/mcp_servers#file-support
If you are a Pi user it may be better to just ask your agent to explain https://github.com/earendil-works/pi/pull/10040
If we fail to explain it, then we need to do a better job explaining it :)
But pi doesn't have security. Pi is full yolo, it's on the user to run it in an environment that minimizes the blast radius if the LLM goes haywire.
So I'm not sure what the plan is here. Will pi support running certain tools like bash as a different OS user than owner of the pi process?
I find this approach is easier to debug and I can also use the tool myself to ensure it's working well.
And you didn’t remember that when you said no to MCP?
No, no MCP for now?
So even with a local Qwen and Pi you can now say things like:
Set up Clop to optimise any PNG that I drop in my website assets folder and convert to a webp with the same name near it
Get Crank to start Time Machine backups immediately when I connect my HDD and notify me when the backup is done.
I want to be able to hold rcmd and fuzzy search and focus cmux agent panes
BetterTouchTool has a great MCP which can create native SwiftUI views and bind them to hotkeys, trackpad gestures etc. It can leverage its immense macOS automation tools and private APIs to let agents do Computer Use.You would need a much more capable coding model to code those tools from scratch and get the same fail-safe logic that the apps have honed over the years.
[0] https://reddit.com/r/macapps/comments/1wkv0dy/mcp_in_macos_a...
A direct quote from March, 2026[1]:
> If you’re still not convinced that a lot of this discourse [regarding the death of MCP] lacks nuance and is just hype, congrats on buying into the current AI-influencer FOMO hype cycle; see you in 6 months when the influencers move on to the next revelation of the moment to stay relevant and get your eyeballs and dollars.
It was fairly obvious why MCP would be needed once AI engineering and uptake moved beyond the solo developer and single harness stack of "what works for Me" versus "what works for My Team", particularly in an enterprise context. The key mistake people made was thinking in terms of their own workflows and own local stacks instead of a team's workflow and a team's operational stack. There was also an ignorance of MCP's stateless HTTP mode (yes, it was already a thing in March; the 2026-07-28 revision of the spec just prioritizes it as the primary focus moving forward) versus local `stdio`.My biggest complaint right now is that OpenAI has still refused to implement the MCP Prompts spec[2] and in general, the major clients have spotty implementation for some of the features in the spec.
[0] https://news.ycombinator.com/item?id=47380270
[1] https://chrlschn.dev/blog/2026/03/mcp-is-dead-long-live-mcp/
It is only going to continue to proliferate in usage and adoption.
It’s suboptimal for the reasons the author outlines: but so is USB-C. So is NVME, so is HDMI.
We use these hugely successful technologies in spite of their flaws because they’re widely compatible and easy for the end user.
That’s why MCP is everywhere. It might not be performant, robust and uniform but it WILL get better over time.
And I’d much rather have the broad MCP ecosystem that we have now than seven or eight different “optimal” ways of plugging in an LLM to something useful.
I hope they'll do the same and eventually add native support for ACP (https://agentclientprotocol.com/get-started/introduction) which, on the contrary, I use quite.
Then again, I don’t even know if general adoption is what Pi/Earendil is going for.
I suspect that a smart model driving multiple dumber models for work and then using sub-agents with the same smart model for adversarial review will be a pretty common pattern.
Personally, I got a bit confused about Pi having most of that stuff as plugins since I remember how much of a mess Eclipse was where so much was just loosely fitting together plugins and just went with OpenCode since it covers most of my needs out of the box. Guess that might also be a sign of me getting older, because my IDEs and desktop environments are all closer to stock too.
Treating MCP as a part of OpenAPI rather than a tool connector is a direction in which we're heading. It is important for the users to have the flexibility of deciding the model, work to be done and the tool call in one prompt. The framework sets up the configuration and gets the output.
Another perk is it allows me to run tools in the exact same way as agents instead of treating MCP as a special way to call on services. Super valuable when debugging.
Another retrospect note, "No MCP" appears to be the first icon on their front page - not sure how I missed that.
Imagine my surprise reading this!
It's pretty effective because of the reasons you noted, but there's a composability problem since each MCP has its own sandbox and can't call into the other ones.
IIUC Pi offer a workaround for this, the harness runs the sandbox and populate it with the MCP tools, that way the composability problem is solved and every MCP do not have to implement their own sandbox.
Codemode is a way for the LLM to orchestrate harness level tools. The reason this happening now, is because the models by the labs are increasingly trained on this. Codex for instance in responses lite requires codemode to even perform parallel tool calling.
* speed - much fewer hops back to the LLM
* fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.
* repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.
If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.
However- in my testing, mcp is really quite fast, and its pretty much free at this point- with frontier models. Context rot is, from what ive tested, not as much of a concern now. I genuinely was not able to hillclimb skills/extensions to beat out the speed of mcp in some cases I've been testing.
The first tool execution runtime in harnesses are direct tool calls with JSON or XML, such as the Read and Edit tools. As an escape hatch, we have Bash tool that allows arbitrary code execution on the host running the agent. The downsides of using bash (on the host) as the main tool execution runtime are:
- Syntax and obvious errors only surface at runtime
- Unergonomic orchestration of parallel and background tasks
- Verbose command output cluttering context
- Dependent on the host environment, packages versions, etc.
- No security measures by default.
To me the last point is the biggest inherent weakness, usually mitigated by creating a dedicated unprivileged user or running bash in a sandbox.
Note that direct tool calling is kind of the polar opposite on these points: syntax errors are caught early, orchestration can be done with some wrapping tools, command output is controlled, and most importantly they are more sandboxed. On the flip side, they obviously have way less power, necessitating Bash tool in the first place.
Codemode is the middle ground between these two extremes. It actually can be derived simply by one idea: what if we replace Bash by another language that can be checked for obvious errors, i.e. type checked?
Everything else falls out from there:
- Any language would do, but I think TypeScript fits the balance between safety, speed, conciseness, and popularity in training data.
- If we use TypeScript, might as well run it in a sandbox as JS runtimes have been designed with this in mind for 20 years
- Orchestration comes for free from the JS runtime. It's not more powerful, just more ergonomic.
- Since the tools are controlled by the harness and not dependent on the host, cloud agent becomes easier.
- With this in place, MCP are not very different from a tool provided to this sandboxed runtime.
Overall I find the benefits compelling enough, but we'll see if the heavily-RLed models these days will use it effectively.
I will use this to talk to my manager about the project status
There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.
It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.
With a skill, updates depend on whatever channel delivered it to you. Whichever channel that is, it's out of my hands as a provider.
So, MCP solves the problem of coordinated distribution of updates to a larger subscriber base. Think inside of a company, for example. I don't have to go around and tell people to `git pull` their skills folder.
It's typed so you can build some governance around it, by allowing only some tools or parameters for your org (this is a pretty weak point, but still)
A skill has one giant description from the frontmatter loaded into the context, where MCP loads a smaller one for every tool. Not necessarily better, the skill approach is often better actually, but sometimes the MCP approach fits more
The idea of using jev as a cheaper faster subagent for specific use cases is interesting. Will have to experiment with that!
I'm actually pretty happy that they did it, since 90% of what I have to integrate in enterprises is MCP-driven (it's a security and auth boundary that has become pretty much mandatory for any third-party agents wanting to reach into corporate data) and this lets me use Pi directly. Am just being cautious about the first version, because, well... it's a first version, and I like my tools stable.
(I actually played around with the idea of using QuickJS myself for codemode, but since I rely on Bun that gives me the ability to use other things... never got around to do it though.)
If you feel that Pi has been drifting away from its original vision, try hax (https://usehax.dev/) - you might like it.
Thank you. I've been frustrated by harnesses hijacking the terminal and breaking basic features such as scrolling and text selection.
It even sends BEL when the agent completes, which makes so much sense, yet Pi never implemented it.
I'm definitely going to use it over the next few days and hopefully make the switch.
Codemode isn't replacing MCP; it's fixing MCP's biggest flaw—its lack of composability.
In this very post:
> While a lot of things have improved about MCP, quite a few have not. The biggest issue with MCP continues to be that it’s hard to compose. Even with codemode, which is just a neat little sandbox to allow composing of tool calls, MCP doesn’t fully deliver on this. But that at this point is less the problem of MCP but the MCP servers out there and different approaches of harnesses to work with them.
Can we get a bit more clarity on this? With codemode, what's the gap? I've also been investigating the search+execute MCP server pattern evangelised by Cloudflare (it uses codemode inside the MCP server to bypass the need to expose a large number of individual tools), but i've seen people say that doesn't compose well either.
Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:
/autocompact 420k
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/
https://github.com/can1357/oh-my-pi
I haven't tried it much though, can't vouch how well it works.
On top of that, somewhat unrelated I’ll agree but still, it has support for vim keybindings
I did create some extensions where it spawns sub agents for specific tasks, especially when I want to keep the context clean or when I really want to offload a piece of work to a cheaper model. And for that I have a high degree of control over, I know which model is being used for each subtask.
I find Claude Code too unwieldy for my tastes. Pi's philosophy of being very light on features nut highly flexible for customization, clicked very well for the way I work.
Like sub-agents, you could just instruct pi/any harness with a user prompt/system prompt to start new invocations of itself, if you share what the exact command is, and pi or any other harness will do their own poor man's version of sub-agent via standard unix programs.
Pi is primarily a coding agent, so yeah, code mode makes sense, but I've found that better MCP design saves everyone a lot of trouble and would also probably have improved the thing's reputation overall (I personally am not fond of the line protocol, would rather have protobuf and more typing, but it is what it is).
I many scenarios, e.g. running the harness server-side, as is the case for chat interfaces, you don't really want to expose OS shell access as that opens up a huge security attack surface.
It does, but a restricted user account mitigates the large majority of those issues. A sandbox mitigates even more.
The number of remaining exploits left is probably going to be the same as the number in the harness. More, in fact, as many of them have no human review anyway.
That's been a trivially solved problem for decades.
Rootless immutable containers without shell access, or SaaS products from multiple vendors with WebAPIs as the only touch point.
Anecdotally, I find the auto compaction (or what I assume is happening when the context magically drops) to be hit or miss. I do like how easy it is to use my work cursor sub and business chat gpt at the same time. Then I use nearly free cursor models for dumb shit and Sol for real problems.
GPT context window is way too low for me and my last experience with it (GPT 5.6 Sol) was so awful and I hit limits way too fast that I cancelled it (and at least got my money back).
I'm no longer using Pi since it got worse IMHO and Claude subs can only be used in Claude Code but I miss the /tree feature which is perfect for first letting the model read & cache the important bits of the codebase and then start your plan from there (as long as you stay in the Cache TTL). Claude Code has /rewind but it's not as good.
I'm only using the 20$ plans.
For that is it not better to have separate sessions for planning stuff and doing actual work? Pi is super flexible with session management, and a lot of that can be automated by its extension system.
But they'd still inevitably get to long in the tooth, and context poisoning meant they'd just eventually not be able to stay in the preferred context size, which for me is 64k-128k. So, I extended it with an eviction command and required a ratio. So instead of a summary of work, it now just places a waypoint. The waypoint basically means the context has a semi-coherent context but without all the baggage.
I'm on like day 3 of a single session with 3m tokens removed and still in the sweet spot. So it evicts to beneath the lower limit, compresses to the upper limit, then evicts again.
It's amazing how resilient it is if you give it a good plan. The work flow has basically been:
1. Write up an implementation document for some new set of features.
2. Rewrite the implementation as a TDD document
3. Set it to work.
The only thing I haven't figured out is it likes to stop when it hits the finish line of the subparts, but likely we're going to end up with the master of puppets monitoring these things and just set them to evaluating what they've done.
Doesn't that destroy the cache? I find that caching significantly sped up my Qwen, especially on said larger contexts.