Pi's Minimalism Is Its Advantage(earendil.com) |
Pi's Minimalism Is Its Advantage(earendil.com) |
If YOLO MODE causes pi to destroy my workstation (it hasn't yet) I'll just nuke the thing from orbit and spin up a new one.
I launch with `srt pi` and get file system and network isolation. There is a seemingly infinite risk surface area to protect, but I think does a reasonable job of balancing security and convenience.
I've also heard good things about nono [1] from colleagues, but I haven't personally tried it out yet.
[0] https://github.com/anthropic-experimental/sandbox-runtime
Its very easy to use and has a pre-made profile for pi. Just do something like `alias pi="nono run -v --profile pi --allow-cwd -- /opt/homebrew/bin/pi"` in your shell.
opencode has never done that.
I'm allured by the minimalism, so I didn't quit there, but I'm not keen on letting it loose with vague instructions, that's for sure.
For example, it comes with a bash tool built that you cannot disable. This is not minimal, it's the full kitchen sink. If I want to build a custom agent I have to literally stop using pi.dev and switch to something else.
So yeah, I fully disagree with the title. "Pi’s Minimalism Is Its Advantage" No. full stop. It's too bloated for me already. It's not minimal enough. If it's minimalism was its strength. it might not even need a sandbox, because it can't run bash commands or update files to begin with.
Think about why a sandbox is needed: Your permissions have been too loose. You now need to deal with the fallout of your decision externally. If all the agent was allowed to do is read your files and run cargo test, you wouldn't need a sandbox at all, the agent is the sandbox.
Now you might say, but what if it needs to modify files? If you wanted to build a sandbox or approval workflow here, you'd put it right into your custom write tool. It could be an extension you just download so you can pick your favorite write tool. Instead, the authors of pi.dev chose the worst possible defaults.
it feels like there's still a strong layer of "for us" vs "for you" within opencode, that i hope, over time can get chewed away at. plugins to rebuild history, to re-title are just impossible, for example. none of these changes, these freedoms are hard to release. the patches i juggle are easy. but whether or not my agentic software serves as a substrate for my desire, or whether it allows me to augment agency: tis the question.
Dax (opencode lead) has such humble takes, is so forthwith about trying failing trying again on and on. about iterative improvement. and it feels like the guts are so in line to deliver, to allow such freedom now in OpenCode. but i don't see the product (anti-product) alignment, where opencode understands that it's competition isn't cc or codex, which can't and won't ever really compete, but pi, that the competition is to be the putty, to deliver the agency, to be a substrate. really hoping, because i love opencode, and these internals in v2 are sick.
the "devtools must be open sourced" debate comes screaming into the fore on this. it certainly argues similar to the post here: that it is minimalism, it is adaptability, programmability, it is directability that unlocks and unleashes us:
> Imagine the convoluted misery it would be trying to plug that into the VS Code extensions API! Or trying to get it into vimdiff. It would certainly be possible, but the machinery to start pre-processing the commits as soon as they appear would be nigh-on impossible. - https://blog.exe.dev/devtools-must-be-open-source https://news.ycombinator.com/item?id=49156111
i don't even fully agree! today more than ever, why not cut a VS Code extension? why not cut some wild coop.nvim async extension that runs whatever subprocesses, talks to whatever system daemon? dream it up and do it; the llm's will cut through the mechanicals. but the core point, about finding software that doesn't obstruct, that accelerates the human agency: it's so Douglas Engelbart. to Augment Agency is so close akin to Augment Intellect, the grand passion for human interest engagement envolvement constructivism fucking-around-and-finding-out. and my how unhindered we can be now. if only our tools/systems/softwares let us be. here's to you, soft software!
Him and Armin especially, have been involved in OSS for years and have worked hard to build communities around projects they’ve built and maintained. We’ve all seen enough projects become what you’re afraid of, Mario included. I’m optimistic that they’ll keep true to their goal of keeping pi open while building their other products around it. I think they understand the community dynamics necessary to keep a project like pi going. And they want it to succeed that way.
Disclaimer: No first hand knowledge
Only when you genuinely find a feature is missing should you write one yourself or have Pi write an extension or skill for you.
Don't use extensions written by others, after all, those extensions were also vibe-coded.
It really doesn't take much for it to be useful, maybe something to search the web?
If you want you can just tell it to write you extensions too, like one to wrap curl so it can get to the web easier for instance. Or just tell it to use curl, really up to you. Just an example to point out: like, do whatever, it's very flexible.
I'm doing quite well with just a "todos" extension, and a "plan" prompt template (doesn't enforce read only tool usage, but prompts to build todos and discuss before doing anything).
I will probably try out some subagent systems soon, but I'm doing surprisingly well without them.
If you want a more tricked out "starter pack" there's oh-my-pi or lazypi and maybe a few others. But worth being careful what you install because 1. A full pack of extensions can destroy the minimalism of pi 2. Random extentions are a security nightmare.
You could try ohmypi but it sort of misses the point.
Why not just install nicopreme's pi-subagents and pi-web-access and then install whatever else you need when you find it's missing?
I realised how use case dependent harness behaviour is when I tried to use my customised-for-a-side-project pi config at work and realised I needed to tweak it significantly to be useful - I would not be surprised if tools like Claude Code needing to be all things for all people is hurting their peak usefulness.
Doesn't make sense to me, but to each their own.
I had a lot of success running it on my Mac with Qwen3.6-35B-A3B model
I ended up on VS Code. I'm very critical of Microsoft generally, but VS Code is a very good editor and my favorite agent harness.
For headless, Pi might be the way.
become ~~ungovernable~~ ungooglable
at least with GPT 5.6 Sol fwiw
https://smolenv.com/t/nested-template-includes-60636/
sh is all you need
If I have agents.md or other context I want it to read I mention it at the beginning of the session
re MCP: I am not using an MCP with smol
but there are ways to convert MCPs into CLI tools or typed js
I imagine that would work well/more token efficient with smol (or most harnesses actually)
The best thing I found so far re smol is that it fits into the context window with plenty of room to spare
So it is easy to adapt (and add stuff to it, even stuff you only need specifically for just 1 project)
Whereas adapting a more complex harness is more error prone
I also see codex do it that way quite often
and at the same time Opus struggles with using the edit tool in Claude Code even though model and harness are by the same company
The result is that the tool gradually morphs into the thing I need rather than me having to adapt myself to whatever new thing Anthropic or OpenAI comes up with.
I can also feel confident that the thing it becomes is what I actually need and not what maximizes token usage...
Please provide evidence of relationship of pi/earendil to peter thiel.
I also suggest that if you believe peter thiel relationship with anything is a nuisance then - maybe - try avoiding gut-assigning links of anything in the world to him. you maybe see things in more positive light and it would be fairer to those things.
Peter thiel doesnt own tolkien work. The world is bigger than one mans bias
#teampi
There are also a ton of small mechanical things a harness can handle that make the whole process much smoother. A really simple example is auto balancing parens. Even frontier models like Claude still struggle with this. Often the model will end up writing a python script to figure out where the mismatch is, and then generate a new version of the code. All of that simply wastes tokens and eats up context on a task that could've been accomplished completely mechanically.
The approach I took with dirge, is to put the model in a loop where it has clearly defined tasks, and the harness handles any repairs that can be done automatically. And I used Janet to provide a plugin system based on what Pi is doing. You get a batteries included experience out of the box, and you can customize it to fit a specific project using plugins if needed.
I do like the grug-brain approach of keeping things extremely simple and easy to reason about.
One thing that I don't like about Pi is that it's almost too extensible, in the sense that I can add a lot of shit into it without really understanding what a given extension is doing. And both from a security and token efficiency standpoint, I like the premise of converting things like MCPs into CLIs. It might be worth investing in tooling that works nicely with the agent harness, but that is not directly integrated with it. I'd be glad to work on that for smol if I can get a workflow going.
Also happy with how much love codex gets from OpenAI.
That said: I was looking at existing agents to find one to build upon and to me they were all too complex and were leaning too heavily into 3rd party dependencies.
Nothing I could understand comfortably in an afternoon (that's also on me I guess). Pi was closest to what I was looking for but still too big and too modular.
(It's hard to come up with good abstractions that work well across all major models + keep up with new concepts that come and go all the time with new releases.)
The more complex agents err on the side of supporting many models 'ok' instead of focusing on taking advantage of a specific model.
With a tiny implementation it is easier to adapt it.
Adding new stuff, removing stuff again, changing it from working well specifically with GPT 5.6 Sol to working with the exact model I want.
Same with any of the off the shelf harnesses.
Go build something useful instead of tweaking the minutiae of the harness.
For a program that's minimal it sure takes a long time to start up, the standard C-p and C-n bindings don't work, it doesn't follow the XDG Base Directory Specification and just pollutes my $HOME directory.
I think there's still space for another harness that's 1/ open-source, 2/ written in a fast compiled language (Rust, Go, etc.) and scriptable in a simple (aka non-JS) scripting language (Lua, etc.), 3/ less opinionated and more sensible so things like XDG isn't a WONTFIX.
But I agree with you, it's biggest weakness is that for a real long time the tagline of it was "there are many harnesses, this one is MINE" (That being Mario's)
I have a lot of respect for Mario and his team, but there's things like you've pointed out that deviate from standards, and other issues that I've seen get posted, only to get knocked down by the team as WON'T FIX because, even though the new owners changed the tagline from MINE to YOURS... It's still very much Mario's.
I do like opinionated things. Truly. But I'm also of the opinion that standards exist for a reason.
That said. I like Pi so much that it's my daily driver, and I've created an ecosystem of plugins to do everything I want, having them all tie together and communicate through the shared bus. Pi is really a good harness.
It's just, well. I don't agree with some of the opinions.
If I'm going to add another thing here... Whilst you cannot get everything you need from the openAI API spec, you can get a surprising amount to get a model config. That said. Versions of Pi are still shipping with model configs for certain inference providers. I do hope that gets decoupled at some stage. I see the groundwork being laid.
So the work is being done in the right direction. I applaud the team but I do get the feeling that a lot of this is because people want to contribute, but the team really wants to hand craft this. And that's great
Being reflexively contarian is not what HN is for.
https://news.ycombinator.com/item?id=45530593
> For a program that's minimal it sure takes a long time to start up
It took the same amount of time (3 seconds) as cursor-agent and codex on my machine, which isn't particularly high-spec.
> the standard C-p and C-n bindings don't work
It has programmable keybindings and you can ask the agent to remap them in five minutes, if not 30 seconds.
> it doesn't follow the XDG Base Directory Specification and just pollutes my $HOME directory
You can make it put .pi/agent anywhere with $PI_CODING_AGENT_DIR in 10 seconds.
> scriptable in a simple (aka non-JS) scripting language
This doesn't make any sense. The very point of Pi is that instead of implementing all of your features directly in the harness, you make the harness minimal and then implement the functionality you want as extensions, which necessarily means that you have some powerful and expressive extension language, ideally the one the harness was written in.
I suspect that if the harness was written in Lua (which I love and is probably the least bad choice of "scripting" language), you would have far more issues with it.
If these are your complaints, then this is one of the strongest endorsements of Pi that I've ever seen. I think I'm bookmarking this comment.
This drives me mad, I believe Ollama and Claude Code also do this. Seems to be rife in the LLM world. IMO there's no excuse for new software sticking dotfiles in my homedir in 2026.
This 11 year old, open issue is very symptomatic of this IMHO: https://github.com/rust-lang/cargo/issues/1734
Agreed, and also, tinfoil hat time:
I believe they opt for this so that state and config files don’t need to be distinguished (it all goes into ~/.appname the same). It’s still not an excuse, but maybe laziness is the reason?
I’m not a fan of config directories being in different locations on different platforms because it’s now one extra thing everyone needs to handle.
(Disclaimer: I work on Pi but I dislike XDG in all settings)
Written in C, tiny footprint, minimalist approach to system prompt and tools (yet essential batteries are included, for example - it has subagents with presets, and background bash tasks out of the box), high quality polished presentation, inspectable (usable transcript view), does not mess with terminal scrollback, respects XDG directory spec, etc.
Open source, MIT-licensed, no commercial agenda. A tool that I myself wanted, so I built it.
Yes, I can probably inspect that but I do think installing through package managers is the best practice.
It looks better than pi with XDG and not being JS but that is it's own red flag for me.
And then I installed it and found all the problems you mention. Not only that but I read Github Issues about the XDG problem and was a bit taken aback by the reaction of the developer.
It's one of the best agents I have used so far but I'm still looking for a very lightweight, token efficient agent NOT written in a JS framework and which respects XDG
I know it's little, but this was the first thing I noticed and it made me think, "Maybe this app isn't for me."
EDIT: FWIW, I just complained to pi and it added the keybinds for me in 10 seconds.
It starts up for me in under a second on my M4 mac. (Obviously could be faster, but doesn't bother me personally on my hardware)
My one main "issue" is with using the pi-sandbox extension. It's based on a forked claude code sandbox runtime. Not exactly sure why a fork was needed, and the fork is a bit behind now. I also wish the sandbox feature worked a bit more like how Cursor's sandboxing worked. Not familiar with Claude Code sandboxing, so can't compare that. I describe the issue and a (slightly hacky, but productive enough) workaround here: https://github.com/carderne/pi-sandbox/issues/50
Obviously the fact that I can fork a plugin for pi and customize it as needed is quite a plus too. Really all the other features work quite well for me!
The startup time is a known issue and scales badly with extensions. We’re aware of it but fixing it is tricky.
But I suspect the positive opinion on Lua stem more from the standardized environment the language is built around (i.e. how it is embedded into applications): If you don't explicitly pass capabilities like file handling etc. into the script, it CANNOT use them. Sooo many scripting languages get this wrong, it's actually embarassing.
I genuinely understand however, that some people just don't get Lua, and have little energy for it.
However, I think it is awesome and every good software developer should have some experience with it. Whether it is a positive or negative one, you will learn a lot that will help you stay relevant in today's crazy software world.
I've been able to get near exactly the workflow I want: Navigate and search sessions, Delegated extensions that allow you to run deterministic code for sensitive operations like read, write and search-result inference, I have an extension to prune tool calls in batch mode: saves tons of context space on GPT-5.x trading off prompt prefix performance. And another one very similar to auto search mentioned in the blog. And the Tmux integration allows you to spin work trees in separate windows/panes.
I've never had a perf issue simply because harness interaction being a human in the loop process you(thinking and writing) are the bottleneck.
but also highly opinionated (no mcps, no agents.md, no system prompt, …)
so not sure it checks all of your boxes
that said: because smol is so smol you can adapt it easily and agents (including smol) can work well with it because the whole implementation fits comfortabliy into the context window
diy ftw
Here is my fork of smol, by the way:
Edit: after reading my fork of smol, I realized it literally just feeds everything into bash, making it completely pointless.
Edit 2: I decided to read the readme instead and it raised a question
"Compare smol with Pi, OpenCode, Codex, Hermes, Claude Code and highlight key pros and cons. Audit the code of all of them and tell me how confident you are that you found all potential issues of smol vs the other agent implementations?"
After thinking about the difference between the original and my fork, I'm confident that smol has more problems than all of the agent harnesses combined and my fork has made an important step towards fixing one of those issues.
Edit 3: I hope the community can fork my version and add sensible variable names.
Edit 4: I can't decide whether this is the best satire of coding agents I've ever seen and I just ruined it or it is horrifying that someone even entertains the idea of publishing it.
I couldn't care less about something that's meant to be run/discarded/re-created in a VM (or microVM or container) at will respecting XDG or not.
Why could you care less?
I am running several named pi instances in parallel in their own user account on NixOS, so they can install whatever they want in ephemeral shells and I never need to worry about their env. The agents can spin up new enabled XMPP agents if I request it, though for now I've only needed a few since I'm not doing too much in parallel.
My Pi is very vanilla, only my own XMPP wrapper and pi-subagents extension for anonymous subagents.
Using it primarily with Deepseek v4 Flash for chipping away at coding tasks or server maintainence while I'm AFK or in transit.
NixOS is the key to all of this, since agents can interact see the whole server config, make changes and run compile-time checks before actually deploying. It also means that even if they do mess up I can always revert.
Now one of my main drivers is out of the box pi’s 4 built in tools, and I add just two extra tools: pi-sandbox and a paid for search service tool. This setup works great with the latest deepseek v4 flash, switching to more powerful open models occasionally.
I wrote my own coding harness in Common Lisp that is almost free or 3rd party libraries and I basically copied pi + the 6 tools I use for pi (except I have two search tools using different vendors in my Common Lisp code).
Everyone (and every company?) should run their own tests and experiments. I find it sad when I talk with people who default to the most complex and the most expensive tools without even trying to evaluate alternatives.
I’d argue that there’s a minimal set of functions that a coding harness needs to just enable a model to get stuff done, and they shouldn’t be an extra effort to set up.
(Oh-my-pi exists for those of a similar persuasion.)
I'd suggest the default should be anything that makes the model more efficient or effective to a reasonable current level.
There are some experimentations by Igor Warzocha to extract Codex shapes and put it in Pi: https://github.com/IgorWarzocha/howaboua-pi-stuff/tree/main/...
I'm expecting every model will have a fine tuned Pi extension at some point.
migrating to the same compaction and exact tools as codex uses will make it at the same level as codex so what benefit will it have over codex? sure you can customize tui to your liking and add something on top, but the efficiency gains will be gone
The result: 38% fewer startup tokens, 17 tools exposed through just three schemas, and 19 skills loaded only when needed.
IMO, I view it more than a coding agent, it's a coding agent platform with powerful extensibility.
If you're willing to put in the work to master the learning curve and push through the issues, it can be a great tool: https://i.sstatic.net/7Cu9Z.jpg
I am slowly shaping it to the way I like to work. It's a bit of a bumpy start, but there are two benefits I see from it:
First is that you get to understand better what goes on behind the scenes. Since it is minimalistic, you have to think about your workflow, about what you want your agents to do, how you want them to do it, etc. It's a completely different way to work since tools like ClaudeCode do a lot of heavy lifting behind the scenes and just force on you their established way of doing things. By adding the building blocks yourself, you end up with a better understanding (and in many ways, control) of what is happening there.
Second, I adore how I can juggle sessions in Pi. /tree, /clone and /name quickly became second nature for me to manage agents. I am totally abusing prompt templates in my development workflows.
Also, it is an awesome harness for the models I use. DS and MiMo feel very snappy now, especially after I started to adapt Pi to my way of working. Lastly, even though those models are already cheap even on Claude Code, it feels like it is even cheaper now.
- outfitter compose different agent profile w, eg skills, mcp, context and wrap pi,Claude,codex - agent-operator run these agents into kubernetes - actions run pi during investigations on failed ci or weekly updates - channels give pi agents an inbox for slack/email/websocket access, calendar wake ups - deepwork build long running workflows with verification gates
The best example of different profiles is having a prod/nonprod bot that has grafana mcp for incident investigation
I'm really enjoying the advisor mode (you can have a second model monitor the output of the primary model and have it "steer" the primary when it makes a mistake or goes off the rails) and the automatic fallback to a second provider if the primary one has issues (Deepseek had some issues yesterday).
I need to dive a bit into the system prompt to see how much context OMP actually adds. I think the system prompt is still 2-3k tokens but that probably depends on bells and whistles.
I was expecting an article about Pi constant or maybe raspberry Pi computer. "My Disappointment Is Immeasurable And My Day Is Ruined!"
To me there's pi, the constant. Then there's "a pi": a Raspberry Pi. Now there's "pi, the agent" too.
It gets confusing.
Especially if you use pi on a Pi to write code that uses pi.
Anyone here using the Rust rewrite of pi.dev as their daily driver? It's endorsed by the author of pi.dev and looks pretty attractive to me being both minimalist and not having npm attached. Any info "from the trenches" are appreciated (setup with sandboxing, extra niceties etc).
First indication of business strategy around Pi's reverse acquisition of Earendil, perhaps?
My advice: focus on getting work done and slowly adapt Pi with small augmentations as you go. You can start getting work done on vanilla setup. When the right idea comes along, try it. Be ready to refine it, and most importantly, rollback the addition. I've rolled back a bunch.
Many "batteries" that are "included" come from speculative and half-baked ideas, from people who were excited about something at some point in their journey. In practice, those ideas may not bring the desired results, and their creator may've moved on already. So it's better to either learn very well established tools, or mold your own slowly.
For example, many automatic memory systems are not helpful. I built a small extension that asked me whether it should remember something (and write it down to a properly scoped SKILL or AGENTS file). Turned out I accepted less than 5% of suggestions. Most were useless one-offs that would pollute the context. Can't imagine how much crap would accumulate if I wasn't in the loop.
I found this one a little bit better and they do support Extensions like Pi. But comes with all features like codex, claude code and it's open-source.
So far API usage is a lot more expensive than subscription and if you need better models than Deepseek and Kimi, and you are not wealthy, I don’t see a way around this.
One game changer when it comes to tweaking configs that are optimized for your use case is that you can easily use a "more powerful" cloud model to identify a good enough config for your local server/pi settings combination [2] in a pattern that applies pretty much anywhere.
- [1] https://huggingface.co/Qwen/Qwen3.5-35B-A3B
- [2] https://alexhans.github.io/posts/find-the-loop-story-first.h...
But regarding coding quality, task completion, over engineering, all related to the final result that the LLM is delivering trough the harness, is there any benchmark that shows the ups and downs of each one?
I'm struggling rotating harness because of a lack os a way to truly compared what is good or not.
- which env do you provide?
- which model(s)?
- subagents?
- system prompt (default or custom?)?
- agents.md
etc etcalso you kinda have to look at many runs and study their traces, if you look at too few runs the outcome variability you're drawing from is too high
The headless pi + xmpp wrapper ended up working much better because the XMPP bridge is the only interface and I get full control over its capabilities.
This is my wrapper: https://github.com/zachpmanson/pi-msg
We added native support for nixos for the same reason - malleability and debugging becomes easier (also because one of our customers asked us to). I think we might be the only sandbox provider to add this in warm pools.
That said, I really don’t like “developing” over chat. I’d much rather wait until I’m really available to inspect diffs properly and watch all the thinking and tool use, real time.
https://github.com/pkulak/nix/tree/main/modules/features/ope...
Funny that the XMPP clients has taken more of my time than the pi XMPP wrapper itself.
I've been able to do 80% of what I needed to do remotely with that setup, so wondering if a more complex setup would really add much...
Want I've been doing for transit is using termux on android with magisk for root access, allowing me to install Nix home manager and harnesses on my phone, and then drive everything from my phone, including eternal terminal sessions to my homes erver, where it spins up agents there as well.
Incidentally it's also great for debugging issues with my phone.
1) system prompt in pi is quite small (way smaller than the one from OpenCode)
2) when your agents.md file changes pi does not re-spam it (preserves cache, good trade-off!)
3) only 4 tools, every tool comes with a description for how to use it and causes reasoning overhead (fewer tools is good)
all of these things add up
here are pi, opencode and smol working on the same tasks in 9 fresh runs
https://smolenv.com/t/nested-template-includes-60636/
you can step through the traces and see how the system prompt + tools steer the agent in a certain way
with GPT 5.6 Sol you can even get away without a system prompt (see smol) and only 1 tool (sh)
it also have soft/soft compaction limit, it tries to compact on turn boundary when possible. with combining with above this can get you about 35% more context (at least it looks like this with the sol)
codex when shell command is executed, will pull output with hard cap at max 30s, so for running compilation it will burn tokens without any benefit.
I have some tasks where agent will have to run some suite that can take over an hour, and codex burns about $20/h just waiting and reasoning every 30s "yep, that's still running". And what is going to happen after compaction, when whole context was just waiting? it will loose the plot and when I'm back it just does completely different thing that I asked it to do.
codex also have a bug, that opening refuses to resolve that adds your last steer after compaction, so imagine that you asked it to cleanup some tmp files or refactor/simplify something. it will do that again and again after each compaction, best case it just burns tokens and figures out, this is already done, or worse do it again and mess up everything and forget about it's task
Basically, it doesn't handle the context "better", it barely does anything special to it, which can actually be better for cost efficiency.
I don't want to spend 2 hours prompting, configuring and fixing features that I need which are standard in every other harness. I don't really want to be wasting my tokens to make an application function like every other harness. I don't really want to have to repeat the cycle on every machine I want to work with. Every VM, every server, every laptop.
For instance the XMPP integration someone mentioned allowing agents to talk to each other and to you remotely; or custom extensions to enable workers to be tmux aware; or adding whatever memory system you’d like; and so on
Also, you can version control your tweaks and easily sync with other machines, just like any code.
I understand both world views and both are legitimate. But I do feel like LLMs are advancing so quickly that it's not a good use of my time to optimize harnesses at this point. I have actual work to do, so sharpening my tools needs to be selective and time-boxed. Personally I'm staying agnostic on harness, not locking into Codex or Claude Code, but also not prematurely optimizing things that tens of thousands of other tools-focused developers are going deep on across the ecosystem. My goal is not to be an early adopter but to reap the benefits of all that experimentation.
Even more so, Pi even has opinionated forks and "distributions" like oh-my-pi that are like LazyVim/AstroVim.
Yeah maybe Claude/OpenCode/KiloCode/Hermes/whatever are not as minimal as Pi but they also work right now.
And if you try vanilla Pi, you will also find out that it works right now.
I think for indie hackers and people that build their own stack is great, but real scenario and people with money Enterprise likes the idea of batteries included.
I found this one a little bit better and they do support Extensions like Pi. But comes with all features like codex, claude code and it's open-source.
as for workflows, its all skills/agents based, here's my dotagents folder
You absolutely can have tables indexed at 0 - what you cannot do, is fail to take responsibility for the use-patterns you apply to those tables, if you do so - and more specifically you have to take responsibility for the requirement that you use 0-based tables, instead of more optimal methods.
The table is an extraordinarily flexible type. You will gain immensely from using it properly - whether its the newbie dilemma over pairs()/ipairs(), or whether its the professional metatable manipulations - the power of this type is undeniable.
However, if you cannot get past the fact that you must learn it, and that it is applicable to your requirements in every single case, then you are for sure going to have a hard time.
Too much power + too little attention to important details = burnt fingers = endless whining. This is my personal stance having used Lua for decades now, professionally and personally, to do amazing things.
To me it's great how minimal the system prompt and tool set is, and I doubt those features are worth the tokens for every model. (Who knows if they improve performance for SOTA models, and they probably harm performance for small local models.)
I would add "better sandbox" support, which I think should be included out of the box. Not having a very simple way to get out of "yolo" mode is kinda crazy. Sure, there are plugins, but they do have some issues.
pi-bash-approval is good, but manual approval plus allowlist is a "bad" way to run coding agents.
The "best" way in my current opinion, is where commands run by default in a sandbox, but commands can be ran unsandboxed as needed, requiring approval or allow-list in that case. Cursor was pretty good at this (when I used it). For example I don't want to configure my sandbox with access to docker, which would present easy jailbreaks, but I do want to allowlist certain docker commands or approve them to run on my host as needed. pi-sandbox is good at allowing me to configure sandbox access, but this feature where somethings can run unsandboxed is missing. I write more about this and about (hacky, but productive enough) workaround here: https://github.com/carderne/pi-sandbox/issues/50
Even worse are those that seemingly follow XDG but not really, like a lot of Electron apps, that just shove config, state and cache in ~/.config.... sigh
One thing you can do in pi that you can't in Cursor: Have a 5-prompt conversation, jump back to prompt 3 and have a new conversation [call this convo2], then jump back to the original point 5, then jump back to convo2.
So fork is really only needed for when you need to interact with the conversation tree in two separate processes.
cargo install --locked --git https://github.com/tontinton/maki.git makiI find it a bit shocking that someone working on an agent harness can't be bothered to spend 5 minutes to research this with the help of an LLM and holds such rigid and uninformed views.
And if you don't want to respect platform standards, just respect XDG on all platforms. The .app solution is the laziest one possible.
Just follow XDG everywhere and create .config/app & co everywhere, at least that way there's a chance more apps end up in subfolders instead of ending up with a million folders in the user directory on BOTH Linux and non-Linux.
Glad we can still use OpenAI subscription through Pi.
So unless it’s been fixed or someone knows a work around, Pi is DOA - I’ve found that on a MULTI tool call (ie one prompt firing off multiple tool calls until it prompts you again) that’s close to hitting the auto-compaction limit (default compactor or extension) it will either keep going until your context spills over and you OOM, or it interrupts itself to compact but then loses the context.
From reading issue after issue on GitHub, I think it’s because Pi doesn’t let extension writers (nor the built-in compactor) hook in between each tool call and so the only place to check if it can compact is when it finishes a request and is about to wait for the next prompt - too late by then
I've found them to be extraordinarily helpful, because they allow me to much more carefully control context and reduce token spend by using a smart model for the parent agent and cheap models for the subagents. Do you just have a big token budget?
I didn't notice any significant change in context usage, and tasks were completed faster. That surprised me, I'm still not sure (not an expert on this), but maybe the handoff boundary was the problem. When the main model gives an isolated task to the subagent, the latter goes wild producing a comprehensive report, trying to satisfy every possibility. Without the handoff, the main model does the job much more precisely and conservatively, checks only specific/narrow things, and stops sooner.
Recently I decided to reintroduce 2 subagents to see how it goes. First was to have a cheaper model drive my real Safari browser instead of using agent-browser and the like. Second, to see if having a cheaper model navigate/search my file system helps in any way.
I think there's some benefit to having a cheap model drive Safari, because there's so much unavoidable garbage produced in that interaction. The filesystem one I don't think I see any benefit, just a lot of unnecessary work that (albeit cheap) wastes more time.
Of course I'm eyeballing this, not benchmarking formally, but I see so many people just onboard these mindlessly. Are you sure that you saw a real improvement in the produced outcomes/timing, or was it based on seeing subagents do a lot of stuff and assuming that the main model would've been doing the same at higher cost?
I admit that subagents may have great benefits, but I wouldn't treat it as just out-of-the-box basic feature that always improves your outcomes.
My root level CLAUDE.md has pretty much just "use a lower tier agent when relevant".
Then I daily-drive Opus, it automatically offloads simpler stuff to Sonnet or even Haiku based its own reasoning because it "knows" their capabilities.
It's so much more cost/token efficient to do it like this. Opus writes the exact implementation plan for Sonnet and then waits for it to complete. After that it checks the work and fixes any issues itself.
In Codex, for example, this doesn't work because the whole system doesn't know about agent tiers and barely can use subagents. So I'm just running Sol all the time.
.cargo's placement is a historical mistake that can't be undone now, but ecosystem participants are generally good participants.
here is a typical bench run with 9 runs and traces for opencode, pi and smol
you can look at every step and which tools are used and how
https://smolenv.com/t/nested-template-includes-60636/
sh is pretty versatile and composes well
pi also only has 4 tools (good!)
I mean, you can't be serious, right? Think about the unintentional commentary you're making here.
You're saying it can be understood in an afternoon but it is intentionally obfuscated.
It might as well be proprietary code, but if it is proprietary code, you're actually making fun of all the coding agents for being a walking security nightmare, because in the end all they do is run bash and no amount of sandboxing or regexes will make bash secure.
So why not drop the pretenses and just expose the coding harnesses for what they are? Inscrutable bash executors that have the potential to go out of control.
I'm confident you can understand what it does (and does not do) and why in an afternoon
(probably in 20-30 minutes actually, even if you are not familiar with Go)
the same is way more difficult with larger agent implementations even if you only want to understand the direct implementation ignoring all the 3rd party dependencies that come with them
no need to start from smol (even though I think it makes a decent starting point)