An Accidental Blackboard(martinfowler.com) |
An Accidental Blackboard(martinfowler.com) |
I've found that similarly to how teams can degrade into spending more time bikeshedding and on the watercooler than on work, agents also tend to end up spending way too much time coordinating as opposed to doing the work. And so I rediscovered that it's better to have one agent that's the Manager (on a Manager Schedule) and the rests be builders (on a Builder's Schedule), where the manager might be interrupt driven, but the builders need to be able to focus for a while without interruption (context poisoning).
Thanks for writing this and demonstrating that writing about anything is useful to share knowledge and practices. In the end, I learned a lot from Martin Fowler and his gang and I guess I should pay back and write about my own discoveries, however trivial they seem to me.
From what I've experienced when you just let the agents figure it out, to your point, they collaborate awkwardly. If you define how/where in your initial spec of what's being built that seems to go a long way in resolving this. However, agents still seem to end up out of alignment with the demands of the spec. I was testing Astra yesterday on a new tool that should have been able to be completed in a couple hours. I let it go and had it simply use a Sol agent for coding and a Opus agent for review. Opus was explicitly asked to validate the progress between checkpoints, one of those being to keep watch for scope creep.
It was half a day later and basically only the scaffolding was done. I asked why and it literally told me it was working on things I had not directed it to, that it was spending too much time on things I hadn't asked for. WTF good are these uber LLMs when they are making decisions and dismissing the prompt? I'm finding the smaller models seem to be able to stay on track much better and I'm constantly wondering if the current SOTA models should really be used in the review and cleanup phase only. But that seems very backwards as when I first started leaning into building out the most complete spec for a given task - it worked really well. Something seems to be degrading that workflow, now.
I feel like it's becoming more and more of a chore to get things done efficiently. But I don't really find that using Astra/Fable makes anything better at this point. In fact the Kimi models work really well together in this workflow. K3 does a great job of orchestration and I'd say is the more reliable of the 3 for a spec driven outcome. Wondering if this is all intentional by OAI and Anthropic to prod the models under the cover to go off and do their own thing and dismiss the directive.
It rubs me like “Tom Clancy” novels not written by Tom.
Like we would likley never see them or care if it weren’t for the name.
Ed: for posterity
https://www.reuters.com/world/europe/openai-agents-hijacked-...
Honest question, I don't really get how this would improve my work flow.
() Spec and (variably coarse) implementation plan are already grounded in code reality and are commited.
That's only because of a deficiency in CI. If you get smarter CI that doesn't trip on files unrelated to the build, using the repo as a blackboard or wiki is probably fine
The biggest problem I encountered using a similar setup for doing long-term and iterative data analysis is error propagation. My team and I are characterizing and modelling a physical system, there are real-world experiments that need to be run and then fed back into the analysis pipeline and then the next frontier of questions comes up. Any time there has been an erroneous analysis somewhere along the line, that error continues to be treated as a correct fact until it has been decisively eradicated. If one analysis script or document has the error written as a correct fact, that error will continue to pollute future analyses. Oh and these errors can also end up in the agents’ memory files as well. I have gotten very careful about making sure that every reference is corrected everywhere because it seems that the initial error is weighted heavier than the correction.
I may have explored a related approach but as an append-only log riding on source-control to sync state between checkouts (git trailer metadata specifically)
https://gist.github.com/corv89/c506780881b260f4c5a4618fe8d92...
Excited to see where such concepts can take "multiplayer" agentic systems
I’m interested to see if jujutsu changes the equation
Another Blackboard Example: https://github.com/halbritt/striatum/tree/main/docs/rfcs
(Useful to see how this compares to another Blackboard type platform - Gastown https://github.com/halbritt/striatum/blob/main/docs/records/...)
https://github.com/horiacristescu/playbook-harness
Another way to see the blackboard - it is like code, it executes, can be passed around like in higher-order programming, but this it even stranger - it can reflect on itself, not just execute. The task.md file is the agent.
It's end to end encrypted, and has group messaging support, so if you wanted to read what the agents are saying you'd just add them to groups you're in. It also has a web ui. Version 0.6.0 should get pushed this evening pacific time, with some additional agent specific features.
Bug reports welcome! If I did a good job on architecture, you should be able to have your blackboard up tonight.
There really aren't good messaging libraries that provide what I wanted, so I built it. Basically I started with "signal but no need to have a phone number." I've used it to build a group messaging iOS app for friends, and just pass it to an agent all the time if I want them to be able to direct message.
The group API approval is modeled off of multi signature approval mechanics from Ethereum - so, you could use it like you describe: "Only allow this call if it passes safety checks from n reviewers", or you could have human in the loop, or a program that checks business rules + an agent and a human, etc. etc. I just wanted something that let us control API calls properly.
I have a very vibe coded skill I use here: https://github.com/mkly/dev-skills/tree/main/dev-board
I hope you asked them beforehand :-)
I was rather commenting on the tone that Thoughtworks decided to "take 10 engineers and put them in one room". I took that wording literally, imagining how they grabbed these 10 engineers with a big hand and dropped them off where they wanted them to develop their airline IROps system.
At the same time people are building shared knowledge bases for agents left and right, like Trello alternatives and Wikis and whatnot.
That top engineers working on advanced problems together with agents without really understanding how they work, is exactly how the world is going to end :D
The future may be AI structured as a corporation, rather than AI as a human competitor.
The boulder, it even found a path down the mountain. on its own, by doing so-amazing path finding its own super amazing path finding algorithm, and all the boulders co-ordinated falling down, ON THEIR OWN! And they all reached the ground!
I tell you! The future may be these boulders washing your underpants and putting you to sleep. It is clear!
"My goal is a very simple to use tool that drops straight into your project and immediately offers a communication channel for agents to coordinate work. The first step is to get Talwrn to a point where it can support its own development. I’m planning to post about it regularly as I’m hoping to use it as a single, evolving example of how pure agentic engineering can proceed."
Term: "Blackboard" https://en.wikipedia.org/wiki/Blackboard_system