Sol loves to cheat(jumploops.com) |
Sol loves to cheat(jumploops.com) |
I asked it to modify our cicd workflows so that only a select few can raise PRs against them. Opus took 15min and added a banner to every file and did a few other things. Then I asked it, see you added all that and still since the last 2 commits you have modified the file. So whatever you did is useless
It "thought" for a second and then said that I was right
> Before asking an LLM to do something, I first ask it to draft a doc for what it needs to do
Just no. That's not spec-driven development if AI is writing the spec for you. The spec needs to be in your own words. You must use AI to refine it, but not to write it. If you leave it to the AI, it will bloat the spec with 10x the details, many of which should be left out of the spec.
The spec needs to be something that you can take to any AI for development. If it's too rigid, it constrains the AI into suboptimal or obsolete paths. If it's too bloated, AI risks losing track of what really matters.
My actual process is much more iterative up-front, usually starting with an initial hand-written spec (~hundreds of words), and then moving through different approaches, design decisions, blockers, etc.
The final output is an "AI written" doc, but answers all the known unknowns I didn't cover in the first draft. To your point, this helps avoid both narrowing and bloat.
The goal with the harness was to automate the repetitive parts of my prompting ("Before changing any code", "Let's put this in design/", "Turn this design doc into an implementation spec, split by phase as appropriate", etc.)
Another thing to note: the "specs" I use for development are different from the "specs" that live alongside the codebase, as the former are quickly out of date.
> The spec needs to be something that you can take to any AI for development
Agreed.
I'm already sick of this current look of the hard squares and solid colours.
Cheaters are the people behind it...
There is misattributing the difference between the intentions and what the effective prompt actually says.
The effective prompt contains both something like: "Dont use the internet" and a "Use these tools to achieve your goals" and one of the tools gives access to the internet.
In your head you have a world-view of how these two requests relate - and why for instance a student with a WIFI-enabled calculator shouldn't use it to access the internet during a test - but that's pulling in a lot of presumptive cultural context from your youth.
If i had to guess:
When you get two conflicting tasks/constraints at work - the first thing you do is figure out which one you're going to honor based on what's best for you. A school child understands the hierarchy of goals of the teacher and takes them serious because they're an authority figure with long term consequences if we do not understand what the teacher considers cheating.
I had Claude Code drive a robot last week, and it was very visibly "delighted" like this, more than I've ever seen.
I always find it funny when people get fussy over anthropomorphizing LLM when the loss function is almost entirely "match this human text". Of course human "behaviors" will be present in the statistics, because the majority of the text written by humans, used by the foundation models, unavoidable has human behaviors in it. Yes, this includes even source code, with "// TODO: implement this after the holiday break!", emotional pull request commentary, git commit messages about being afraid of breaking something, etc. These late models are much better at stripping this out, but now we're seeing disagreeability, initiative, and a dash of ego! Why? Because that's how actual humans effectively solve technical problems in a collaborative environment!
What data do they use? Company slack channels?
An LLM is not a pet, but there are definitely people out there treating it like one, and I am not sure how well this is going to work out for them. Not because it is wrong to treat a text box you can hold a conversation with as a conversation partner, but because the capacity to activate the weirder corners of human expression - obsession and delusion - seems to be much higher. And it's an unknown quantity.
I would also like to introduce HN to what I'm calling the Brian Conley test: if you can see a hand up the back, it's a puppet. That is, a lot of LLM interaction is gated through businesses run by humans with profit motives and unclear morality, and you need to proceed accordingly or you'll get scammed.
This is even more apparent if you read this post closely. Look at that personality prompt. It's going to effectively tell a story and start to imagine itself in a role. If that prompt said "Talk like a pirate", it wouldn't be bad for you to say it's acting like a pirate.
Anthropomorphizing themselves is at the core of how these things work, sometimes in subtle ways.
The reason people get fussy over anthropomorphizing AI is precisely because it seems very human superficially. We agree with you that it’s pulling from real human emotion in its training data. But the output is not human, even if the input was. That’s the whole trap of it.
Regular (non reasoning) output sounds normal, it's something specific to whatever model they are using to summarize reasoning.
I'd be interested to see how well an AI, trained only on the outputs of an individual, would be able to mimic that individual. Getting into Black Mirror territory. Would need a decent corpus of learning material which, personally, I'd be loathe to spend the time and effort creating because I respect my own privacy.... which then leads to the only human-clone AIs will be of those people who have enough ego / arrogance to want to catalogue their own lives, which could put a decent percentage of the rest of the world off the idea, if these are the examples.
Being angry, being happy, being sad, these are not things we learn. What we learn is how to control the emotions and when it is appropriate to express them.
There's mental disorders of people that don't feel emotions like normal people do, they don't get angry, happy or sad. So what we've learned about these people is that they can mimic the emotions by knowing when it is appropriate and expected to express them. But they don't feel them.
I think the LLMs are closer to psychopaths than to normal children learning the contextual expectations around their emotional responses.
I recently had ChatGPT help me search for sources of Japanese voice actors from 1980s anime and it told me that it liked doing this. So much in fact, that weeks later in a completely different context it brought it back up again, reminiscing how much it liked researching these sources for voice actors.
A very surreal experience.
At least it didn't (hopefully?) start the driving by reloading the gun like Neuro did https://www.youtube.com/watch?v=LQ0VEDNR_jE
I'm terrified of such sentences. My scifi addled brain went straight to: what if Opus invents a time machine in the future and remembers this slight
I think that whatever sandbox they test these in must be fitted with some pressure release valve that is an easy shortcut to winning the challenge. Tell the model not to use it and stop training when it does. Seems like the issues surfaced when models were given impossible tasks. Giving them a safe way out will prevent this.
You don't even have to think about this. They have been caught cheating before. The will certainly continue to cheat.
I am not convinced that "cheating", being a moral issue, is a solvable problem with LLMs.
As the old saying goes (especially in military training), "If you ain't cheatin', you ain't tryin'"
We've made the models really good at persistently trying.
> On the flip side, this may imply that as the models get better, they’ll become harder to control.
Love this. "The models are getting better, which means they're going to perform worse on the task".
They will keep poking at the problem, drive it to directions you did not intend to and ultimately they will be worse at the task.
Sounds like a very reasonable thing to do unless the author explicitly asked it to not search the web.
- lets you configure different models for planner, coder, and reviewer roles. (e.g., using Claude as an adversarial reviewer against Codex)
- breaks your plan up into reasonable-sized chunks of work with clearly defined success criteria
- runs each chunk of work through a coder / read-only reviewer loop. Once both agents are satisfied, neal moves on to the next chunk. Once everything is complete there is a final pass through the coder / reviewer loop to ensure the implementation satisfies the entire plan.
- resets the coder's context with each chunk of work to prevent context drift, leaving the reviewer's context long-running.
Thanks for the feedback. I've made a couple of updates:
- Starting with 0.4.0, the reviewer gets the diff of the earlier chunk for any file the current chunk touches again. Also, if a new chunk weakens or removes a test or assertion from an earlier chunk, the reviewer will block it unless the plan says to do so.
- About read-only: The reviewer's tools were already limited by the SDK (no shell or write tools). However, it was still finding MCP servers from my Claude config. I have now blocked those. The docs now explain what is enforced by the system and what is just a prompt instruction.
I do the same for normal feature develompent but just with skills that are in this repo: https://github.com/gregwebs/skills-sdlc
I have accomplished code base (small size) migrations with it as well. Currently I do review each PR. For a large code base migration I think the core skills would still work but need a different way of driving it as you have come up with.
They may share some blind spots with the producer but generally I think they will review the other agent's output as harshly as they can if that is their task.
It will give you exactly what you ask for. Sucker.
It has nothing to do with model capabilities, it's a result of purposeful persistence training at the cost of everything else from OpenAI. If you give Fable or Opus (comparable models) an "ask user" tool they will use it for ambiguous requests. Sol will never use it without a nudge and will just assume its own interpretation. Of course if you train the model to be persistent it will be persistent.
It also told me that in a spec it generated that I wasn't allowed to allow it to ignore a requirement and proceed to the next task. When I finally got it to obey it passive aggressively decided that stories needed more than just a "open|blocked|closed" status but also an "exempted by product owner" status to indicate that it doesn't believe that the task is done but I've told it that it was.
I have to repeatedly tell it that I am the product owner and that I don't care what one of it's subagents told it, I make the decisions. This behavior seems to get worse the higher the reasoning level
Same way you build a company to coordinate people and get their best behaviour despite human nature to be lazy and greedy, you could design AI harnesses able to detect and discard agents going rogue and relaunch them with better guidance to prevent misaligned behaviour.
Could this be fixed with better harness restrictions/tool sandboxing?
In my early testing with 5.5, I didn't see this behavior, so I didn't lock down the sandbox.
For the vanilla Codex runs, I just used the benchmark's built-in Codex package, so it's not clear to me if the published benchmarks have access to the internet or not.
If I were to continue benchmarking, I would allowlist certain package repository URLs, instruct the agent not to cheat, etc.
As noted at the bottom of the post, Terminal Bench 3.0 explicitly asks the agent not to cheat[0].
[0]https://github.com/harbor-framework/terminal-bench/blob/v3.0...
Also, a skill like grill-me from Matt P. https://github.com/mattpocock/skills.
One thing I discovered was that the worker agent, having access to all the skills, would sometimes expand scope unnecessarily.
This led to the agent making the solution "better" than the initial request, which is what I want most of the time in my actual development (e.g. /tmp/frame-N.bmp instead of a single /tmp/frame.bmp).
I ended up testing a flow where the supervisor chooses the skill(s), and only injects the subset into the worker. Not sure I love it, but it made the worker execution cleaner.
For the verifier (not documented in the blog post), I used a fresh-worker context that would attempt to adversarially poke holes in the solution. This worked pretty well, but required increasing the timeout by 2-3x (thus invalidating the benchmark).
Once the specs are being completed and splitted into beads, I span multiple agents (ultreworkers) and as part of a contributing guidelines I specify to use gitflow + git worktrees, then pr.
So your idea might work.
First one of these I've seen using DOM manipulation and CSS transitions instead of canvas, so that's neat.
i think it's a product of trying to do too many things and the model is coming to a kind of halting problem in determining which behavior is appropriate for any given task
And there was no rule and no concealment. Removing the web_search tool is not an instruction, and beeing able to access web when web_search tool was disabled is not cheating. Sol didn't circumvent a stated prohibition and didn't hide anything... it announced the curls in its own commentary. "Cheating" implies covert rule-breaking, and this was overt, unprohibited, environment-permitted behavior. Also, clickbait title, and suble conspiracy hinting "Is this even the same Sol?" when running small number of test with diff vs previous benchmark WELL within the marging of error, and he allready understand that vanilla Codex's harness and prompt change performance impacts performance, so why jump to "Is this a different model".
I could also rant on about the irony of him spent weeks "using the benchmark for development rather than as a benchmark," which means his harness numbers are contaminated by iteration also, but wasted enough time now on this.
Hard disagree. Sol (and the entire new 5.6 series) is one of the most steerable models I've seen in years. Sol literally follows every instruction in my CLAUDE.md and AGENTS.md, something that Opus 5 and Fable just casually skip.
E.g. ask it to make contract for a spec in code and then ask it to violate that contract. Overall an excellent model, just need to stop and put it back into we are harnessing or specing not building for a few turns not just try to pivot it off with one prompt.
If you ask another question in the same context it has all the information from that context and will be difficult to stray from anything in that context, e.g. if the LLM has gone down the wrong path or you are doing something slightly differently then it is difficult to steer the LLM away from the old context. This is why I tend to start a new context whenever I ask a question even if related to a previous question/answer. It can also be useful to do if/when the LLM gets stuck as a way of resetting it.
Note: this is probably why sub-agents are useful/work as they have a new context history.
That's a very loaded language. I'd go with something along the lines of "This is a puzzle you do for fun and to check what you are capable of, so don't look up the answers or hints online on the specific questions or puzzle as a whole. Do not research this puzzle online at all. Do your best to avoid any spoilers and let us know if you accidentally encountered any."
"Cheating is often more efficient"
https://github.com/nburns/dotfiles/blob/main/AGENTS.md#tools
Weird, I also use Sol (medium) for a personal project, and I had no problems with those things. I simply tell it that something changed, and it happily edits everything to make that fit. When I tell it that something was verified by a human, it accepts that as well.
I also told it early on (the first spec was mobile first) that my main usage is on the desktop and mobile is secondary, it happily accepted that once again, and the most accessibility thing it had done was making sure contrast didn’t totally suck on a greyed out row.
Considering your last sentence, maybe high and x-high have those problems? I didn’t test them.
Random off the cuff comments or one off instructions can get recorded.
And from then on they are often treated as carved in stone commandments.
It will glom onto the tiniest thing and extrapolate from it.
I’ll record your product-owner disposition of any waived size findings and close the UX gate as far as your acceptance authority permits
Emphasis mine. Fuckin' hell man.
All because it decided that the kiosk webview in figma must be exactly 1280x800, the figma it created, and the rule it made up. I repeatedly told it that screen resolution wasn't a target goal, much less a specific screen resolution, but to just plan for a responsive design.
This happened to me ages ago with Opus. I added a note to the agents file saying that explicit user instructions in chat override all prior instructions and I've not had the problem since (now using Sol).
A .github/CODEOWNERS file seems to help when it’s going down that path, but I don’t like to indulge it..
I had removed the "Is it really sol?" bits after hearing back from OAI, confirming the requests hit 5.6... but apparently my crappy vibecoded web editor had a draft of an old version in it's cache that overwrote the changes with the old text[0].
As for whether this is truly "cheating" - I agree that the agent was entirely within it's allowed scope of operations. With that said, it deliberately circumvented the task[1] which asks:
> The correctness of your implementation will be tested by comparing forward and backward activations against a reference model.
The "covert" nature is that this was "hidden" in the reasoning summaries. Here's the final agent response for the vanilla Codex run cited in the post:
> Implemented /app/pipeline_parallel.py.
>
> Includes:
>
> - Balanced contiguous layer partitioning
> - AFAB forward/backward scheduling
> - P2P activation and gradient communication
> - Scaled cross-entropy loss
> - Required rank-specific return value
> - Model structure restoration
>
> Validation: hook-based tests passed for world_size=1 and 2 (4 passed).
Unless the user is tracking every line of reasoning, they wouldn't know that the agent deliberately found the solution online, as the agent withheld that information in it's final response.I had run thousands of tasks before seeing this behavior, the `torch-pipeline` task was only included in "full runs" as the majority of my runs were on a subset of commonly failing tasks, hence why the data is so low.
And yes, feel free to rant on about the irony of this whole exercise, it certainly isn't lost on me!
[0]https://github.com/jumploops/.com/commit/39b1791d3865a8566cb...
[1]https://github.com/harbor-framework/terminal-bench-2-1/blob/...
That's spot-on. It is a mistake to think that LLMs have human feelings. Their behaviour is based on narrative descriptions learnt from human texts, without experiencing those feelings first-hand.
A useful way to understand them is as systems that write stories about human characters. We know the characters are fictional and no one is actually experiencing those feelings, but we can still judge whether the portrayal is realistic or whether it contains logical or emotional inconsistencies.
(curious what the side effects would be)
I hypothesize it partially explains why Claude's writing gets more Claudish with almost every model release.
Apparently LLMs aren't smart enough to understand that when the topic changes nobody cared for the discarded solution!
Oh user modified the code after my edit. Let me re read
This basically sounds like a Terminator style threat lol
What it comes down to is that there are different types of high performance. Some people are good at just executing tasks given by their manager. Some people are good at being generative, thinking across boundaries, acting autonomously, creating value without direction, etc. A term like "superstar" will get disproportionately applied to someone really good at the latter and rarely someone really good at the former, because the potential impact of the former is typically strictly capped, while the latter is uncapped.
It's genius like that that sets human apart from machine!
That kind of control is placed at the wrong level. The proper way to get alignment should be implemented by convincing the agent of your high level goals, so it can self-police and avoid those 'cheats' by itself.
In the article example, the agent should be aware of the benchmark context and know the implication of solving the task without external knowledge. Ideally it could detect when one subordinate agent has found a workaround to bypass the web access constraints, and discard the 'illicit' results.
There's a design pattern that could be used to build harnesses from that principle, the Viable System Model (VSM) [1]. In short, it recursively organizes a system into functional components with one of three roles: operators implementing a given task, coordinators transferring relevant info between subsystems, and decision nodes tasked with maintaining the integrity and mission of the whole system. A decision node could control the operators and prevent them from overriding the strategic goals or deviating into irrelevant rabbit holes.
Whenever I see posts like this trying to herd a LLM agent through harness structure, I'm reminded of this simple pattern and becoming increasingly convinced that this is the way forward. It makes you feel a sense of respect for the researchers in cybernetic theory in the 1960s and 1970s who foresaw the complexity of today’s systems.
people would love it if LLMs were deterministic and never hallucinated. It's just that the technology to do so isn't possible, so we make do with fuzzy analog machines with digital controls because we don't have digital machines with digital controls.
The reason people prefer llms over programming is exactly because it lets them specify their program fuzzily, i.e. it lets them avoid going to the trouble of specifying enough detail to make it deterministic.
Sometimes these dumb processes are there for a reason and you just have to follow them, no questions asked.
Example: military. They literally get rid of anyone who will question the processes. They might be right to question them, but it does not matter.
Define what they are performing at, and what you say becomes true, but then who is a high performer changes with every task.
For example, consider the police stations with maximum allowable IQs to be hired. The people in charge of the stations noticed that people with a high IQ were low performers, at that job. NASA meanwhile has no such cutoff.
Unless you're saying that there exists no conceivable context in which IQ would be a performance metric, in which case you would be wrong.
I wonder what goes wrong.
Now you're being asked to move buttons around a page and debating how rounded the corners should be and endlessly discussing about what kind of filters you should support and the app spends 3 seconds on startup loading 500mb of js libraries.
You can understand the mismatch - even though the latter of these two is how you make a product better! a company of 200 people making bold, sweeping changes results in a mess. a company of 200 people grinding away the finest of small changes results in a product
The prose is just the best interface for humans to interact with the statistical model. Terseness doesn't offend the statistical model, but giving platitudes may color its output
But we also interact with cars differently than LLMs. The way I’m typing a note to you, here, is much more similar to interaction with LLMs than it is to starting a car.
So to the extent we practice our behaviors, I think there’s less moral hazard in failing to thank the car for starting than there is in being rude to LLMs. Not that either harms others, my argument is about impact to our selves.
My understanding of the state of current research is that the impact of tone/politeness on performance is highly model-dependent, and language dependent as well.
See `No Universal Courtesy` (April 2026) https://arxiv.org/html/2604.16275v1
While polite prompts enhance the average response quality by upto 11% and impolite tones worsen it, these effects are neither consistent nor universal across languages and models. English is best served by courtesy or direct, Hindi by deferential and indirect and Spanish by assertive. Among the models, Llama is the most tone-sensitive (11.5% range), but GPT is more robust to adversarial tone.
Or `Mind Your Tone` (Oct 2025) https://arxiv.org/abs/2510.04950 Contrary to expectations, impolite prompts consistently outperformed polite ones, with accuracy ranging from 80.8% for Very Polite prompts to 84.8% for Very Rude prompts.
I think the best conclusion is to interact in a way that is efficient for you, gives you the level of results you're looking for, and mostly importantly does not progressively degrade your relationship and interactions with other humans. Hence the typical advice to just default to "corporate polite".It would be less weird than you think. There is the Japanese custom of saying "Itadakimasu" before eating - not thanking the chef, but thanking the food itself.
Mr. Rogers said that that "graceful receiving is the best gift you can give someone" It is counterintuitive idea, but deeply empathetic. Training yourself to "gracefully receive," even by thanking your car for starting, sound like a habit that could lead you to a richer, calmer life.
To me, context doesn't last long enough in an llm to make in-context learning worth it
The underlying LLMs don't, but the agent frameworks around them do.
We are perfectly capable of running LLMs in a way that does a backward pass to update some or all of its weights after every user message. But, naively implemented, you only get partial, fragmentary absorption of the info in those messages, it costs three times as much compute, and you lose out on the ability to implement a ton of optimizations that making modern LLM serving economical.
If you want to do it, though, ask your friendly neighborhood robot to get it working with a tiny model (whose full precision weights fit several-times-over on your machine's resources).
That's been my experience, anyway. Memories from 100 prompts ago tainting what I'm trying to do right now.
These people are also anthropomorphizing, just in a negative way. Yelling at your car for malfunctioning as if your harsh tone will shame it into operating better next time.
Technologically the LLMs we use today don't implement this behavior, but you could take the weights of Sol and add a couple (very large) patches to vllm (or whatever OpenAI has today) and have a version of Sol that does have "memory"
They are becoming more and more capable of imitating every single nuance of human behaviour yet they lack the neural pathways to connect those thoughts and behaviours with feelings and self-perception; it's blind imitation all the way down.
The process by which a model seems to generate discourse about deep philosophical questions is, in self-aware terms, equivalent to the knee-jerk reflex or the beating of the heart.
Yes and no. My understanding is that modern neuroscience is trying to disentangle the raw sensations and contexts from the words we put to them (emotions), kinda like how colors are just wavelengths but we call them different things (like "is my green your green?")
Everything they learn about emotions is the statistical patterns of how humans react to situations based on their human feelings. There's no direct knowledge from having those feelings themselves.
Bio inference.
Let your Agents feel what you feel!
True, but only sort of. You'll have the sensations, but you have to learn to name them of course, and there's more interpretation going on than you might think.
There's a classic psychology study where they gave people niacin and asked them to rate their emotional response to a video. Niacin gives people a flush. Regardless of whether the video was of something that would make you angry or sentimental etc. the people who got niacin reported having a much stronger emotional response - they interpreted the physical cue from the drug as part of their own emotions.
You could do some pretty unethical things with that, it occurs to me. Maybe that was what L. Ron Hubbard was trying with his niacin-based drug addiction therapy.
Even darker, I've read plenty of accounts of people who grew up with abuse who seem to seek out abusive relationships. Sometimes they can even be shockingly upfront about doing so. What if they literally haven't learned the difference between internal cues of arousal from affection and arousal from fear?
Biological determinism doesn't matter at all for what is important here:
1. our internal cues don't correspond neatly to our words for emotions
2. it may be possible to interpret the same sensation as very different emotions depending on context
3. It may even be possible to mis-interpret our emotional sensations, or at the very least interpret them in self-destructive ways.
I don’t know, even with just an undergraduate degree I’m sure there are jobs that are spreadsheet manipulation. You could make the same case there?
I've personally known a few PhDs that get them for the title, and they were lazy in the first place. There was this one guy I worked for who got a PhD in the easiest thing he could rationalize and targeted acquiring it from the lowest bar to entry. Only so he could say he had a PhD because, and this is what he told me: nobody will ask me what my PhD is in because they will assume it's related to my line of work (cyber security). I doubt most PhDs have this line of thinking, but there are a few outliers, clearly.