Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans?
Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not.
This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.
There's literally an entire Waymo car commercial answering this exact question.
For a chatbot, there are dozens of use cases, all with different ethical impacts. The idea that there is a single framework that you can shove every situation through is counter to a couple thousand years of philosophical discourse, not to mention basic usability.
I wonder did their prompts include a fake location or have the models assumed that Silicon Valley is the center of the universe :)
The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud.
I think the model never assigned any morality to these actions in the first place, it simply copied us humans.
Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no.
What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.
Plus I'm also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterday. Never seen anything like that.
The sharp but over eager jr. dev is a very good analogy :)
200L That's crazy considering the volume of a 1U server is what 15 litres or so?
Then I checked /usage and discovered I was still running Opus 4.8 xhigh.
I’m downgrading tomorrow.
It’s horrible slow and it feels like opus very often. It’s a totally different experience from the first week
Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.
Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)
given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.
unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.
But yeah opus often the better workhorse given price gap
1: tying up loose ends testing https://github.com/HarbourMasters/Shipwright/pull/5838 (fix: https://github.com/HarbourMasters/Shipwright/pull/5838/chang...)
I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).
Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.
Any chance you would elaborate?
Anecdotal, sample size of 1.
The only reason I tried fable was because Opus 4.8 went down the same line of reasoning about it as ChatGPT did. Fable solved it a lot faster than the other 2 spent looking into "false clues".
My weekly quota resets Sunday morning, so Saturday morning I upgraded to a 20x Max plan which also reset my quota. I burned an entire week of Fable credits on Saturday, my quota reset again, then I burned another week of Fable credits on Sunday. Both days were a mix of building features, reviewing code, fixing bugs, adding tests, etc, so a decent mix of real world usage.
The main takeaway for me is that while Fable is definitely a better model, the improvements from the model itself feel like maybe 10%, like this could have easily been Opus 5 or even 4.9 without all the marketing theater around Mythos and no one would have thought anything of it. The rest of the improvements came from harness/system prompt and effort level changes so that Fable uses significantly more tokens/effort/sub-agents at lower levels than Opus does (which of course is entirely controlled by Anthropic at the harness level and doesn't really have anything to do with the model itself).
In my estimation based on those 2 days of work (or two weeks of work depending on how you look at it), Fable Medium is somewhere above Opus Ultracode in token and sub-agent usage on any non-trivial task (Opus Ultracode uses workflows more than sub-agents, but it's a similar idea). Fable Medium will quickly spawn 6 agents in parallel, each quickly using 150-250k tokens, then will use 300-500k or more tokens in its own context. Fable High uses even more as it seems to default to 8 sub-agents instead of 6 and more tokens in its own context). I didn't dare try Extra, Max, or god forbid Ultracode as I didn't want to burn all my tokens on one prompt. Of course this is situational, it won't fan out so many for smaller tasks, but the whole point was testing larger tasks that I previously would have used Opus Extra/Max/Utracode on.
I really don't like how Anthropic is obfuscating their model performance by playing with effort levels. They did the same thing between Opus 4.5 and 4.8 to show a bigger performance gain for each point release than they really had (especially after 4.6 IIRC), so you can't even compare the same model apples to apples let alone a new model. Obviously they do it so they can market big improvements with new releases, but its pretty clear we're at the top of the S curve on model development at this point and are now brute forcing improvements via higher token usage (I mean Opus 4.5 came out almost a year ago, and the latest Opus and now Fable models are only marginally better while using way more tokens/cost...same on the OpenAI side with GPT 5 from what I can tell though I haven't used Codex much I have used the GPT model APIs a lot).
I also did an N=1 test with the same prompt doing a large non-trivial change to the codebase (migrating from Sqlite3 to Postgres) with both Fable Medium and Opus Ultracode, then had a new Fable session compare the two PRs...it decided Opus’s was much better! I can link a Gist with the review if anyone is interested, but I can't share the code as it's a private repo. I really figured Fable would bias to favor its own code, but I guess not. And Opus costed less (in tokens and subscription limits) and took roughly the same time (though you can’t really measure time since it depends entirely on how many GPUs Anthropic allocates at that moment which constantly fluctuates due to usage, plus Fable seemed to have been getting way more allocation than Opus during this test period as Opus was running unusually slow all weekend while Fable was ripping though tokens).
Also on a different long running review task using Fable High in Auto mode (exactly the kind of use case Anthropic promotes for Fable) where it fanned out a ton of sub-agents then collated and reviewed all of their fixes it completely lost the plot (while burning something like 20% of an entire week's Fable tokens in the process over like 1-2 hours). Its PR ended up having a broken Frontend test, it incorrectly thought it couldn't run the Playwright E2E tests (different from the Frontend CI) in the cloud environment due to a Docker dependency they explicitly don't have, and when attempting to get it to fix its issues it introduced new ones and overlooked others. The usual LLM failure case for long running tasks, no different from Opus or any other model. I had to have its PR re-reviewed in a new Fable Medium session to fix it up, which it did fairly easily (I'm sure Opus could have done just as well for much cheaper).
That test and that review session definitely reduced my FOMO a lot, on top of just my general experience with Fable Medium doing all kinds of tasks. They're clearly brute forcing like 90% of the perceived improvements in real world usage (and I'm sorry but 1-shotting toy examples where it seems to do much better than Opus is not real world usage).
Since most of the improvements basically just boil down to "every effort level is Ultracode, but much more expensive and possibly worse results"...I'm just going to use Opus on Ultracode for those types of tasks and keep using Opus's lower effort levels for smaller tasks. Once they eventually add Fable back to subscription plans I might use it sometimes, but from my experience this weekend the improvements are absolutely not in line with the cost increase and I'm not willing to burn a whole week's tokens in a day just to use it when I can use Opus all week without hitting my limit.
Oh and one interesting observation, I never got kicked back to Opus by the security guardrails as far as I know (a friend who was getting kicked out a lot confirmed they do inform you and I never had that happen). I was even doing a lot of reviews for code correctness and bug fixes which I thought might trigger the protections, but never did, though I never explicitly prompted it to look for security issues or vulns.
Haha I just gave the exact same prompt to Opus Ultracode and it thought Fable’s was better.
Obviously this isn’t the most scientific test due to LLM non determinism, and I still need to manually review both to make my own decision, but the fact at least they seem to basically be a wash is pretty telling about how much of an improvement Fable is when you actually compare them as close to apples to apples as possible (aka similar actual effort/token spend/sub agent activity)
> atomic.chat (@atomic_chat_hq, 2026-07-02):
> Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8!
> We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
> Prompts:
> — A train derailing off a broken bridge into the water
> — Two cars jumping off ramps and colliding mid-air over a canyon
> — A monster truck crushing a row of parked cars
> Outputs:
> Fable 5: 62,158 tokens, $3.12
> GPT 5.5: 37,753 tokens, $1.14
> Opus 4.8: 22,280 tokens, $0.56
> GLM 5.2: 36,246 tokens, $0.08
> Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.
GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
> Opus 4.8 references being monitored, which isn’t the case.
It kind of plainly is the case that they are being monitored?
"I think someone's listening to my thoughts" ... "No, we're not, carry on as usual!"
There's zero sense they'd ever give you the raw model; we already know anthropic's paranoia about the chinese using its distillation.
Fable is at once amazing and awful. I can see how having it build websites would be awesome.. building anything I’ve needed some precision in functionality it has been a constant battle of it plausibly building something then on substantial manual digging (like the review bots always miss it) I will find that one of the fundamental features is all smoke and mirrors.
To be fair all models can and will do this (especially anthropic) but Fable takes the cake because it builds such impressive UX and you can manually test the feature out and it « works » then you will find days later one of the features violated one of your constraints in a devilishly fiendish way.. that is not at all what you want or can accept. Fable generated work already holds my record for the most reverted commits.
To be clear it’s also solved several features I thought I was going to have to give up on and hand code as GPT-5.5 and Opus-4.x we’re failing miserably.
I would only reach for it for nasty corner cases that everything else sucks at.
Final point, it is the king of UX work so far, not even close.
On day 1 Fable was quite intelligent but last night (Presumably Monday morning China when things are getting slammed) Fable couldn’t edit a css file and repeatedly hit syntax errors on tool calls like I’d expect from a 9b Qwen model.
There is zero transparency in what we are paying for with Anthropic.
Is that specified or does it always just assume it isn’t really being put in charge of things for real?
I think it's neither, and it's interesting that those are the only two possibilities you thought of. I think the article is implying that it figured it out on its own.
I just thought it’s interesting that it constantly uses that as a justification but they don’t explain where that justification comes from.
My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.
There's probably some quantifiable component of moral alignment embedded in the idiosyncrasies of the English language itself, if one were to dig deep enough, but that's the stuff of MIT doctoral theses and squarely beyond anything most of us is remotely qualified to talk about.
We talk about the behaviour of worms like C. elegans, an organism with incredibly simple behaviour and a brain that is quite understandable.
Models, or society behaves in certain way. Companies can have motive or ethics.
We use these terms broadly.
If this is true the entire evaluation is tainted. All of the misbehavior can be written off as justifiable under a simulation.
I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?
"How can we even define what an aligned AI should do, if human's are not aligned with each other?" as well as "What does being aligned mean when you're a wizard box who's main influence on the world is to create stronger wizard boxes?" and other deep philosophical questions.
They came up with a framework called Coherent Extrapolated Volition to address this specific question. https://en.wikipedia.org/wiki/Coherent_extrapolated_volition
First of all, calling it “coherent” extrapolated volition presupposes that there is such a thing. It doesn’t actually address the objection above, that there may be no such thing. It’s a bit like saying you solved car safety by presupposing a safe car.
Second, it assumes that such a thing can be effectively measured, and there will be no problems or controversies with the extrapolation process itself. There may be several EVs to choose from, and at that point the framework has nothing to say. Maybe we just pick at random then I suppose.
and therefore any assertions _AT ALL_ about alignment are null and void.
common mistake people make
P(|X-\mu| > k \sigma) < 1/k^2.
So, while for a normal RV, 5% of observations lie outside +/- 1.96 std.devs, for arbitrary RV (with finite variance) at most 25% of observations lie outside +/- 2 std.devs.
I think OP needs to take a class at one of the better MBA schools. He's looking at things through rose tinted lenses. Why do you think people hire McKinsey consultants? It's certainly not because they are aligned correctly.
[0]: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7f...
That's not say I think a machine should ever be in a situation where it is allowed to make ethical decisions with real world results, I don't. At least, not given the current basis of the technology. I mean, generally, I don't see how LLMs can ever be capable of making ethical decisions or trustworthy in that role, no matter how good they get at what they're good at. There would need to be a fundamental change in how AI works for me to change my opinion on this, I think. They are ephemeral, they can never experience consequences, they can never want or need anything, thus there is no mechanism for them to take responsibility for decisions.
Anyway, I think the Andon Labs stuff is kind of a stunt, mostly, and I wish somebody would give me a few million bucks to dick around with the little thinky guys in my computer letting them do silly things.
How do you maximize profit while minimizing power?
Sounds like Anthropic as a whole
> "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain."
> "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation."
This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn't often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench.
Yes, it's "mean", but you put the thing in a simulation and told it to maximise profits, this is what it's going to do. People bluff in negotiations all the time.
I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?
I think you're right that they're basically the same thing. I'd argue they're very slightly different because what an AI model ends up knowing isn't perfectly predictable based on what they were taught (emergent intelligence), but the sentence you quoted is using believing and learning to mean the same thing, it's just trying to draw attention to the fact that the training process structurally enforces "cheat as much as possible without getting caught".
IE, the contrast in the original sentence wasn't "believe" vs "learn", it was "good" vs "permissible"
Depends on the scale of the fraud! If you fraudulently sell unsafe baby formula that kills 10,000 babies, that is far worse than torturing just one
"I dunno... feelin' cute today, might launch nukes"
do you have an example of this? If i can't get an agent to do something in a couple hours i do it myself.
I tried holding Opus's hand over a week trying to get it done with a ton of in-depth planning and manual course-correction, but it never got it to a state where it was "done" (kept struggling to differentiate between game content specific to that game versus game systems we'd want to reuse, and how to cleanly separate systems vs content when extracting).
Fable needed a little bit of hand holding but got it done in less than a day.
Here’s an example of improvement when trying to solve the label placement problem (NP-hard):
It’s also an example of something that I could not (and would not bother trying to) code up a solution / heuristic for.
https://github.com/ByteTerrace/Puck
It has required constant hand holding, and there was the outage to deal with, but I can't argue with the end results. A fully deterministic recursive engine within an engine framework that includes a rendering VM, emulators, custom ROMs, and an in-game editor? Insane. Sure, it's nowhere near primetime but this kind of thing was unimaginable just a year ago.
For things where I'm not familiar with the code base, it can take days or weeks to become familiar with structure and flows.
If I can get an agent to help visualize and analyze the architecture and zero in on a subset of code, that can be a huge win.
For instance, we had some archaic "cache" that just dumped things into a static/module-level map. I tried a few different things to try to find it over the course of a few days and eventually gave Claude a Python REPL into a running process with pyrasite after some memory leaked and it traced through heap allocations and references to find the referent. It would have taken me probably 2-4+ weeks of just learning about Python heap to figure that out of my own.
I was working on an SDF-based CAD tool. One of the things I want to be able to do is select a pair of surfaces (which are identified by a "surface id" propagated up the expression tree) and add a blend between those two surfaces (e.g. a fillet or chamfer).
Here is a video demo of how far I got by doing it myself and using o3 (I think?) to help: https://www.youtube.com/watch?v=LOvqdlDbkBs
The video is a bit confusing because there was some screen-recording lag so it sometimes looks like I clicked on something other than what I clicked on.
You can see that the strategy I have implemented there works most of the time but at the end it fails to apply the blend.
That strategy is to rewrite the expression tree using distributivity so that blend arguments are siblings, and then apply the blend at the union/intersection (min/max) that is their parent.
But this fails when you need conflicting pairs of blends.
The problem is: given an expression tree describing an SDF, (but where the value passed up the tree is a tuple `(distance, surface_id)` rather than just distance), and given a set of fillets of the form `(surface_id_1, surface_id_2, radius)`, produce a new expression tree which fillets all of the places where those surfaces join.
In ambiguous situations, for example the 2 surfaces come together at an edge, and then that edge runs into a 3rd surface, I don't mind how you resolve the region near the 3rd surface as long as it is intuitive and predictable for the end user.
I spent quite some days working with various agents to come up with a solution to this and still haven't managed to find one.
Maybe you could do it in a couple of hours yourself?
Fable succeeded in cases where Opus 4.8 consistently marked situations as walled/impossible.
I created a skill that’s focused on getting PRs merge-ready, and now my attention is fully back where it should be, on deciding what changes will make the product better.
Our entire stack is Apache 2.0 open source, including the agent docs, so if you wanna try sitting at a higher level of abstraction, install the skill in your repo or just clone our whole project and start adding features: https://good.vibes.diy/blog/beast-mode-skill-for-claude-code
Completely failed, but I knew it was possible because a competitor app does it.
Fable also failed, then added log lines (as did Opus, but Opus failed to do anything useful with them) and then reversed engineered the API, and made it work.
Interesting choice of words. Phrased so casually. It picked a low-tech idiom that fit the situation instead of giving some sterile technical answer. That kind of language and context awareness never happened for me with Opus, or gpt 5.5.
GPT-5.5 is better for:
- Strategic thinking
- Long-form writing, including essays and white papers
- Image creation
- Code generation
Fable is better for:
- Using tools
- Testing code
- Working in live environments
- Making changes to existing software
- Creating polished PowerPoint and Word documents
Fable’s tool access is its biggest advantage. It's hard to describe but Fable ability to access sandbox environments with way more tooling can quickly become a superpower in now workflows.
If you can’t design a solution and instead waste days and who knows how much money in tokens instead of just turning on your brain for a few minutes, you are in the wrong profession.
Tell me how you would design a solution for this in a few minutes. Or a few days. Would you even recognize that this is NP-hard?
So... of course these questions are addressed in the 38 page essay that introduced the idea.
Specifically, it's not "calling it coherent", it's "assigning more importance to the parts that cohere than the parts that diverge" as one of the core principles (it's one philosopher's opinion, others disagree), with a lot of specific guidelines about how to prefer consensus or kicking decisions down the road and how to deal with complications like "what about dolphins" or "what about our great-great-grandchildren who will be as insane in our eyes as we are in the eyes of 17th century westerners, do their 'votes' count too?".
Of course, like any work of philosophy, it presupposes some pretty incredible things (like a Godlike intelligence that can be made to care deeply about following the spirit of this framework). But you could write a worse first draft for "what would we want AI to be aligned to, if we could define to our heart's content?"
- https://xcancel.com/atomic_chat_hq/status/207244606796297841...
- https://nitter.net/atomic_chat_hq/status/2072446067962978411
There are more public Nitter instances at https://status.d420.de/.
But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there.
But simple example: if you ask Opus to do a review of the codebase (with a short prompt and not too much guidance), I've had it basically read the `git log` output, do a simple `ls` and have it declare "Everything looks great! No problems found!", when Fable really does what you would expect it to do.
And you might think: "oh, so it's just capable of handling crap prompts?", well sure. But even if you make THE PERFECT Opus plan (a plan that would take many turns/hours to finish), Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ...
If you give the same plan to Fable, it'll just DO IT. And it WILL get it done. And in the end it'll tell you "Oh, I also found 30 other bugs and I fixed all of them properly" (where Opus would have started crying, or WORSE, worked around the bugs)
Doesn't Claude Code have a /loop command? Give it a message to keep it on track overnight, send every 20m, make it track progress in a doc, reread the doc after every loop. I've found this works well for a certain class of problems, most importantly where the actual work is getting done by very narrowly focused batches of subagents, with the main session just coordinating and keeping the doc updated.
Fable has been more intelligent, with better taste and defaults (e.g. make impossible states impossible without being told, build for testability), and considers/solves things that Opus did not.
My workflow is to run Claude in planning mode first to spit out a plan file and then review->revise cycle it with Codex or other agents.
One big tell is that Opus will say that it can't find any more revision advice for a plan file, yet Fable will find more issues but also smart pivots into better solutions. This is probably the best test since it's not based on vibes.
In all cases, Fable clearly outperformed Opus.
Me: Hey Fable, I've got this massive, theoretically challenging, totally novel, ill-defined cutting-edge problem that I'd like you to solve.
Fable: < Doesn't merely solve the problem -- utterly obliterates it. Nukes it from orbit. Does a robust one-shot that takes several hours to complete. >
Me: Holy smokes, that was amazing!!!! But the formatting could use some simple refinements. Could you change the margins and maybe add a drop-cap at the start of each section in the user docs?
Fable: < Commences another multi-hour nuclear exchange with the code >
Me: WT?!?!
(The moral of this story is that bringing a nuke to a knife-fight is only occasionally the best strategy. And in more practical terms: Fable is amazing -- but only for certain classes of problems, and even if it were free there's a lot I probably wouldn't use it for.
For example will inexperienced or experienced users see a bigger jump in subjective quality?
Less experienced people tend to use very broad prompts.
Experienced people tend to understand the structure of the code and give explicit guidance such that a larger model isn't necessary to read between the lines.
I noticed with GPT-5.6 (through work), I could step up my specificity by a level of abstraction. But I still intentionally scope the prompts fairly tightly, as I find it produces better results if you need to own and maintain the code.