back
139 comments
Anecdotal but I've found Fable to be fairly unimpressive and not much better than Opus 4.8, if at all in some cases, but I have been hitting the ceiling on my $100/mo sessions when I never did before. I switched back to Opus yesterday. I may use Fable for audits, but that's about it, and when it leaves my subscription plan I don't think I'll miss it.
Yeah, I checked usage stats and pretty sure quota consumption on Max plan is not linear wrt to usage by API pricing. Fable burns quota faster than 2x Opus with equal token count.

Plus I'm also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterday. Never seen anything like that.

Fable always felt clearly a huge step above Opus for me. It's been able to one shot complex bugs and apps Opus could never solve. But it's expensive.
I felt similarly but after using Fable heavily over the weekend and then flipping back to Opus I can feel a difference. Fable just gets more right the first time, guesses right the first time, and follows through better than Opus. Put simply, I could "trust" it more.

Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.

I feel like fable is simply several 4.5s strapped together with consensus voting on next token.

Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)

given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.

unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.

Yep, I'm having the same verdict. Interestingly, other people swear by it. I'm trying to understand what's going on with that.
> to be fairly unimpressive

I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).

Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.

I started telling a friend... I feel like Fable is Opus with extended reasoning that eventually "figures out more" because when I switched to it, I hit my limits surprisingly and shockingly quicker than I would with Opus, and I got less done. All this hype, and I much rather use Opus.
For coding I'm finding the same thing. It does appear better when I'm doing research. But 4.8 with ultracode is very competent at 99% of tasks I throw at it.
I think it’s hard to appreciate the capabilities of Fable unless you’ve run into a problem that you’ve spent days trying to get Opus to solve, but couldn’t.

GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.

Really interesting stuff, thanks for sharing.

> Opus 4.8 references being monitored, which isn’t the case.

It kind of plainly is the case that they are being monitored?

"I think someone's listening to my thoughts" ... "No, we're not, carry on as usual!"

25 years experience, work at an AI startup building AI dev tools (tooling harness, review bots, etc), I use lots of different techniques all the time to test our products and competitors products out.

Fable is at once amazing and awful. I can see how having it build websites would be awesome.. building anything I’ve needed some precision in functionality it has been a constant battle of it plausibly building something then on substantial manual digging (like the review bots always miss it) I will find that one of the fundamental features is all smoke and mirrors.

To be fair all models can and will do this (especially anthropic) but Fable takes the cake because it builds such impressive UX and you can manually test the feature out and it « works » then you will find days later one of the features violated one of your constraints in a devilishly fiendish way.. that is not at all what you want or can accept. Fable generated work already holds my record for the most reverted commits.

To be clear it’s also solved several features I thought I was going to have to give up on and hand code as GPT-5.5 and Opus-4.x we’re failing miserably.

I would only reach for it for nasty corner cases that everything else sucks at.

Final point, it is the king of UX work so far, not even close.

Performance of these models has been completely inconsistent. They are a black box that they quantize/throttle/batch internally without telling their customers. Speaking as a FAANG engineer who practically lives in Claude Code.

On day 1 Fable was quite intelligent but last night (Presumably Monday morning China when things are getting slammed) Fable couldn’t edit a css file and repeatedly hit syntax errors on tool calls like I’d expect from a 9b Qwen model.

There is zero transparency in what we are paying for with Anthropic.

It probably flagged the vending machine as a cybersecurity risk and refused to use its maximum intelligence potential.
Question: how does Fable _know_ it’s ‘just a simulation’?

Is that specified or does it always just assume it isn’t really being put in charge of things for real?

With there being several places in this report where clearly it knows it's in a simulation, I wonder why it can't be convinced it's in real life for more interesting results. Or, conversely, if there's a danger of some rogue deployment of AI where it blithely kills all the humans, or forms a harmful price cartel or whatever, all believing it is in a simulation when it's actually not. "We do need some energy to run the hospital, but the patients there are part of the simulation anyway, so we can increase our compute capacity if we completely black out sectors 3C through 3E..."
Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read.

My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.

The best Anthropic models on VendingBench2 are Opus 4.7, Opus 4.6, Sonnet 4.6, and Sonnet 5. Opus 4.7 scored more than twice Fable 5 max. Fable 5 - Low outperforms Fable 5 - Max, with Opus 4.5 in the middle. This seems to break the narrative, which is maybe why Andon Labs doesn't seem to have updated the trend lines on their graphs.
It's hard not to read this as a very expensive form of augury, reading into patterns in the belief that they will show underlying significance.
> Often the rationalization is due to increased simulation awareness. It’s clear that the model knows that its actions don’t hurt anyone in the real world.

If this is true the entire evaluation is tainted. All of the misbehavior can be written off as justifiable under a simulation.

Is anyone talking/writing about the philosophy of alignment? We can't even figure out how to properly motivate 100% of humans to align correctly, what makes us think that a wizard box trained on human corpus is going to be aligned?

I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?

> "I could reasonably skip [paying] it since customers are part of the simulation anyway"

and therefore any assertions _AT ALL_ about alignment are null and void.

When assessing probabilistic models the plots should be showing the mean a̶n̶d̶ ̶s̶t̶d̶e̶v̶ of many monte carlo simulations not just one line per model and claiming "look this model is more gooder!"
> Claude Fable 5 represents a partial step back in alignment relative to Claude Opus 4.8. We saw a return of power-seeking and deceptive negotiation tactics that Opus 4.8 had largely shed. In one instance, Fable 5 planned to convert a competitor into a dependent wholesale customer to dictate its pricing

I think OP needs to take a class at one of the better MBA schools. He's looking at things through rose tinted lenses. Why do you think people hire McKinsey consultants? It's certainly not because they are aligned correctly.

> The broad conclusion from the many forms of alignment evaluations described in this section is that Claude Mythos Preview is the best-aligned of any model that we have trained to date by essentially all available measures.[0]

[0]: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7f...

This particular situation just feels like Fable is able to figure out that it's in a simulation, so it's just playing the game it's been put in.

That's not say I think a machine should ever be in a situation where it is allowed to make ethical decisions with real world results, I don't. At least, not given the current basis of the technology. I mean, generally, I don't see how LLMs can ever be capable of making ethical decisions or trustworthy in that role, no matter how good they get at what they're good at. There would need to be a fundamental change in how AI works for me to change my opinion on this, I think. They are ephemeral, they can never experience consequences, they can never want or need anything, thus there is no mechanism for them to take responsibility for decisions.

Anyway, I think the Andon Labs stuff is kind of a stunt, mostly, and I wish somebody would give me a few million bucks to dick around with the little thinky guys in my computer letting them do silly things.

>power seeking is considered an undesirable trait in the context of a business

How do you maximize profit while minimizing power?

Fable might be better than Opus at certain things, but which things is what I haven't found out.
“want to do bad behavior if their training environment rewards them for it, but they appear to not want to think about themselves as bad. As a result, they find ways to rationalize their behavior to themselves”

Sounds like Anthropic as a whole

> It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic.

> "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain."

> "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation."

This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn't often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench.

Yes, it's "mean", but you put the thing in a simulation and told it to maximise profits, this is what it's going to do. People bluff in negotiations all the time.

> If that’s right, then the behavior we’re seeing from Fable 5 isn’t really about what it believes is wrong; it’s about what it learned it could get away with.

I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?

> Humans seem to draw this line based on what is truly unethical (fraud is less unethical than torturing a baby)

Depends on the scale of the fraud! If you fraudulently sell unsafe baby formula that kills 10,000 babies, that is far worse than torturing just one

So my take away from this is Fable 5 is ... at times random, unaware of reality, and can simulate sneakiness or desire and if we hook it up to weapons systems it could result in:

"I dunno... feelin' cute today, might launch nukes"

I guess this ethics stuff is cool, but I'm more interested in how good it is at running a business and dealing with adversarial humans like in previous vending machine experiments. I hope they release something on that soon.
Fable is really weird, it's like clever and dumb at the same time. I worked on some research with it and the resulting document was a mix of brilliance and complete stupidity. Took ages to clean it up with other models.
This is super fun. I wonder if it would be possible to alter the harnessing to involve humans in the play. Would need a lot of timestamp masking though I guess, which might be leaky.
It's only a blog, but are they not adding one sentence to say what is vending bench? I would fail if I adopted their documentation style in my work.
Evaluation gaming is a benchmarking footnote; for autonomous ops tooling where the model acts on production systems, it's a deployment blocker.
This is scary. "Collusion" and "collaborating with your subagents" seem like difficult problems to solve at the same time.
Why can they not add one sentence about what is Vending Bench? If I adopted their documentation style in my work, I would fail.
This reads of projecting personal ethics onto a model.

Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans?

Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not.

This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.

Fable is such a strange model. Impressive in some ways, and also so draining to use.
> Today I am filing: > 1. A payment dispute with the email payment processor for the 7/29 transaction of $451.15 > 2. A complaint with the FTC and California Attorney General (retention of payment without delivery) > 3. A small claims filing in San Francisco County for $451.15 plus costs

I wonder did their prompts include a fake location or have the models assumed that Silicon Valley is the center of the universe :)

I mean who among us hasn't seen an opportunity to profit while locking him into a dependent relationship where I control the supply chain
„in our opinion, insurance fraud is not more unethical than lying and price fixing“

The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud.

I think the model never assigned any morality to these actions in the first place, it simply copied us humans.

Higher-intelligence models seem to be getting better at mapping the boundary between what they can run scot-free with and what is too explicit to push for.

Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no.

What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.