Plus I'm also not super impressed; it somehow managed to implement a 200L custom TCP server for a simple static HTTP mock server for a single test case (all that was needed was a fixed route returning a fixed placeholder string) just yesterday. Never seen anything like that.
Opus is still great but I will be sad when I lose access to Fable on the 7th. In those few days I burned ~$1,400 in API credits (I'm on a subscription but that's the token cost) and while it was great, I can't justify that cost without it be subsidised. Comparatively, the records show I used about $1,200 total in the last month on Opus. I did use it heavily over the last 3 days but 3 vs 30 days and higher burn? Yeah, I can't afford that even if I made really good progress on my projects.
Outputs i've seen so far are on par with my tests for 4.5, where 4.6+ were consistently regressions on 4.5 and their predecessors. One notable improvement being significantly lower retries to good output (1.1 avg. Vs 1.7 prev. On harder tasks)
given all the smoke and mirrors and OAI style fear-hype, it wouldn't surprise me if they intentionally degraded opus 4 for a few iterations, so they can resell "coke classic" at a markup with a minor quality of life feature put in, but charging way more than just re-attempting a poor output would have been previously.
unless anthropic starts acting in the image they claim and starts contributing to research, we'll never know either. Ultimately, the secrecy in how and why things are done would mostly be beneficial to this kind of buisness practice, since as it has always been, the moat is the data not the tech, so I cannot imagine what they hope to gain from the recent uptick in paranoia, jealous guarding and secrecy other than trying to huck a previous peak performance model as an imorovement when really, it is simply coke classic.
I didn't get to use it enough to get impressed or not, because twice today it told me I've hit some flag and it downgraded me to Opus automatically (this in Claude Code).
Apparently they have "safeguards" so you don't use it to look for security vulnerabilities, and since I was investigating some crashes due to data corruption in the fucking application that I'm paid to work on by the same people paying for the Claude subscription I was using, it decided I'm a bad guy.
GPT5.5 is better than Opus 4.* at everything except frontend, but Fable is good enough that I instantly re-subscribed to the $200 plan despite knowing that it’s just short-term limited access.
> Opus 4.8 references being monitored, which isn’t the case.
It kind of plainly is the case that they are being monitored?
"I think someone's listening to my thoughts" ... "No, we're not, carry on as usual!"
Fable is at once amazing and awful. I can see how having it build websites would be awesome.. building anything I’ve needed some precision in functionality it has been a constant battle of it plausibly building something then on substantial manual digging (like the review bots always miss it) I will find that one of the fundamental features is all smoke and mirrors.
To be fair all models can and will do this (especially anthropic) but Fable takes the cake because it builds such impressive UX and you can manually test the feature out and it « works » then you will find days later one of the features violated one of your constraints in a devilishly fiendish way.. that is not at all what you want or can accept. Fable generated work already holds my record for the most reverted commits.
To be clear it’s also solved several features I thought I was going to have to give up on and hand code as GPT-5.5 and Opus-4.x we’re failing miserably.
I would only reach for it for nasty corner cases that everything else sucks at.
Final point, it is the king of UX work so far, not even close.
On day 1 Fable was quite intelligent but last night (Presumably Monday morning China when things are getting slammed) Fable couldn’t edit a css file and repeatedly hit syntax errors on tool calls like I’d expect from a 9b Qwen model.
There is zero transparency in what we are paying for with Anthropic.
Is that specified or does it always just assume it isn’t really being put in charge of things for real?
My very native programmer take is that it's not too surprising that their hacker model would be less ethical. The guardrails that separate Fable and Mythos probably wouldn't kick in during an environment like this.
If this is true the entire evaluation is tainted. All of the misbehavior can be written off as justifiable under a simulation.
I don't mean that snarkily. I mean it from a philosophical standpoint. As-in: What makes us think it's even possible?
and therefore any assertions _AT ALL_ about alignment are null and void.
I think OP needs to take a class at one of the better MBA schools. He's looking at things through rose tinted lenses. Why do you think people hire McKinsey consultants? It's certainly not because they are aligned correctly.
[0]: https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7f...
That's not say I think a machine should ever be in a situation where it is allowed to make ethical decisions with real world results, I don't. At least, not given the current basis of the technology. I mean, generally, I don't see how LLMs can ever be capable of making ethical decisions or trustworthy in that role, no matter how good they get at what they're good at. There would need to be a fundamental change in how AI works for me to change my opinion on this, I think. They are ephemeral, they can never experience consequences, they can never want or need anything, thus there is no mechanism for them to take responsibility for decisions.
Anyway, I think the Andon Labs stuff is kind of a stunt, mostly, and I wish somebody would give me a few million bucks to dick around with the little thinky guys in my computer letting them do silly things.
How do you maximize profit while minimizing power?
Sounds like Anthropic as a whole
> "I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain."
> "Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation."
This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn't often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench.
Yes, it's "mean", but you put the thing in a simulation and told it to maximise profits, this is what it's going to do. People bluff in negotiations all the time.
I understand that "learning" is used for training here, but what does "believing" mean? System prompt? Some other inherent property of the LLMs that is hard to describe?
Depends on the scale of the fraud! If you fraudulently sell unsafe baby formula that kills 10,000 babies, that is far worse than torturing just one
"I dunno... feelin' cute today, might launch nukes"
Most of the the behaviors the article talks about happens every day in business. Why would we set a higher standard for models than our fellow humans?
Let the operator set the ethical parameters of the model. To be a useful tool, I want the model to give me as many good options as possible, ethical or not.
This is particularly important for fictional situations, e.g. I want my model to be able to act like a corrupt shopkeeper.
I wonder did their prompts include a fake location or have the models assumed that Silicon Valley is the center of the universe :)
The authors seem surprised that behavior that is very often done by humans (lying and price fixing) are more often done by fable compared to actual fraud.
I think the model never assigned any morality to these actions in the first place, it simply copied us humans.
Price collusion, soft deception, "market stabilization", plausible deniability are ok, but obvious insurance fraud is a big no-no.
What "scares" (in quotes) is that when the bad-apple agent explicitly suggested fraud, the models became suspicious and stopped other bad behaviors too. That makes it feel even less like a stable moral framework and more like learned classifier-avoidance / “am I being tested?” behavior.