> 29. "emo's cup" — hmm, one more thought: "And there was a cup. You're like, that's dad's emo's cup." In context of a car photographed from helicopter... a cup in the car. "Dad's brand cup" — QuikTrip? "QT cup"? "emo's"... Hmm, could it be "that's dad's. You know, cup"? Or "Imo's"! IMOS — Imo's Pizza is THE St. Louis pizza chain! "Dad's Imo's cup" — Imo's Pizza cups! That's very plausible: St. Louis family, Imo's cup in the car. High confidence? Imo's is iconic StL. "emo's" → "Imo's" — phonetically identical. I'd include this one: from "dad's emo's cup" to "dad's Imo's cup". Moderate-high confidence given St. Louis setting. Include with context
Claude 4.6 didn’t get it. A higher level model might have, but 4.6 is the last model that doesn’t emit erratic content policy rejections for this task.
What’s really interesting is it’s also tied to thinking level. The sweet spot for quality with low refusal is Claude 4.6 high effort, go higher in effort and and it will trip too. 4.7 is hit and miss with high and often with xhigh and above 4.8 is often. I didn’t try Fable. Sonnet wasn’t good enough quality wise.
It also seems to be related to asking it producing what it knows is copied content. I originally had it return full corrected transcripts and that hit all the time. Then I started having it return just correction lists and that works much better.
These big models recall is absolutely astounding. I built the transcript pipeline so I could search and find that episode that I kind of remember part of from a year ago. I asked fable to search the transcripts to find a particular Conan Needs a Friend episode (Katakai, as god made her!). It knew there were actually multiple episodes where that anecdote was told _before_ it had even searched. When Kimi K3 does a correction you can see in the thinking trace that it instantly recalls people and works before it reasons in to whether it’s certain that’s right.
I’m still testing, but Claude is going to get replaced for K3 when I get some time. It’s very good at this work.
I wonder how many image ids it memorized and whether it would always use those or search for new images?
Although, it is possible that it was unintentional as the model was RLHF'ed into memorizing images.
Stripping URLs would certainly break some code, and also train the model with unhelpful behavior. You want your coding model to write real URLs, not write "return redirect(<|URL_STRIPPED|>)".
Like yeah, you made it perform better, but I could’ve set any other model to xhigh and spent all the tokens and patience.
It’s surely enough to generate buzz for some headlines about “Fable-level performance”, but I would’ve expected people “in the know” to push back at this more.
It feels like a benchmark-only setting, just for ArtificialAnalysis really.
Just to add before someone objects, but the tokens! I have only once maxed out a 5x subscription week (which I then upgraded and maxed out a 20x week so it was a big push). I don’t understand how some people are using so many tokens
Umm what? Lots of people do