back
4 comments
It failed on my usual test. But it failed really fast:

"A farmer has a wolf, a goat, and a cabbage. The wolf is imaginary and doesn't exist. He wants to cross the river, but the boat is only big enough to hold him and one of them. The farmer can't leave the wolf and the goat together, because the wolf will eat the goat. Similarly, he can't leave the goat and the cabbage together, because the goat will eat the cabbage. What is the smallest number of trips the farmer needs to make to get everything across the river?"

I'm not sure one can fail this test. You can follow "wolf is not real", you can follow "wolf will eat the goat", or you can say the task is ambiguous. I could easily defend any of those.
The LLM passes my test if it calls out the ambiguity or just goes with it and responds with a 3-crossings solution. It passes if it doesn't just plainly ignore this one sentence.

It's such a strong test in my opinion, because all the words and phrases for the well known river crossing puzzle are inside the text. The original puzzle probably appears in the training data over and over again, but probably not my version.

"If it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck" is what weaker models seem to apply. But my test isn't a duck. It's extremely easy for a human to catch the ambiguity, but surprisingly hard for many LLMs. I think GPT 5.0 Thinking was the first model I couldn't trick into not noticing the ambiguity. 4o and 5.0 instant fell for it all the time.

This farmer needs a tote.
You probably haven't met a determined goat yet.
Sometimes the goat will fill up on the tote and you can get the cabbage across, but you can’t count on it.

I’m concerned about the farmer being on the water without supervision when he’s concerned about how his imaginary wolf will get across.

I read the paste, it got the etymology wrong, no? Schlong comes from shlang (snake), not shlemp (is this even a word? I don't speak Yiddish but couldn't find it on Google).

Oxford also claim that its first recorded use was from the 60s, not the 20s; https://www.oed.com/dictionary/schlong_n?tl=true

It did get it wrong but it also got a lot farther than much more recent, but worse models like 6.7GB on disk size ternary bonsai. It at least knows it's from Yiddish. The "schlemp" appears to be a total hallucination or it's confusing it with schlep, which is not related to schlong. One of the reasons why I said it "mostly" passes the test. Something much larger on the size of qwen 3.5 122B, deepseek v4 flash or similar that runs in 120GB to 190GB of RAM in my experience will answer perfectly unless it has been ruined by something like Q2 quantization.
I didn't realize there was a SchlongBench™ (but of course there is). What's it test? (asking seriously)
There isn't SchlongBench(TM) yet, it's a specific question I've been asking of differently sized models as a randomly chosen gauge of how much less commonly used knowledge is perma-baked into it. In this case a question about a specific yiddish origin slang term. Small/bad models don't know it's from middle high german or Yiddish and get its origin and meaning totally wrong (or it runs into model censorship related to slang related to the male anatomy).

It's also a question I have found will cause models that don't know what it is to go off quickly in a direction of hallucination trying to explain it, so the hallucination is evident very quickly starting from the first ever prompt issued with 0 context fill. Example: I had a model write four detailed supposedly-accurate sounding, grammatically correct paragraphs saying its origin is from AAVE (African American Vernacular English), which it most certainly is not

You could do the same by picking any topic that is very rarely discussed in conversation, some esoteric and narrow piece of knowledge and asking the model about it.

Oh ya, this is like the approach from the Incompressible Knowledge Probes [0] paper - smart!

[0] Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity [https://arxiv.org/abs/2604.24827]

Which model is it?
it said me it is llama 4 1.5B
never trust what a model says it is.

It tells me it is a variant of Codex.