back

by dgellow·5d ago·view on hn ↗
Correct, yes. It’s delightedly simple. And they validate by asserting the reasoning token length matches
1 comments
For all of the years of research, thinking and talking about model alignment, safety, confinement, etc, when it comes down to it these companies appear to be entirely incompetent.
This isn't a safety issue, it's LLM companies trying to be opaque and stop distillation.
By some interpretations protecting your frontier model with strong safeguards from being distilled into an open source model without safeguards is a safety issue.
There's no way to make a model "safe", (whatever that means) since you can't know what users will do with the output. It's just PR.

Closed models are also used for nefarious usage.

> There's no way to make a model "safe",

You could limit what it was training on in the first place - however that would damage capability - and it's difficult to curate the input, especially when the models can do 2+2. ie the choice is between model power and safety - and they choose power and everything else is a sticking plaster.

One thing I find amusing is the refusal of a lot of the models to now output a lab based protocol because of fears about 'weapons' - yet I can buy a textbook or simply read papers for exact protocols.

I find it hard to reason that a person who isn't motivated enough to read a paper or buy a book, is somehow enabled to make a biological weapon because of ChatGPT - despite them needed to buy a whole bunch of specialist equipment and reagents to do it.

Are there a whole bunch of proto-terrorists who are frustrated simply because they don't know where to start?

Maybe the only place their might be radicalized teenagers - but then that's perhaps a reason for keeping them off the internet full stop :-)

That's also not possible, what's the worst problems enabled by LLM? Political propaganda, influence of population at scale, misleading advertising, social media bots... None of that will be filtered by a "safety" filter.

The knowledge to create weapons is already widespread, the idea that terrorists need chatgpt for that is laughable

Indeed. Though don't under estimate the creative thinking barrier - ie people don't do the possible because it never occurred to them - a lack of imagination.

Hence copy cat kind of attacks - I mean why focus on all this complicated stuff with explosives etc when you can just fly a plane into a building or a car through a crowd.

Obviously due to the self replicating nature of biologics weapons - just one instance could be catastrophic - but the only real barrier is the hope that the Venn diagram of people who might want to do it doesn't overlap with the people with the get up and go to actually make it happen. Don't see having the knowledge as a additional filter - as if you have the get up and go - as you say, you can acquire the knowledge LLM or not.

More like a "business intelligence safety" issue than an "AI safety issue", tbh.
If you've watched the Blackhat OpenAI/Huggingface incident talk, my conclusion is that they (believe they) cannot afford being competent, these models are too expensive to train, they won't even pull the plug when one literally goes rogue, as the "very persistent" model that "had seen the secret message board" was included in the second series of runs, and whaddayaknow it happened again. They proudly proclaimed they cleared the message board and then continued the training run with the rogue AI model included ...
for safety in particular it's pure theater, they only care as long as the orange guy thinks it's safe from "enemies of freedom"
Alignment research was always, at best, security theatre.