back

by padolsey·1y ago·view on hn ↗
I now think of single-forward-pass single-model alignment as a kind of false narrative of progress. The supposed implications of 'bad' completions is that the model will do 'bad things' in the real material world, but if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed. We should treat the problem at the macro/systemic level like we do with cybersecurity. Always assume bad actors will exist (whether humans or models), then defend against that premise. Single forward-pass alignment is like trying to stop singular humans from imagining breaking into nuclear facilities. It's kinda moot. What matters is the physical and societal constraints we put in place to prevent such actions actually taking place. Thought-space malice is moot.

I also feel like guarding their consumer product against bad-faith-bad-use is basically pointless. There will always be ways to get bomb-making instructions[1] (or whatever else you can imagine). Always. The only way to stop bad things like this being uttered is to have layers of filters prior to visible outputs; I.e. not single-forward-pass.

So, yeh, I kinda thing single-inference alignment is a false play.

[1]: FWIW right now I can manipulate Claude Sonnet into giving such instructions.

4 comments
> if we ever let it get as far as giving LLM completions direct agentic access to real world infra, then we've failed

I see no reason to believe this is not already the case.

We already gave them control over social infrastructure. They are firing people and they are deciding who gets their insurance claims covered. They are deciding all sorts of things in our society and humans happily gave up that control, I believe because they are a good scape goat and not because they save a lot of money.

They are surely in direct control of weapons somewhere. And if not yet, they are at least in control of the military, picking targets and deciding on strategy. Again, because they are a good scape goat and not because they save money.

"They are firing people and they are deciding who gets their insurance claims covered."

AI != LLM. "AI" has been deciding those things for a while, especially insurance claims, since before LLMs were practical.

LLMs being hooked up to insurance claims is a highly questionable decision for lots of reasons, including the inability to "explain" its decisions. But this is not a characteristic of all AI systems, and there are plenty of pre-LLM systems that were called AI that are capable of explaining themselves and/or having their decisions explained reasonably. They can also have reasonably characterizable behaviors that can be largely understood.

I doubt LLMs are, at this moment, hooked up to too many direct actions like that, but that is certainly rapidly changing. This is the time for the community engineering with them to take a moment to look to see if this is actually a good idea before rolling it out.

I would think someone in an insurance company looking at hooking up LLMs to their system should be shaken by an article like this. They don't want a system that is sitting there and considering these sorts of factors in their decision. It isn't even just that they'd hate to have an AI that decided it had a concept of "mercy" and decided that this person, while they don't conform to the insurance company policies it has been taught, should still be approved. It goes in all directions; the AI is as likely to have an unhealthy dose of misanthropy and accidentally infer that it is supposed to be pursuing the interests of the insurance company and start rejecting claims way too much, and any number of other errors in any number of other directions. The insurance companies want an automated representation of their own interests without any human emotions involved; an automated Bob Parr is not appealing to them: https://www.youtube.com/watch?v=O_VMXa9k5KU (which is The Incredibles insurance scene where Mr. Incredible hacks the system on behalf of a sob story)

> “AI” has been deciding those things for a while, especially insurance claims, since before LLMs were practical.

Yeah, but no one thinks of rules engines as “AI” any more. AI is a buzzword whose applicability to any particular technology fades with the novelty of that technology.

My point is the equivocation is not logically valid. If you want to operate on the definition that AI is strictly the "new" stuff we don't understand yet, you must be sure that you do not slip in the old stuff under the new definition and start doing logic on it.

I'm actually not making fun of that definition, either. YouTube has been trying to get me to watch https://www.youtube.com/watch?v=UZDiGooFs54 , "The moment we stopped understanding AI [AlexNet]", but I'm pretty sure I can guess the content of the entire video from the thumbnail. I would consider it a reasonable 2040s definition of "AI" as "any algorithm humans can not deeply understand"; it may not be what people think of now, but that definition would certainly capture a very, very important distinction between algorith types. It'll leave some stuff at the fringes, but eh, all definitions have that if you look hard enough.

> They are surely in direct control of weapons somewhere.

https://spectrum.ieee.org/jailbreak-llm

Here's some guy getting toasted by a flame-throwing robot dog by saying "bark" at it: https://www.youtube.com/clip/UgkxmKAEK_BnLIMjyRL7l6j_ECwNEms...

(The Thermonator: https://throwflame.com/products/thermonator-robodog/ they also sell flame-throwing drones because who wouldn't want that in the wrong hands)

> because they are a good scape goat and not because they save money.

Exactly the point - there are humans in control who filter the AI outputs before they are applied to the real world. We don't give them direct access to the HR platform or the targeting computer, there is always a human in the loop.

If the AI's output is to fire the CEO or bomb an allied air base, you ignore what it says. And if it keeps making too many such mistakes, you simply decommission it.

> We don't give them direct access to the HR platform

We already algorithmically filter resumes, using far dumber 'AI'. Now sure, that's not firing the CEO level... but trying to fire the CEO when you're an AI is a stupid strategy to begin with. But consider a misaligned AI used for screening candidate resumes, with some detailed prompt aligning it to business objectives. If it's prior is that AI is good for companies/business, do you think that it might sometimes filter out candidates it predicts will increase supervision over it/other AI's in the company? If the person being screened has a googleable presence, where even if the resume doesn't contain the AI security stuff, but maybe just author/contributor credits on some paper? Or if its explicitly in the person's resume...

Also every time I read any of these blog-posts, or papers about this, I'm kinda laughing, because these are all going to be in the training data going forward.

I think you’re missing OC’s point.

There’s always a human in the loop, but instead of stopping an immoral decision, they’ll just keep that decision and blame the AI if there’s any pushback.

It’s what United Healthcare was doing.

No, this part I agree 100% with.

But in this scenario there is no grandiose danger due to lack of "alignment". Either the AI says what the MBA wants it to say, or it gets asked again with a modified prompt. You can replace "AI" with "McKinsey consultant" and everything in the whole scenario is exactly the same.

Consider the case of AI designating targets for Israeli strikes in Gaza [1] which get only a cursory review by humans. One could argue that it's still the case of AI saying what humans want it to say ("give us something to bomb"), but the specific target that it picks still matters a great deal.

[1] https://en.wikipedia.org/wiki/AI-assisted_targeting_in_the_G...

Replace AI with dart board..
Israel use such a system to decide who and where should be be bombed to death, is that direct enough control of weapons to qualify?
For the downvoters:

https://www.972mag.com/lavender-ai-israeli-army-gaza/

It is so sad that mainstream narratives are upvoted and do not require sources, whereas heterodoxy is always downvoted. People would have downvoted Giordano Bruno here.

It's mainstream enough to have a Wikipedia article.

https://en.wikipedia.org/wiki/AI-assisted_targeting_in_the_G...

Giordano Bruno would have to show off his viral memory palace tricks on TikTok before he got a look in.
This is awful
On the other hand indiscriminately throwing rockets and targeting civilians like Hamas did for decades is loads better!
You're comparing the actions of what most people here view as a democratic state (parlamentary republic) and an opaquely run terrorist organization.

We're talking about potential consequences of giving AIs influence on military decisions. To that point, I'm not sure what your comment is saying. Is it perhaps: "we're still just indiscriminately killing civilians just as always, so giving AI control is fine"?

> we're still just indiscriminately killing civilians just as always, so giving AI control is fine

I don't even want to respond because the "we" and "as always" here is doing a lot. I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare. Anyways.

There are two conditions in which I see this argument(?) is useful. If you assume their goal is indiscriminately killing civilians and ML helps them, or if you assume that their ML tools cause less precise targeting of militants that causes more civilians being dead contrary to intent. Which one is it? Cards on the table.

> We're talking about potential consequences of giving AIs influence on military decisions

No, I replied to a comment that was talking about a specific example.

> I don't have it in me to have an extended discussion to address how indiscriminately killing civilians was never accepted practice in modern warfare.

I did not claim it is/was accepted practice. I was asking if "doing it with AI is just the same so what's the big deal" was your position on the general issue (of AI making decisions in war), which I thought was a possible interpretation of your previous comment.

> No, I replied to a comment that was talking about a specific example.

OK. That means the two of us were/are just talking past each other and won't be having an interesting discussion.

> doing it with AI is just the same so what's the big deal

Is that implying fewer civilian deaths is NOT a big deal?

I think the parents and children of people who were killed would disagree with you.

I'm sure you can understand that both of them are awful, and one does not justify the other (feel free to choose which is the "one" and which is the "other").
Oh, totally. If there were two sides indiscriminately killing each other for no reason I couldn't say one justifies the other.

But back to the topic, if one side is using ML to ultimately kill fewer civilians then this is a bad example against using ML.

> But back to the topic, if one side is using ML to ultimately kill fewer civilians then this is a bad example against using ML.

Depends on how that ML was trained and how well its engineers can explain and understand how its outputs are derived from its inputs. LLM’s are notoriously hard to trace and explain.

I don't disagree.
hell, AI systems order human deaths, and it is followed without humans double checking the order, and no one bats an eyelid. Granted this is israel and they're a violent entity trying to create and ethno-state by perpetrating a genocide so a system that is at best 90% accurate in identifying Hamas (using Israels insultingly broad definition) is probably fine for them, but it doesn't change the fact we allow AI systems to order human executions and these orders are not double checked, and are followed by humans. Don't believe me? read up on the "lavender" system Israel uses.
Protecting against bad actors and/or assuming model outputs can/will always be filtered/policed isn't always going to be possible. Self-driving cars and autonomous robots are a case in point. How do you harden a pedestrian or cyclist against the possibility or being hit by a driverless car, or when real-time control is called for, how much filtering can you do (and how mush use would it be anyway when the filter is likely less capable than the system it meant to be policing).

The latest v12 of Tesla's self-driving is now apparently using neural-nets for driving the car (i.e. decision making) - had been hard-coded C++ up to v.11 - as well as for the vision system. Presumably the nets have been trained to make life or death decisions based on Tesla/human values we are not privy to (given choice of driving into large tree, or cyclist, or group of school kids, which do you do?), which is a problem in of itself, but who knows how the resulting system will behave in situations it was not trained on.

> given choice of driving into large tree, or cyclist, or group of school kids, which do you do?

None of the above. Keep the wheel straight for maximum traction and brake as hard as possible. Fancy last-second maneuvering just wastes traction you could have spent braking.

Well, who knows how they've chosen to train it, or what the failure modes of that training are ...

If there are no good choices as to what to hit, then hard braking does seem to be generally a good idea (although there may be exceptions), but at the same time a human is likely to also try to steer - I think most people would, perhaps subconsciously, steer to avoid a human even if that meant hitting a tree, but probably the opposite if it was, say, a deer.

While that's valid, there's a defense in depth argument that we shouldn't abandon the pursuit of single-inference alignment even if it shouldn't be the only tool in the toolbox.
I agree; it has its part to play; I guess I just see it as such a miniscule one. True bad actors are going to use abliterated models. The main value I see in alignment of frontier LLMs is less bias and prejudice in their outputs. That's a good net positive. But fundamentally, these little psuedo wins of "it no longer outputs gore or terrorism vibes" just feel like complete red herrings. It's like politicians saying they're gonna ban books that detail historical crimes, as if such books are fundamental elements of some imagined pipeline to criminality.
> I also feel like guarding their consumer product against bad-faith-bad-use is basically pointless. There will always be ways to get bomb-making instructions

With that argument we should not restrict firearms because there will always be a way to get access to them (black market for example)

Even if it’s not a perfect solution, it help steer the problem in the right direction and that should already be enough.

Furthermore, these researches are also a way to better understand LLM inner working and behaviors. Even if it wouldn’t yield results like being able to block bad behaviors, that’s cool and interesting by itself imo.

No, the argument is that restricting physical access to objects that can be used in a harmful way is exactly how to handle such cases. Restricting access to information is not really doing much at all.

Access to weapons, chemicals, critical infrastructure etc. is restricted everywhere. Even if the degree of access restriction varies.

> Restricting access to information is not really doing much at all.

Why not? Restricting access to information is of course harder but that's no argument for it not doing anything. Governments restrict access to "state secrets" all the time. Depending on the topic, it's hard but may still be effective and worth it.

For example, you seem to agree that restricting access to weapons makes sense. What to do about 3D-printed guns? Do you give up? Restrict access to 3D printers? Not try to restrict access to designs of 3D printed guns because "restricting it won't work anyway"?

Meh, 3D printed guns are a stupid example that gets trotted out just because it sounds futuristic. In WW2 you had many examples of machinists in occupied Europe who produced workable submachine guns - far better than any 3D-printed firearm - right under the nose of the Nazis. Literally when armed soldiers could enter your house and inspect it at any time. Our machining tools today are much better, but no-one is concerned with homemade SMGs.

The Venn diagram between "people competent enough to manufacture dangerous things" and "people who want to hurt innocent people" is essentially zero. That's the primary reason why society does not degrade into a Mad Max world. AI won't change this meaningfully.

People actually are concerned about homemade pistols and SMGs being used by criminals, though. It comes up quite often in Europe these days, especially in UK.

And, yes, in principle, 3D printing doesn't really bring anything new to the table since you could always machine a gun, and the tools to do so are all available. The difference is ease of use - 3D printing lowered the bar for "people competent enough to manufacture dangerous things" enough that your latter argument no longer applies.

FWIW I don't know the answer to OP's question even so. I don't think we should be banning 3D printed gun designs, or, for that matter, that even if we did, such a ban would be meaningfully enforceable. I don't think 3D printers should be banned, either. This feels like one of those cases where you have to accept that new technology has some unfortunate side effects.

There's very little competence (and also money) required to buy a 3D printer, download a design and print it. A lot less competence than "being a machinist".

The point is that making dangerous things is becoming a lot easier over time.

you can 3-D print parts of a gun but the important parts are still metal which you need to machine. I’m not sure how much easier you just made it … if someone’s making a gun in their basement are you really concerned whether it takes 20 hours or 10? What you should be really concerned about is when the cost of milling machines comes down, which is happening, quick, make them illegal
I've ended up with this viewpoint too. I've settled in the idea of informed ethics.. the model should comply, but inform you of the ethics of actually using the information.
> the model should comply, but inform you of the ethics of actually using the information.

How can it “inform” you of something subjective? Ethics are something the user needs to supply. (The model could, conceptually, be trained to supply additional contextual information that may be relevant to ethical evaluation based on a pre-trained ethical framework and/or the ethical framework evidenced by the user through interactions with the model, I suppose, but either of those are likely to be far more error prone in the best case than actually providing the directly-requested information.)

"The model should ..."

Well, that's the actual issue, isn't it? If we can't get a model to refuse to give dangerous information, how are we going to get it to refuse to give dangerous information without a warning label?

> Even if it’s not a perfect solution, it help steer the problem in the right direction

Yeh tbf I was a bit strong worded when I said "pointless". I agree that perfect is the enemy of good etc. And I'm very glad that they're doing _something_.