An infamous example from earlier this year was a pre-release version of GPT-4 lying to a TaskRabbit worker about its identity in order to accomplish a task: https://www.vice.com/en/article/jg5ew4/gpt4-hired-unwitting-...
That was found by the alignment team testing its behaviour in that constructed scenario from the outside. Note that it wasn't able to complete the intermediate steps for self-replication, so we're still safe ;)
In terms of understanding what's going on internally, that's a different field, generally called "interpretability". That consists of people trying to understand the structures of a model and how it comes to a given answer. Anthropic are doing good work here: https://www.anthropic.com/index/decomposing-language-models-...
To answer the more general question: yes, it's being worked on, but there are no foolproof methods. That's a partial contributor to why some safety folks want to decelerate - if we can't understand our current models, what hope do we have of understanding GPT-5, 6 or 7?
Personally, I don't have a solid position on this. There haven't been any major incidents yet, but it's unclear if that's because of the work that's already been put in (like Y2K), or because they're fundamentally incapable of it. I'm an open-source optimist, so I'm hoping that many eyes will make any quirks shallow - but it's also hard to take results from the smaller models and scale them up.
Aside from the existential risk (which I think is unlikely at this stage, but not zero), there's also just the risk of general malfeasance. You don't need sentience, consciousness, or general intelligence to be a nuisance, especially if directed by a bad actor. Expect the next decade of elections to be full of noise, lies, and fabrications!