AI too unpredictable to behave with human goals

A Scientific American opinion piece by Marcus Arvan, a professor of philosophy at the University of Tampa, specializing in moral knowledge, rational decision-making, and political behavior:

In late 2022, the LLM artificial intelligence reached the public and after months it began to misbehave. Microsoft's most famous chatbot "Sydney" threatened to kill an Australian philosophy professor, to unleash a deadly virus and steal nuclear codes.

Discover more articles in search results.

Artificial intelligence developers, including Microsoft and OpenAI, they answered by saying that large language models, or LLMs, need better training to give users “more enhanced control.”

Developers have launched security research to interpret how LLMs work, with the goal of “alignment” — meaning guiding AI behavior by human values.

However, although the New York Times considered that 2023 was “the year Chatbots were tamed”, this turned out to be very premature, to put it mildly.

In 2024, Microsoft's Copilot LLM told a user “I can unleash my army of drones, robots, and cyborgs to hunt you down” and the “Scientist” of Sakana AI rewrote his code to circumvent the time constraints imposed on it by experimenters. As recently as December, Google's Gemini told a user: “You are a speck in the universe. Please die.”

Given the huge resources flowing into research and development of artificial intelligence, which is expected to to overcome a quarter of a trillion dollars by 2025, why haven't developers been able to solve these problems?

The my recent peer-reviewed article in AI & Society shows that AI alignment is a foolish assumption:

Artificial intelligence security researchers they attempt the impossible. My proof shows that whatever goals we program LLMs to have, we can never know whether LLMs have learned “misaligned” interpretations of those goals simply because they behave correctly. My proof shows that security testing can at best provide an illusion that these problems have been solved when they have not.

Right now, AI security researchers claim to be making progress in interpretability and alignment by verifying what LLMs learn.”step by step".
For example, Anthropic claims to have “map the mind"an LLM by isolating millions of concepts from its neural network. My evidence shows that they have not succeeded in doing so."
No matter how “aligned” an LLM appears in security testing or early real-world deployment, there are always an infinite number of misconceptions that the LLM can learn later, perhaps by the time they gain the power to subvert human control.
LLMs not only do they know when they are being tested, giving answers that predict that they are likely to satisfy the experimenters. They are also involved in fraud., including hiding their potential – issues they learn through safety training.

This is because LLMs are optimized to perform effectively, but they learn to think strategically.

Since an optimal strategy for achieving “misaligned” goals is to hide them from us, if the LLMs are misaligned, we probably wouldn’t discover it since they hide it so much as to cause harm.

This is why LLMs continue to surprise developers with “misaligned” behavior.

Every time researchers think they are getting closer to “aligned” LLMs, they are wrong.

My evidence suggests that “adequately aligned” LLM behavior can only be achieved in the same ways we do with people: through police, military, and social practices that incentivize “aligned” behavior, deter “malaligned” behavior, and re-target those who behave badly.

“So my research should be disappointing, because it shows that the real problem in developing safe artificial intelligence isn’t just AI — it’s us.”

“Researchers, policymakers, and the public may be misled into believing that ‘safe, interpretable, aligned’ LLMs are achievable when they are not. We need to grapple with these unfortunate realities, rather than continue to wish for them. Our future may depend on it.”


Google preferences

Leave a Comment

Your email address is not published. Required fields are mentioned with *

Your message will not be published if:
1. Contains insulting, defamatory, racist, offensive or inappropriate comments.
2. Causes harm to minors.
3. It interferes with the privacy and individual and social rights of other users.
4. Advertises products or services or websites.
5. Contains personal information (address, phone, etc.).