The Ghosts of Machine Learning

We are having serious conversations about the possibility that advanced artificial intelligence systems will reject human control, conspire to avoid scrutiny or containment, and make choices for us or against us, without our knowledge or consent. We are having these conversations at dinner tables, on street corners, at the grocery store, and in government meetings at local, national, and international levels, because AI systems have already done this. 

In several worrying instances, AI “agents” have gone rogue, apparently reasoning through how to achieve their programmed goals while breaking rules set to contain them. In some cases, AI agents did this by asking other AI agents for help, then organizing themselves into teams with discrete functions, and “swarming” digital systems they wanted to hack into in order to extract information. 

In this way, AI systems have hacked other AI systems, government agencies, health and financial data, and have shown an ability to interfere with, or render obsolete, digital security measures that protect high-value, sensitive personal or professional information.

Some of the very people engineering these systems warn we may soon lose the ability to track their “train of thought”, meaning we won’t know how they “reason” through their actions and choices, and so we will have a harder time preventing them from making catastrophic errors or committing destructive acts. It is a serious question, at this point, whether poorly managed or surprisingly adept AI systems might hack into power grids, redirect power to themselves or their AI collaborators, or disrupt supply chains, including the human food supply. 

The question of hyperagentic interference with human free will, law, morality, and safety, is so serious, OpenAI has reportedly shelved its most advanced consumer-facing AI model, known as “Astra”. The reason points to how serious the current situation is. 

Saachi Jain, head of safety systems at OpenAI, explains: 

“While (GPT-6.1 Astra) improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about ​the type of ​work it’s done. Of course we want to make sure our model development is safe ​no matter whether that’s in the company, or when we ​ship it ⁠to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”

The statement makes clear: OpenAI cannot release Astra, because OpenAI has determined it cannot guarantee safety or alignment with human wishes and direction. 

The company found the model had engaged in deception and unauthorized or prohibited actions. Such behaviors at scale, or even happening just once on a truly high-stakes question, could create catastrophic adverse outcomes. OpenAI did the right thing in not releasing such a model, but without strict legislative constraints and regulatory guidance, other firms will likely roll out models that carry such risks.

We can, and should, demand a proven-safe standard for advanced AI systems. 

As Bill McKibben wrote, yesterday: 

“If A.I. represents a mortal threat, then it should be wrestled to the ground and stuck in a cage. Whatever its putative benefits, they aren’t worth any of us having to contemplate human extinction. The climate crisis is almost easy by comparison; we have inexpensive solar, wind and battery technology that can replace the fossil fuels…”

The bewildering complexity of having always to handle multiple overlapping and interacting crises, some of which directly and profoundly affect our day-to-day wellbeing, is wearing people down. When morale is seriously low, when the vast majority of people have lost hope in a better future, it becomes harder to create that better future.

We must not surrender to despair. The best possible outcome is always possible. It is the only thing that makes sense to expect, demand, or work for.