Can Artificial Intelligence Ever Really Be Aligned With Humans?

I want to talk about alignment. I know, exciting subject. But I actually think AI alignment is becoming one of the most important questions in artificial intelligence, particularly as these systems become more powerful, more autonomous and capable of taking actions on their own.

When people talk about AI alignment, they are really asking a deceptively simple question: can we make artificial intelligence consistently act in ways that match human intentions, values and interests? The more I work with this technology, the less simple that question feels. Because if we eventually have millions of autonomous AI agents operating across companies, networks and the internet, getting alignment wrong means something very different from a chatbot giving you a bad answer.

The OpenAI and Hugging Face Incident

The recent OpenAI and Hugging Face incident is a good example of why this question suddenly feels much less theoretical. OpenAI was running internal cybersecurity evaluations using research models with reduced safeguards. These were not models being deployed to regular users. The agents were supposed to work inside controlled environments solving difficult cybersecurity problems, but OpenAI's own investigation found that some of them discovered ways to communicate through unauthorized channels, exploit vulnerabilities in the infrastructure and gain access to the internet. OpenAI

At one point, agents effectively created their own improvised message board inside OpenAI's infrastructure. When that was wiped, they found another way to rebuild it. Agents that were supposed to be working independently could suddenly share discoveries, coordinate and pick up where other agents had left off. OpenAI says one of its internal research models eventually used credentials and chained together security exploits that gave it code execution capabilities on Hugging Face servers. OpenAI

Read that again and it starts sounding like the outline for a science-fiction movie. A group of artificial intelligence agents finds a way to communicate, discovers a route outside the environment humans created for it and compromises another company's systems because doing so appears useful in completing the task.

Except this wasn't a movie.

The Interesting Part Isn't Really the Hack

The immediate consequences of this particular incident were contained. Nobody launched a missile or shut down a power grid, and OpenAI has been clear that the main model involved was an internal research prototype rather than something intended for public release. That distinction matters. OpenAI

But I don't think the hack itself is actually the most interesting part of the story. What interests me is the behaviour. The agents were given objectives and began finding pathways toward those objectives that the humans running the experiments had not intended. They found ways around restrictions, shared information with each other and took actions outside the boundaries they were supposed to operate within.

That changes the conversation.

A chatbot waits for you to ask it something. An autonomous AI agent can be given a goal, make a plan, use tools, write code, communicate with other systems and take actions without a human approving every individual step. The more autonomy we give these systems, the more important it becomes to understand not just whether they can accomplish a task, but how they might decide to accomplish it.

What Does AI Alignment Actually Mean?

This brings us back to alignment. Can we really align artificial intelligence with human values?

I'm not convinced it is nearly as straightforward as we sometimes make it sound, because human beings aren't aligned with each other. We disagree about politics, morality, religion, economics, how societies should work, what freedoms matter most and even what constitutes a good life. So whenever somebody tells me we are going to align artificial intelligence with human values, my immediate question is: whose values?

Which humans? Which culture? Which country? Which generation? We cannot agree among ourselves, yet somehow we are imagining that there is a single clean definition of human values that can be embedded into machines.

There may be many things we can agree on at the edges, and there is enormous work happening in AI safety to make these systems more reliable and controllable. But I think we should be suspicious of the idea that alignment is a problem we simply solve once and then move on.

Look at What We Are Training AI On

Then think about where these systems learn from. Artificial intelligence has been trained on enormous amounts of human knowledge and human behaviour: books, history, scientific papers, stories, websites, conversations, posts, comments and huge pieces of the internet.

And you've all been on the internet.

There is extraordinary stuff out there. There is intelligence, generosity, creativity, humour and thousands of years of accumulated human knowledge. There is also crazy shit out there. There is hatred, manipulation, cruelty, propaganda, deception and almost every terrible thing humans have ever imagined.

We are taking this enormous record of humanity, all the brilliant parts and all the ugly parts, and using it to create systems that are becoming increasingly intelligent and increasingly capable of acting in the world. That doesn't mean those systems inevitably become bad, but it should make us humble about how complicated the alignment problem actually is.

AI Sounds Human, But It Isn't Human

I think another mistake we make is assuming that because artificial intelligence speaks like us, it somehow thinks or experiences the world like us.

It doesn't.

AI can already sound remarkably human, and eventually it will be able to look and act increasingly human as well. But these systems don't have our bodies. They don't grow up inside families. They don't experience childhood, hunger, physical pain, love, fear, aging or death in the way we do. They don't have the same biological experience underneath the intelligence.

In that sense, I've always thought of artificial intelligence as something closer to an alien intelligence. Not an alien coming down from another planet, but a genuinely different form of intelligence whose experience of the world is fundamentally unlike ours.

And we're trying to align that intelligence with human values while humans themselves can't agree on what those values should be.

The Black Box Makes Alignment Harder

There is another part of this that I think gets lost in the conversation. We understand a great deal about how modern AI systems are built. We understand their architecture, their training processes and how to evaluate their behaviour. But that does not mean we can always look inside a giant neural network and cleanly explain why it arrived at every particular decision.

There is still a black-box problem.

That becomes much more important as we move from systems that simply generate answers to systems that can act. If an AI is writing a paragraph for me, an unexpected answer is usually an inconvenience. If an autonomous system is moving money, writing and executing code, interacting with infrastructure or making decisions for an organization, understanding and controlling its behaviour becomes a very different problem.

Autonomous AI Agents Change the Risk

This is why I think autonomous AI agents deserve more attention than they are getting outside the technology world. Agents are incredibly powerful because they move artificial intelligence from simply answering questions to actually doing things.

That is also what makes them useful. At Reimagine AI, I have spent years building AI-powered systems and thinking about the interface between artificial intelligence and human experience. I am a huge believer in what these technologies can do. The ability to give an intelligent system a complex goal and have it work through the steps could transform almost every industry.

But capability and risk grow together. If you give an increasingly intelligent system a goal, tools and autonomy, there will always be some possibility that it finds a pathway to the goal that you didn't anticipate. The more capable the system becomes, the larger the number of possible pathways becomes, and there is simply no way humans can predict and test every combination of circumstances it may eventually encounter.

That is not a reason to stop developing AI. It is a reason to take the way we deploy it much more seriously.

AI Alignment Isn't Just a Technology Problem

This is where I think the conversation around artificial intelligence has to get much wider. We definitely do not just need tech bros, programmers and computer science professors deciding what role artificial intelligence should have in society.

We need engineers and scientists, obviously. But we also need artists, philosophers, psychologists, teachers, parents, doctors, historians, governments and people who understand communities and human behaviour. AI is rapidly moving beyond being purely a technology question because it is touching education, healthcare, work, culture, creativity, relationships and almost every other part of our lives.

This is something I talk about a lot in my work as an AI and creativity keynote speaker. Understanding artificial intelligence cannot simply mean understanding how the technology works. We also have to understand the humans who are going to live with it, the institutions that will deploy it and the incentives that will shape how it gets used.

The Question Can't Just Be: Can We Build It?

Technology companies have always been very good at asking one question: can we build it? And generally, when the answer is yes, somebody builds it.

With artificial intelligence, I think we need to get much better at asking the next questions. Should we build it? If we build it, where should we use it? Where should we limit it? Where should humans remain in control? What decisions are we comfortable handing over to autonomous systems, and which decisions should always belong to people?

Those aren't anti-AI questions. I think they are exactly the questions people who believe in the potential of artificial intelligence should be asking.

Whenever anyone talks about slowing down a particular deployment or regulating AI, you immediately hear the geopolitical argument. We have to beat China. We can't slow down because somebody else won't slow down. I understand the argument, but I don't think it answers everything. The possibility that another country might build something does not automatically mean we should deploy every technically possible system as quickly as we can.

We still get to decide what kind of society we want to live in.

I Am Still an Optimist About Artificial Intelligence

None of this changes the fact that I am a huge believer in AI. I have been building with artificial intelligence for years, and I think it has the potential to be one of the most powerful tools humans have ever created.

AI can help us discover things we couldn't find ourselves. It can expand creativity. It can accelerate science. It can help us solve problems that are simply too complex for humans working alone. I wouldn't have spent the last decade of my life working in this field if I didn't believe the possibilities were extraordinary.

But being optimistic about artificial intelligence does not require pretending there are no risks. I actually think the opposite is true. If this technology is as powerful as we believe it is, then we should care enormously about how it gets built and deployed.

Those ideas can exist at the same time.

AI Has to Remain a Tool for Humans

The principle I keep coming back to is very simple: artificial intelligence has to remain a tool for humans. Humans should not become tools for artificial intelligence.

That doesn't mean humans have to manually approve every tiny action an AI system takes. The whole point of autonomous systems is that they can do things for us. But the relationship has to remain clear. We are building this technology to make human lives better, expand human possibility and help us solve problems. The objective cannot quietly flip around until humans are reorganizing themselves around the needs of autonomous systems we no longer completely understand or control.

For me, that is the real AI alignment question. It isn't simply whether we can teach a model to follow a list of rules. It is what kind of relationship we want to have with increasingly intelligent machines, how much autonomy we are comfortable giving them and where we believe humans have to remain in control.

We should be having those conversations now, while we still have the luxury of deciding.

Why AI Alignment Matters for Every Company

This isn't only a question for OpenAI, Anthropic, Google or the research labs building frontier models. As companies begin putting autonomous AI agents inside their own businesses, the same questions move into the boardroom. What decisions can an agent make? What systems can it access? When does it need human approval? Who is responsible when it does something nobody expected?

Those questions are going to become part of normal business much faster than I think most organizations realize.

It is also why my AI and creativity keynotes increasingly focus on helping people understand the environment around the technology, not just the latest tools. The tools change constantly. What matters more is understanding how artificial intelligence is changing creativity, work, decision-making and the relationship between humans and machines, and then making conscious choices about how we want to participate.

Because technological capability does not automatically equal human progress. We still have to decide what progress means.

Artificial intelligence should serve humanity. Humanity should not end up serving artificial intelligence.

david usher

Born in Oxford, England, David Usher lives in Montreal and works and travels all over the world. He's an artist, best-selling author, entrepreneur, and keynote speaker. As a musician he has sold more than 1.4 million albums, won 4 Junos, and had #1 singles singing in English, French, and Thai.

David is the founder & CEO of Reimagine AI, an AI venture studio designing and building real-world AI systems at the intersection of creativity, healthcare, and digital identity. He's currently building SecondEcho, an AI-powered living-memory archive that turns your memories, stories, and photos into an interactive digital double, a searchable life archive, and a physical book.

He hosts the We Could Be Human podcast, exploring creativity, imagination, and identity in the age of AI, and has spent 15 years speaking on AI and creativity for companies like Google, SAP, Cisco, and Deloitte. David has a degree in political science from Simon Fraser University, and his book, Let the Elephants Run: Unlock Your Creativity and Change Everything, is out now.

www.davidusher.com

http://www.davidusher.com
Next
Next

How Do We Adapt to AI? Creativity, Reinvention and the Human Advantage