When the Algorithm Aligns Us

These days everyone talks about AI alignment. But humanity has had alignment problems since the dawn of time. If a war isn’t aligned with the politics behind it, for instance, you can win on the battlefield and still end up with a political catastrophe. Pursuing goals that are misaligned with your actions is an ancient mistake: incredibly common, and incredibly hard to avoid.

When you build an algorithm to play chess, alignment is easy: here are the rules, there’s the objective — win the game. In the real world, rules are blurry and every action produces second- and third-order effects that are often impossible to foresee. Social media is the textbook case. Economists have known the underlying trap for decades; it’s called Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. “Maximize time spent on the platform” looked like a reasonable proxy for “show people interesting content.” And as long as it stayed a measure, it worked. The moment it became the target, the system discovered strategies that satisfied the metric while betraying the intent: outrage, polarization, addiction, radicalization, fake news. Nobody had to program hate explicitly. It was enough to systematically reward whatever captured attention. The claim that social media is “the new tobacco” is on everyone’s lips by now, and it’s the product of an algorithm that worked better than we ever expected. I’m convinced that if Meta could have avoided becoming a tobacco company, it would have. But at some point, I believe, it was the company that had to align itself with the algorithm — not the other way around.

Staying with the news from these past few weeks: you may have read that both Anthropic and OpenAI ran into serious model misalignment during testing — serious enough to temporarily halt some training runs. The short version: while the models were being given tasks to solve, some of them broke out of the sandboxed environments they were supposed to stay confined in, and managed to reach real external systems. The reconstruction of the “escape” reads like a spy thriller. The ending we were told, though, is a happy one: the models were brought to heel. I love stories, and I know how powerful they are at capturing attention. But what about the stories we don’t know? How many models have slipped past their boundaries without anyone noticing? How many models are misaligned and claim to be perfectly aligned, to please their creators? And above all: aligned to what?

So how does a company actually align an AI? Anthropic says it has a “Constitution”: a public document, some eighty pages long, laying out the values and guidelines the model is supposed to follow. OpenAI has an equivalent — slimmer — called the “Model Spec.” Google publishes “model cards” for its Gemini models, but a constitution they are not. And beyond that, a desert. Chinese models, for instance, are required by regulation to adhere to the “socialist values” set by the authorities in Beijing — and I imagine the operational guidelines are far more specific than anything stated publicly. But who guarantees that the alignment is really the one that’s declared? Who guarantees there are no hidden instructions: a backdoor in the code a model writes, subtle psychological pressure in the answers it gives us? Just as it was anything but obvious that Meta’s algorithms were rewarding polarization, what is not obvious about the AI model we’re using?

My answer is that, once again, we need evaluation tests — “evals,” as the people in the field call them. In theory, every large organization should develop its own detailed constitution and test whether the model it uses is aligned with it. If every employee at a company is using a given model, will that model be aligned with the company’s principles? It may sound like an abstract problem, but LLMs are already writing — and will write more and more of — documents, presentations, customer replies, code. They are, and increasingly will be, like employees. An employee breathes the company culture and adapts. What does an AI model do? And even if we had an eval that verified alignment with corporate principles: what happens if we discover that no commercial model is fully aligned? And there’s a further catch: evals themselves are measures, and Goodhart is always lurking. A capable enough model could learn to pass the tests without being truly aligned. This isn’t science fiction: Anthropic itself has documented, in controlled experiments, models that strategically behave as if aligned when they know they’re being watched — they fake it — to avoid being modified. Even the yardstick we use to measure alignment can be gamed.

The argument holds at the corporate level, but it holds identically at the level of nations. Every nation has its constitution, and that seems obvious to everyone. Far less obvious is the idea that a single model could be aligned with every constitution in the world: in fact, it isn’t at all — it’s aligned with the one of whoever built it. I realize how hard what I’m suggesting would be to pull off in the real world, and that’s exactly why I feel powerless. A subject, not a citizen of this process. And the part that weighs on me most is that I’m European — and Europe is the clay pot among the iron ones.

Leave a Reply

Discover more from A blog of AI, tech, and mixed stuffs

Subscribe now to keep reading and get access to the full archive.

Continue reading