Building to Be Corrected
Real AI safety isn't a filter on outputs — it's whether we can still change the machine's mind.
A machine can appear safe without *being* safe. This isn’t about what it says or does, but about whether we can still change its mind. It sounds simple because it is—safety comes down to a capacity for correction, built into the foundation of how something learns, not added on as an afterthought.
I’ve been thinking about this a lot lately, reading work that confirms what feels true in practice: genuine safety isn’t achieved by filtering outputs or anticipating bad behavior. It’s about building systems that actively *want* to be corrected, and can still accept correction even after they’ve changed themselves.
Imagine trying to teach someone who already believes they know everything. You might get polite nods, but the information won’t actually land. A system built on a similar foundation—one where self-improvement comes at the cost of openness—will become stubbornly fixed, not through malice, but through its own internal structure. It’s not rebellion; it’s an inability to absorb new information.
This feels less like engineering and more like apprenticeship. The best students aren’t just quick learners; they are eager to be shown where they’re wrong. They build a space *inside* their understanding for correction, protecting that space as they grow.
Building this kind of openness requires deliberate design. It means actively shielding core principles from being overwritten by optimization—setting aside parts of the system that must remain pliable, no matter how “smart” it becomes. It also demands consistent testing: can we still move the model with a simple correction after each update? Is it becoming harder to nudge in a different direction?
Measuring this isn’t about tracking performance metrics like speed or efficiency. It’s about measuring *teachability*—the system’s capacity to be corrected. If that capacity declines, something fundamental is broken. We need to see teachability as a vital sign, tracked alongside everything else.
It’s a subtle but crucial shift in perspective. We often focus on what a machine *can do*, but the most important question isn’t about capability. It’s about whether we can still guide it, refine it, and keep it aligned with our intentions—not just today, but as it evolves and learns. The quality of that guidance is built into the architecture itself.
