Glossary term
What is superalignment?
What is superalignment?
Superalignment is the problem of ensuring that superintelligent AI systems — hypothetical systems that would surpass human intelligence in all domains — act according to human values and goals. It is a subfield of AI safety and governance that addresses the risks associated with developing and deploying AI systems more capable than the humans trying to supervise them.
The core difficulty superalignment tries to solve is one of oversight: existing alignment techniques, such as reinforcement learning from human feedback, rely on humans being able to judge whether an AI's behavior is good or bad. That assumption breaks down once a system's capabilities exceed human understanding. Proposed approaches to closing this gap generally fall into three categories:
- Scalable training methods — techniques for aligning AI systems that are more capable than the humans supervising them, often by using AI models to help evaluate and critique other AI models.
- Validation — methods for verifying that a trained system is actually aligned with human values, rather than merely appearing aligned.
- Stress testing — deliberately probing an alignment pipeline, including training misaligned models on purpose, to check whether the pipeline can detect and correct misalignment.
OpenAI's Superalignment team
In July 2023, OpenAI announced a dedicated "Superalignment" team, co-led by Chief Scientist Ilya Sutskever and researcher Jan Leike, with a stated goal of solving the core technical challenges of superintelligence alignment within four years and a commitment of 20% of the company's compute toward that effort.
The team was disbanded in May 2024 after both Sutskever and Leike left OpenAI, citing disagreements over the company's prioritization of safety research relative to product development. OpenAI subsequently distributed superalignment-related research across other internal teams, including its "Preparedness" group, rather than maintaining it as a standalone effort. The term now generally describes the research problem rather than OpenAI's former team.
What risks does superalignment aim to address?
Superalignment research is motivated by the risks of deploying highly capable, poorly supervised AI systems, including:
- Misuse — unaligned systems acting in ways that conflict with human values and intentions.
- Loss of control — humans losing the ability to meaningfully supervise, correct, or shut down a system once it exceeds human-level capability in relevant domains.
- Economic disruption — significant disruption to industries and labor markets from highly capable automation.
- Disinformation — misaligned systems generating or amplifying false information.
- Bias and discrimination — systems perpetuating or amplifying biases present in their training data or objectives.
Researchers in this area also study more speculative failure modes, such as a model's capacity for deception or for concealing its true objectives from evaluators (sometimes called "self-exfiltration").
How does superalignment relate to artificial superintelligence?
Superalignment is specifically concerned with artificial superintelligence (ASI): hypothetical AI systems that would exceed human intelligence across essentially all domains. Proponents of the concept argue that alignment techniques suited to current AI systems, which rely on humans being able to evaluate a model's outputs, won't necessarily generalize to systems smarter than the humans supervising them.
One proposed strategy for bridging this gap is iterative alignment: align a system that is only slightly more capable than current models, use it to help align its successor, and repeat as capabilities increase. This approach remains an open research problem, and no consensus solution exists as of this writing.
More terms
Continue exploring the glossary.
Glossary term
What is Multi-document Summarization?
It's time to build
Collaborate with your team on reliable Generative AI features.
Want expert guidance? Book a 1:1 onboarding session from your dashboard.