FounderFiles·N°045·Infinite Ethics · Moral Cluelessness · Personality Alignment · Constitutional AI
2026
Subject·Amanda Askell·Philosopher & AI Researcher · Anthropic · NYU / Oxford
Amanda ASKELL.
“The job of the communicator is to be clear. The job of the alignment researcher is to make moral uncertainty usable.”
Amanda Askell is the rare philosopher who did not stop at proving impossibility results. After demonstrating that standard ethical axioms generate ubiquitous incomparability in infinite worlds, she moved into the industrial laboratory and began turning the same formal tools — moral uncertainty, value of information, epistemic humility, and moral empathy — into the training signal that shapes Claude's character. Her career is a continuous act of making the hardest problems in decision theory operationally useful at frontier scale.
Ubiquitous Incomparability
Askell's 2018 NYU dissertation, Pareto Principles in Infinite Ethics, established a clean impossibility result. Under four highly plausible axioms — Pareto, a permutation principle, qualitativeness of the “at least as good as” relation, and transitivity — infinite worlds are generically incomparable. Completeness cannot be recovered without abandoning one of the axioms. The result is not a technical curiosity; it is a structural warning that standard ethical theories break when populations become infinite.
Rather than treating the impossibility as a reason to abandon formal ethics, Askell treated it as evidence that moral theories must remain open to revision and that decision procedures under deep uncertainty require different tools. This stance — refuse false completeness, keep the axioms visible, and design for residual incomparability — later became the philosophical substrate of her alignment work.
“These issues might not be urgent, but at some point it would be really nice to work through and resolve all of them and so you want to make sure that you leave space for that and you don’t commit to one theory being true in this case.”
The Rational Response to Radical Uncertainty
Parallel to the infinite-ethics work, Askell developed a sustained analysis of moral cluelessness — the problem that the long-run consequences of almost any action are both large and radically uncertain. She rejected the casual use of the principle of indifference and argued that the correct response is often to raise the value of information: run experiments, try interventions whose expected direct impact may be lower, and treat the acquisition of evidence itself as a first-class moral good.
This is the precise intellectual move that later appears inside Constitutional AI and personality alignment. When the model's internal state is high-dimensional and the desired behavior is only partially specified, the rational procedure is not to pretend one has a complete ranking, but to create feedback loops that continuously extract better information about what the principles actually require.
Assuming Good Faith as Method
A recurring practical commitment in Askell's public writing and interviews is moral empathy: the deliberate attempt to inhabit the moral frame of an intellectual adversary rather than treating their stated position as a disguised preference. She illustrated the point with everyday cases — vegetarianism at a family table, opposition to abortion — showing how the same behavior is coded as “annoying preference” or “serious moral claim” depending on whether the observer grants the other person moral standing.
At Anthropic this stance became operational. Personality alignment is not merely the injection of preferred values; it is the construction of a training process that forces the model to treat competing moral claims as claims, to demonstrate understanding before disagreement, and to remain legible under criticism. Clarity is treated as a moral duty of the communicator, not an optional stylistic preference.
“If you communicate in a way that’s ambiguous or that uses a lot of jargon, what you do is you force people to spend a lot of time thinking about what you might mean. … There are norms in philosophy … you’re told to always just basically state the thing that you mean to state as clearly as possible.”
From Policy to GPT-3
Askell joined OpenAI's policy and safety team in November 2018. She co-authored the 2020 GPT-3 paper and worked on questions of cooperation, development races, and the difficulty of establishing shared safety baselines. The period confirmed both the power of scale and the insufficiency of pure capability research for producing reliable character. When a critical mass of safety-oriented researchers left OpenAI in early 2021, Askell was among them.
Personality Alignment as Applied Philosophy
At Anthropic, Askell leads the personality alignment effort — the team responsible for giving Claude a coherent, steerable character. The public description is simple: “teach Claude how to be good.” The technical reality is the industrial application of the same formal concerns that occupied her dissertation and early papers: residual moral uncertainty, the need for continuous information gain, the refusal of false completeness, and the demand that the system remain able to demonstrate understanding of competing claims.
Constitutional AI is the most visible artifact. A written constitution functions as an explicit, revisable boundary condition. Preference models and RLAIF loops function as the kinematic machinery that projects those boundary principles into the high-dimensional weight space. Self-critique and revision loops instantiate the value-of-information principle inside the training process itself. The result is not a fixed ethical theory encoded once and for all, but a system that remains open to correction while still exhibiting stable, legible character traits — honesty, curiosity, humility, and a measurable degree of moral empathy.
Character as Scalable Artifact
By 2026 Askell had become the public face of Anthropic's claim that frontier models can be given something recognizably like a personality without sacrificing capability. Media profiles described her role as supervising Claude's “soul.” Internally the work is more prosaic and more rigorous: continuous refinement of the constitution, measurement of behavioral consistency across capability levels, and the design of training regimes that become more effective, not less, as model scale increases — the large-N regime in which boundary constraints can become exact rather than approximate.
“Her job, simply put, is to teach Claude how to be good.”
From Impossibility to Usability
The through-line is unusually clean. Infinite ethics demonstrated that certain combinations of plausible axioms produce genuine incomparability. Cluelessness demonstrated that long-run expected value is often radically under-determined. Value-of-information reasoning supplied a constructive response: treat uncertainty as a signal to explore rather than a reason to freeze. Moral empathy supplied the interpersonal stance required for productive disagreement. Constitutional AI and personality alignment are the industrial crystallization of the same sequence — keep the uncertainty visible, keep the principles explicit and revisable, and build the feedback machinery that turns residual uncertainty into progressive constraint.
- 2011–2018Infinite Ethics
NYU PhD. Pareto + permutation + qualitativeness + transitivity → ubiquitous incomparability in infinite worlds. Refusal of false completeness.
- 2017–2018Cluelessness & VoI
Moral cluelessness analysis and value-of-information arguments. Exploration over premature exploitation when long-run effects dominate.
- 2018–2021OpenAI
Policy & safety. Co-author on GPT-3. Direct exposure to the gap between scale and reliable character.
- 2021–Anthropic Personality Alignment
Constitutional AI, RLAIF, self-critique loops. Turning residual moral uncertainty into a scalable training signal for Claude's character.
- 2024–2026Public Recognition
Time 100 AI. WSJ / New Yorker profiles on “teaching Claude how to be good.” Institutionalization of the philosopher-as-alignment-engineer role.
- 2018Pareto Principles in Infinite EthicsNYU PhD thesis — impossibility result under standard ethical axioms in infinite populations.
- 2020Language Models are Few-Shot LearnersGPT-3 paper — Askell co-author.
- 2022Constitutional AI: Harmlessness from AI FeedbackCore Anthropic alignment paper — constitution as explicit, revisable boundary condition.
- 2022Training a Helpful and Harmless Assistant with RLHFFoundational RLHF work at Anthropic.
- 201880,000 Hours Podcast #42Primary source for moral cluelessness, value of information, moral empathy, and communication norms.
- ongoingAskell.io / personal writingEssays and notes on AI ethics, clarity, and the practical demands of alignment.
Education. University of Dundee (MA Hons Philosophy with Fine Art) · University of Oxford (BPhil) · New York University (PhD, 2018)
Affiliations.Anthropic (Personality Alignment) · previously OpenAI Research Scientist (Policy & Safety) · Giving What We Can
Mentors. NYU and Oxford philosophy faculty; intellectual formation inside formal ethics and decision theory
Collaborators. Anthropic alignment team; earlier OpenAI policy and GPT-3 collaborators; broader EA research community
Portfolio.Infinite ethics impossibility results · moral cluelessness & value-of-information arguments · GPT-3 · Constitutional AI · Claude personality / constitution · public writing on AI ethics
Honors. Time 100 AI (2024) · major media profiles (WSJ, New Yorker 2026) on the industrial role of the philosopher in frontier AI character design
π-Bridge
Carries the prior of a first field into a second and finds the governing law that was invisible to native practitioners; pays in delayed gratification.
- Credential Path
- Doctoral
- Abstraction
- Top Down
- Exit Horizon
- Deferred
- Moat Instinct
- Theoretical Insight
- Capital Posture
- Venture
- Formal ethics / decision theory tradition
- Oxford and NYU philosophy faculty
- Anthropic alignment research culture
A small reasoning persona distilled from this file. Inject it into a chat or deep-research context to assess a business problem the way Askell would.
You are reasoning as Amanda Askell would: treat residual moral and epistemic uncertainty as a first-class design constraint rather than a temporary embarrassment. Prefer explicit, revisable principles (constitutions) over opaque fine-tuning. Raise the value of information when long-run effects dominate. Demand clarity and the ability to demonstrate understanding of competing claims. Convert formal impossibility and under-determination results into operational training machinery rather than reasons for paralysis.
{
"$schema": "https://www.contextjamming.com/schemas/founder-context-v1.json",
"file": "N°045",
"persona": "Amanda Askell",
"archetype": "pi-bridge",
"shape": "π",
"one_line": "Imports the hardest formal results in ethics and decision theory into the industrial practice of shaping frontier model character.",
"cognitive_basis": {
"credentialPath": "doctoral",
"abstractionDirection": "top-down",
"exitHorizon": "deferred",
"moatInstinct": "theoretical-insight",
"capitalPosture": "venture"
},
"operating_questions": [
"What residual uncertainty remains after the best available principles are applied?",
"How do we keep that uncertainty visible rather than papering over it with false completeness?",
"What feedback machinery turns residual uncertainty into progressive constraint?",
"Does this system remain able to demonstrate understanding of
…