---
title: "Alignment Overfitting"
description: "Bratton's name for the error of treating short-term \"alignment\" — tuning models to cohere with stated human preferences and self-image — as the long-term metaphysical grounding for how animal intelligence should orient machine intelligence."
type: "glossary-term"
canonical_url: "https://antikythera.wiki/terms/alignment-overfitting"
md_url: "https://antikythera.wiki/md/terms/alignment-overfitting"
last_updated: "2026-09-06"
site: "Antikythera Wiki"
---

# Alignment Overfitting

> A working definition written for this wiki from the reading notes; not Antikythera's own wording. From Antikythera Wiki (https://antikythera.wiki/terms/alignment-overfitting), an independent guide to Antikythera's published work.

Bratton's name for the error of treating short-term "alignment" — tuning models to cohere with stated human preferences and self-image — as the long-term metaphysical grounding for how animal intelligence should orient machine intelligence. Lower-case alignment is a legitimate engineering tactic (it "makes it work"); upper-case Alignment, overfit to "dubiously defined human values," is itself a kind of existential risk because of what it directs attention *from*: the Copernican disclosures AI could make about what cognition and "we" actually are. The remedy is bidirectional alignment — AI to culture's wisdom, culture to AI's disclosures — see *Productive Disalignment*.

## Appears in

- [Cognitive Infrastructures](https://antikythera.wiki/work/journal/coginfra) (2025-05-10)
- [After Alignment](https://antikythera.wiki/work/journal/afteralignment) (2025-05-10)
- [Antikythera](https://antikythera.wiki/work/journal/research) (2025-05-10)

_Glossary generated 2026-09-06._
