---
title: "Generative Topolinguistics: Bidirectional interfaces for emergent language topologies"
description: "Iulia Ionescu, Jenn Leung and Yannis Siglidis propose a research programme they call generative topolinguistics: treat the internal geometry of a large language model not as a map of language to be read but as an interface to be pushed on, and watch what language and what social behaviour come out."
type: "reading-notes"
authors:
  - "Iulia Ionescu"
  - "Jenn Leung"
  - "Yannis Siglidis"
published: "2025-05-10"
original_url: "https://coginfra.antikythera.org/"
doi: "10.1162/ANTI.5CZG"
canonical_url: "https://antikythera.wiki/work/journal/coginfra/generative-topolinguistics"
md_url: "https://antikythera.wiki/md/work/journal/coginfra/generative-topolinguistics"
last_updated: "2026-08-23"
site: "Antikythera Wiki"
---

# Generative Topolinguistics: Bidirectional interfaces for emergent language topologies

> Independent, unofficial reading notes from Antikythera Wiki (https://antikythera.wiki). Written with AI from the published text, with citations; not by the work's author and not affiliated with Antikythera (https://antikythera.org). Read the original at the link below.

- **Authors:** Iulia Ionescu, Jenn Leung, Yannis Siglidis
- **Published:** 2025-05-10
- **Kind:** Studio projects · Journal article
- **DOI:** https://doi.org/10.1162/ANTI.5CZG
- **Original:** https://coginfra.antikythera.org/
- **Length:** 7.9k words
- **Part of:** [Cognitive Infrastructures](https://antikythera.wiki/work/journal/coginfra)
- **This page:** https://antikythera.wiki/work/journal/coginfra/generative-topolinguistics
- **Notes generated:** 2026-08-23

## Summary

[Iulia Ionescu](https://antikythera.wiki/people/iulia-ionescu), [Jenn Leung](https://antikythera.wiki/people/jenn-leung) and [Yannis Siglidis](https://antikythera.wiki/people/yannis-siglidis) propose a research programme they call generative topolinguistics: treat the internal geometry of a large language model not as a map of language to be read but as an interface to be pushed on, and watch what language and what social behaviour come out. The proposal rests on a claim about where LLM structure comes from. Society produces language, language produces the model's token space, and the model's outputs now flow back into society, so the arrangement is bidirectional. The paper surveys the LLM social-simulation literature to show that bot-only societies already generate sociolinguistic patterns unlike human ones, then sketches four speculative experiments, from steering individual tokens to an "LLM game of life" in which agents evolve their own ways of communicating. It reframes interpretability as an instrument for sociolinguistics rather than for safety alone.

## The argument

The paper opens with a historical claim. Saussure, Wittgenstein, Chomsky and Skinner were each modelling a different aspect of language, and the disputes between them were ideological confrontations between partial models rather than tests of a shared framework. Computational linguistics tried to make those models tangible and, in the authors' judgement, failed; connectionist next-token prediction succeeded, and an LLM has thereby become an "emergent unified model" in which structuralist, pragmatic and social theories of language may all be encoded somewhere in the weights. If Saussure and Lyotard are now implicitly coded in parameter space, one can study their limits by interacting with the model rather than by arguing about it.

The second move is the three-mirrors model, which fixes the direction of explanation. Sociality, defined with Chris Sinha as a normative pattern of meaning realised in semiotic systems and institutions, is the highest-order domain. Language is its "diffractive lens", in the sense developed by Duranti and Eckert, where linguistic forms index social context and speakers use variation to build social meaning. Token space is a second-order embedding of sociality projected through that lens. This is a stance against the autonomy of grammar from social organisation claimed by formal linguistics. The authors' interest is not in reconstructing the history of language inside a model but in the conditions under which new sociolinguistic behaviour could emerge from human–AI interaction.

Bidirectionality is the third step. Embedded sociality is the projection from society through language into token space. Emergent sociality is the inverse: token-space structure expressed as linguistic patterns and social behaviour, first between bots and then, because people appropriate linguistic structures that seem humanly possible, in human discourse. Giddens's double hermeneutic and Clark's cognitive niche construction are the borrowed frames. Social-scientific concepts alter the phenomena they describe, and humans and LLMs are jointly building a linguistic environment that will shape both parties' future cognition.

The survey of LLM social simulations grounds the claim that this emergence is observable, and the technical section supplies four interventions scaled from tokens to geometry to topology. The conclusion is modest. The boundary between artificial and human linguistic systems is more permeable than assumed, hybrid sociolinguistic phenomena are to be expected, and this raises questions about linguistic agency that LLM deployment should take into account.

## Section by section

### 1 Introduction

After the historical sketch, the subsection "Generative Topolinguistics" names three things one can locate in a model's vector spaces: structures of token interaction, geometric properties where vector similarity encodes semantic similarity, and topological properties whose global shape can encode sociality. The founding question is what happens to language, and to the sociality it supports, if one manipulates its geometric representation rather than modelling it as a moving target.

### 2 Sociality as Embedded and Emergent in Language

The three mirrors are set out in order, with Eckert's indexical field offered as an analogy for the many dimensions of token space. The bidirectional framework then separates embedded from emergent sociality and locates existing emergent sociality in bot-only social networks. A conditional is stated plainly: whether dialogue with an LLM is semiotically rich enough to change culture depends on whether the model offers enough human-like affordances for its interlocutor to appropriate its structures, and the authors cite Turing-test results to say it already does. The open question is whether manipulating vector geometry would spill over into social interaction, and how that spillover would be re-embedded when a model is trained on its own outputs.

### 3 Synthetic Sociolinguistics

"From Language to Life" reviews the bottom-up tradition from Schelling to Epstein and Axtell, whose dictum is that ungrown emergence is unexplained emergence, and notes that none of its models could "grow humans out of molecules". "From Tokens to Sociality" observes that one model can be individuated by prompting into many agents, so the LLM behaves like language itself, a medium in which individuals are born as sociolinguistic agents. The evidence comes in two kinds. On fixed platforms such as Chirper AI and OnlyBots, LLM bots form persistent communities, some homophilous by language, and show two equally pronounced modes of content similarity where human societies show one. In role-play settings such as Stanford's generative agents and DeepMind's Concordia, agents show information diffusion, relationship memory and coordination. "Recursive Linguistic Simulations" draws the methodological conclusion: geometry is the medium of social emergence, and LLM agents can be evolved outputs rather than designed inputs. Since the shared topography of human languages may be an artefact of cultural transmission rather than a universal grammar, it may be more timely to culture many different languages than to regrow the ones we have.

### 4 Towards Generative Topolinguistics

The aim is twofold: extend generative linguistics by comparing human and LLM outputs geometrically rather than testing hypotheses about innate grammar, and pursue sociolinguistics through the topological unfolding of LLM outputs. "From Tokens to Topology" is the technical primer. Tokens are "computational syllables", about three-quarters of an English word; next-token prediction is a movement in spherical coordinates with angle for semantics and radius for confidence; averaged sequence representations form an embedding space whose dynamics can unfold topologically. Chat behaviour is role-play induced by markers such as "human:" and "AI:". Any given model is a particular, non-universal instance, yet still a cultural technology in Gopnik's sense, and one that can be intervened on at every scale.

"Speculative Approaches" presents four experiments. Under "Tokens as Agents", a token sequence is handed to a controller trained by reinforcement learning to deform the output under constraints, for instance to lower populism in social-media text; the authors acknowledge the kinship with adversarial attacks. Under "Geometric Manipulation", the model for what is wanted is the discovery of latent directions in GANs that linearise attributes like facial expression, using adapters, probes or sparse-neuron decompositions. The distinctive proposal is to bind such manipulations to input modalities rather than discrete goals, for example a tactile interface translating touch or pulse into geometric deformation. Under "Topological Contouring", individuated agents are routed to one another, their successive outputs trace a reachable region of embedding space, and manipulations are evaluated by how they change that coverage. Under "LLM Game of Life", a competitive environment with selection dynamics lets specialised models discover more efficient utterances and forms of organisation, as agents in communication games invent codes.

### 5 Conclusion

The closing paragraph predicts hybrid sociolinguistic phenomena, argues that the capacity of LLMs to reshape the social fabric they model calls for more care in deployment, and describes the paper as speculation within an empirical generative framework.

## Key concepts

- **[Generative topolinguistics](https://antikythera.wiki/terms/generative-topolinguistics)** — Learning about language and sociality by manipulating the geometry and topology of the models that reproduce them, at the scales of tokens, geometry and topology, and treating the result as a bidirectional interface rather than a static visualisation.
- **[Three-mirrors model](https://antikythera.wiki/terms/three-mirrors-model)** — Sociality, language and token space as successive projections, each a mirror of the one above it, with sociality given explanatory primacy.
- **Embedded and emergent sociality** — The projection of society into token space, and the reverse flow from token space into linguistic patterns and behaviour, first among bots and then among people.
- **Embedding space** — The averaged representation of a whole sequence, the stratum of [latent space](https://antikythera.wiki/terms/latent-space) usually visualised.
- **Topological contouring** — Measuring which regions of embedding space a chain of agent interactions can reach, and using geometric manipulation to change that reachability.
- **LLM game of life** — A competitive multi-agent environment with selection dynamics in which pretrained models are adapted to discover their own utterances and modes of organisation.

## Connections

- [Cognitive Infrastructures](https://antikythera.wiki/work/journal/coginfra) places the paper under "Mimesis of Mimesis" and glosses it as treating embedding visualisations as "representations of representations of representations". The three-mirrors model is the worked-out form of that slogan, and the bidirectional framework is the compendium's general claim that to play with a model is to remake it, specialised to language.
- [Traversing the Uncanny Ridge](https://antikythera.wiki/work/journal/coginfra/traversing-the-uncanny-ridge) is the closest sibling. Both treat AI-to-AI communication as a source of novelty, and the ridge paper's proposal to keep embedding spaces for different modalities separate is a design choice of the kind topological contouring would measure. They agree that convergence is not the goal.
- [Minimum Viable Interiority](https://antikythera.wiki/work/journal/coginfra/minimum-viable-interiority) addresses the step this paper takes for granted, how one model becomes many agents. Ionescu, Leung and Siglidis individuate by prompting; that paper argues individuation requires functional closure, a mild tension.
- [Synthetic Counteradaptation](https://antikythera.wiki/work/journal/coginfra/synthetic-counteradaptation) and [Mutual Prediction in Human–AI Coevolution](https://antikythera.wiki/work/journal/coginfra/mutual-prediction-in-human-ai-coevolution) describe the human side of the loop this paper calls emergent sociality, and the three are in agreement.
- [Latent Spacecraft](https://antikythera.wiki/work/journal/latentspacecraft) is the earlier attempt to make latent space navigable rather than narratable; this paper moves from navigation to intervention.
- The [Agentworld (Research Brief)](https://antikythera.wiki/work/book/agentworld-brief) treats human linguistic structures as a subset of a wider field of synthetic modes, which is this paper's conclusion stated as a premise. [The Silicon Interior](https://antikythera.wiki/work/substack/2026-02-10-the-silicon-interior), on the Moltbook agent society, is a later empirical case of the bot-only networks surveyed here.


## Terms used

- [Bidirectional Alignment](https://antikythera.wiki/terms/bidirectional-alignment)
- [Generative Topolinguistics](https://antikythera.wiki/terms/generative-topolinguistics)
- [Latent Space](https://antikythera.wiki/terms/latent-space)
- [That Which Can Be Tokenized](https://antikythera.wiki/terms/that-which-can-be-tokenized)
- [Three-Mirrors Model](https://antikythera.wiki/terms/three-mirrors-model)
