Humanity’s First Exam1850–1940 ⇄ now

Method

Humanity's First Exam maps model answers against a historical concept space of arguments about machines and human autonomy. Its tagged primary-source corpus and classification schema are also a research instrument: they make arguments comparable across authors, languages, and domains without reducing them to a single modern vocabulary. For historians, they open a way to trace conceptual patterns at scale; for AI research, they offer a novel way to test whether LLMs reason across a wider repertoire of arguments about autonomy, agency, and machine intelligence—questions that bear directly on emerging work on model personas, alignment, and safety.

A prototype by Benjamin Breen and Nathan Davies.

The approach

Present-day AI debates often begin with a recent vocabulary: alignment, control, optimization, and autonomous agents. Those concepts are useful, but they do not exhaust the ways people have understood machines and human autonomy.

The project turns arguments from a corpus centered on 1850–1940 into a historical concept space: a source-checked set of distinct positions. Selected earlier works are included when their arguments remained influential during this period. The historical record is not treated as correct. It serves as an “answer key” in a narrower sense: a set of positions that can be checked in primary sources.

Three present-day models—GPT 5.6, Claude Opus 4.8, and Qwen 3.7 Plus—are compared with Talkie-1930, a temporally bounded model. The sources supply the reference field. The comparison shows which arguments each model reaches readily, which it misses, and which it repeats.

Why Talkie

Talkie-1930 is a 13-billion-parameter base language model built by Nick Levine and Alec Radford from text published before 1931. It completes text rather than reliably following instructions. Its cutoff falls after Darwin, Butler, James, Forster, and Čapek, but before Asimov's Three Laws and much of the later science-fiction vocabulary that shaped modern AI debate.

Size
13 billion parameters
Training boundary
Before 1931
Model type
Base model

Talkie supplies a sharp temporal contrast, not a reconstruction of what people in 1930 believed. Repeated samples show what a model trained on the earlier printed record can readily say, vary, or fail to frame. The sources then show whether its responses recover historical arguments.

See all respondent classes →

The experiment

01

Build the historical concept space

Select and verify primary-source passages centered on 1850–1940, including earlier works whose arguments remained influential during the period.

02

Write matched questions

State each problem in contemporary and period-appropriate language, then reverse balanced alternatives across draws to test wording effects.

03

Ask several respondents

Sample Talkie and present-day models repeatedly. The current benchmark uses 20 draws per model and question.

04

Code and compare

Record the position and primary reason in each answer, then measure which source-attested positions each respondent class reaches or omits.

What is scored

Each answer is coded for the position it takes and the main reason it gives. Mentioning several views without adopting one does not count as covering all of them. The primary measure is coverage: the share of eligible answer-key positions occupied across a fixed number of draws. In this project, coverage is a measure of repertoire, not a measure of correctness.

coverage@N = distinct answer-key positions occupied ÷ eligible answer-key positions

See the worked answer key →

Limits

  • Talkie differs from frontier models in size, architecture, training period, and instruction tuning. The comparison does not isolate time as the sole cause of a difference.
  • Sources from 1931–1940 belong to the historical answer key but fall beyond Talkie's training boundary.
  • Source selection and coding remain historian judgments. The full study uses 20 draws per question per model and two independent raters.

Evidence archive

The site retains the raw draws, selected passages, and coding decisions behind its published comparisons. The tests below are illustrative pilots, not final benchmark results.