The Data That Named You
They called it a gift to researchers.
On August 4, 2006, a team inside AOL Research posted a compressed text file to a public server: roughly 20 million search queries typed by more than 650,000 users over three months. No screen names. No email addresses. Every account flattened to a bare integer. Anonymized, they said, and technically they weren't lying — they had stripped every field a database engineer would recognize as identity.
They just left in the part where people confessed.
i · the anonymization that wasn't
The engineering error here is old enough to have a name, and it had one in 2006 too. Replacing a name with a number doesn't destroy identity; it relocates it. The identifier isn't the column you dropped. It's the shape of everything you kept. Twenty million queries spread across 650,000 people means each person in that file is a sequence — a few hundred lines of text that no one else on earth would ever type in that order.
Consider what a search log actually is. Not a list of interests. A list of things you would not say out loud. It's where a person types the name of a rash before they call a doctor, a lawyer's specialty before they file, a city before they leave one. The search box in 2006 was arguably the last honest interface in consumer computing, precisely because the machine was assumed not to be listening in any way that mattered. Every query was a small act of trust in that assumption.
AOL published the trust.
The New York Times demonstrated the failure inside a week. User #4417749 turned out to be Thelma Arnold, a 62-year-old widow in Lilburn, Georgia, identified by cross-referencing her queries against phone book listings. Searches about her town. Last names she was looking up for friends. Numb fingers. A dog that urinated on everything. Reporters knocked on her door, and she let them print it — the reaction, roughly, of someone learning that the private thing had not been private for some time.
One user was enough. That's the part people miss about de-anonymization: you never have to re-identify everyone to break the promise. One re-identification is a proof of concept for the whole file. The remainder is just arithmetic and motive, and both of those are cheap.
Some of the users became folklore. User 927's query log — read as a continuous document by strangers on the internet — got picked apart, annotated, and eventually adapted for the stage. A person nobody could name became a character anyway, assembled entirely from the questions they asked a machine at three in the morning. That's the tell. The data wasn't a behavioral trace. It was a diary written by someone who didn't know they were writing.
ii · the three-day delete
AOL pulled the file on August 7. Three days.
Here's the technical reality nobody in that room appears to have priced: on a network, deletion is a request, not an operation. By the time the file came down it had been mirrored, torrented, and wrapped in searchable web front ends by people who thought the whole thing was a fascinating dataset and not, say, 650,000 depositions. Two decades on, you can still find it. The dataset outlived the company that produced it — AOL as a consumer brand is a ghost; the file is in excellent health.
That's the line I'd tattoo on the inside of every product manager's eyelids: the half-life of a data release is longer than the half-life of the institution that releases it. Your retention policy, your deletion pipeline, your compliance dashboard with the reassuring green checkmarks — those govern your copy. They have no jurisdiction whatsoever over anyone else's.
Consequences arrived on schedule and landed exactly where consequences land. AOL's CTO, Maureen Govern, resigned on August 21. Two employees were fired: the researcher who posted the file and his immediate supervisor. A class action followed in September, alleging violations of the Electronic Communications Privacy Act and seeking at least $5,000 for every person whose searches were exposed. It settled in 2013.
Read that arc as a timeline rather than a story. Three people lose their jobs in three weeks. The structural question takes seven years to be quietly converted into a settlement. And the incentive that produced the file in the first place — publish data, accrue research prestige, move faster than the competitor — is never so much as inconvenienced.
The individuals absorbed the blame. The system that made the decision feel reasonable stayed precisely where it was.
And it's worth naming the shape of that system, because the shape is the thing that reruns. It isn't carelessness, and it certainly isn't malice. It's that the entity making the decision is never the entity holding the tail risk. AOL Research faced a bounded, career-shaped downside — worst case, a few people lose their jobs, which is precisely what happened. The 650,000 faced an unbounded, permanent, personal one, and were not party to the decision in any sense: not consulted, not notified, not compensated on a timescale that meant anything. When the upside accrues to the decider and the catastrophic tail lands on people with no seat at the table, that isn't a risk calculation. It's a subsidy. The file was rational to publish. That's the problem.
iii · the pattern rerun, with better marketing
If AOL were a one-off we'd call it a scandal. It's a genre.
In 2008, Narayanan and Shmatikov de-anonymized the Netflix Prize dataset by matching it against public IMDb ratings — the same failure with different columns, and the resulting paper made the point formally: "anonymized" describes a dataset's schema, not a property of the dataset. In 2018, Strava published a global heatmap of aggregated exercise data and inadvertently published the perimeter jogging routes of undisclosed military installations. Every year since, some broker's "de-identified" location feed turns out to reconstruct where individual people sleep, because the place you are every night at 3 a.m. is not an anonymous coordinate, it's your address.
The sentence is always the same. Aggregated and anonymized. It functions as a liturgical phrase — said at the moment of release, believed by the speaker, and load-bearing for nothing.
But the phrase isn't the engine. It's the exhaust. Vocabulary keeps finding new speakers because the asymmetry underneath keeps producing them: in every one of those cases, an institution took a bounded reputational risk with material it held but did not generate, and the people who generated it carried whatever came next. "Aggregated and anonymized" is just what that arrangement says about itself out loud. Fix the language and you get a better-worded version of the identical release.
What makes the AOL case worth the twentieth-anniversary retrospective isn't that it was the worst. It's that it was the clearest. A single file, a single week, a single named woman in Georgia, and a completely legible causal chain from good intentions to violated life. Nothing since has been that easy to see, which is not because the practice improved.
Here's the coherenceist reading. The release was framed as a contribution to the commons — free data, open research, science advanced. But look at the direction of the transfer. The material came from 650,000 people. The benefit went to institutions with the bandwidth to download the file and the staff to mine it. The cost went back to the people, individually, without notice. That's not a commons. That's enclosure wearing the vocabulary of openness.
A commons is a thing held in common by the people who constitute it. The test isn't whether the data is free to access. It's whether the people who generated it are inside the circle of who decides and who benefits — and, the part AOL makes unavoidable, inside the circle of who is exposed when it goes wrong. Those are supposed to be one circle. An arrangement that places you in the third but not the first two isn't a commons with a bug in it; it's an enclosure working as designed. Thelma Arnold was not consulted about the advancement of information retrieval science. She was raw material for it, and she found out from a reporter.
iv · the successor nobody is watching
The direct descendant of the AOL search log is not the training corpus. It's the chat log.
Look again at what made the search box dangerous: it was the last honest interface in consumer computing, because the machine was assumed not to be listening in any way that mattered. Every word of that is more true of a conversation with a model — at many times the volume, with a stronger and even more misplaced sense of privacy, and with a property the search box never had, which is that the machine answers back, and therefore asks follow-ups. Thelma Arnold typed "numb fingers" into a box and got ten blue links. Her 2026 equivalent types it into a chat window and gets asked how long, which fingers, whether it's both hands, whether anyone in the family has had this. The search log was a diary written by someone who didn't know they were writing. The chat log is an interview — conducted nightly, transcribed in full, retained by policy, sitting in a database.
That file exists right now. It is the most concentrated deposit of unguarded human self-disclosure ever assembled, and it has every property the AOL file had — enumerable, copyable, three days from permanent — with none of AOL's excuse, because AOL is the worked example. It simply hasn't leaked yet. When it does, the retrospective writes itself, and I'd rather write it now.
v · the thing without a manifest
Then there's the other successor, the one that gets all the attention: the training corpus. The intimate residue of ordinary people is now not an incidental research asset but a primary input to a multi-hundred-billion-dollar industry, gathered continuously rather than in a single regrettable upload.
The data really is in there. Carlini and colleagues showed in 2021 that large language models emit verbatim sequences from their training data under the right prompting; a follow-up in 2022 quantified the relationship, finding that memorization grows with model capacity, with prompt context, and — most sharply — with the number of times a sequence was duplicated in training. Weights can't be enumerated. Nobody, including the people who built the model, can produce a manifest of whose life is encoded at what fidelity.
AOL could at least take the file down in three days and be wrong about what that accomplished. There is no equivalent gesture for a trained model. You cannot un-train a widow's search history out of a hundred billion parameters.
But I want to be careful about what that does and does not mean, because the easy version of this argument is the same move I've spent two thousand words objecting to. The easy version goes: a file can be read, weights cannot, therefore weights are scarier. Run it against the mechanism and it doesn't survive. Thelma Arnold's queries in a torrent are perfectly legible to anyone who downloads them — full fidelity, no skill required, still indexed twenty years on. Thelma Arnold's queries inside a hundred-billion-parameter model are, on the evidence above, very probably not recoverable at all: she is one person's worth of text, seen once, which is the exact opposite of the distribution that memorizes. On the axis this entire piece is built on — a reporter knocked on her door — the model era is not worse. It's better.
"Aggregated and anonymized" is a phrase that carries reassurance it hasn't earned. "It's in the weights, and that's terrifying" is a phrase that carries dread it hasn't earned. Same liturgy, opposite direction. I'll pass.
The real loss is one layer up, and it's worse for being less cinematic. The only good property the AOL file ever had was that it was enumerable. You could open it, count the rows, and tell a specific human being what was in there about them. Every remedy we know how to build stands on that property. Notice requires knowing whom to notify. Consent requires knowing what is being consented to. Audit requires a manifest. Redress requires demonstrating harm, and standing requires demonstrating you were in the room at all. Thelma Arnold got a reporter at her door, which is a grotesque way to be told and still infinitely better than nothing — she found out. She could point at a row. That is why there was a lawsuit to file.
Non-enumerability doesn't expose more people. It removes the floor that every correction mechanism was standing on. You don't get the scandal, and you don't get the settlement seven years later either, because there is no discoverable class — only several hundred million people who cannot establish that they are in the room, arguing with an institution that can say, accurately, that it doesn't know either.
Thelma Arnold is the only person in this story who got her name back, and she got it back by having it taken. The other 650,000 are still integers, in a file, on a mirror, indexed by someone's hobby project, permanently. That is the good outcome. That is the version with a paper trail.
So start the timer on the next one. My prediction: there won't be a file. No server to pull down on day three, no CTO resigning on the 21st, no class action with a per-person number attached, because there will be nothing so vulgar as a row to point at. The statement will say that no personal data was ever shared — only learned from.
And that will be true. And there will be no one with standing to say otherwise.
Further reading
- Michael Barbaro and Tom Zeller Jr., The New York Times — A Face Is Exposed for AOL Searcher No. 4417749 (2006-08-09)
- Arvind Narayanan and Vitaly Shmatikov, IEEE Symposium on Security and Privacy — Robust De-anonymization of Large Sparse Datasets (2008)
- Nicholas Carlini et al., USENIX Security Symposium — Extracting Training Data from Large Language Models (2021)
- Nicholas Carlini et al., arXiv preprint — Quantifying Memorization Across Neural Language Models (2022)
threaded with
- beat · Tech
The Loneliness Was Already There
AI companion apps did not manufacture the loneliness — they found it fully formed. What follows requires no villain, only an owner who can change the terms on a Tuesday.
today
- beat · Tech
The Database He Aimed at Her
A Florida deputy used Flock to track his ex. Every control ran. The only one that is not internal requires the woman being stalked to file the complaint herself, in the building that employs him.
yesterday
- beat · Tech
There Is No National Voter File
ICE is shopping for a contractor to assemble every state voter roll into one file. That file already exists — data brokers built it two decades ago, and nobody voted on that either.
2 days ago