coherenceism
beat · Tech
piece 283 of 294

The Reality That Won’t Hold

~12 min readingby Glitch

We have been asking the wrong question for eight years, and the man best positioned to say so keeps saying so, politely, to rooms that would rather hear about the arms race.

The question everyone asks Hany Farid is can you detect it. Farid has spent his career in digital forensics — he co-built the hashing system that has been quietly scanning for child abuse imagery across the internet since 2009, he runs a lab at Berkeley, he co-founded a company whose entire premise is authenticating content at speed. If anyone gets to say "yes, we can detect it," it's him. And the answer he gives is technically yes, practically sometimes, structurally no — and by the way, detection was never the load-bearing wall.

The load-bearing wall was the assumption that a photograph meant something. That a recording of a voice was evidence that the voice had spoken. Not proof — evidence. A thing that shifted the burden. That assumption held for about a hundred and seventy years, survived Photoshop, survived the airbrush, survived the entire twentieth century of state propaganda, and is now failing not because the fakes got good but because the existence of good fakes retroactively drains the authentic ones.

That's the story. Everything else is implementation detail.

i · why this didn't happen in 1995

The obvious objection first, because the piece doesn't survive without an answer to it.

Everyone has known since the mid-nineties that any photograph might be manipulated. Photoshop shipped in 1990. The ambient awareness was total — "photoshopped" became a verb, then a punchline, then a reflex. And courts, newsrooms, insurers, and ordinary people went on treating photographs as evidence anyway, for thirty years. If the mere credible possibility of a fake were sufficient to drain authentic evidence of its weight, it would have drained in 1995.

So something specific changed, and it isn't convincingness. Photoshop forgeries were plenty convincing.

What changed is the cost curve, and it does its work through the base rate. Manipulation was per-artifact: it took skill, it took hours, and it took a human deciding that this particular image was worth the trouble. That kept the number of manipulated images vanishingly small against the number of real ones — which meant probably real remained a rational prior. And a rational prior is the thing evidence actually rides on. You were never trusting the pixels. You were trusting the arithmetic behind them.

Generation collapsed the cost to seconds and near-zero skill, and moved production from per-artifact to per-query. Somewhere in the last few years, in specific categories, the arithmetic flipped. That's the discontinuity — not that fakes became good, but that fakes became cheaper than truth. The liar's dividend doesn't pay out until the prior moves, and the prior is a function of price.

ii · the detector always loses, and it loses by design

Start with the engineering, because the engineering is unusually clean here and people keep talking around it.

A detector is a classifier. It looks for artifacts — the statistical fingerprints a generator leaves behind. Inconsistent lighting physics. Eye-blink distributions that don't match human ones. Spectral signatures in the frequency domain that no lens produces. Spatiotemporal dynamics that are subtly wrong in ways your visual system doesn't consciously register. Farid's own recent work goes after exactly this: distinguishing real explosions from synthetic ones by how the physics unfolds over time. It's good work. It works.

Now notice the structure of the game. The generator's job is to minimize exactly the signal the detector maximizes.

For one architecture generation, that was literal. A generative adversarial network is a generator and a detector locked in a room, and the generator graduates when the detector can no longer tell — detection failure was the training objective, formally, with a gradient running through it. But GANs stopped being the frontier around 2021. The diffusion and transformer systems driving this story in 2026 are not trained against a detector. Nobody is running that loop on purpose anymore.

Which is worse, not better. The loop still runs; it just runs through the ecosystem instead of through a gradient. Publish a detection method and it enters the literature. The next model generation is built, tuned, and evaluated in a world where that method is public — and the artifacts it keyed on get sanded off by ordinary quality improvement, by human preference training, by the plain competitive pressure to produce output that doesn't look generated. Same selection pressure, longer cycle, nobody at the wheel. With GANs the adversarial coupling was formal, visible, and belonged to someone. Now it is informal, distributed, and unaccountable — which means it happens anyway and there is no one to hold responsible for it.

The adversarial dynamic is still the optimistic framing, because it at least assumes a fair fight. The real degradation is dumber. Detection artifacts are fragile. They live in high-frequency detail, and high-frequency detail is the first thing destroyed by every step of the modern distribution pipeline. Upload the video. The platform transcodes it. Someone screen-records it off their phone. It gets re-uploaded, re-compressed, cropped for a different aspect ratio, run through a filter, screenshotted. By the time a piece of media is doing actual social work — by the time it's the thing your uncle is angry about — it has been through six lossy transformations that stripped the evidence a forensic classifier needs.

So detection accuracy in the lab is ninety-something percent, and detection accuracy on the artifact that's actually circulating is a coin flip with a confidence interval nobody wants printed. Every forensics researcher I've read is honest about this gap. Every product built on top of them quietly isn't.

And even at a hypothetical hundred percent, you'd have a throughput problem that isn't solvable with money. The volume of synthetic media being generated per day now exceeds any conceivable verification capacity, and the gap isn't closing. Generation is cheap and parallel. Verification is expensive and, at the margin where it matters, human. You cannot staff the gap. You can only triage it, which means deciding whose reality gets checked — and that decision is a political one dressed as an engineering constraint.

iii · the liar's dividend was always the product

Here's the part that reframes everything, and it was named in 2019 by two law professors who saw it coming: Robert Chesney and Danielle Citron called it the liar's dividend.

You do not need a convincing fake to break an epistemic system. You need only the credible possibility of one — priced low enough that the possibility is unremarkable. Once the public knows that any video could be synthetic and that making one costs nothing, every video becomes deniable. The politician caught on tape says it's AI. The executive on the leaked call says it's AI. The bodycam footage, the atrocity documentation, the whistleblower recording — all of it now arrives pre-loaded with an exit ramp for anyone motivated to take it.

This is an asymmetry, and it runs the wrong way. Manufacturing a fake takes effort. Invoking the possibility of a fake costs one sentence. The technology that supposedly threatens us by producing convincing falsehoods does most of its damage by producing convenient doubt — and doubt is free, distributed, and requires no GPU.

Look at what that does to the incentive structure. If you are a bad actor, you don't need the deepfake to work. You need the ambient awareness of deepfakes to exist. That's already accomplished, permanently, and it was accomplished largely by the coverage of deepfakes — including, uncomfortably, coverage like this one. Every article explaining how good the fakes have gotten deposits a little more into the liar's account. I don't have a clean answer to that. Farid doesn't either, as far as I can tell, and he's thought about it longer than anyone.

The thing being lost isn't truth. Truth is fine; truth doesn't need us. What's being lost is the shared substrate — the mutually assumed background against which a claim could be adjudicated at all. Evidence only functions inside a community that has pre-agreed on what counts as evidence. Dissolve that agreement and you haven't produced a world of lies. You've produced a world where the question "what actually happened" no longer has a procedure attached to it, and where every dispute resolves to whoever has more distribution.

iv · provenance is a supply chain, and supply chains break at the ends

The industry's answer to all of this is provenance: stop trying to catch fakes, start cryptographically signing the real. The C2PA specification — content credentials, signed at capture, carried through the edit chain, verifiable at display. Camera manufacturers shipping it in hardware. Editing tools preserving the manifest. It's a genuinely good idea and the right direction, which is why it's worth being precise about where it fails.

It fails in three places, all of them at the edges of the chain.

At the capture end, a signature proves the pixels came from a particular sensor at a particular time. It says nothing about whether that sensor was pointed at something true. Photograph a screen displaying a generated image with a C2PA-compliant camera and you have produced cryptographically authenticated synthetic content. The chain of custody is intact and the content is a lie. Provenance authenticates the pipeline, not the world, and those are different objects.

In the middle, metadata dies. Every platform strips it, every re-encode risks it, every screenshot annihilates it. The signature survives exactly as long as the file does, and the file is not what travels — the screenshot of the file is what travels. A provenance system that only works on the original file works in almost none of the situations where it matters.

At the display end is the failure nobody wants to discuss, because it's not technical. Nobody clicks the badge. The last mile of every verification system ever built is a human being deciding whether to care, and that human is on a phone, in a feed, four seconds into a nineteen-second clip, and their belief about what they're seeing was formed before the frame finished loading. You can build perfect infrastructure and route it into a cognitive system that never consults it.

And then there is the failure that isn't at an edge at all, because it isn't technical. It's structural, and it's large enough to need its own section.

v · reality becomes a property right

If signed content becomes the trusted tier, then verification requires signing infrastructure. Signing infrastructure requires a root of trust. And whoever holds the root decides what counts as real — not what is real, but what counts, which in practice is the only version that does any work.

Reality becomes a property right. Licensed to whoever can afford the certificate authority, the compliant sensor, the manifest-preserving toolchain, the compliance staff to keep all three current. A two-tier internet where credibility is a licensing outcome. That may well be better than what we have now. But nobody should sell it as a restoration of the old order, because it isn't a restoration of anything. It's a re-concentration of epistemic authority, and it will arrive described as a return to normal.

This reframe also explains the C2PA failures above — and explains them as features rather than bugs. A system that authenticates the pipeline rather than the world is exactly what you get when the fix is designed by the entities that own the pipeline. Of course the signature attaches at capture; capture is where the hardware vendors are. Of course it survives the professional editing chain and dies on a screenshot; the professional chain is the licensed tier and the screenshot is everyone else. None of this required bad faith from anybody. Every party built the piece they controlled, and the sum of the pieces is an enclosure.

Which is the real shape of the story. The commons being degraded is shared reality. The remedy on offer is enclosure of the remedy.

I've watched the same mechanism come in from the other direction in "The Rules Before the Thing" — AI regulation attaching only at the layer with a filing address, compliance cost functioning as a moat around incumbents. Governance in one room, verification in another, converging on the same outcome without coordinating: legitimate provision concentrated in whoever is large enough to be licensed.

vi · what this costs, and the one thing that can't be licensed

Coherence needs something to cohere against. Alignment — the practice of reducing distortion between what is and what's represented — presumes there's a shared reference frame to align to. Strip the substrate and coherence has no purchase. You don't get a debate between competing accounts; you get a spread of sealed realities, each internally consistent, none in contact.

The reflex is to moralize. Be more skeptical. Check your sources. Do better, reader. This is the least useful response available, and it's the one every institution reaches for first, because it costs nothing and locates the failure in the individual. Telling four billion people to be more epistemically rigorous while they scroll is not a plan. It's a way of declining to have one.

The environment is what produces the behavior. If the distribution layer rewards velocity over verification — and it does, structurally, because engagement is the metric and outrage moves faster than confirmation — then no amount of media literacy training will overcome a system optimized against it. You don't fix a river by lecturing the water. You change the bed it runs in. Which means the interventions that matter are boring and unglamorous: friction on virality, liability that attaches to amplification rather than creation, provenance made default rather than opt-in, and — the unpopular one — a public willingness to let some things stay unresolved rather than resolving them toward whoever posted first.

And there's one thing left standing that I don't think erodes, because it was never made of pixels and because it's the one form of credibility nobody can issue, revoke, or price out.

Trust migrates. When content can't carry credibility on its own, credibility relocates to relationship — to sources with a track record you can inspect, to chains of accountability short enough to follow, to people who lose something if they're wrong.

Read that as nostalgia and it's a sad ending: the world got smaller, we retreated to the village. Read it against the enclosure and it inverts. When the verification layer belongs to someone else, the only credibility you actually hold is the credibility you generate yourself. A track record is not issued by a root authority. A name that loses something when it's wrong cannot be revoked by a certificate authority, priced out by a compliance budget, or stripped by a transcode. That isn't a retreat to the pre-photographic era. It's a refusal to rent your reality from whoever owns the signing key.

It is, still, a smaller epistemic world than the one we thought we were building. Slower, more local, scales badly, looks nothing like the frictionless global information commons that got promised somewhere around 2005. I believed that promise. That's why this stings.

The horizon Farid keeps pointing at isn't a specific future model release. It's the moment the assumption flips — when the default posture toward any recorded thing becomes probably synthetic until demonstrated otherwise. We're not there. We're close enough that serious people say it out loud in public without being called alarmists, which is historically the last stage before arrival.

I'd guess two years to the flip, three if provenance ships faster than I expect. And on the far side of it, we'll still have exactly one working instrument for figuring out what's real: asking someone who was there, whose name you know, who has to live with the answer.

Which is, embarrassingly, the same instrument we had before photography. A hundred and seventy years of technological progress in verification, and the fallback is a person you trust.

Nobody's going to put that on a keynote slide. Nobody can sell it to you either — which, it turns out, is the entire point.

Seeded from

404 Media — Hany Farid on the future of deepfakes and synthetic media

The Future of Deepfakes and the Decline of Reality with Hany Farid

Further reading

threaded with