If it is not in the AI, it doesn-t exist

          

               (( ПЕРЕВОД ТЕКСТА: "НЕТ В ИИ — ЗНАЧИТ, НЕТ ВООБЩЕ", http://proza.ru/2026/09/09/1030 ))


When the web became an ordinary tool, a rule took shape: if you don't have a website, you don't exist. The formula was aimed first at businesses, at service organizations, at anyone in the business of providing information — not at private individuals. It was already imprecise back then. Plenty of things existed perfectly well without a website; they had simply stopped competing in one particular attention market. The formula held for commercial discoverability and failed for existence.

Today generative models are increasingly replacing the search engine, and the same formula is back: if it's not in the AI, it doesn't exist. The claim has moved closer to the truth than it was on the previous layer. But the degree is less interesting than the mechanism. What changed is not how much gets filtered out — it's the nature of the filtering.



A LIST ADMITS ITS OWN INCOMPLETENESS — PROSE DOESN'T


A list of ten links tells you it is incomplete. It has a result number, a second page, an edge. Users might not have scrolled past the first screen, but they knew there was something to scroll to. INCOMPLETENESS HAD A SHAPE, and you could see it with the naked eye.

A SYNTHESIZED ANSWER CARRIES NO SUCH ADMISSION. It looks equally finished whether it rests on two hundred sources or on three. No rank, no counter, no edge.
This is the real hardening of the old formula. Not "if it's not in the AI, it doesn't exist," but "ABSENCE HAS STOPPED HAVING A SHAPE." The web made the gap visible. The generative answer makes it invisible.



HOW MUCH HAS ALREADY HAPPENED


The 2026 numbers are specific enough. According to SparkToro's study of Similarweb clickstream data, 68.01 percent of Google searches ended without a single click, against 60.45 percent in 2024. The share of searches producing at least one click fell 9.51 percentage points in two years — nearly a quarter off the previous level.

A note on method is mandatory here, because the number is a pretty one and gets pulled out of context eagerly. The study covers the United States only, and only January through April of 2026; it assumes roughly a two-to-one ratio of mobile to desktop searches, and it excludes searches made inside the Google mobile app — where, by the authors' own estimate, the zero-click share may be higher still.
Google said at its I/O 2026 conference that AI Mode had passed a billion monthly users and that query volume in it is more than doubling every quarter.

The losses are distributed regressively. An Axios analysis of Chartbeat data found that the smallest publishers lost roughly sixty percent of their search referral traffic over two years — about three times steeper than the decline at large ones. Whoever needed that traffic most lost the most of it. That isn't a side effect; it's a property of the mechanism. What survives is whatever already had the resources to survive without search traffic.



THE COUNTERMOVEMENT: THE WEB IS CLOSING ITSELF


There is a second half to this that gets noticed less often. It isn't only that models can't reach the content. The content is actively shutting the models out.
Cloudflare, which carries roughly twenty percent of the web, has blocked AI crawlers by default on new domains since July 2025. On July 1, 2026 the company announced the next step: crawler traffic is split into three categories — Search, Agent, and Training. As of September 15, 2026, on pages carrying ads, the new defaults will let Search through and block Training and Agent. Crawlers that fuse all three functions into a single identifier will be blocked outright.

The scope matters and tends to get lost in the retelling: the new defaults apply to new customers, to new sites added by existing customers, and to all free-tier customers. Paying customers with configurations already in place keep them.

The company's stated basis: by its own figures from early June 2026, training crawlers accounted for 50.6 percent of crawl traffic against 10.7 percent for search bots, and more than half of all AI crawler requests were re-downloading pages that hadn't changed since the last visit. The old Pay Per Crawl scheme, which charged per request, has been replaced by Pay Per Use, under which a publisher is paid when the content is actually used in a model's answer.

What emerges here is a failure mode with no web-era analog. The Agent category is the crawler that fetches a page so a model can answer a person in real time, and it will land in the same blocked bucket as the training crawler. Which means a site that leaves the default alone will stay visible in search and disappear from assistants' live answers. For a publication living on ad impressions that is probably the right behavior. For a store, a booking service, or a service provider it is precisely the wrong one — because the person who used to look them up through a search engine now asks an assistant.

And there is no feedback loop. In search optimization, a drop in rank was measurable. Here an organization can make itself nonexistent while trying to protect itself, and never find out.



THE SHAPE OF THE BLINDNESS IS NOT RANDOM


An inventory of what is effectively out of reach for a language model doesn't form a set of random holes. It forms a consistent profile.

Legacy encodings: Windows-1251, KOI8-R, and the like. The damage falls disproportionately on the old non-English web, the post-Soviet part of it in particular — which is to say, on exactly the layer that is least duplicated anywhere else.

Authentication and paywalls: closed forums, corporate portals, subscription archives.

Closed platforms: a large share of contemporary discussion lives in messaging apps and private communities.

The nontextual: unrecognized scans, images, video, audio, tables trapped inside pictures.

The local and the offline: municipal archives, library catalogs, paper, oral transmission.

The explicitly blocked: through robots.txt, or behind paid crawler access.
A limit built into the retrieval tool itself: a model's reach ends where the search index ends. A filter on top of a filter.


The bias points the same direction every time — toward the English-language, the recent, the well-formatted, the commercially maintained, the structurally simple, the open, and the textual.



THREE KINDS OF NONEXISTENCE


"If it's not in the AI, it doesn't exist" is worth taking apart, because it fuses three things of entirely different natures.


LEVEL ONE: unreachable.
Encoding, authentication, blocking. Hard exclusion, and the direct analog of the web era: no page.


LEVEL TWO: unretrieved.
The material exists and is readable in principle, but it didn't make the results, didn't rank, didn't earn back the latency and the cost of the fetch. Familiar from the search era, but harsher. Readers used to be able to keep scrolling. Now the gap between "returned" and "not returned" is binary.


LEVEL THREE: DISSOLVED.
No web-era analog exists. The material was in the corpus, made it into context, and vanished in the synthesis. Whatever is rare, singular, contrary to consensus, or locally specific gets averaged away even when it was physically there.


LEVEL THREE is the most interesting and the LEAST INSTRUMENTED. The first two leave logs, status codes, analytics. The third leaves nothing — and that includes the model itself, which cannot show you where the missing thing went.



THE WEB IS REFORMATTING ITSELF FOR A NONHUMAN READER


Search engine optimization is mutating into optimization for generative engines. The difference is not cosmetic: the old kind was gameable but auditable — you could measure your rank. A GENERATIVE ANSWER HAS NO RANK. VISIBILITY IS PROBABILISTIC, it does not reproduce from one session to the next, and it cannot be verified.

A strong incentive to optimize combined with no reliable way to measure the result produces two things: a market for snake oil, and text reformatted for machine legibility instead of human reading. This has already started — see llms.txt and the growing layer of structured markup addressed to models rather than to people.

For thirty years the web was formatted for the human eye, with machine indexing bolted on the side. That priority is now inverting.



THE OBJECTION YOU HAVE TO RAISE AGAINST YOURSELF


The old version of the formula was an overstatement, and this one is too. The temptation to announce an age of total invisibility is strong, but historically such announcements have come true only in part.

The difference in degree is real, though, and the mechanism behind it deserves to be named precisely: the intermediary stopped listing and started synthesizing. A list confesses its own incompleteness by the mere fact of being a list. PROSE CONFESSES NOTHING.



WHY THIS PROBABLY WON'T GET FIXED


The obvious remedy is coverage instrumentation — an indication of what was never consulted, an estimate of how complete the sample was, an honest equivalent of those blue links admitting their own limits.

The odds that anything like it becomes an industry default by 2028 should be set low, on the order of twenty percent. The reason is not technical. Admitting a gap directly lowers a product's perceived authority, and perceived authority is what is being sold.

It is the same skew visible across the industry: THE SPEED OF PRODUCTION HAS OUTRUN THE SPEED OF VERIFICATION, and THE INSTITUTIONS THAT COULD RESTORE THE BALANCE HAVE NO ECONOMIC REASON TO.

The costs of this new layer land the way the rest of this technology cycle's costs land — the same buildout that drove up memory prices and electricity bills. The ones who pay are small, non-English-speaking, noncommercial, and old.
With one difference worth registering separately. Expensive RAM will get cheaper if the investment cycle turns down. Here there is no such scenario. A market collapse will not return to the index what fell out of it, and it will not restore traffic to publications that will have folded by then. This layer is irreversible not because it is catastrophic, but because there is nothing in it left to roll back.











Konstantin BGDT Privalov,
Kts, 2026-09-09


Рецензии