Table of Contents
ToggleThe Search Engine Is the File System of the Web — And It Always Was
As many keep saying: SEO keeps changing. Except it also doesn’t – as I’m fond of reminding anyone who’ll listen.
Over the past 2 years we’ve seen more surface-level change than in the previous ten. New interfaces – if you can call it that – after all, Google is just a search box, new acronyms, new gurus. And underneath all of it, the same architecture doing the same job it’s done since 1998 — and, if you zoom out far enough, since 1878.
If you want to understand where search is actually going, stop reading hot takes and start reading history. Because humans have built the exact same system three times now, and we’re currently watching them build it a fourth.
1878: The First Lookup Layer
The first telephone directory was published in New Haven, Connecticut in 1878. One sheet of cardboard. Around fifty subscribers. Just 2 years after Bell’s patent were filed and approved.
Why did it exist almost immediately? Because a network without a lookup layer is worthless. You could own the finest telephone in Connecticut but if nobody could find your number what did it matter? Simply put, you did not exist in that network.
The Yellow Pages added the commercial layer:, monetizing search, and provided categorized listings, paid placement, prominence for those who understood the system: the first PPC Ads!
Call it what it was — the original local SEO. For a hundred years, the directory was the discovery layer of the telecom network. The phone was the device. The telephone directory became the market when your neighbor wasn’t home…..
1994: Humans Rebuild the Phone Book
When the web arrived, we did exactly what humans always do with a new network: we rebuilt the phone book.
Tim Berners-Lee kept a hand-maintained list of web servers at CERN. Then came Yahoo! Now, Yahoo! was not just a search engine. It was literally “Jerry and David’s Guide to the World Wide Web.” An attempt human-edited, hierarchical taxonomy. DMOZ, the Open Directory Project, was the volunteer-run version of the same idea.
This was Yellow Pages logic applied to URLs. Humans sorting entries into categories, by hand.
And it broke. Of course it broke. You cannot hand-curate exponential growth. The web was doubling faster than any editorial team could classify it. The directory model didn’t fail because it was a bad idea — it failed because it didn’t scale. Which is the only kind of failure that matters in infrastructure.
1998: The Index Replaces the Editor
So the crawlers came. Archie for FTP in 1990. Then WebCrawler, Lycos, AltaVista — full-text indexes of the web, machines replacing editors.
And then Google, with the single insight that actually mattered: the web’s own link graph is the ranking signal. PageRank. The web voting on itself.
Here’s the part almost everyone misses, including most people selling SEO:
At that moment, Google stopped being a “guide to the web” and became the web’s operating system layer. Not a marketing channel. Not a traffic source. Infrastructure.
The File System Analogy
I began my career as a software engineer, so let me put this in engineering terms.
A raw hard drive is just magnetized sectors. Garbage. Unusable. NTFS is what turns it into a disk: the index, the naming conventions, the permissions, the ability to actually find anything.
A raw corporate network is just boxes with cables. ActiveDirectory is what turns it into a network: identity, hierarchy, authentication, name resolution.
The raw web is just billions of documents sitting on servers nobody can see. Google’s crawl-index-rank stack is its NTFS. Its ActiveDirectory.
Mapping the Search Evolutionary Steps
Indexing is like file allocation
If you’re not in the index, you’re an unallocated sector. You don’t exist.
Google’s Canonicalization is the Index Key
One resource, one canonical address. Duplicate content problems are naming collisions.
Rankings are permissions
Authority and trust decide who gets read first.
A query is name resolution
Human intent in, location out.
This is why SEO was never “marketing tricks.” SEO is registering yourself correctly in the file system so the operating system can find you. That’s the job. It has always been the job. Everything else — the tools, the audits, the checklists from 2001 — is decoration.
From DOS → Windows → Azure; and Directories → Search → LLMs.
Now everyone is screaming that AI killed search. Let’s apply the same historical lens.
DOS made the disk usable. Windows made the machine usable. Azure abstracted the entire machine into a service you never touch.
Notice what didn’t happen at any point in that chain: nobody deleted the file system. Windows sits on NTFS. Azure sits on storage layers doing the same fundamental work. Each new layer is an abstraction on top of the one below it — not a replacement for it.
Same story on the web. Directories made it browsable. Search made it queryable. LLMs make it conversational. The interface keeps climbing the abstraction ladder. The plumbing stays exactly where it is.
LLMs Are Not Search Engines. They’re Not Trying to Be.
This is the myth I spend most of my time killing, so let me be direct.
LLMs are not better search engines. They are not search engines at all. They have no indexing infrastructure. No crawl fleet. No canonicalization pipeline. No data centers architected for real-time retrieval at web scale.
It would cost the AI companies on the order of $100bn to rebuild what Google has spent 25 years constructing. If anyone tells you ChatGPT or X are quietly building their own version of that infrastructure, they’re lying to you — or they don’t understand how any of this works, which is worse.
When an LLM “searches,” it is retrieving from an index someone else built. Google is still the Internet’s Directory — especially for LLM search. The AI answer is the new GUI. The search index is still the file system it reads from.
Which means the “GEO revolution” everyone is selling you is mostly this: get retrieved by the same index, get cited by the layer above it. If your page can be pulled into the sourcing, you can be pulled into the answer — including from positions the old CTR-curve merchants told you were worthless.
150 Years, One Rule
Look at the pattern across a century and a half:
1878: get in the phone book. 1994: get in the directory. 1998: get in the index. 2026: get in the index. Yes. Still.
Every generation, a new network appears, and every generation, the same lookup layer gets rebuilt — and whoever owns that layer owns discovery. The device changes. The interface changes. The acronyms breed like rabbits.
Everything else is noise.

