(In today’s internet, clicking a link and actually seeing the original webpage has become a luxury.)
It’s about 11 degrees in Shanghai today. The clouds are thin, but sitting by the window still feels a bit chilly. Back at my computer, I casually clicked on a thread summarizing a historical event sent by a friend. Upon clicking the source link inside, the screen punctually popped up that blood-pressure-raising prompt once again—”Hmm…this page doesn’t exist. Try searching for something else.”
Even more absurdly, right above this 404 error hovered a highly mocking, bold line of text: “Don’t miss what’s happening. People on X are the first to know.”
The Severed Nerve Endings
I don’t know if you’ve had similar experiences recently, but in my daily life, these dead ends are appearing with increasing frequency.
Yesterday, to verify an original statement from a few months ago, I used an OSINT (Open Source Intelligence) tool specifically designed to fish for data in major caches. The result? Even the caches are failing on a massive scale; running it yielded nothing but red error messages. Because ever since X began fully restricting unlogged access, even normal search engine crawlers struggle to completely index its webpages.
There is actually a very serious technical phenomenon behind this, called “Digital Decay” or “Link Rot.” I recently looked at a set of monitoring data released by the Pew Research Center in 2024, which is fascinating. They found that up to 38% of webpages that existed in 2013 have completely vanished today. Even looking at a closer timeframe, a quarter of all webpages generated in the decade between 2013 and 2023 are currently inaccessible.
(Looking at this downward curve, the speed at which webpages are disappearing is far more drastic than we imagine.)
But X has absolutely hit the fast-forward button in this wave of internet amnesia.
In the past, a tweet was an open node. Breaking news, an industry leader’s casual rant, or even a functional snippet of open-source code—you could just copy the link, paste it into a blog or chat group, and anyone could click to view it directly. And now? If you don’t log in, or if you happen to be on a new device without a saved login state, all you see is a cold wall.
To put it bluntly, that phrase “People on X are the first to know” carries an exclusive subtext: “If you’re not on X, you don’t deserve to know anything.” A platform that once touted itself as the “global digital public square” has essentially locked its doors, standing on the wall shouting at passersby: “The view is great in here, hurry up and hand over your phone number and privacy data to register an account ╮(╯▽╰)╭.”
Ostensibly Anti-Crawler, Actually Anti-Passerby
If you dig into Elon Musk’s previous explanations for these kinds of changes, his reasoning has always been quite stern—to prevent “data pillage.”
Put plainly, large language model (LLM) companies are frantically scraping high-quality corpus data from the platform to train their AI, and he is unhappy about it. Why should my data be given to you for free? Thus, the platform immediately raised the highest level of alert.
What does a normal anti-scraping strategy look like? You could implement more granular API rate limiting, set dynamic threshold limits for unlogged user access frequencies, utilize browser fingerprinting, or even introduce slightly milder CAPTCHA mechanisms. However, X chose the lowest-cost and most brutal approach: a one-size-fits-all, forceful login wall.
Any request without a valid Cookie, even if it’s just requesting the basic Meta information of a page (like the title and description), will be directly intercepted and forcibly redirected to that welcoming page that simply says “Let’s go.”
(Behind this blue “Let’s go” lies a fiercely aggressive “pay-to-pass” logic regarding data.)
This reflects an extremely arrogant product logic. They presume that every unlogged visitor is a machine coming to steal data. To intercept a small fraction of high-concurrency malicious crawlers, they do not hesitate to shut the door directly on the vast majority of real human users clicking through external search engines or article links. This technical trade-off is like sealing up all the doors and windows of a house with concrete just to catch a few flies.
Did you stop OpenAI’s crawlers? Honestly, they probably bypassed it long ago using hundreds of thousands of residential IP proxies and simulated login accounts. The only ones truly blocked are the ordinary people who occasionally want to read a native link from the news but give up because they forgot their passwords.
Everyone is Building Walls, Who is More Decent?
To put it harshly, it’s not just X doing this nowadays. The entire internet environment is undergoing a severe “privatization” movement.
Take Reddit, for instance. Also envious of AI companies using their data to train LLMs, they raised their API calling fees to a level that general developers simply couldn’t afford, directly shutting down a host of excellent third-party clients like Apollo. But even with Reddit, at least when you open a post in a desktop browser, it lets you see a few top comments. Although the page is full of pop-ups urging you to download the app, it still retains a tiny shred of “publicly visible” decency.
Looking at China, people are actually much more familiar with this playbook. An impenetrable high wall has long been built between a certain national-level social APP and a veteran search giant. You want to find an in-depth article from a certain official account using a general search engine? Absolutely impossible. All information flows are tightly cut off within their respective clients, forcing you to download, register, and labor for their Daily Active User (DAU) and Monthly Active User (MAU) metrics.
(Today’s various social platforms are essentially locking themselves in cages.)
But there is one problem you must understand: the direct consequence of this infrastructure-level closure is the fragmentation of the internet’s underlying architecture.
When we talked about Hypertext in the past, its core lay in the “link.” A single URL could transport you seamlessly from one corner of the world to another; this is the soul of the World Wide Web. Now, more and more links lead to “This page does not exist” or “Please download the app to read the full text.” When links no longer lead to knowledge and content, but rather to a series of login walls, that idealistic internet exists in name only. It is regressing into countless disconnected, self-congratulatory local area networks (LANs).
Some Things I Sometimes Think About
Following this logic, I sometimes wonder, how will the digital history of our generation be preserved?
Twitter or blogs from a decade ago were essentially the “first draft of history” for human society. The release of a policy, the real-time recording of a social movement, or even an SOS signal during a sudden disaster—these naturally settled into that massive network. Any subsequent researcher could cite them and corroborate their arguments through a simple URL.
Now, these URLs are turning into a pile of dead code at a visible speed.
If a digital historian in the future wants to study a certain internet controversy from the early 2020s and clicks on a reference link in a news report from that time, they might only face a vast wasteland of inaccessible redirect error pages. Those once-vivid public dialogues, filled with intense emotion, evaporate instantly into thin air just because a certain company changed its server rules.
A few days ago, over tea with a friend researching internet anthropology, we discussed this, and he was actually very pessimistic. He said that today’s tech giants no longer view the data on their platforms as “humanity’s shared memory”; it is solely seen as a string of numbers waiting to be monetized on a balance sheet. The more data is lacking in the LLM era, the higher these giants will build their walls.
Maybe I’m overthinking it. Perhaps in four or five years, everyone will realize that this highly closed ecosystem simply cannot attract fresh blood, or perhaps something new based on decentralized protocols (like the current AT Protocol) will emerge to reconnect these artificially created islands. Some things really do need to be destroyed to their absolute limits before a true rebound can occur.
After typing these words, I looked up out the window.
The clouds seemed to have parted slightly, but the sun still hasn’t quite broken through. The coffee on the desk has completely gone cold, with a thin layer of oil forming on the surface. Never mind, I’ll go brew another cup.
References:
- When Online Content Disappears – Pew Research Center
- Twitter has apparently disappeared behind a login wall | SpaceBattles
- All Twitter content seems to be behind a login wall today – Hacker News
- Data Decay Rate Statistics: 20 Critical Facts Every GTM – Landbase
—— Lyra Celest @ Turbulence τ.
