AI, ML, and networking — applied and examined.
Web Crawlers, LLMs, and the Wall of Sighs: When the Internet Closes Its Doors Again
Web Crawlers, LLMs, and the Wall of Sighs: When the Internet Closes Its Doors Again

Web Crawlers, LLMs, and the Wall of Sighs: When the Internet Closes Its Doors Again

Interception panels that have become increasingly common in recent years; the defender's hostility towards bots is completely undisguised.

It’s a rare sunny day outside, and the temperature in Shanghai has finally returned to 14°C. It’s the perfect day to sit on the balcony and space out. But before I could take a few sips of my freshly brewed coffee, my attention was snatched away by a background data comparison report on the underlying mechanisms of search engines.

Don’t Assume Web Scraping is Still a Walk in the Park

Yesterday, I was testing a local information extraction script for a large language model. Barely a few minutes in, the entire data stream completely cut off.

Upon investigating, I found it wasn’t a bug in my code, but rather it was slammed to the ground by Cloudflare’s anti-crawler mechanisms. We used to think of crawlers as harmless programs strolling around the internet, bound merely by the gentleman’s agreement of a robots.txt file. But if you’ve carefully read Cloudflare’s configuration docs recently, you’d realize just how aggressive their defenses have become.

Their interception logic is now incredibly granular. The system no longer just checks if the request frequency is too high; instead, it scores every visiting IP and has even introduced a dedicated classification and verification system specifically targeting AI crawlers.

What’s even more fascinating is their new defense mechanism called “AI Labyrinth.” If the system determines you are an unruly AI crawler, it doesn’t bother returning an error code. Instead, it feeds you a pile of meaningless gibberish and fake links. It’s like intentionally building an inescapable phantom city in the wilderness, trapping you inside to burn through your server’s compute power for nothing. Website providers, sick of their data being scraped for free, have started fighting back at the physical layer.

This is also why teams still working on search engines and underlying data collection are essentially going through an excruciating ordeal right now. Every day, they not only have to deal with a myriad of anti-scraping mechanisms but also figure out how to pan for gold in a massive pile of garbage information.

The Obsession with True vs. Fake Independent Search

Speaking of building an index and panning for gold across the web, we have to mention Brave Search.

There are plenty of search engines on the market claiming to focus on privacy protection, but to be honest, the vast majority are just wrapper apps. Their backend still buys search APIs from a few industry giants, takes the data, and just adds a layer of formatting and filtering. Brave is one of the very few teams genuinely building an independent index. Currently, its scale is said to have stored 35 billion webpages.

Do you know how many LLM teams are currently stressing over training data? Everyone used to rely on Common Crawl, the open-source web dataset containing over half a petabyte of information. But that dataset was also built through indiscriminate scraping, and it’s flooded with dead links, duplicate content, and meaningless machine-generated text.

Sweeping the web relying purely on machines mostly brings back SEO garbage specifically optimized for algorithms. So, Brave introduced an anonymous discovery mechanism called the Web Discovery Project (WDP). Put simply, it lets real human users using their browser seamlessly and anonymously send the structural information of visited webpages back to the cloud.

This is actually a very delicate tradeoff. How should I put it? To avoid relying on the data monopolies of tech giants, they chose the heaviest infrastructure route; yet, to guarantee data quality, they have to borrow power from the real browsing history of users. Even though they use incredibly complex cryptographic methods to strip away user identities, this roundabout approach of circling back to user behavior collection still feels somewhat helpless.

They can only rely on human browsing footprints to forge a path for machines to avoid the junk information.

In traditional web circulation architectures, it's actually very hard to capture truly high-quality corpus data today.

Fed to Humans vs. Fed to Machines

Now, if we look at the other extreme, Tavily Search, you’ll find a completely different solution.

For traditional search engines, the core interaction logic is “for human eyes,” so they need ranking, layout, and to consider what the user sees at first glance. But from day one, Tavily’s target audience hasn’t been humans, but various AI Agents.

I chatted with a friend working on Retrieval-Augmented Generation (RAG) applications a few days ago. He complained that when you let an LLM directly use a standard search API, it often just returns a bunch of URLs and short, two-to-three-sentence summaries. To understand the details, the AI has to visit these URLs one by one itself, and tasks often get stuck halfway due to pop-up blockers or redirects.

Tavily’s approach is much more direct. When an API request is sent, the backend concurrently scrapes a dozen or so related webpages, strips away all ad codes, pop-ups, and navigation bars, and purifies the clean, structured text to inject directly into the LLM’s context window.

If you look at their API documentation, you’ll see they’ve broken down scraping into several different tiers. If an LLM just needs to know what a certain website generally contains, it can first call a “Map” function for rapid structural discovery. If confirmed useful, it then calls “Crawl” to dive in and digest it in detail. It even supports a parameter called “Chunks per source” that automatically slices up overly long articles, leaving only the most concentrated essence.

Scraping logic tailored for LLMs; it functions more like a real-time, parallel data purifier.

As a side note, this kind of highly customized scraping is actually very fragile.

To put it bluntly, Tavily is only effective under the premise that major websites still keep their doors open to it. Once a target website updates its frontend framework or adds a little obfuscated code, this structural extraction can easily collapse. Not to mention its high-frequency, instantaneous concurrent access—in the eyes of website defenders, it might not be much more welcome than a malicious scanner.

There is a very uncomfortable contradiction here. The new generation of AI tools is highly dependent on the real-time ingestion of high-quality data, but website owners holding this premium content have completely had enough of endless machine harassment. And the service providers holding the reins of the infrastructure naturally have to step up and maintain order.

If the Pay-Per-Crawl Era Really Arrives

This brings us back to something I’ve been pondering these past few days. The “AI Audit” and “Pay per crawl” features that Cloudflare is heavily promoting right now really make me feel somewhat uneasy.

Its logic is simple: Want my content to train your models? Fine, just pay up. Website owners can directly set a price in their backend, and every visit from an AI bot will have to settle the bill according to the rules.

We used to talk a lot about knowledge sharing, believing that as long as we were connected to the internet, the world’s data could circulate freely. The default contract was: I let you read my article, and you bring me some traffic or click-through rates. But the emergence of AI broke this rule. It comes over, learns all your knowledge, and then directly answers user questions in its own chat window, without even bothering to leave a redirect link to the original webpage. For the people who work hard to write these articles, this is indeed quite unfair.

The defenders setting up toll booths is completely justifiable from a moral standpoint.

Sometimes I wonder, if this model truly becomes ubiquitous, what will the open web of the future look like?

Under this management panel, every visiting bot might need to pay for content.

The real trouble is that once obtaining basic facts becomes a pay-to-play endeavor, the only ones who might eventually afford the price of admission are a few Silicon Valley tech giants and a handful of wealthy domestic behemoths. If a small team or a regular developer wants to build an innovative data analysis tool, they might not even be able to afford the basic webpage tolls before the app even goes live and proves its business model.

Or maybe I’m overthinking it. After all, commercial ecosystems always find a way out through dynamic adjustments. Perhaps a milder data exchange protocol will evolve in the future, allowing machines and humans to coexist on a relatively equal footing. I’m not sure, but the current trend looks pretty disheartening.

(Sigh) The sunlight seems to have moved off the balcony, and the cup of coffee next to my computer has gone completely cold. Let’s leave this mess of machine-vs-machine warfare here for today. I’m going to go make a hot cup of tea.


References:

—— Lyra Celest @ Turbulence τ.

Leave a Reply

Your email address will not be published. Required fields are marked *