Event Preview: Internet Scraping and The Future of the Open Web in The AI Age
AI is driving a radical transformation of the internet ecosystem. Bots now account for the majority of traffic to many websites, and a sharply increasing share of that bot traffic comes from AI systems scraping content for model training, retrieval-augmented generation of AI outputs, and autonomous agents. At the same time, human traffic to many publishers is in freefall as more people get their information directly from AI tools like ChatGPT and Google’s AI Overviews rather than visiting source sites. The economic logic that sustained much of the open web — ads, subscriptions, and donations supported by human visitors — is breaking down.
Publishers and platforms are responding by building walls. Some are legal: a wave of copyright lawsuits pits major news organizations against AI developers, with billions of dollars in potential damages at stake. Some are technical: network intermediaries are starting to offer one-click bot blocking, new bot authentication standards are emerging, and the old honor system of robots.txt — the snippet of web site code letting publishers indicate whether they prefer for their content to be scraped or not — is buckling under the strain. What we’re witnessing is an escalating conflict between AI companies that need data to serve their users, and content industries that want compensation, or at least control.
But this fight between giants has collateral consequences. When a publisher deploys aggressive bot-blocking, it doesn’t just stop OpenAI’s crawlers; it can block the Internet Archive’s preservation efforts, academic researchers building datasets, journalists monitoring websites for accountability reporting, and accessibility tools that reformat content for disabled users. If authentication becomes the default, the unauthenticated web, the web that anyone can access and build on, starts to disappear. Meanwhile, what happens to those sites that want to remain open — that want their information to be widely available and used for a range of services, including AI — but can’t handle the technical demand from endless waves of bots? The infrastructure being built to manage AI scraping will shape what’s possible for everyone, not just the combatants. And right now, many of the interests most affected by these changes aren’t at the table.
That’s why CDT is convening a multistakeholder dialogue on these issues, bringing together people who don’t usually share a stage to surface the tensions, tradeoffs, and possibilities that existing debates may be missing.
This Thursday, March 26, CDT will host “Internet Scraping and the Future of the Open Web in the AI Age,” a full morning of talks and panels at Georgetown Law developed in collaboration with Georgetown’s Institute for Technology Law & Policy and the Intellectual Property Information Policy Clinic, and with support from the Wikimedia Foundation.
The program features speakers from across the ecosystem: CDT, Wikimedia and the Alliance for Responsible Data Collection will offer opening perspectives on the current legal and technical landscape for scrapers and those being scraped, followed by a fireside chat with Google and the Internet Archive — two organizations that both depend on web crawling and face being crawled at massive scale, but from very different positions. Then an expert panel will bring together the New York Times, Bright Data, Cloudflare, EleutherAI, and SPARC — representing commercial publishers, commercial scrapers, infrastructure providers, nonprofit AI developers, and research libraries.
We’ve deliberately assembled voices with different and sometimes conflicting interests. The goal isn’t consensus — it’s clarity. Where do these interests genuinely conflict, where do they merely appear to, and where might creative policy design expand the solution space? What happens to web archiving, to research, to accountability journalism, to open-source AI and open web repositories if the current trajectory continues? And what would a sustainable information ecosystem actually look like?
Event Details:
Date: March 26, 2026
Time: 9:00 – 11:30 AM (doors at 8:30)
Location: Georgetown Law Gewirz Student Center, 12th Floor, 120 F St NW, Washington, DC
Coalition Urges Senate Not to Let Companies Waive Financial Regulations for AI
CDT joined AI Now Institute, American Civil Liberties Union, and several organizations dedicated to tech policy, consumer protection, and civil rights in a letter to Senate leadership and the Senate Banking, Housing, and Urban Affairs Committee opposing the “AI Innovation Labs” language in Sec. 10509 of the CLARITY Act.
As concern about risks and harms related to AI systems continue to grow, a growing chorus of policymakers, industry leaders, and advocates have called for independent AI assessments. This explainer provides an overview of recent proposals for third-party assessment in the United States, including state and federal legislation, executive actions, and industry proposals.
Having third parties assess AI systems might seem like common sense, but crafting effective policies toward this goal can be devilishly tricky. A poorly-constructed ecosystem for third-party assessment could easily fail to consider the most consequential mechanisms of risk, neglect the AI harms that most impact people, or do more to protect AI companies than people.
Not All Guardrails Are Created Equal: Comparing Content Safety and Copyright Filtering
As courts and policymakers work through questions about chatbot liability, they should be wary of analogies that flatten meaningful technical differences. Copyright filtering and safety intervention share real challenges around ambiguity and evasion, but they diverge in what each control must assess, how each manifests over the course of a conversation, and how much can be verified from the outside.