For thirty years, a website had exactly two postures toward a bot: let it in or block it. robots.txt was a request, not a lock. If a crawler ignored it — and plenty do — there was no real recourse short of an IP ban or a lawsuit. That binary is starting to break down. A new mechanism called pay-per-crawl gives publishers a third option: let the bot in, but only if it pays first. The plumbing for that option is, improbably, a status code that has sat unused in the HTTP specification since 1997.
That matters to anyone who runs a website. AI systems increasingly read your pages, either to train models or to answer a user's question directly inside a chat interface, and often send little or no traffic back in return. Blocking them outright can cost you visibility in AI-driven answers; allowing them for free means giving away the content that pays your bills. Pay-per-crawl, built on HTTP 402 Payment Required, is the first widely deployed attempt to put a price on that access at the protocol level.
This guide explains what pay-per-crawl is and how the 402 exchange works step by step, why the status code was dormant for so long, why the shift is happening now, how it compares with robots.txt, bot blocking, and licensing deals, how to set an AI crawler policy for your own site, and what it means for AI companies whose retrieval and training pipelines now face a real per-request cost.
What pay-per-crawl actually is
Pay-per-crawl is a metered access model for web content, aimed squarely at AI crawlers — the bots that fetch pages to train large language models, echoing the same uncompensated-extraction dynamic playing out in open source software, or to answer live queries via retrieval. Instead of a flat "allow" or "disallow" in robots.txt, a site fronted by pay-per-crawl responds to an unrecognized or unpaid crawler with a request for payment. If the crawler (or the company operating it) has an account and agrees to the price, the request is retried with proof of payment attached, and the content is served. If not, the crawler gets a rejection instead of the page.
Cloudflare, which sits in front of a large share of the web as a reverse proxy and bot-management layer, built the first widely deployed version of this. Because it already fingerprints and classifies bot traffic for security purposes, it was well positioned to add a billing step to that pipeline: identify the crawler, check whether it's associated with an AI company, and either wave it through, block it, or ask it to pay, depending on the site owner's settings.
The pitch to publishers is straightforward. Search engines historically crawled a site in exchange for sending it traffic — an implicit trade that justified letting Googlebot in for free. AI crawlers that scrape a page to train a model, or to answer a user's question directly inside a chat interface, often return no traffic at all. The page's value is extracted and the visitor never arrives. Pay-per-crawl reframes that extraction as a transaction with a price attached, rather than something a site either tolerates or blocks outright.
How HTTP 402 fits in
HTTP status codes are how a server tells a client what happened to its request: 200 for success, 404 for not found, 403 for forbidden. Status code 402 was reserved in the original HTTP/1.1 specification under the label "Payment Required," with a note that it was set aside for future use. For nearly three decades, nothing used it. There was no standard way to define what "payment" meant in an HTTP exchange, so browsers and servers alike left it dormant.
Pay-per-crawl activates it. The mechanics are simple enough to describe as a short exchange:
- A crawler requests a page. The server checks whether the request carries valid payment credentials.
- If it doesn't, and the site owner has pay-per-crawl enabled for that content, the server returns
402 Payment Requiredalong with metadata describing the price and how to pay it. - The crawler's operator — an AI company with a funded account on the payment layer — either accepts the price or abandons the request.
- On acceptance, the crawler resends the request with a signed payment token attached.
- The server verifies the token, logs the charge, and returns
200 OKwith the actual content.
The verification step depends on being able to tell crawlers apart reliably — separating a legitimate, self-identified AI agent from something spoofing its user-agent string to dodge the toll. That relies on cryptographic bot verification (signed requests tied to a known operator's identity) rather than the easily-faked User-Agent header the web has leaned on for years. Without that layer, pay-per-crawl is trivially bypassed by any crawler willing to lie about who it is.
Why 402, specifically
The web already had ways to gate content — paywalls, login walls, IP allow-lists. What 402 adds is a machine-readable signal that fits inside the normal request/response cycle a bot already understands. A crawler doesn't need to parse a paywall's HTML or run JavaScript to discover it owes money; the status code itself carries that meaning, the same way 404 carries "not found" without any prose. That's what makes it plausible as a protocol rather than a one-off business feature: any client, not just one company's crawler, can be built to recognize and respond to it.
Why this is surfacing now
The near-term forcing function is a policy change on Cloudflare's platform: starting September 15, 2026, new domains onboarded to Cloudflare will default to blocking AI training crawlers, with paid access offered as the alternative route in rather than an opt-in extra bolted on afterward. That's a meaningful shift in default behavior. Today, a site owner who wants to restrict AI crawling has to actively configure it — write robots.txt rules, turn on bot-blocking features, or negotiate directly with an AI company. After that date, new sites on Cloudflare start from "blocked, unless paid," and the site owner has to actively choose to allow free crawling if they want it.
Defaults matter enormously on the web, because most site owners never touch settings past whatever ships out of the box. A default of "open" produced the web AI companies trained on through the early 2020s: crawlers largely operated under the assumption that anything not explicitly disallowed was fair game, and few robots.txt files anticipated LLM training as a use case worth blocking. Flipping the default for new sites — even just on one platform — changes the shape of what's crawlable going forward, without requiring millions of individual publishers to take action.
It also signals where the infrastructure layer of the web thinks this is heading: not toward every publisher hand-negotiating licensing terms with every AI lab, but toward the reverse-proxy and CDN layer handling identification, pricing, and payment as a built-in utility, the way it already handles TLS termination or DDoS mitigation.
Benefits of Pay-Per-Crawl
Pay-per-crawl is new and imperfect, but it adds options that neither robots.txt nor outright blocking could provide.
A middle option between open and blocked
Before, a site owner worried about AI crawling had two choices: allow it for free or block it and risk disappearing from AI-driven answers. Pay-per-crawl adds a third: let the crawler in on terms. That lets publishers keep some presence in AI products while being compensated for the content those products rely on, instead of choosing between visibility and giving their work away.
Compensation for sites too small to negotiate
Large publishers can sign direct licensing deals with AI companies. Millions of smaller sites cannot, because no AI company will negotiate individually with each of them. Building pricing and payment into the infrastructure layer gives those sites a route to compensation without lawyers or business development teams. The amounts may be small at first, but the mechanism exists where previously there was nothing.
Enforcement rather than requests
robots.txt depends on crawler goodwill. Pay-per-crawl, paired with verified bot identity, enforces a decision at the network layer: an unpaid crawler simply does not receive the content. That shift from advisory to enforced control matters most for crawlers whose operators have little incentive to respect voluntary rules, which is exactly the gap that made AI crawling contentious.
Machine-readable terms bots can act on
Because the price arrives as part of a standard HTTP response, a crawler can discover what access costs and decide automatically whether to pay, with no human negotiation and no page rendering. That makes the arrangement workable at web scale, where millions of sites and thousands of requests per second rule out anything that needs a person in the loop.
Cleaner provenance for AI companies
For AI developers, a paid, logged crawl leaves a record of where content came from and on what terms it was obtained. As questions about training data provenance continue in courts and policy debates, that audit trail can be useful. It turns part of data collection from an anonymous scrape into a documented transaction. Paying also gives AI companies predictable, sanctioned access to sources that might otherwise be blocked outright, which is easier to plan around than an escalating contest between crawlers and bot-blocking rules.
Pay-Per-Crawl Use Cases
The mechanism is general, but some kinds of sites and content stand to gain more than others. These are the situations it fits best today.
Mid-sized publishers and niche media
Specialist publications, trade media, and regional news sites produce distinctive content that AI products draw on, yet they lack the leverage for bulk licensing deals. Enabling pay-per-crawl on their articles lets them charge training and retrieval crawlers while still serving human readers normally. Uptake in this segment is the clearest test of whether the model works, since these sites are the long tail it was built for.
Reference and expert content
Sites that answer specific questions well, such as technical guides, recipes, product comparisons, and specialist explainers, are particularly exposed to AI answers that satisfy a query without a click-through. For them, charging retrieval crawlers, or at least training crawlers, offers a way to recover some value from content that would otherwise be summarised elsewhere with little traffic coming back.
Premium or members-only sections
Many sites mix free marketing pages with premium material behind a paywall. Pay-per-crawl can apply to the premium sections specifically, so crawlers pay for access to content human visitors also pay for, while public pages stay open for discovery. Segmenting by section keeps the site visible in AI answers for introductory material without giving away the work that funds it.
Newly launched sites on default-blocking platforms
For new domains on platforms that block AI training crawlers by default, paid access becomes the route by which AI companies reach that content at all. Site owners in that position do not need to configure blocking themselves; their main decision is whether to price access and at what level. This makes the default change a practical entry point for owners who would never have set up crawler controls manually.
AI companies budgeting for data access
On the other side, AI developers and retrieval-based products can use pay-per-crawl as a structured way to reach content they would otherwise be blocked from. Teams set budgets, decide which sources are worth paying for, and keep records of paid access. Early adopters are likely to treat it like any other data acquisition cost, weighing price against the quality the content adds.
The broader landscape of AI access control
Pay-per-crawl doesn't exist in isolation. It's one entry in a small but growing set of mechanisms sites use to manage AI crawler access, each with different tradeoffs.
| Mechanism | What it does | Enforcement strength | Payment involved |
|---|---|---|---|
robots.txt | Publishes crawl rules bots are expected to honor voluntarily | Weak — advisory only, unenforced | No |
| Bot-management blocking | Fingerprints and blocks traffic identified as an unwanted crawler | Strong at the network layer, but an arms race against spoofing | No |
| Pay-per-crawl (HTTP 402) | Requires a payment token before serving content to an identified crawler | Strong, contingent on reliable bot identity verification | Yes, per-request |
| Direct licensing deals | Publisher and AI company negotiate a bulk content-access agreement | Strong, but slow and only viable for large publishers | Yes, negotiated |
| Machine-readable licensing terms (e.g. RSL-style tags) | Embeds licensing terms and price in page metadata for bots to read | Advisory, similar to robots.txt unless paired with enforcement | Optional |
The pattern across all five is a shift from advisory controls that depend on crawler goodwill toward enforced controls that depend on infrastructure. robots.txt was written in an era when the crawlers reading it were mostly search engines with a commercial incentive to stay in publishers' good graces. AI training crawlers don't share that incentive structure in the same way, which is a large part of why enforcement mechanisms like bot blocking and pay-per-crawl gained traction rather than a renewed push to get everyone to simply respect robots.txt more strictly.
Where this sits relative to licensing deals
Large publishers with negotiating power — major news organizations, big reference sites — have generally preferred direct licensing negotiations with AI companies, because a bulk deal can be worth far more than metered per-request fees and comes with contractual guarantees around attribution or usage limits. Pay-per-crawl is aimed more at the long tail: the millions of sites too small to get an AI company's licensing team on the phone, for whom a self-serve, automatically-priced toll is the only realistic form of compensation. The two models aren't competitors so much as they cover different ends of the publisher-size spectrum.
Practical implications for businesses and builders
For site owners, the calculus is now more granular than "allow or block." A few things worth working through before deciding on a stance:
- Traffic value vs. content value. A site that earns most of its revenue from ad-supported pageviews loses more from AI answers that satisfy a query without a click-through than from training scrapes that happen once. Pricing (or blocking) decisions should reflect which of those two problems actually hurts more.
- Segmenting by crawler purpose. Not all AI crawlers do the same thing. Some scrape for one-time model training; others fetch pages live to answer a specific user's question (retrieval-augmented generation). A blanket policy treats both the same, but many site owners will want different rules — for example, allowing live-retrieval bots that plausibly send users back to the source, while charging or blocking pure training scrapes.
- Verification dependency. None of this works if crawler identity can be faked. Site owners adopting pay-per-crawl are implicitly betting on the underlying bot-verification layer staying ahead of spoofing attempts, which is an ongoing security problem, not a solved one.
- Platform lock-in. Today's pay-per-crawl implementations are tied to the CDN or bot-management vendor sitting in front of a site. A site not using that vendor doesn't get the feature. That's a meaningful lever for CDN vendors and a dependency site owners should weigh consciously as part of their broader web infrastructure planning.
For AI companies and builders of crawler-based products, the implications run the other direction:
- Crawling costs become a line item. A pipeline that used to fetch content for the price of bandwidth now needs a budget line, a payment integration, and a policy for which sites are worth paying for versus skipping.
- Data provenance gets easier to audit. A paid, logged crawl leaves a cleaner trail than an anonymous scrape, which could matter for companies trying to demonstrate where their training data came from — a question that's come up repeatedly in copyright litigation over AI training data.
- Retrieval products face real-time cost pressure. A chatbot that fetches a live page to answer a user's question now potentially pays per fetch. At scale, that turns retrieval-augmented answers into a cost model that looks more like a paid API call than a free web request, which changes the unit economics of that class of product.
Pay-Per-Crawl Best Practices: Setting an AI Crawler Policy
Whether or not you enable pay-per-crawl, it's worth making a deliberate decision about AI crawlers rather than inheriting whatever your platform defaults to.
- Measure current bot traffic. Use your CDN, server logs, or analytics to see which AI crawlers already visit, how often, and which sections of the site they fetch.
- Decide what you want from AI systems. Some sites want to appear in AI answers and accept crawling as a form of distribution; others mainly want compensation or protection for premium content.
- Separate crawler purposes. Treat search indexing, AI training, and live retrieval for answers as different categories with different rules.
- Segment your content. Public marketing pages, documentation, and paywalled or premium content may each deserve a different policy.
- Write the rules down in
robots.txtfirst. Even though it's advisory, a clear robots.txt file documents your intent and is honored by many well-behaved crawlers. - Add enforcement where it matters. Use bot management or pay-per-crawl for the sections and crawler types where advisory rules aren't enough.
- Set and review prices. If you charge, start with a simple per-request price, watch which crawlers pay and which leave, and adjust.
- Monitor referral traffic and AI visibility. Track whether your choices change traffic from AI products and search, and revisit the policy quarterly.
- Keep search indexing separate. Confirm that whatever you block or charge does not catch the search crawlers that send you organic traffic. Check how each operator identifies its different bots, since one company may run both a search crawler and AI products.
- Record decisions and their reasons. Note which crawlers you allowed, charged, or blocked, and why. When a new AI product appears or your traffic shifts, a short written rationale makes it much faster to adjust the policy without undoing choices that were working.
Common Pay-Per-Crawl Mistakes
Blocking every bot with one blanket rule
The quickest reaction to AI crawling is to block everything that looks like a bot. That can also catch search indexers and live-retrieval bots that send visitors back, cutting organic traffic and AI visibility at the same time. Distinguish search, training, and retrieval crawlers and set rules for each, rather than treating "bots" as a single category.
Relying on robots.txt alone
robots.txt documents intent, and many well-behaved crawlers honour it, but it enforces nothing. Site owners who write careful rules and assume the matter is settled are often surprised by what their logs show. Use it as the statement of policy, then add bot management or pay-per-crawl where enforcement actually matters to you.
Pricing without measuring
Setting a price before knowing which crawlers visit, how often, and what they fetch is guesswork. Too high, and every crawler walks away, removing the site from AI answers with no revenue in return; too low, and the charge barely registers. Measure current bot traffic first, start with a simple price, and adjust based on which crawlers pay and how referral traffic changes.
Ignoring the verification dependency
Pay-per-crawl only works if paying crawlers can be told apart from impostors. Owners who enable it and stop paying attention may not notice when spoofed traffic slips through or when a scraping service ignores the toll altogether. Review bot traffic periodically and treat verification as an ongoing security concern rather than a setting that stays solved.
Expecting it to undo past scraping
Some owners enable pay-per-crawl expecting it to address content already used in model training. It governs future requests only. Content collected before the mechanism existed is outside its reach, and questions about that material are being argued elsewhere. Set expectations accordingly and treat pay-per-crawl as a forward-looking control.
Limitations and open questions
The idea is coherent; the execution has real gaps. A few are worth naming plainly:
- Bot identity is still an arms race. Pay-per-crawl only works if a server can tell a paying, legitimate crawler apart from one impersonating it, a verification challenge closely related to content authenticity more broadly. Spoofing user-agent strings and IP ranges is a well-worn tactic, and cryptographic verification schemes, while stronger, are neither universal nor immune to compromise.
- Fragmentation risk. If every CDN vendor ships its own flavor of payment metadata and its own verification scheme, AI companies face a patchwork of incompatible integrations rather than one protocol — undermining the value of standardizing on a shared status code in the first place.
- Pricing has no established market yet. There's no consensus on what a single page fetch should cost, how it should scale with content value or crawler volume, or who arbitrates disputes. Early pricing will likely be arbitrary and volatile until usage data accumulates.
- Small AI builders may simply route around it. A well-funded lab can absorb crawl fees or negotiate bulk deals. A solo developer or small startup running a crawler is more likely to skip paid sites entirely, or to route requests through less scrupulous scraping services that ignore the toll altogether — which would concentrate high-quality paid data with the largest players while starving smaller ones.
- It doesn't retroactively cover past training data. Pay-per-crawl governs future crawling. It has no bearing on content already scraped and incorporated into models trained before the mechanism existed, which is the crux of most ongoing legal disputes over AI training data.
What to watch next
A few signals will indicate whether this becomes durable infrastructure or a niche feature:
- Whether other CDN and bot-management vendors ship comparable 402 support. A single vendor's implementation is a feature; multiple interoperable implementations start to look like a protocol.
- Uptake among mid-sized publishers, who have the technical means to enable it but, unlike major news organizations, lack in-house teams to negotiate direct licensing deals — they're the segment pay-per-crawl was built for.
- How AI companies respond in practice — whether major crawlers pay routinely, negotiate around it, or simply avoid paid domains and rely on the remaining open web.
- Standards-body involvement. A mechanism that started as one company's product feature gains durability if it moves toward a formally documented, vendor-neutral specification that others can implement independently.
- What happens after Cloudflare's September 15, 2026 default change lands — whether it visibly shifts the composition of what's crawlable, and whether other platforms follow with similar default-blocking policies of their own.
Teams weighing how pay-per-crawl, bot management, and AI-crawler policy should fit into their own site architecture can get hands-on help from Woyce Technologies.
FAQ
What is pay-per-crawl?
Pay-per-crawl is a mechanism, pioneered by Cloudflare, that lets website owners charge AI crawlers a fee for each request before serving them content. When an identified AI crawler asks for a page, the server can respond with the HTTP 402 "Payment Required" status code and a price instead of the content. If the crawler's operator has a funded account and accepts the price, it retries with proof of payment and receives the page. It gives publishers a third option between allowing all AI crawling and blocking it.
How does HTTP 402 work in this context?
When an unpaid or unrecognized crawler requests a page, the server responds with a 402 status code and pricing details instead of the content. The crawler's operator decides whether to pay. If it does, it resends the request with a signed payment token, the server verifies the token and records the charge, and the page is served with a normal 200 response. Because the signal is part of the standard HTTP exchange, crawlers can handle it automatically without parsing a paywall or rendering a page.
Is pay-per-crawl the same as a paywall?
Not exactly. A paywall is built for human visitors and typically involves a login, a subscription, and a flow rendered in HTML and JavaScript. Pay-per-crawl operates at the HTTP protocol level, is aimed specifically at automated crawlers, and charges per request rather than per subscription. It is designed to be machine-readable, so a bot can discover the price and pay without any page rendering. A site can run both at once: a paywall for people and pay-per-crawl for AI bots.
Will this stop AI companies from training on my content for free?
It can meaningfully raise the cost and difficulty of doing so if your site is protected by verified bot identification and payment enforcement, but it isn't airtight. Crawlers that spoof their identity, rotate IP addresses, or route through third-party scraping services can still attempt to bypass it. It also does nothing about content that was already collected before you enabled it. Treat it as one layer of control alongside robots.txt, bot management, and, where relevant, legal terms of use.
Do I need to use Cloudflare to enable pay-per-crawl?
Currently, pay-per-crawl implementations are tied to the infrastructure vendor sitting in front of a site, and Cloudflare's is the most widely deployed. Sites on other CDNs or hosting providers without a comparable bot-management layer don't currently have equivalent built-in support, although they can still use robots.txt, server-level blocking, and direct licensing. If other vendors adopt interoperable HTTP 402 support, the choice of CDN will matter less.
What happens on September 15, 2026?
New domains onboarded to Cloudflare after that date will default to blocking AI training crawlers, with paid access offered as an explicit alternative route, rather than requiring site owners to opt in to blocking after the fact. Existing sites keep their current settings unless their owners change them. Because most site owners never change defaults, the shift is expected to reduce how much newly published content is freely available for AI training over time.
Does this affect search engine crawlers too?
Implementations generally aim to distinguish AI training and retrieval crawlers from traditional search-indexing crawlers, since blocking search indexing would cut off organic traffic. Site owners can typically configure different rules for different categories of bot rather than applying one blanket policy. The line can blur when one company operates both a search crawler and AI products, so check how each crawler identifies itself and what its operator says the data is used for.
Is pay-per-crawl worth it for a small website?
It depends on your goals. If your content is distinctive and AI products are answering questions with it while sending you little traffic, charging or blocking training crawlers can make sense, and pay-per-crawl is one of the few compensation routes open to small publishers. If you rely on discovery and want to be cited in AI answers, blocking or pricing too aggressively may reduce visibility. Many small sites start by measuring AI bot traffic, blocking pure training crawlers, and allowing retrieval bots.
Conclusion
For decades a website could only ask bots to stay away. Pay-per-crawl changes that by giving publishers a third choice — let AI crawlers in, but only if they pay — and by using the long-dormant HTTP 402 status code to make that request machine-readable inside the normal web exchange.
The central insight is that AI crawling isn't one thing. Training scrapes, live retrieval for answers, and search indexing affect a site's traffic and revenue very differently, and a good policy treats them separately. Pay-per-crawl is mainly useful for the long tail of sites too small to negotiate licensing deals, while large publishers continue to sign direct agreements.
The caveats matter. Everything depends on reliable bot identification, implementations are currently tied to specific infrastructure vendors, pricing has no established market, smaller AI builders may route around paid sites, and nothing here affects content already collected for past training.
The practical next step is to measure which AI crawlers already visit your site and decide what you want from each category. If you'd like help designing a crawler policy, bot management setup, or AI-ready site architecture, our web development team can help.
