There is a setting somewhere in your infrastructure stack that determines whether AI systems can read your website. Someone configured it. It was probably not anyone in marketing.
For most B2B companies this happened without a decision being made in any meaningful sense. A CDN or edge provider shipped a bot management feature, AI crawler blocking moved from opt-in to on-by-default, and an infrastructure or security team accepted the recommended configuration during a routine review. The reasoning was sound on its own terms: unidentified automated traffic consumes bandwidth, scrapes content, and creates no measurable value. Blocking it is hygiene.
The problem is that “unidentified automated traffic” now includes the systems that decide whether your company appears in a buyer’s consideration set. Over the past two years, discovery has shifted decisively toward AI-mediated paths—answer engines that synthesize responses from retrieved sources, assistants that research vendors on request, and increasingly the autonomous purchasing agents we discussed in July, which build shortlists before any human reviews them. All of these depend on being able to fetch and read your content.
Which means a control designed to reduce scraping has become, in effect, a distribution policy. And it is being set by people who were never told that.
Three Kinds of Crawler, One Blunt Control
The core failure here is categorical. Most bot management configurations treat AI-associated traffic as a single class, when at least three distinct things are arriving at your origin with very different implications for marketing.
Training crawlers collect content at scale to build or refine models. There is no direct, attributable return for you in this exchange—your content becomes part of a system’s general capability with no citation, no referral, and no visibility. This is the category where blocking is most defensible, and where the licensing conversations of the past two years have concentrated. It is also, notably, the category most B2B marketing content is least valuable to: a model has limited use for your product positioning page.
Retrieval fetchers pull specific pages at query time to answer a live question, and typically surface a link or citation. This is the mechanism behind AI search results and assistant answers. Blocking these is functionally equivalent to deindexing yourself from a growing share of buyer research. Teams that spent 2026 investing in answer engine optimization while their edge configuration denied retrieval fetchers have been running two initiatives against each other.
Agent traffic is a request made on behalf of a specific person—someone’s assistant reading your pricing page because they asked it to, or a procurement agent checking whether you meet a compliance requirement. Functionally this is a prospect visiting your site, with an intermediary handling the reading. It is the highest-intent traffic arriving at your domain, and it is frequently the traffic most likely to be challenged, because agents often present unfamiliar user agents, arrive from datacenter IP ranges, and cannot solve interactive challenges designed to verify humanity.
Any policy that cannot distinguish between these three is not a policy. It is a default. And the practical consequence in many B2B organizations right now is a site that is open to the crawler with the least commercial value to it and closed to the two with the most.
Why This Is Harder Than “Just Allow Them”
The obvious correction—allow anything AI-shaped—is wrong for reasons worth being precise about.
Identity verification is genuinely difficult. Well-behaved commercial crawlers publish their IP ranges and support reverse DNS verification, so allowlisting them is straightforward. But user-agent strings are trivially forged, and a meaningful volume of traffic claiming to be a major AI assistant is not. Some of it is competitive scraping, some is content harvesting for spam sites, and some is reconnaissance. Allowing by user-agent string alone is not access policy; it is an open door with a sign on it.
Cost is real and asymmetric. Aggressive crawlers can generate request volumes far in excess of human traffic, and they concentrate on exactly the dynamic, uncached surfaces—search pages, filtered resource libraries, parameterized URLs—that are most expensive to serve. Infrastructure teams pushing back on unrestricted access usually have a bandwidth invoice supporting their position.
And the licensing question is unresolved in a way that rewards nobody’s confidence. A handful of large publishers have negotiated paid access arrangements. Per-crawl payment mechanisms exist at the edge layer. Machine-readable licensing standards have emerged. But the honest assessment for nearly every B2B company is that this is not your leverage. Your content is not a scarce corpus that a model developer needs; your content is marketing material whose entire purpose is to be encountered. Treating it as a revenue-generating asset to be metered misreads what it does for you. The few B2B organizations with genuine leverage here are those sitting on proprietary datasets, original research programs, or technical documentation with no substitute—and for them the right move is usually to segment that material rather than to meter the whole domain.
Building an Access Policy That Reflects Marketing Reality
What follows is the operational work, and most of it is unglamorous.
Find out what your current posture actually is. Not what you assume, and not what the marketing team decided—what the edge configuration does. Ask your infrastructure team for the current bot management rules, the list of blocked and challenged categories, and whether AI crawler blocking is enabled. Then verify independently: check your robots.txt as served in production, and pull edge logs for the major AI user agents to see whether requests are being answered, challenged, or refused. The gap between assumed and actual posture is where most of the surprises live.
Segment your content by what access to it is worth. Three tiers cover most cases. Open and actively promoted: positioning, product, pricing, documentation, comparison content, customer proof—material you want retrieved, cited, and read by agents, because being absent from those answers costs you consideration. Gated by exchange: original research, benchmarks, proprietary data—valuable enough that a form or an explicit license is reasonable, and where wide unattributed ingestion genuinely dilutes an asset you paid to create. Not worth serving to anyone automated: internal search results, faceted filter permutations, session-parameterized URLs, and the long tail of low-value dynamic pages that inflate crawl cost without ever appearing in an answer.
Differentiate rules by crawler purpose, not by whether the word “AI” appears in the user agent. Verified retrieval fetchers and agent traffic should reach your open tier without friction. Training crawlers can be a deliberate choice—allow, deny, or restrict to a subset—made by someone who understands the tradeoff rather than inherited from a vendor default. Unverified traffic claiming a known identity should be treated as unverified traffic, which usually means rate limiting rather than a hard block.
Check that your gates degrade gracefully. Interactive challenges, JavaScript-dependent rendering, cookie walls, and consent interstitials all block legitimate agent access as effectively as they block scrapers, and they do it silently. If a buyer’s assistant cannot read your pricing page because a consent banner intercepts the request, you will never see the failure—you will simply not be in the answer. Test your key pages the way a retrieval system would: fetch them without JavaScript, without cookies, and see what comes back.
Get agent traffic into your reporting. Standard analytics either filters this traffic as non-human or misclassifies it as direct, which means the fastest-growing segment of your site’s audience is invisible in the dashboards your team reviews. Server-side log analysis segmented by user agent is currently the practical answer. It is not elegant, but it is the only way to know whether the volume is growing, which pages draw it, and whether a configuration change helped or hurt.
Put the policy in writing and give it an owner. The failure mode here is not a bad decision; it is an unowned one that gets silently re-defaulted at the next vendor upgrade. A one-page document naming which categories are allowed, which are blocked, why, and who reviews it prevents your distribution strategy from being reset by a patch note.
What This Is Actually About
It is tempting to file this under technical SEO and hand it to whoever owns the site. That undersells it.
For most of the last two decades, the question of who could read your marketing content had one answer: everyone, and please do. Discovery ran through search engines whose access you actively courted. There was no version of the strategy in which restricting readership made sense, so no one built the muscle for deciding.
That era is over, and what replaces it is a genuine allocation problem. Access has costs. Some readers return attributable value and some do not. The systems arriving at your domain have materially different relationships to your pipeline, and treating them identically—whether by blocking everything or allowing everything—guarantees you get the tradeoff wrong in one direction or the other.
The teams handling this well have made it a joint decision rather than a handoff. Infrastructure brings the traffic data and cost picture. Security brings the threat model. Marketing brings the one input neither of the others has: which of these visitors represent demand. That input is not optional, and in most organizations it has not yet been offered.
The reason to act now is straightforward. The share of B2B discovery mediated by AI systems is rising through the second half of 2026, and the buyers behind those systems are not going to tell you when they failed to find you. A blocked retrieval fetcher does not generate an error report or an angry email. It generates an answer about your category that does not include you—and a shortlist you were never on.
Somebody at your company made that decision months ago. It is worth finding out what they chose.