How to Set Search, Agent and Training Access Without a Blanket AI Bot Block
How should a small business choose different access policies for Cloudflare’s Search, Agent and Training crawler categories?
Make a separate decision for Search, Agent and Training. Preserve Search access when it matches the intended public availability of the site; assess Agent access according to how the business wants automated assistants to use public content; and set Training to disallow when the business wants to state a no-training preference. Record the reason and approval for each choice. Do not start with a blanket AI-bot block simply because one category raises a concern.
Make a separate decision for Search, Agent and Training. Preserve Search access when it matches the intended public availability of the site; assess Agent access according to how the business wants automated assistants to use public content; and set Training to disallow when the business wants to state a no-training preference. Record the reason and approval for each choice. Do not start with a blanket AI-bot block simply because one category raises a concern.
Why one AI-bot rule is too blunt
Cloudflare lets a site state what it wants to do about Search, Agent and Training traffic. Because those choices can be expressed separately, begin with three business decisions rather than a single yes-or-no rule for everything described as AI-related.
Sources: Say it once: introducing Bot Preference Sync | Cloudflare Blog.
Bot Preference Sync aligns Cloudflare’s configured AI bot policy with the public robots.txt instructions through the Search, Agent and Training categories. This gives the policy review a clear structure: decide the intended treatment of each category, document the reason and obtain approval before configuration.
Sources: Cloudflare Bot Preference Sync and AI Crawlers.
- Treat Search, Agent and Training as three approval decisions.
- Record why each category should remain available or be restricted.
- Do not let concern about Training silently become a block on Search or Agent traffic.
Decide whether Search access should remain available
For Search and Agent traffic, site operators can allow access, block access on pages serving advertisements or block access entirely. For Search, choose among those approaches according to the intended availability of the public site, without assuming or promising a measurable visibility outcome.
Sources: Cloudflare Bot Preference Sync: Automated Robots.txt Rules.
A practical policy question is: “Do we intend crawler-based search services to access this public content?” If the answer is yes, preserve the corresponding access unless another approved requirement justifies a narrower restriction. If the answer is uncertain, investigate before choosing an entire-category block.
- Record whether the site is intended to be publicly available through crawler-based search.
- Identify any page-specific concern separately from the category-wide decision.
- Do not claim that allowing Search guarantees traffic, citations or commercial results.
Decide how Agent traffic should access the site
Cloudflare’s policy model also permits a separate choice for Agent traffic. Site operators can choose to allow Agent access, block it on pages serving advertisements or block it entirely. Keep this decision separate from Training so that a no-training policy does not automatically determine how automated assistants may access public content.
Sources: Cloudflare Bot Preference Sync: Automated Robots.txt Rules.
Ask whether Agent access fits the business’s intended use of its public website. Record the selected treatment, the reason, the approving person and any review trigger. Avoid inventing Agent use cases or crawler examples that are not part of the approved evidence.
- Assess Agent access independently of Search and Training.
- Choose an approach based on the intended use of public content.
- Record a review trigger if the business requirement may change.
Set the Training preference separately
Training can be set to disallow as a no-training preference while Search access remains available for crawlers that clear the transparency bar. This supports a category-level policy in which the business communicates a Training preference without automatically closing Search access.
Sources: Cloudflare Bot Preference Sync: Robots.txt Meets Enforcement.
Category-wide controls can allow or block Search and Agent bots while setting Training to disallow with transparency checks for mixed-use crawlers. The evidence does not define the transparency criteria, so do not invent them or claim that your own informal check establishes compliance.
Sources: Cloudflare Bot Preference Sync: Robots.txt Meets Enforcement.
Describe Training disallow accurately: it states a no-training preference. If the business requires access to be prevented rather than merely discouraged, document that as a separate enforcement requirement for authorised review.
- Keep the Training decision separate from Search and Agent.
- Label disallow as a preference, not proof of blocking.
- Escalate any access-prevention requirement to a separate enforcement review.
Complete the category decision matrix
Before changing settings, complete one row for each category. Record the business purpose, desired access, reason, whether the requirement is a preference or enforced restriction, the approver and the event that should trigger review. Finish by checking that no concern about one category has produced an accidental blanket restriction.
- Every category needs an explicit desired treatment.
- Every restriction needs a plain-language business reason.
- Every access-prevention requirement needs separate identification.
- Every undecided category should remain open for review rather than being silently grouped with another.
Search–Agent–Training access decision matrix
Complete this matrix before changing Cloudflare settings. It keeps the three policy choices distinct and exposes any accidental move from a targeted concern to a blanket restriction.
| Category | Business question | Desired treatment | Control meaning | Review trigger |
|---|---|---|---|---|
| Search | Should crawler-based search access this public content? | Allow, restrict on the available basis or block as approved | Record whether this is an availability decision or access-prevention requirement | Site purpose, content model or policy changes |
| Agent | Does Agent access fit the intended use of public content? | Allow, restrict on the available basis or block as approved | Keep this decision separate from Training | Approved use or risk assessment changes |
| Training | Does the business want to state a no-training preference? | Set the approved Training preference separately | Disallow states a preference; do not call it proof of blocking | Training policy or enforcement requirement changes |
| Blanket-policy check | Did one concern alter unrelated categories? | Return unintended changes for review | Use category-level decisions unless broader blocking is deliberately approved | Any proposed category-wide expansion |
This planning matrix does not prescribe a dashboard path or test enforcement. Confirm current product details through authoritative Cloudflare information before implementation.
Frequently asked questions
Why should Search, Agent and Training be decided separately?
Cloudflare provides category-level policy choices, so a concern about one activity does not need to determine the treatment of all AI-related traffic. Separate decisions reduce the risk of an unintended blanket restriction.
Can Training be set to disallow while Search remains available?
Yes. The accepted evidence supports a no-training preference while preserving Search access for crawlers that clear Cloudflare’s transparency bar. The evidence does not define those transparency criteria, and the preference is not proof of enforced blocking.
What choices are available for Search and Agent traffic?
The supplied evidence says operators can allow access, block access on pages serving advertisements or block access entirely. Choose according to approved business intent without promising traffic, citation or sales effects.
How should a business decide Agent access?
Assess whether Agent traffic fits the intended use of the site’s public content. Record the desired treatment, reason, approver and review trigger independently of the Training decision.
What should happen if the business cannot decide on a category?
Mark it for investigation rather than applying a blanket block. Clarify the intended availability and whether the requirement is merely a preference or genuine access prevention before changing controls.
Related guidance
What follow-up questions matter most?
- Why should Search, Agent and Training be decided separately?
- Cloudflare provides category-level policy choices, so a concern about one activity does not need to determine the treatment of all AI-related traffic. Separate decisions reduce the risk of an unintended blanket restriction.
- Can Training be set to disallow while Search remains available?
- Yes. The accepted evidence supports a no-training preference while preserving Search access for crawlers that clear Cloudflare’s transparency bar. The evidence does not define those transparency criteria, and the preference is not proof of enforced blocking.
- What choices are available for Search and Agent traffic?
- The supplied evidence says operators can allow access, block access on pages serving advertisements or block access entirely. Choose according to approved business intent without promising traffic, citation or sales effects.
- How should a business decide Agent access?
- Assess whether Agent traffic fits the intended use of the site’s public content. Record the desired treatment, reason, approver and review trigger independently of the Training decision.
- What should happen if the business cannot decide on a category?
- Mark it for investigation rather than applying a blanket block. Clarify the intended availability and whether the requirement is merely a preference or genuine access prevention before changing controls.
What steps does this workflow follow?
Create a Search, Agent and Training access policy
- List the three categories: Create separate decision rows for Search, Agent and Training rather than one combined AI-bot rule.
- Define the intended availability: For each category, state how its access fits the intended use of the business’s public website.
- Choose the desired treatment: Record whether the category should be allowed, restricted through an available option or covered by a declared preference.
- Separate preference from prevention: Identify whether the decision only communicates policy or requires an enforceable restriction to prevent access.
- Record the reason and approver: Write a plain-language reason, name the approval role and capture any event that should trigger review.
- Check for a blanket restriction: Confirm that a concern about one category has not unintentionally changed the approved treatment of the other two.