NO KEYNOTES • NO BREAKOUTS • NO VENDOR PITCHES • NO CANVA TUTORIALS • 30 SEATS • SAN DIEGO • FEB 2027 •NO KEYNOTES • NO BREAKOUTS • NO VENDOR PITCHES • NO CANVA TUTORIALS • 30 SEATS • SAN DIEGO • FEB 2027 •NO KEYNOTES • NO BREAKOUTS • NO VENDOR PITCHES • NO CANVA TUTORIALS • 30 SEATS • SAN DIEGO • FEB 2027 •NO KEYNOTES • NO BREAKOUTS • NO VENDOR PITCHES • NO CANVA TUTORIALS • 30 SEATS • SAN DIEGO • FEB 2027 •NO KEYNOTES • NO BREAKOUTS • NO VENDOR PITCHES • NO CANVA TUTORIALS • 30 SEATS • SAN DIEGO • FEB 2027 •NO KEYNOTES • NO BREAKOUTS • NO VENDOR PITCHES • NO CANVA TUTORIALS • 30 SEATS • SAN DIEGO • FEB 2027 •

Guides / AI Search & Church Discovery · Part 6 of 8

AI Crawlers and Church Websites: What Your Web Team Needs to Know

By Chris Mefford · 7 min read · Published

Contents
  1. Separate search, training, and user-requested access
  2. Check the page before checking the bot policy
  3. Understand what robots.txt does
  4. Review page-level indexing and preview controls
  5. Check rendering without assuming every crawler is Google
  6. Inspect security and hosting layers
  7. Give developers a testable ticket
  8. Validate access before looking for citations
  9. A technical review checklist
  10. Frequently asked questions
  11. Make the public information retrievable
  12. Sources

A church website needs accessible public content, appropriate crawler permissions, and working pages before that content can reliably participate in search. Diagnose access, indexing, and rendering separately. Make deliberate decisions about AI search and model training rather than applying one blanket rule to every bot.

If the communications team can see a page in a browser, that does not settle what a crawler receives. A security challenge, login requirement, rendering problem, or contradictory directive may produce a different experience.

The right first step is a documented check of important pages. Hand your web team a short list: the homepage, visit page, each campus page, beliefs page, and a few priority public ministry pages.

Separate search, training, and user-requested access

OpenAI distinguishes three relevant agents: OAI-SearchBot supports ChatGPT search; GPTBot relates to potential model-training use; and ChatGPT-User handles certain user-initiated visits. Their controls are not interchangeable. OpenAI explicitly permits allowing its search crawler while disallowing its training crawler. User-initiated access has separate behavior. [1]

For Google, Google-Extended controls certain Gemini training and grounding uses. Google states that this token does not determine inclusion or ranking in Google Search. [2]

Use a policy table before changing configuration:

PurposeDecision to makeEvidence to retain
Search discoveryWhich documented search crawlers should access public pages?Current vendor guidance and approved configuration
Training useWhat training policy has the organization chosen?Separate decision and applicable controls
User-requested retrievalHow does the vendor describe visits initiated by a person?Current documentation and relevant access behavior
Private contentWhich areas require authentication?Access controls and an owner

For vendors not covered here, consult their current official documentation. Do not infer a crawler’s purpose from its name or from another vendor’s terminology.

Check the page before checking the bot policy

Start with the URL someone should actually visit. Does it return the intended content? Does it redirect to another page? Does it show a login prompt, an error, or a challenge page?

Record the final URL, response status, and whether the main text is present. Check both the ordinary visitor experience and the diagnostic tools appropriate to the search system.

A response marked successful is not enough by itself. A server can return a 200 status for a generic app shell, an empty template, or an error message. Inspect the content too.

Test representative pages individually. A homepage can work while campus or ministry routes fail on direct loading.

Understand what robots.txt does

Robots.txt communicates crawl instructions to compliant crawlers. It is not a password, and blocking crawling is not a reliable method for keeping a URL out of search results. Google explains that blocked URLs can still be indexed through information found elsewhere. [3]

Keep private member records, pastoral information, and administrative tools behind real access controls. Their protection should not depend on a crawler voluntarily following a text file.

When inspecting robots.txt, ask:

  • Are the directives intended for this live hostname?
  • Is a broad rule unintentionally blocking public content?
  • Does a crawler-specific group change the behavior expected from the general group?
  • Are necessary resources treated consistently with the intended rendering strategy?
  • Has the organization separated search and training decisions?

Do not paste a generic “allow all bots” file over an existing configuration. Have the owner of the current rules review the proposed change, and keep a copy of the previous version.

Review page-level indexing and preview controls

Inspect relevant meta tags and HTTP headers. A page may be crawlable yet intentionally marked to stay out of an index. A staging setting may also have been carried into production by mistake.

Check the canonical URL as part of the same investigation. If several distinct campus pages all point to the homepage as canonical, ask the web team whether that is intentional and appropriate.

For Google’s AI search features, ordinary Google Search eligibility applies. Its documentation identifies Googlebot and search preview controls as the relevant mechanisms for those surfaces. [4]

This is a reason to diagnose the actual configuration, not a reason to remove every restrictive directive. Some pages should remain private or excluded. Document which pages are meant to be discoverable.

Check rendering without assuming every crawler is Google

Compare the initial HTML response with the rendered page. Can the intended crawler obtain the service schedule, location, and ministry description? Does essential content require clicking, logging in, or interacting with a script before it is available?

Google documents its JavaScript rendering process and associated limitations. That capability should not be assumed for all retrieval systems. [5]

Where feasible within your platform, deliver important public information in the initial or prerendered HTML. Treat this as an accessibility and retrieval engineering choice, not a magic citation technique.

Keep a concise evidence record: URL, initial content observed, rendered content observed, tool used, and the proposed fix. A screenshot alone does not establish what the response HTML contains.

Inspect security and hosting layers

Robots.txt may permit access while a firewall or content delivery network blocks requests. Ask the web team to inspect relevant access and security logs for the affected URLs.

Look for patterns such as forbidden responses, repeated rate limiting, challenge pages, or redirects that do not occur for normal visitors. Verify crawler identity using the vendor’s current documented method before treating a request as authentic.

A user-agent string can be imitated. Testing a request with a crawler’s name may help reproduce a rule, but it does not prove a real vendor crawler can reach the site.

If a legitimate search crawler is blocked, adjust the narrow rule causing the problem. Preserve the site’s protection against abusive traffic. Do not disable the entire firewall to improve discovery.

Give developers a testable ticket

Here is a hypothetical ticket format:

FieldExample
PageThe live Bay Street campus URL
Expected resultPublic campus schedule and directions available to the intended search crawler
Observed issueA verified crawler request receives a challenge instead of the page
EvidenceTimestamp, request identifier, response status, and relevant log excerpt
Proposed investigationIdentify the security rule and confirm crawler identity
Acceptance conditionThe intended public content is retrievable without weakening unrelated protections
Follow-upRecheck logs and repeat the same page test after deployment

Do not paste credentials, private visitor details, or unrelated logs into a public issue tracker. The example describes the structure of a ticket, not a finding about your church.

Validate access before looking for citations

After a fix, repeat the technical check first. Confirm that the intended page and content can be retrieved under the tested conditions.

Then check the relevant index or webmaster tool where available. Finally, repeat a small set of AI search prompts from the visibility audit.

These are separate outcomes. A successful fetch proves that a particular request succeeded. It does not prove indexing, selection for an answer, a citation, or a visit from a real person.

Use the measurement guide to keep those outcomes separate in your report.

A technical review checklist

  • Priority public URLs return the intended page on direct load.
  • The team has checked content, not only response codes.
  • Search and training policies are documented separately.
  • Robots.txt and relevant page directives match the intended policy.
  • Canonical URLs identify the correct pages.
  • Essential facts are available through the intended rendering path.
  • Security rules have been checked using verified evidence.
  • Private information remains protected through access controls.
  • Changes have an owner, a rollback record, and a follow-up check.

Frequently asked questions

Does allowing a crawler guarantee an AI recommendation?

No. Access is a technical condition, not a promise of selection or recommendation.

Should we add llms.txt before doing this audit?

Do the access and content checks first. Google says special AI text files are not required for its AI search features. Evaluate any additional file against documented use by the systems you care about. [4]

Make decisions by purpose and vendor. Search access and training permission are different policy questions.

Make the public information retrievable

Ask your web team to review five priority pages and return a short evidence-based findings list. Fix a demonstrated barrier, verify the fix, and record what remains unknown.

Explore Not Another Church Marketing Conference to connect technical work with the content church visitors need.

Sources

[1] OpenAI: Overview of OpenAI crawlers.

[2] Google: Common crawlers and Google-Extended.

[3] Google Search Central: Introduction to robots.txt.

[4] Google Search Central: AI features and your website.

[5] Google Search Central: JavaScript SEO basics.

Sources reviewed September 27, 2026. Vendor controls can change; recheck the official documentation before editing production settings. The diagnostic workflow is an editorial recommendation.

// Next step

How findable is your church?

Take the free 2-minute audit: 10 questions, an honest score out of 100, and the 3 fixes that matter most.

30 seats · Application only · Feb 18–19, 2027

Get my score →