Contents
- Build a report with four separate layers
- Keep the test set stable enough to compare
- Define each measure before using it
- Show counts beside percentages
- Track accuracy at the level of individual claims
- Measure identifiable referrals without calling them the total
- Define useful actions carefully
- Ask visitors how they found you
- Keep a change log so the trend has context
- A monthly report leadership can use
- Frequently asked questions
- Report progress people can understand
- Sources and methodology
Measure church AI search performance in separate layers: what AI tools say in controlled tests, whether the facts are accurate, what identifiable website traffic arrives, and what visitors report or do next. Keep the underlying counts and test conditions visible. None of these measures alone captures the entire path from discovery to attendance.
“ChatGPT mentioned us” is a useful observation. It is not a complete performance report.
The practical question for a church marketing team is whether its work is making useful information easier to find and helping people take meaningful next steps. A small, consistent report can help answer that without pretending to measure activity you cannot observe.
Build a report with four separate layers
| Layer | What to measure | What it establishes |
|---|---|---|
| Visibility tests | Mentions, recommendations, and citations for a fixed prompt set | What happened in those tests |
| Accuracy | Verified wrong claims and unresolved issues | Whether tested answers help or mislead |
| Website activity | Identifiable AI referrals and relevant actions | Observable activity under your analytics setup |
| Visitor outcomes | Inquiries, registrations, and self-reported discovery | Specific actions or accounts of discovery |
Do not merge the layers into one score. A higher mention rate alongside more incorrect service times is not unqualified progress.
Use the AI visibility audit to establish the initial test set and save the answers behind the report.
Keep the test set stable enough to compare
Write down the prompts, platforms, modes, locations, and frequency before measuring change. Preserve the exact wording of your core questions for successive rounds.
Separate branded questions from unbranded questions. Being found when someone names your church tests a different task from appearing when someone asks for a relevant type of church or program in your city.
If you add new questions, report them as a new or expanded set. Do not compare a ten-question baseline with a twenty-question follow-up as though the composition stayed constant.
Keep platforms separate too. A change in one system can be obscured when it is averaged with several others. When a platform or mode changes materially, record that in the reporting notes.
Define each measure before using it
Here are practical definitions your team can adopt.
Mention rate: completed responses that name your church divided by all completed responses in the specified prompt set.
Recommendation rate: completed responses that explicitly recommend your church as an option divided by completed responses in the specified set. Retain the sentence supporting the classification.
Owned-site citation rate: completed responses that link to a page on your church’s website divided by completed responses in the specified set.
Verified errors: a count of specific claims that conflict with approved current facts. Report their substance and status, not just their total.
Useful next-step coverage: completed branded responses that provide a relevant, working contact or destination divided by completed branded responses.
Define treatment of failures. Exclude a failed request from completed-response rates and show its count separately. Keep a completed answer that omits your church in the denominator.
For Google AI Overviews, record feature appearance separately. If no overview appears, there was no AI Overview response in which to be mentioned. Report both the number of queries tested and the number producing the feature so the context remains visible.
Show counts beside percentages
The following is an invented example for one platform and an unchanged prompt set. It illustrates reporting arithmetic, not results from an actual church.
| Measure | Baseline | Follow-up | Interpretation |
|---|---|---|---|
| Unbranded mentions | 4 of 20: 20% | 7 of 20: 35% | Increase of 15 percentage points in the test set |
| Owned-site citations in unbranded answers | 2 of 20: 10% | 5 of 20: 25% | More tested responses linked to the church’s site |
| Branded answers with a useful next step | 6 of 10: 60% | 8 of 10: 80% | More tested answers provided a usable action |
| Verified wrong claims | 3 | 1 | Review which errors were corrected and which remain |
An increase from 20% to 35% is 15 percentage points, or a 75% relative increase. The percentage-point description and raw counts are usually clearer for a small test set.
Do not label the 35% figure as market share. It does not represent the proportion of local residents seeing your church or the share of all AI church searches.
Track accuracy at the level of individual claims
Maintain an issue register beside the metric table. Give recurring errors stable identifiers so the same wrong service time does not look like a new issue every month.
For each issue, record the claim, correct fact, source of truth, affected platform, first observed date, last observed date, and owner. Mark the source as unknown when no citation establishes its origin.
Distinguish source correction from answer correction. Your website may be fixed while a tested answer still repeats the old information. Both statuses belong in the record.
Use the church information guide to coordinate updates across pages and profiles.
Measure identifiable referrals without calling them the total
In GA4, the Traffic acquisition report uses session-scoped acquisition information. Session source and medium can help identify visits attributed to referring AI services when that information is available. [1]
Start with the sources actually present in your data. Verify which ones belong to relevant services and maintain a reviewed grouping or report filter. Do not assume one permanent domain list captures all AI-related visits.
Compare AI referral sessions, landing pages, engagement, and meaningful actions over consistent date ranges. Document changes to consent, tags, filters, and site behavior that could affect measurement.
Some discovery journeys will not provide an identifiable AI referral. A person may read an answer and later reach the church another way. Report what your analytics can identify without assigning all unattributed activity to AI.
Google includes AI search feature traffic within Search Console’s Web reporting. That aggregate should not be presented as an isolated AI referral total. [2]
Define useful actions carefully
Choose actions that correspond to what the page helps someone do. Possible examples include a successful visit inquiry, a completed public-program registration, or a click to a directions provider.
These actions have different meanings. A directions click does not prove arrival. A registration does not prove attendance. An inquiry can be valuable even if it is not followed by a visit.
Have your analytics owner document each event’s trigger and verify that it fires as intended. A form submission event should represent a successful submission rather than merely a button click or validation error.
Avoid collecting names, email addresses, prayer requests, or sensitive form contents in analytics parameters. Count the relevant action without copying the person’s message into your reporting tools.
Ask visitors how they found you
Offer an optional discovery question on an appropriate form or in a welcome conversation. Include an AI assistant or AI search option alongside other relevant channels and allow an open-ended response.
People may remember several influences: a friend’s invitation, a search, a video, and an AI-generated suggestion. Let them describe that path rather than forcing one precise source when they do not know it.
Keep self-reported discovery separate from analytics attribution. If someone says ChatGPT helped them find the church, that is useful reported evidence. It does not justify assigning every similar visit to ChatGPT.
Watch for duplication when the same person completes more than one form. Where your existing process can responsibly reconcile inquiries, use it; do not collect more personal data simply to improve a marketing dashboard.
Keep a change log so the trend has context
Record meaningful changes alongside the measures:
- Website page published or substantially revised.
- Service schedule or campus change.
- Profile or directory correction.
- Technical access issue resolved.
- Analytics configuration change.
- Seasonal campaign or major event.
- Prompt set, platform, or testing-mode change.
If mentions rise after you update pages and profiles, say that the increase followed those changes. Without a design that isolates the cause, do not claim that one schema field or paragraph produced the result.
The crawler guide can help document technical fixes precisely. “Verified search access restored for these pages” is a stronger report than “AI optimization completed.”
A monthly report leadership can use
Use this structure for a one-page summary:
| Section | What to report |
|---|---|
| Coverage | Dates, platforms, modes, prompt counts, repeat rounds, and failures |
| Visibility | Branded and unbranded mention and citation counts by platform |
| Accuracy | New, resolved, and remaining errors, with visitor impact |
| Website activity | Identifiable referrals and defined actions, with measurement caveats |
| Visitor evidence | Self-reported discovery and meaningful inquiries, kept separate from session totals |
| Work completed | Specific pages, profiles, or technical issues addressed |
| Next actions | Three priorities, named owners, and review dates |
The summary should lead to a decision: maintain what is working, correct an error, investigate a gap, or improve the next page.
Frequently asked questions
How often should we report?
Choose a cadence that matches your capacity and change volume. A monthly core review is a practical starting point, with faster checks for serious factual errors or major schedule changes. It is a workflow choice, not a platform requirement.
Can we compare ourselves with nearby churches?
You can record which organizations appear in the same unbranded tests. Describe the sample and context. Do not present the result as a complete market ranking or infer the other churches’ visitor outcomes.
Should we buy a visibility tracking tool?
Consider it when maintaining consistent tests becomes burdensome. Evaluate whether it exposes prompts, modes, dates, source URLs, and actual answers. A score without inspectable evidence is difficult to act on.
What if the numbers do not improve?
Review the specific results. You may still have corrected visitor-impacting errors or made a program easier to access. If discovery remains weak, investigate relevant content and technical gaps rather than changing the test until the score rises.
Report progress people can understand
Start with a stable test set, a short accuracy log, and a few defined website actions. Keep the claims proportional to the evidence and use each report to choose the next useful improvement.
Explore Not Another Church Marketing Conference to turn measurement into practical work with other church marketing leaders.
Sources and methodology
[1] Google Analytics Help: GA4 Traffic acquisition report.
[2] Google Search Central: AI features and your website.
Sources reviewed September 27, 2026. The reporting framework and formulas are proposed operational measures. The numerical example is fictional, not research or church performance data. Manual tests do not estimate population-wide exposure.