The short answer
You cannot control what a search engine or AI assistant says about your organization, but you can make the accurate answer easy to find. Four things do most of the work:
- A consistent identity: the same name, description, service area and contact details everywhere you appear.
- Readable content: pages that explain what you do, for whom and where, in plain text that crawlers can access.
- Search eligibility: pages that can be crawled, indexed and shown with a snippet.
- Deliberate crawler choices: allowing the crawlers that power search and answers, and deciding separately about crawlers used for AI training.
Google states that there are no additional requirements or special markup for its AI Overviews and AI Mode, and that Google Search ignores llms.txt files. Be wary of anyone selling a shortcut.
1. A consistent identity
AI assistants and search engines build a picture of your organization from many sources. Conflicting information makes that picture blurry.
- Use one form of your organization's name and one short description of what you do.
- State where you serve plainly, such as "organizations across the US" or specific states and cities.
- Keep contact details identical on your website, business listings and social profiles.
- Have a clear About page that explains who you are and what you offer, without puffery.
- If customers visit you or you serve a local area, keep your Google Business Profile complete and accurate. Google says businesses with complete and accurate information are more likely to show up in local results.
- Organization structured data can help Google understand details such as your logo and contact information. Google lists no required properties and does not promise that the information will be displayed.
2. Readable content
- One page per service or topic, with a descriptive title and headings, rather than everything on one long page.
- Answer the questions buyers actually ask: what it is, who it suits, how it works, what affects cost, what happens next.
- Put important information in text, not only in images, videos or downloadable PDFs.
- Use semantic HTML (real headings, lists and tables), which Google's generative AI guidance recommends for readability.
- Write for people first. Google's guidance on helpful, people-first content favors original, useful information from people who know the subject.
- Offer pages in other languages if you serve customers who prefer them, as separate, properly linked pages rather than automatic translation widgets.
Avoid mass-produced pages written mainly for search engines, such as near-identical "service in every city" pages. Google's spam policies treat scaled content abuse and doorway pages as spam.
3. Search eligibility
Google's guidance says that to appear in its AI features a page must be indexed and eligible to be shown with a snippet. Check that:
- important pages are not blocked by robots.txt or marked
noindexby mistake, - pages you want kept out of search use
noindexrather than a robots.txt block, because robots.txt manages crawling, not indexing, - pages load reliably and their main content does not depend on actions a crawler cannot take,
- each page has a unique, descriptive title and meta description,
- your sitemap lists the pages you care about.
4. Search crawlers and training crawlers are separate choices
Several AI companies use different crawlers for different purposes, so you can allow one and block another in robots.txt.
| Company | Used for search or answers | Used for model training | Notes |
|---|---|---|---|
| OpenAI | OAI-SearchBot (ChatGPT search) | GPTBot | ChatGPT-User visits pages when a user asks; each setting is independent |
| Perplexity | PerplexityBot (search results, not training) | None listed | Perplexity-User handles user-initiated visits (Perplexity crawlers) |
| Googlebot (Google Search, including AI features) | Google-Extended token | Google says Google-Extended does not affect inclusion or ranking in Google Search |
Blocking a search crawler usually means that service cannot cite or link to your pages. Blocking a training crawler is a policy decision about how your content may be used, and does not by itself remove you from search. OpenAI notes that changes to robots.txt can take about a day to take effect for its search results. Check each provider's current documentation before changing your file, because user agents and behavior change.
5. What you can and cannot measure
Measurement in AI search is limited, and it helps to say so plainly to anyone expecting precise reports.
- Google Search Console includes a generative AI performance report that shows impressions in AI Overviews and AI Mode. It does not report clicks.
- Bing Webmaster Tools has an AI Performance report in public preview that shows citations, cited pages and grounding queries.
- Other assistants provide little or no reporting to site owners. Some visits arrive as referrals you can see in analytics; others show up only as direct traffic or not at all.
- Nobody can promise citations or rankings. Treat any tool that claims to report your "AI share" precisely with caution, and ask how it collects its data.
A practical approach: track impressions where they are reported, watch referral traffic from AI services in your analytics, and ask new inquiries how they found you.
Checklist
Identity
- One organization name and one short description, used everywhere.
- Service area stated plainly on the site.
- Contact details identical on the website and all listings.
- Business listings claimed, complete and accurate where relevant.
Content
- A page for each main service, with a clear title and headings.
- Key facts in HTML text, not only in images or PDFs.
- Buyer questions answered on the relevant pages.
- Pages in other languages where you serve customers who prefer them.
Eligibility
- Important pages indexable and not blocked by robots.txt.
noindex, not robots.txt, used for pages kept out of search.- Sitemap submitted in Google Search Console and Bing Webmaster Tools.
Crawlers and measurement
- A written decision on search crawlers and, separately, training crawlers.
- robots.txt reviewed against current provider documentation.
- AI impression and citation reports checked monthly where available.
- "How did you hear about us?" recorded for new inquiries.
Limitations
This guide covers foundations; it cannot promise results. AI products change quickly, and each decides independently what to show and cite. The official documentation linked above was current when this article was reviewed; check it again before making significant changes.
Next step
For how buyer research is shifting, see how AI assistants are changing the way buyers find information. If your website needs clearer structure, faster pages or better content to support all of this, see web development.
Sources and further reading
Product capabilities and guidance change. These are the primary sources this article relies on, checked on the review date above.
- AI features and your website, Google Search Central
- Optimizing your website for generative AI features on Google Search, Google Search Central
- Creating helpful, reliable, people-first content, Google Search Central
- Organization structured data, Google Search Central
- Introduction to robots.txt, Google Search Central
- Google's common crawlers (Google-Extended), Google Crawling Infrastructure
- Overview of OpenAI crawlers, OpenAI
- Perplexity crawlers, Perplexity
- Generative AI performance report, Google Search Console Help
- Introducing AI Performance in Bing Webmaster Tools (public preview), Microsoft Bing Webmaster Blog
- Tips to improve your local ranking on Google, Google Business Profile Help
- Spam policies for Google web search, Google Search Central
This article is general information, not legal, accounting or security advice for your specific situation. Examples are hypothetical unless stated otherwise.