Glossary
Thirty terms, plainly defined. No marketing spin, just what each word means.
- AEO (answer engine optimization)
- The practice of structuring a website so AI answer engines such as ChatGPT, Perplexity and Gemini can find, parse, cite and recommend it. AEO prioritizes direct answers, clear structure and verifiable facts over the ranking signals that drive traditional search engine optimization.
- Agentic readiness
- Whether an AI agent acting on a user's behalf can complete a task on a site — check pricing, book an appointment, start a purchase — without a human relaying instructions. It depends on clear navigation, structured data, working forms and plainly stated facts.
- AI Overviews (Google)
- The AI-generated summary Google shows above traditional results for many queries, synthesizing several sources into one answer with citation links. There is no separate AI Overviews crawler; eligibility depends on Google's standard Search index, so appearing there requires the same crawlability and indexing as ranking normally.
- Answer engine
- A search tool — ChatGPT, Perplexity, Gemini, or Google AI Overviews — that returns one synthesized answer instead of a ranked list of links. Answer engines draw on crawled or indexed content and often cite sources, which makes earning a citation the answer-engine equivalent of ranking well in traditional search.
- Bing index
- Microsoft's search index, which powers Bing search and supplies grounding data to Copilot. Some AI answer engines have also drawn on Bing's index for web results rather than crawling and indexing the web independently, so Bing coverage can affect visibility beyond Bing's own market share.
- ChatGPT-User
- OpenAI's user-agent for real-time fetches ChatGPT performs when a live user asks it to look at a specific page — distinct from GPTBot's training crawl and OAI-SearchBot's indexing crawl. Blocking ChatGPT-User in robots.txt stops ChatGPT from opening that page on a user's request.
- Citation (AI search)
- The explicit source link an answer engine shows next to a claim, letting a reader verify where the information came from. Getting cited requires content the model can parse cleanly and attribute — the AI-search equivalent of ranking, and often a stronger source of qualified traffic per visit.
- Claude-SearchBot
- Anthropic's crawler that indexes web content to build and improve Claude's search feature, separate from ClaudeBot, used for model training, and Claude-User, which fetches pages live when a person asks Claude a question. Each can be allowed or blocked independently in robots.txt.
- Crawler (AI)
- An automated program that fetches web pages on behalf of an AI company — to train models (GPTBot, ClaudeBot), build a search index (OAI-SearchBot, Claude-SearchBot, PerplexityBot), or fetch a page live for a user (ChatGPT-User, Perplexity-User, Claude-User). Each type is controlled separately in robots.txt.
- DefinedTerm / DefinedTermSet
- Schema.org types for marking up a glossary. A DefinedTermSet is the collection — the glossary itself; each DefinedTerm inside it is one term with a name and description. This structure lets AI systems read exact definitions directly, instead of inferring them from surrounding prose.
- Entity (search)
- A distinct, identifiable thing — a person, place, organization, or product — that a search or answer engine recognizes and connects across sources. A clear, consistently named entity profile helps AI systems associate your business with the correct facts rather than a similarly named competitor.
- FAQPage schema
- A schema.org type marking a page's questions and answers in structured form. Google deprecated FAQ rich results in Search in May 2026, so the markup no longer produces expanded snippets there, but AI answer engines can still parse it to extract clean question-answer pairs when composing a response.
- GEO (generative engine optimization)
- Optimizing content so generative AI systems draw on it when composing an answer, rather than only ranking it in a results list. GEO overlaps with AEO but is used more broadly, for writing and structuring content that any generative model — not only a dedicated answer engine — can use.
- Google-Extended
- A token added to robots.txt that controls whether Google may use content it has already crawled to train Gemini and ground its responses. It has no separate user agent of its own, applies only to pages Googlebot already fetches, and does not affect standard Search indexing.
- GPTBot
- OpenAI's crawler that collects publicly available web content to train and improve its generative AI models. Disallowing GPTBot in robots.txt opts a site out of having its content used for that training; it does not affect whether ChatGPT can search or fetch the page for a user.
- HowTo schema
- A schema.org type for step-by-step instructions, listing each step with text and optional media. Google deprecated HowTo rich results in 2023, removing the visual snippet, but the structured markup can still help an AI system parse a procedure into discrete, orderable steps when answering how-to questions.
- JSON-LD
- JavaScript Object Notation for Linked Data — the standard format for embedding schema.org structured data in a script tag (type application/ld+json). It keeps machine-readable facts separate from visible HTML, so search engines and AI systems can read a page's entities and relationships without parsing its design.
- llms.txt
- A convention, defined at llmstxt.org, for a plain markdown file at a site's root that summarizes its key pages for AI systems in a format simpler than HTML. No major AI crawler has confirmed it reads llms.txt, but publishing one costs little and may help as adoption grows.
- LocalBusiness schema
- A schema.org type for marking up a business's name, address, phone number, hours and category. It is the structured-data backbone of local search and increasingly of local AI answers — an agent asked whether a shop is open now needs this markup to answer without guessing from unstructured text.
- MCP (Model Context Protocol)
- An open standard, introduced by Anthropic and now stewarded under the Linux Foundation, that lets AI applications connect to external tools and data sources through one common interface. It underlies agentic readiness: it defines how an agent can call your systems directly, not just read your pages.
- NAP consistency (name, address, phone)
- Whether your business name, address and phone number match exactly across your website, directories, social profiles and structured data. Mismatches make it harder for search engines and AI systems to confirm which listings refer to the same business, which weakens local trust signals and local answer accuracy.
- OAI-SearchBot
- OpenAI's crawler that builds the index behind ChatGPT's search feature, functioning like a traditional search-engine crawler rather than a training crawler. Allowing OAI-SearchBot in robots.txt is what makes a page eligible to appear as a cited source in ChatGPT search answers.
- Organization schema
- A schema.org type describing a company: legal name, logo, address, social profiles and contact points. It helps search engines and AI systems build a confident, disambiguated profile of who you are, underlying both local knowledge panels and AI answers that need to describe or contact a business correctly.
- Perplexity-User
- Perplexity's user-agent for on-demand page visits, triggered when someone asks Perplexity a question that requires reading a specific page live. It is distinct from PerplexityBot's ongoing indexing crawl, and each can be allowed or blocked independently in robots.txt.
- PerplexityBot
- Perplexity's automated crawler that indexes web content for its search and answer features; Perplexity states it is not used to train foundation models. Its IP ranges are published so site owners can verify requests, and it can be allowed or disallowed independently of Perplexity-User in robots.txt.
- Question readiness
- Whether a page directly and unambiguously answers the specific questions people, and AI agents, ask about a topic — near the top of the content, in plain language. It is one of the six checks Agent Ready Search scores, because answer engines favor content that resolves a question without inference.
- Robots.txt
- A plain-text file at a site's root that tells crawlers, including AI crawlers, which parts of the site they may access. Each AI crawler — GPTBot, ClaudeBot, PerplexityBot and others — is named as its own user agent, so a site can allow some and block others in the same file.
- Schema.org
- A shared vocabulary of structured-data types — Organization, LocalBusiness, FAQPage, Article, Product and hundreds more — maintained jointly by Google, Microsoft, Yahoo and Yandex. Marking up a page with schema.org vocabulary, usually via JSON-LD, gives search engines and AI systems a precise, agreed-upon way to read what a page is about.
- Structured data
- Machine-readable markup, typically JSON-LD using schema.org vocabulary, embedded in a page to describe its content explicitly rather than leaving it to be inferred from prose and layout. It is the mechanism behind rich results, knowledge panels, and much of what lets an AI system quote a fact with confidence.
- Trust and authority (AI search context)
- The signals an answer engine weighs before citing or recommending a source: consistent facts across a site and the wider web, clear authorship, verifiable claims, working citations, and a coherent entity profile. Unlike backlink-based SEO authority, AI trust leans heavily on internal consistency and verifiability.