AI Search & GEO Glossary

Plain-English definitions for the concepts behind AI visibility, citations, answer engines, and generative search.

01

AI Visibility

AI visibility is the degree to which AI answer engines (ChatGPT, Perplexity, Google AI Overviews, Claude) can find, understand, cite and recommend a brand when users ask relevant questions. Unlike search rankings, AI visibility is measured across mentions, citations and sentiment inside generated answers rather than positions on a results page. A site can rank well in classic search yet be invisible to AI engines if its content lacks crawl access for AI bots, machine-readable structure, or quotable evidence. In practice, teams measure AI visibility with three numbers: mention rate (how often the brand appears in sampled answers), citation rate (how often its pages are linked as sources), and sentiment (how the brand is described when it does appear). BrandGEO's audit adds a fourth, upstream layer — whether your site is even retrievable by the crawlers those engines depend on — because a brand cannot be recommended from pages the engines never saw.

Read definition
02

Generative Engine Optimization (GEO)

Generative Engine Optimization (GEO) is the practice of improving how often and how favorably AI answer engines cite and recommend your content. The term was popularized by the 2023 Princeton/Georgia Tech/IIT-Delhi paper "GEO: Generative Engine Optimization" (KDD 2024), which measured how content changes — added statistics, quotations, citations — raise a page's share in generated answers, reporting relative visibility improvements of up to 40% in its benchmarks. In practice GEO spans crawler access, structured data, evidence-dense writing, entity consistency and freshness signals. A working GEO program runs as a loop: audit what engines can see, fix access and structure, rewrite key pages around quotable evidence, then track mentions and citations to verify the changes moved the numbers. GEO complements rather than replaces SEO: search rankings remain one input into what AI engines retrieve, but citation-worthiness decides what they quote.

Read definition
03

Answer Engine Optimization (AEO)

Answer Engine Optimization (AEO) is optimizing content to be selected as the answer by engines that respond with a single synthesized reply — AI chatbots, voice assistants and featured snippets — rather than a list of links. AEO and GEO overlap heavily and are often used interchangeably; when distinguished, AEO emphasizes question-shaped content (FAQs, direct definitions, how-to steps) while GEO covers the broader pipeline including crawl access, schema and entity signals. In practice AEO work concentrates on three page patterns: definition blocks that answer "what is X" in the first eighty words, FAQ sections marked up with FAQPage schema, and step-by-step structures engines can lift wholesale. Teams typically treat AEO as the content-formatting layer inside a broader GEO program rather than a separate discipline with its own toolchain. If your buyers ask AI tools for recommendations, AEO/GEO determines whether you are in the answer.

Read definition
04

LLM SEO

LLM SEO is an informal umbrella term for making content discoverable and citable by large language model applications — chat assistants, AI search and agents. It covers the same ground as GEO/AEO: allowing AI crawlers, publishing machine-readable structure (JSON-LD, llms.txt), writing evidence-first content, and monitoring how LLM products describe your brand. The term appears in searches like "llm seo tools" and "llm seo meaning"; most practitioners now converge on GEO as the standard name for the discipline. Treat searches for LLM SEO as a naming signal rather than a separate method: the audience asking for it usually wants the same checklist — crawler access, llms.txt, structured data, evidence-first rewriting and answer tracking. The practical difference from classic SEO is where the work lands: instead of optimizing for a ranked list of links, you are optimizing to be retrieved, trusted and quoted inside a single synthesized answer.

Read definition
05

AI Citation

An AI citation is a source link or attribution inside an AI-generated answer — the pages Perplexity lists as sources, the links ChatGPT search shows, or the sites named in Google AI Overviews. Citations are the currency of AI visibility: engines cite pages that are retrievable, parseable and contain specific, quotable evidence (numbers, definitions, comparisons). Citation behavior differs by engine: Perplexity cites nearly every claim inline, Google AI Overviews link a rotating subset of supporting pages, and ChatGPT search shows sources for browsed answers. The common thread is that engines prefer pages where a specific claim can be lifted with its evidence attached — which is why evidence-dense sections, named statistics and clearly attributed quotes consistently out-earn general marketing copy. Tracking which of your pages earn citations — and which competitors get cited instead — is the core feedback loop of GEO work.

Read definition
06

Share of Voice in AI

Share of voice in AI is the percentage of AI-generated answers about your category that mention or cite your brand versus competitors. It adapts the classic media metric to answer engines: sample the questions your buyers ask, run them across engines, and count whose products appear. Because AI answers concentrate on a few sources, share of voice tends to be winner-take-most — which is why early GEO investment in a niche can lock in outsized visibility. To make the metric operational, fix three variables before you measure: the prompt set (real buyer questions, not vanity keywords), the engine list, and the sampling cadence — AI answers vary run to run, so single-shot checks mislead. Reported movement should always come with its prompt set attached; a share-of-voice number without the underlying questions is marketing, not measurement.

Read definition
07

Prompt Tracking

Prompt tracking is monitoring a fixed set of questions (prompts) across AI engines over time to measure brand visibility: who gets mentioned, cited and recommended for each prompt. It is the AI-era equivalent of rank tracking. A useful prompt set mirrors real buyer questions ("best X for Y", "X alternative", "is X worth it") rather than vanity keywords. A workable starter set is 20-50 prompts spanning four intents: category recommendations, direct comparisons, problem questions your product solves, and brand verification ("is X legit"). Because engines answer stochastically, track each prompt across repeated runs and report trends rather than single answers — a brand that appears in three of five runs this month and five of five next month is the signal you're looking for. Movement in prompt tracking data — after shipping fixes or content — is the main evidence that GEO work is paying off.

Read definition
08

Query Fan-Out

Query fan-out is the process where an AI engine expands one user question into multiple internal sub-queries, retrieves results for each, then synthesizes a single answer. Google's AI Mode publicly describes this technique. Fan-out means your page can be cited for a question it never targeted verbatim — the engine may pull you in via a sub-query about pricing, a definition or a comparison. The practical implication: a comparison page can be cited for a pricing sub-query, and a glossary definition can be pulled into a buying-guide answer. Structuring pages so each section answers one self-contained question — with its own heading, evidence and definition — turns a single URL into many fan-out targets, which is why section-level structure matters more in GEO than page-level keyword targeting did in classic SEO.

Read definition
09

llms.txt

llms.txt is a proposed plain-text file at a site's root (like robots.txt) that gives large language models a curated, markdown-formatted map of the site's most important content. Proposed by Jeremy Howard in 2024 (llmstxt.org), it lets sites present clean, token-efficient context to AI systems instead of forcing them to parse full HTML. The format is deliberately simple: an H1 with the project name, a short summary, then curated markdown links grouped by section; a companion llms-full.txt can carry expanded content. Adoption is voluntary and engine support varies, so treat llms.txt as a complement to — not a replacement for — robots.txt, sitemaps and schema: engines that ignore it lose nothing, and engines that read it get your best pages in clean form. Publishing one is a low-cost way to control the narrative AI models see first.

Read definition
10

GPTBot

GPTBot is OpenAI's web crawler, identified by the GPTBot user-agent, used to collect publicly available web content that may improve OpenAI's models. Sites control it via robots.txt: allowing GPTBot makes content available for model training and related uses; OpenAI also operates separate agents (like OAI-SearchBot for ChatGPT search) with distinct user-agent strings and purposes, per OpenAI's published bot documentation. For GEO, the practical decision is which OpenAI bots to allow: blocking GPTBot while allowing OAI-SearchBot keeps you out of training data but visible in ChatGPT search results. Verify real GPTBot traffic before drawing conclusions from server logs: OpenAI publishes its IP ranges, and spoofed user-agents are common. The robots.txt decision is also reversible — many publishers allow search-facing bots while restricting training crawlers, and revisit the split quarterly as OpenAI's bot roster changes.

Read definition
11

AI Overview

AI Overviews are Google's AI-generated summaries shown above classic results for many queries, synthesizing information from multiple sources with links. They matter for GEO because they compress the results page: studies across 2024-2026 consistently measure lower click-through to organic links when an AI Overview is present. Appearing inside the Overview (as a cited source) recovers some of that lost visibility — which requires the same crawlability, structure and citability work as other answer engines. For measurement, separate two questions: how often Overviews appear for your target queries (trigger rate), and how often you are cited when they do (inclusion rate). Trigger rates shift with Google's rollouts and vary sharply by query type — informational queries trigger far more often than navigational ones — so a traffic drop should be diagnosed against both numbers before blaming content.

Read definition
12

IndexNow

IndexNow is an open protocol (indexnow.org) that lets websites instantly notify participating search engines — including Bing and Yandex — when URLs are added, updated or deleted, instead of waiting for recrawls. A site hosts a key file and POSTs changed URLs to the API. Implementation is deliberately light: generate a key, host the key file at your site root, and POST changed URLs to api.indexnow.org — one request can carry up to 10,000 URLs. Google does not participate; it relies on sitemaps and its own crawl scheduling. For GEO, IndexNow matters because Bing's index feeds Microsoft Copilot and is one of the retrieval sources in the ChatGPT search stack — fast Bing indexing shortens the path from publish to AI-citable. The protocol only signals that a URL changed; participating engines still decide whether and when to crawl and index it, so IndexNow shortens discovery, not evaluation.

Read definition
13

AI Search Optimization

AI search optimization is the umbrella practice of making a brand discoverable, citable and favorably represented across AI-powered search experiences — including ChatGPT search, Perplexity, Google AI Overviews and Microsoft Copilot. The term is broader than GEO or AEO alone: it covers crawler access, structured data, evidence-dense content, entity consistency, freshness signals and answer tracking as a unified discipline. In practice, AI search optimization work follows a loop: audit whether engines can crawl and parse your pages, fix structural gaps (robots.txt rules, JSON-LD, llms.txt), rewrite key content around quotable evidence, and monitor mentions and citations across engines to verify lift. The term gained traction in 2025 as marketers searched for a plain-English label that avoids the alphabet soup of GEO, AEO and LLM SEO. Teams that already run GEO programs are doing AI search optimization — the difference is naming, not method.

Read definition
14

Answer Engine

An answer engine is any search product that responds to a user question with a synthesized answer rather than a ranked list of links. The category includes ChatGPT (with browsing), Perplexity, Google AI Overviews, Microsoft Copilot, Claude with web access, and voice assistants when they give spoken answers. Answer engines share a retrieval-then-generate architecture: they fetch candidate pages from an index (often Bing or Google), select relevant passages, and compose a single reply with optional source citations. For brands, the shift from link lists to synthesized answers changes the game: visibility is no longer about ranking position but about being retrieved, trusted and quoted. A brand that ranks on page one may still be invisible if its content lacks the quotable specifics — numbers, definitions, comparisons — that engines select for their answers. The term answer engine is increasingly used interchangeably with AI search engine, though purists reserve the latter for engines that also return traditional link results.

Read definition
15

Citation Rate

Citation rate is the percentage of AI-generated answers in which a specific page or domain is linked as a source. It is one of the three core AI visibility metrics, alongside mention rate (brand named but not linked) and sentiment (how favorably the brand is described). Citation rate matters more than mention rate for driving traffic because a citation is a clickable link — Perplexity places citations inline, ChatGPT search shows source cards, and Google AI Overviews link supporting pages. Measuring citation rate requires a fixed prompt set, a consistent sampling cadence, and multi-run averaging because AI answers are stochastic: a single check can show zero citations while the true rate is 40%. Teams typically report citation rate as citations earned divided by total sampled answers for a prompt set, broken down by engine. Movement in citation rate — after deploying structured data, rewriting content, or publishing new evidence — is the primary feedback signal for GEO programs.

Read definition
16

OAI-SearchBot

OAI-SearchBot is OpenAI's web crawler dedicated to retrieving content for ChatGPT search (the browsing feature), distinct from GPTBot which collects data for model training. The user-agent string is OAI-SearchBot; OpenAI publishes its IP ranges for verification. The distinction matters for robots.txt policy: many publishers block GPTBot (training) while allowing OAI-SearchBot (search) — this keeps content out of future model training but visible when ChatGPT users browse the web for answers. If you block OAI-SearchBot, ChatGPT search cannot retrieve your pages in real time, which means you lose citation eligibility for browsed queries even if your content was in the training data. Verify real OAI-SearchBot traffic by checking source IPs against OpenAI's published ranges; user-agent strings alone can be spoofed. Review your robots.txt quarterly as OpenAI may introduce additional search-facing crawlers.

Read definition
17

ClaudeBot

ClaudeBot is Anthropic's web crawler, identified by the ClaudeBot user-agent string. It collects publicly available web content that may be used to improve Anthropic's Claude models. Unlike OpenAI, which splits training (GPTBot) and search (OAI-SearchBot) into separate crawlers, Anthropic currently operates ClaudeBot as a single crawler. Blocking ClaudeBot via robots.txt prevents Anthropic from indexing your content for training purposes. As of mid-2026, Claude's web access features (search and citations) rely on third-party search APIs rather than ClaudeBot's own index, so blocking ClaudeBot does not immediately remove you from Claude's search answers — but this architecture could change. For GEO decision-making, the practical stance for most brands is to allow ClaudeBot: the training-data inclusion increases the chance Claude mentions your brand in answers from its parametric knowledge, and the privacy risk is limited to content that is already publicly accessible.

Read definition
18

PerplexityBot

PerplexityBot is Perplexity AI's web crawler, used to build and refresh the index that powers Perplexity's search answers. Perplexity is among the most citation-heavy answer engines — it links sources inline for nearly every factual claim — making it a high-value channel for brands that produce evidence-dense content. Blocking PerplexityBot via robots.txt removes your pages from Perplexity's proprietary index; however, Perplexity also retrieves results from Bing's index, so a block may reduce but not eliminate your appearance in answers. Perplexity also documents a separate user-triggered fetcher for pages a reader pastes or asks about directly; that agent is distinct from the indexing crawler and should be evaluated on its own line in robots.txt. The pragmatic GEO stance for most brands is to allow PerplexityBot: Perplexity's citation-heavy format means every appearance is a branded, clickable link. Because user-agent strings are trivially spoofed, verify real PerplexityBot traffic by reverse DNS lookup against Perplexity's published documentation before making robots.txt decisions on log data alone.

Read definition
19

AI Crawler

An AI crawler is a web bot operated by an AI company to collect content for model training, retrieval-augmented search, or both. Major AI crawlers include GPTBot and OAI-SearchBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity AI), Google-Extended (Google, for Gemini training), and Bytespider (ByteDance). Each crawler has its own user-agent string and is controllable via robots.txt. For GEO, the critical distinction is purpose: training crawlers collect data that shapes a model's parametric knowledge (what it "knows" without searching), while search crawlers retrieve pages in real time for citation in answers. Blocking a training crawler keeps you out of future model weights; blocking a search crawler removes you from live search answers. Most GEO programs allow search crawlers while making a case-by-case decision on training crawlers. Maintain an up-to-date crawler list and review robots.txt quarterly — AI companies launch new bots faster than the industry tracks them.

Read definition
20

Zero-Click Search

A zero-click search is a query where the user gets their answer directly on the search results page — via a featured snippet, knowledge panel, or AI Overview — without clicking through to any website. Studies consistently find that zero-click queries account for a significant and growing share of Google searches, and AI Overviews accelerate this trend by synthesizing multi-source answers at the top of the page. For brands, zero-click search reshapes the value equation: traditional SEO metrics like click-through rate decline, but being the cited source inside an AI Overview or featured snippet maintains brand visibility even when the user never visits your page. The GEO response is to optimize for citation rather than clicks: structure content so engines can lift a quotable passage with attribution, include your brand name near key facts, and track citation rate alongside traffic. A page that earns zero clicks but is cited in every AI Overview for its category still builds brand awareness and trust.

Read definition
21

Entity SEO

Entity SEO is the practice of establishing your brand, products and key people as recognized entities in search engines' knowledge systems — not just keywords on a page, but distinct things with verified attributes and relationships. Google's Knowledge Graph, Bing's entity index, and AI models' parametric memories all operate on entities: they store facts about named things (founding year, headquarters, product categories, executive names) and surface them in answers. For GEO, entity SEO is foundational because AI engines trust entities they can verify across multiple sources. Practical entity SEO work includes: publishing consistent Organization JSON-LD with @id and sameAs, maintaining accurate profiles on Wikipedia, Crunchbase, G2 and LinkedIn, ensuring NAP (name, address, phone) consistency across directories, and building editorial mentions that use your exact brand name in context. The test is simple: ask an AI engine "What is [your brand]?" — if it returns a confident, accurate description, your entity is established; if it hedges or hallucinates, your entity signals are weak.

Read definition
22

Knowledge Graph

A knowledge graph is a structured database of entities (people, organizations, places, products) and the relationships between them, used by search engines and AI systems to understand the world beyond keywords. Google's Knowledge Graph, introduced in 2012, powers knowledge panels and feeds into AI Overviews; Bing maintains a similar entity index that informs Microsoft Copilot answers. AI language models also build implicit knowledge graphs from training data — they store entity facts in model weights and retrieve them during generation. For GEO, the knowledge graph determines whether an AI engine recognizes your brand as a real, distinct entity or confuses it with similarly named things. Strengthening your knowledge graph presence requires consistent structured data (Organization JSON-LD with @id), verified profiles on authoritative platforms (Wikipedia, Wikidata, Crunchbase), and editorial coverage that connects your brand name to specific attributes. A brand with a strong knowledge graph entry gets cited with confidence; a brand without one gets hedged or omitted.

Read definition
23

Structured Data

Structured data is machine-readable markup added to web pages that explicitly labels content for search engines and AI systems — telling them not just what the page says but what the content means. The dominant vocabulary is Schema.org, typically implemented as JSON-LD embedded in the page's HTML head. For classic SEO, structured data triggers rich results (star ratings, FAQ accordions, how-to steps); for GEO, it serves a different purpose: it gives AI engines high-confidence facts to cite. An Organization block with @id, sameAs and foundingDate tells an AI engine that your brand is a verified entity, not just a keyword match. Structured data does not guarantee citation — engines still evaluate content quality and relevance — but it reduces the chance of misidentification or hallucination. Priority types for GEO are Organization (brand identity), Article (content authorship), FAQPage (question-answer pairs), HowTo (step-by-step instructions), and Product (specifications and reviews). Validate with Google's Rich Results Test and schema.org's validator before deploying.

Read definition
24

JSON-LD

JSON-LD (JavaScript Object Notation for Linked Data) is the recommended format for embedding structured data in web pages. It sits in a <script type="application/ld+json"> block in the page head, separate from the visible HTML — which means it can be added without touching the page layout. Google, Bing and AI engines all parse JSON-LD to extract entity facts and content metadata. For GEO, the three essential JSON-LD blocks are: Organization (brand identity, @id, sameAs links to verified profiles), Article (authorship, datePublished, headline), and FAQPage (question-answer pairs that engines can lift directly into answers). The @id field is particularly important: it creates a stable, unique identifier for your brand entity that search engines and AI systems can reference across pages and sites. Common mistakes include omitting @id, listing unverified sameAs URLs, and duplicating blocks across pages with inconsistent data — all of which weaken rather than strengthen entity signals.

Read definition
25

Schema Markup

Schema markup is the implementation of Schema.org vocabulary on web pages to provide search engines and AI systems with explicit, structured descriptions of page content. Schema.org defines hundreds of types (Organization, Article, Product, FAQPage, HowTo, Review) and properties; markup is typically deployed as JSON-LD, though Microdata and RDFa are also valid. For GEO, schema markup serves two functions: it triggers rich results in traditional search (which can feed into AI retrieval pools), and it provides AI engines with high-confidence entity and content signals. Not all schema types matter equally for AI visibility — prioritize Organization (brand identity), Article (content metadata), FAQPage (direct question-answer pairs), and Product (specifications) over decorative types like BreadcrumbList. The key quality check is truthfulness: schema markup that contradicts visible page content (inflated ratings, unearned review counts, fabricated authorship) is penalized by Google and treated as a negative trust signal by AI engines.

Read definition
26

AI Hallucination (Brand)

A brand hallucination occurs when an AI engine states confidently incorrect information about your company — wrong founding dates, fabricated product features, invented partnerships, or confused identities with similarly named brands. Hallucinations arise because language models generate plausible text from statistical patterns rather than verified facts; when entity signals are weak or contradictory, models fill gaps with confident-sounding fiction. For brands, hallucinations range from embarrassing (wrong CEO name) to damaging (false safety claims, invented lawsuits). The defense is entity strength: consistent structured data (Organization JSON-LD with verified sameAs links), authoritative third-party coverage (Wikipedia, Crunchbase, G2), and content that states key facts in clear, quotable form near the top of pages. Monitoring for hallucinations is part of AI brand monitoring: run your brand name through major engines monthly and flag answers that contain invented facts. When hallucinations persist, the fix is usually upstream — strengthening the entity signals that engines use to verify claims.

Read definition
27

Sentiment in AI Answers

Sentiment in AI answers is how favorably or unfavorably an AI engine describes your brand when it appears in a generated answer — the qualitative complement to mention rate and citation rate. An AI engine might mention your brand frequently but frame it negatively ("Brand X has faced criticism for...") or with hedging ("some users report issues with..."). Sentiment is shaped by the training data the model absorbed and the real-time sources it retrieves: negative reviews on Reddit, critical articles, and complaint threads all feed into how engines characterize your brand. Measuring sentiment requires reading actual AI answers — not just counting mentions — and coding them as positive, neutral or negative. For GEO, sentiment improvement comes from the same evidence-first content strategy that drives citations: publishing specific, verifiable claims with supporting data crowds out vague negative signals with concrete positive evidence. Brands with strong first-party content and third-party endorsements (G2 reviews, press coverage, case studies) consistently earn better sentiment than those relying on marketing copy alone.

Read definition
28

AI Brand Monitoring

AI brand monitoring is the ongoing practice of tracking how AI answer engines mention, cite and describe your brand across a set of buyer-relevant prompts. It extends traditional brand monitoring (media mentions, social listening, review tracking) into the AI answer layer. A working monitoring program fixes three variables: a prompt set (20-50 real buyer questions), an engine list (ChatGPT, Perplexity, Gemini, Google AI Overviews at minimum), and a cadence (weekly or biweekly, with multi-run sampling to account for stochastic variation). Each monitoring cycle captures four signals: mention rate, citation rate, sentiment, and competitor presence — who else appears in the same answers. The output is a trend dashboard that shows whether GEO work is moving the numbers. Manual monitoring is possible with spreadsheets; tools like BrandGEO automate the prompt-engine-sampling loop and surface changes over time. The key discipline is consistency: changing the prompt set between cycles makes trends unreadable.

Read definition
29

lastmod

lastmod is the XML sitemap element that declares when a page was last meaningfully updated. Search engines and AI crawlers use lastmod as a freshness signal to prioritize recrawling: a page whose lastmod moved recently is more likely to be refetched than a stale one. For GEO, accurate lastmod matters because AI engines prefer fresh content for time-sensitive queries — and incorrect lastmod (setting today's date on unchanged pages) erodes trust in your sitemap, causing engines to ignore the signal entirely. Google's documentation explicitly warns against updating lastmod without meaningful content changes. Best practice: update lastmod only when the page's substantive content changes (new data, revised conclusions, added sections), not for cosmetic edits (typo fixes, layout tweaks). Combine lastmod with IndexNow pings for Bing-indexed engines (Microsoft Copilot, ChatGPT search via Bing) to minimize the delay between publish and AI-citable. Automate lastmod in your CMS or build pipeline so it reflects actual content diffs.

Read definition
30

hreflang

hreflang is an HTML link attribute (or sitemap annotation) that tells search engines which language and regional version of a page to show to users in different locales. The tag uses the format <link rel="alternate" hreflang="en" href="..." /> and must be reciprocal: if the English page points to the Chinese version, the Chinese page must point back. For GEO, hreflang matters because AI engines that retrieve results from Google or Bing inherit the language-matching logic: a query in Chinese should retrieve your /zh/ page, not the English version. Incorrect hreflang (missing reciprocal tags, wrong language codes, broken URLs) can cause the wrong language variant to appear in AI answers — or cause both variants to compete and dilute each other's authority. Common mistakes include using country codes ("cn") instead of language codes ("zh"), omitting the x-default fallback, and failing to update hreflang when adding new language variants. For bilingual GEO sites, audit hreflang with Google Search Console's international targeting report and validate that AI engines return the correct language variant for locale-specific prompts.

Read definition