45 terms
AI Visibility GlossaryAI Visibility Terms A–Z
AI visibility describes how often, and how accurately, a brand or source appears in AI-generated answers. This section covers large language models, RAG, embeddings, semantic search, grounding and hallucination from one angle: how AI answers use web content.
The aim is to explain technical terms without hype: what models can and cannot do, and why visibility measurements are based on sampling.
Searches term names and alternate names; on category pages also the short definitions.
A
- AI AgentAn AI agent is a software system that uses a language model as its decision-maker to pursue a goal: it plans steps, calls tools such as search, APIs, code execution or file operations, evaluates the results and chooses what to do next. Unlike a single question-and-answer exchange, it carries out a task over multiple steps while interacting with external systems. Reliable agents depend on permission limits, human approval and logging.Professional · AI Visibility · Software
- AI CitationAn AI citation is a link, footnote or source card in an answer generated by an AI-powered search or answer system that shows which web page the information came from. Citations let users verify claims and are the main route by which websites gain visibility and traffic from AI answers. Being selected as a cited source is never guaranteed.Core · GEO · AI Visibility
- AI CrawlerAn AI crawler is a bot operated by an AI company that visits web pages and retrieves their content. AI crawlers serve three main purposes: collecting content for model training (GPTBot, ClaudeBot), building an index for AI search (OAI-SearchBot, PerplexityBot), and fetching a page on demand when a user asks (ChatGPT-User). Most can be managed with robots.txt rules.Core · GEO · AI Visibility
- AI SearchAI search is a search experience that responds to a question with an answer written by a large language model instead of only a list of links. The system typically breaks the question into several searches, retrieves relevant web pages and presents a synthesized answer together with source links. Google AI Overviews and AI Mode, ChatGPT search and Perplexity are examples.Core · GEO · AI Visibility
- AI Share of VoiceAI share of voice is a measure of how much space a brand takes up, relative to the competitors being tracked, in the answers AI systems give to a defined set of questions. It is usually calculated as the brand's mentions or citations divided by the total for all tracked brands in the same answers. Because answers vary from run to run, it is estimated by sampling rather than measured exactly.Advanced · GEO · AI Visibility
- AI VisibilityAI visibility describes how often a brand is mentioned, how accurately it is described and whether it is cited as a source in answers from AI-powered systems such as ChatGPT search, Google AI Overviews and AI Mode, Perplexity or Claude. Because these answers vary from one run to the next, it is measured by repeatedly sampling a set of prompts rather than by a single ranking.Core · GEO · AI Visibility
C
- ChunkingChunking is the practice of splitting long documents into smaller, self-contained pieces of text before they are embedded and indexed. RAG and AI search systems retrieve these chunks rather than whole documents. Chunk size, the overlap between neighboring chunks and where the boundaries fall all have a direct effect on whether the right information can be found.Professional · AI Visibility
- Context EngineeringContext engineering is the practice of deciding what information enters a language model's context window on each call, in what order and at what token cost, and of managing that information as a task unfolds. It covers system instructions, tool definitions, retrieved documents, conversation history and memory, treating prompt writing as one part of a wider problem alongside retrieval and memory management.Advanced · AI Visibility
- Context WindowA context window is the maximum number of tokens a language model can take into account in a single request. System instructions, conversation history, attached documents, retrieved sources and the model's own response all share this budget. Anything that does not fit in the window simply does not exist for the model during that request.Professional · AI Visibility · Software
E
- EmbeddingAn embedding is a numerical vector that represents the meaning of a piece of text, an image or other data, produced by an embedding model. Content with similar meaning ends up with vectors that sit close together, so related items can be found even when they share no words. Embeddings power semantic search, retrieval-augmented generation (RAG) and recommendation systems.Professional · AI Visibility · Software
- EntityAn entity is something that search engines and AI systems recognize as a single, distinguishable thing: a person, company, product, place or concept. Entities are identified by their identity and attributes rather than by the words used to name them. Consistent naming, structured data and sameAs links help systems merge different spellings into one entity and keep similarly named entities apart.Core · SEO · GEO · AI Visibility
- Entity DisambiguationEntity disambiguation is the task of choosing the correct real-world entity when a name in a text could refer to several, for example deciding whether "Mercury" means the planet, the chemical element or the singer. Systems rely on surrounding context, candidate lists and persistent identifiers such as Wikidata QIDs. It is the core decision step in entity linking and closely related to entity resolution.Advanced · GEO · AI Visibility
F
- Fine-TuningFine-tuning is the process of continuing to train an already pre-trained AI model on a smaller dataset prepared for a specific task, style or domain. It changes the model's weights, which separates it from prompting, where only the input changes, and from RAG, where documents are retrieved at answer time. Typical goals are a consistent output format, a narrow classification task or a house writing style.Professional · AI Visibility · Software
- Foundation ModelA foundation model is an AI model trained on broad, diverse data at scale, usually with self-supervised methods, that can be adapted to a wide range of tasks through fine-tuning, instructions or added components. The term was popularized by Stanford researchers in 2021. Large language models are the best-known examples, but image, speech and multimodal models can be foundation models too.Professional · AI Visibility
- Function CallingFunction calling, also called tool use, is a mechanism in which a language model responds not with prose but with a structured request, usually JSON, to call a function the application has defined, along with the arguments to pass. Each function is described by a name, a description and a JSON Schema for its parameters. The application, not the model, executes the call and returns the result to the model.Advanced · AI Visibility · Software
G
- Generative AIGenerative AI is the umbrella term for AI models that learn patterns from training data and use them to produce new content: text, images, audio, video or code. What separates them from models that classify existing data or predict a value is that their output is new material. Large language models, diffusion-based image models and speech synthesis systems are well-known examples. Their output can be fluent and still wrong, so it needs checking.Core · AI Visibility
- GEO (Generative Engine Optimization)Generative Engine Optimization (GEO) is the practice of helping a brand and its content be understood correctly, mentioned and cited as a source in answers produced by AI-powered search and answer systems. It builds on SEO fundamentals such as crawlability and indexing, and focuses on clarity, consistent entity information and quotable content. No technique can guarantee inclusion in AI-generated answers.Core · GEO · AI Visibility
- GroundingGrounding is the practice of making a language model base its answer on verifiable sources supplied at answer time — web results, documents or database records — rather than only on what it learned during training. The aim is to keep answers current and checkable and to reduce the risk of fabricated information. A grounded answer usually shows the sources it relied on.Professional · GEO · AI Visibility
H
- Hallucination (AI)An AI hallucination is output in which a model states something false, fabricated or unsupported by its sources in fluent, confident language. Typical examples are invented statistics, citations to sources that do not exist, wrong dates, or services attributed to the wrong company. It stems from language models generating likely-sounding text rather than verifying facts.Core · AI Visibility
- Hybrid SearchHybrid search is an approach that runs lexical search, such as BM25 keyword matching, and vector search over embeddings for the same query, then merges both result lists into a single ranking. Lexical search catches exact matches like product codes and precise phrases, while vector search catches the same meaning expressed in different words. The lists are usually combined with Reciprocal Rank Fusion (RRF) or a weighted score combination.Advanced · AI Visibility
I
K
L
M
- Machine LearningMachine learning is the branch of artificial intelligence in which computers learn patterns from example data, rather than following explicitly written rules, in order to make predictions or decisions. A model adjusts its parameters on training data and is expected to generalise to data it has never seen. Spam filters, recommendation systems, demand forecasting, fraud detection and large language models are all applications of machine learning.Core · AI Visibility · Software
- Model Context Protocol (MCP)The Model Context Protocol (MCP) is an open protocol that standardises how AI applications connect to external data sources, tools and workflows. Anthropic open-sourced it in November 2024. Its architecture has a host (the AI application), one client per connected server, and servers that expose tools, resources and prompts. Messages between clients and servers use the JSON-RPC 2.0 format.Advanced · AI Visibility · Software
- Multimodal ModelA multimodal model is an AI model that can take more than one type of data, or modality, as input or produce it as output: text, images, audio or video. Such a model can answer questions about a photo, read a chart or generate an image from a description. Because the different modalities are processed in a shared representation space, the model can connect an object in an image with the words that describe it.Professional · AI Visibility
N
- Named Entity Recognition (NER)Named entity recognition (NER) is the natural language processing task of finding names and similar expressions in text and sorting them into predefined types such as person, organisation, location, date, product or monetary amount. In “Maria Lopez opened the Leeds office in March”, Maria Lopez is a person, Leeds a location and March a date. NER is usually the first step in how search engines and AI systems understand text in terms of entities.Advanced · GEO · AI Visibility
- Natural Language Processing (NLP)Natural language processing (NLP) is the field of artificial intelligence and linguistics concerned with enabling computers to analyse, interpret and generate human language, both written and spoken. Typical NLP tasks include text classification, sentiment analysis, entity recognition, machine translation, summarisation and question answering. The field began with hand-written rules and now relies mostly on machine learning and Transformer-based language models.Core · AI Visibility
P
- Parametric KnowledgeParametric knowledge is the information an AI model absorbed during training and now holds in its weights (parameters). The model reaches it from the inside, without consulting any outside source. Information that arrives at answer time through search, document retrieval or a tool call is called retrieved, or non-parametric, knowledge. Parametric knowledge is frozen when the training data was collected, cannot be traced to a source and can be recalled wrongly.Advanced · AI Visibility
- Passage RetrievalPassage retrieval is a retrieval approach that finds and ranks paragraph- or section-sized passages within documents, rather than whole documents, in response to a query. AI search systems typically build their answers from such passages. Google also documents a passage ranking system that identifies individual sections of a web page to better understand how relevant the page is to a search.Advanced · GEO · AI Visibility
- PromptA prompt is the input text given to a language model to produce a response: a question, a task description, examples or a document to work on. In chat applications, the message the user types is called the user prompt. The model, however, reads more than that: it responds to the whole context, including system instructions, conversation history and any attached or retrieved sources. How clearly a prompt is written directly shapes the quality of the answer.Core · AI Visibility
- Prompt InjectionPrompt injection is a security vulnerability in which instructions placed in the input a language model processes change the application's behavior in ways the developer did not intend. In a direct injection the instructions are in the user's own message; in an indirect injection they are hidden in content the model reads, such as a web page, an email or a document. OWASP ranks it first (LLM01) in its 2025 Top 10 for LLM applications.Professional · AI Visibility · Security & Infrastructure
R
- RAG (Retrieval-Augmented Generation)Retrieval-augmented generation (RAG) is a technique in which a language model first retrieves relevant content from an external source, such as a search index, a document store or the web, and then bases its answer on that content. It keeps answers current and tied to sources, and it underpins most AI search experiences that show citations.Professional · AI Visibility · Software
- Relevance ScoringRelevance scoring is how a search system expresses, as a number, how relevant a document or passage is to a particular query; results are then sorted by that number. Classic methods rely on word statistics such as TF-IDF and BM25, while modern methods use embedding similarity and cross-encoder models. Large search engines combine these scores with many other signals.Advanced · SEO · AI Visibility
- RerankingReranking is the step in which candidate documents or passages found by a fast first-stage retriever are scored again, and reordered, by a slower but more accurate model. The most common model for the job is a cross-encoder, which reads the query and the passage together. The goal is to make sure the few results shown to a user or passed to a language model really are the most relevant ones.Advanced · AI Visibility
- RetrievalRetrieval is the step that takes a query and returns a ranked list of the most relevant documents or passages from a large collection. The component that does it is called a retriever. It is the first stage of a search engine and, in RAG systems, the part that decides which sources a language model gets to read. Retrieval can be sparse (based on word matching), dense (based on embeddings) or a combination of both.Professional · AI Visibility
S
- Semantic SearchSemantic search is a retrieval approach that returns results based on the meaning of a query and of the content, not just on exact word matches. It accounts for synonyms, context, entities and user intent. Today it is often implemented with vector search, where text is converted into numerical vectors (embeddings) and the content closest in meaning to the query is retrieved.Professional · SEO · AI Visibility
- Semantic TripleA semantic triple is a unit of data that records one fact as a subject, a predicate and an object, such as "Albert Einstein – place of birth – Ulm". It is the basic building block of the W3C's RDF data model. Linked together, triples form a knowledge graph in which entities are nodes and relationships are edges. Schema.org markup written in JSON-LD can also be read as a set of triples.Advanced · GEO · AI Visibility
- Source SelectionSource selection is the process by which an AI-powered search or answer system decides which web pages to read, which passages to take from them and which to display as sources when answering a question. It typically involves generating search queries, retrieving candidates, re-ranking them and picking the passages that fit in the model's context window. Providers document some prerequisites but not the weights or the details.Advanced · GEO · AI Visibility
- System PromptA system prompt is the set of instructions an AI application's developer gives a language model before the conversation starts, and which stays in force for the whole conversation. It defines the model's role, scope, tone, available tools and output format, and end users usually never see it. Models are trained to give system instructions priority over user messages, but a system prompt is not a security boundary.Professional · AI Visibility
T
- Temperature (LLM)Temperature is a sampling parameter that controls how sharply or how flatly a language model weighs its options when choosing the next token. Low values concentrate on the most likely tokens and give more consistent output; high values spread probability out and give more varied, less predictable output. Even a temperature of zero does not guarantee identical answers in every environment.Professional · AI Visibility
- Token (AI)A token is the basic unit a language model uses to process text: it can be a whole word, part of a word, a punctuation mark or a string of characters with a leading space. The component that splits text into tokens is the tokenizer, and each model family has its own. Context windows, output limits and API pricing are all measured in tokens. Token counts do not equal word counts, and Turkish text usually needs more tokens than English.Core · AI Visibility
- Transformer ArchitectureThe Transformer is a neural network architecture built on the attention mechanism, introduced in the 2017 paper “Attention Is All You Need”. Instead of reading text one word at a time, it processes all tokens in a sequence together and calculates how strongly each one relates to every other. Today's large language models, and many image and multimodal models, are built on this architecture.Advanced · AI Visibility
V
- Vector DatabaseA vector database is a data management system designed to store high-dimensional vectors, such as embeddings, together with their source text and metadata, and to run fast similarity queries over them. On top of approximate nearest neighbor indexes it provides persistence, updates, filtering, access control and scaling. It can be a standalone product or an extension of an existing database, such as pgvector for PostgreSQL.Professional · AI Visibility · Software
- Vector SearchVector search is a search method that finds, among records stored as numerical vectors, the ones closest to a query vector according to a chosen distance metric. The records are usually embeddings of text, images or products. Because comparing the query with every vector is slow at scale, large systems use approximate nearest neighbor (ANN) indexes such as HNSW or IVF, which give up a little accuracy for a large gain in speed.Professional · AI Visibility · Software
Other categories
- SEO Glossary167 termsCanonical URL · Core Web Vitals · Crawling · Entity
- GEO Glossary63 termsAI Visibility · Entity · GEO · Structured Data
- Web Design Glossary97 termsCore Web Vitals · Responsive Design · Web Accessibility · Above the Fold
- Software Glossary138 termsAPI · LLM · API Endpoint · API Key
- Performance & Analytics Glossary71 termsCore Web Vitals · Abandoned Cart · Bounce Rate · Call to Action
- Security & Infrastructure Glossary82 termsAPI Key · Authentication · Authorization · Backup
- CMS & E-commerce Glossary27 termsAbandoned Cart · Checkout · CMS · CRM

