Contact

What is Source Selection?

Definition

Source selection is the process by which an AI-powered search or answer system decides which web pages to read, which passages to take from them and which to display as sources when answering a question. It typically involves generating search queries, retrieving candidates, re-ranking them and picking the passages that fit in the model's context window. Providers document some prerequisites but not the weights or the details.

Also known as: AI source selection, answer engine source selection, how AI chooses sources

Diagram of an AI system narrowing candidate pages step by step by relevance and trust to the few it cites

The chain of decisions behind an answer

Web-grounded answer systems differ in detail, but public documentation and the research literature describe a broadly similar flow. Source selection is not one decision; it is a series of filters:

  1. Query generation: the user's question becomes one or more search queries. Google calls its version of this query fan-out for AI Overviews and AI Mode: multiple related searches across subtopics and data sources.
  2. Candidate pool: those queries go to a search index or search API, and candidate documents are retrieved. A page that is not indexed, or that blocks the system's crawler, never reaches this stage.
  3. Narrowing: candidates are re-ranked by a model that judges relevance more precisely, and the best passages, rather than whole pages, are picked out.
  4. Fitting the context: because the model's context window is finite, only a handful of passages from a handful of sources are passed to it.
  5. Writing and attribution: the model writes the answer from those passages, and the sources it used are shown as links or cards.

Which search backend, index and models sit behind each step varies between products and is not always disclosed.

What providers do document

The internals are closed, but a few prerequisites are on the record:

  • Google: to be a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to appear in Search with a snippet; there are no additional requirements. Google also says fan-out lets these features show a wider and more diverse set of links than classic results.
  • OpenAI: sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers, though they can still appear as navigational links.

The shared message is simple: access and indexing are the entry ticket. Everything after that is down to each system's own ranking logic.

What nobody publishes: the weights

No major provider explains, factor by factor, why a page was chosen. Most lists of “AI ranking factors” are built on observation and correlation. A reasonable inference is that retrieval and re-ranking reward passages that closely match the question in meaning, and that trustworthiness and source authority weigh more on health, finance and legal topics. How those factors trade off against each other is not something any outside observer actually knows.

Selection is also unstable. Ask the same question at a different time, from a different location or in slightly different words, and the source list can change. Drawing conclusions from a single screenshot, whether “we're in” or “they dropped us”, is therefore misleading.

What a site can and cannot influence

StageWithin the site's control
Candidate poolCrawl permission, indexing, snippet eligibility, fast and error-free responses
Matching sub-queriesCovering a topic's sub-questions in distinct, clearly headed sections
Passage choiceAnswering at the top of each section, paragraphs that stand on their own
TrustSourced data, named authors and publisher, consistent entity information
Final decisionNo direct control; it can only be monitored by sampling

Most of what a site can do amounts to not being disqualified. No method guarantees selection, and any service that promises it deserves a skeptical look.

Related terms

← Back to the glossary