·8 min read

The 4 Types of Pages That Drive AI Citations and How to Get Your Brand Featured on Them

82% of AI brand mentions come from pages you do not own. This workshop breaks down exactly which 4 types of third-party pages AI cites in B2B categories, why each model cites different ones, and how to audit which sources are driving citations for your competitors right now.

Will Leatherman

Will Leatherman

Founder, Catalyst

TLDR

When someone asks ChatGPT, Perplexity, or Google AI Overview about the best tool in your category, the answer is built almost entirely from third-party pages, not your website. 4 page types account for most of these citations: best-of lists (43.8% of ChatGPT citations according to Ahrefs), review platforms like G2, community threads on Reddit, and broad reference sources like Wikipedia. Each AI model pulls from a different source set (86% of top citation sources are not shared across ChatGPT, Perplexity, and Google, per Ahrefs' cross-platform study of 76.7M results). The practical implication: one content strategy optimized for a single platform fails. The fix is to map which sources each model is using in your category and get your brand named on those specific pages. Will ran this audit live in the workshop on a real CRM company and identified exactly which third-party pages were driving competitor citations, and which gaps to close in 30 days.

Most teams improving their AI search presence start with their own website. They add FAQ sections, restructure pages, and optimize metadata. That work matters, but it addresses the smaller half of the problem. According to Geonimo's analysis of 50,000+ AI responses, 82% of brand mentions in AI answers come from third-party sources, pages on other websites that the model trusts and cites. Your homepage plays a minor role in whether AI recommends you. The sources that matter most are the ones you do not control directly.

This workshop walks through how those citations work, which 4 page types produce most of them, and how to run the audit that reveals exactly which sources are driving citations in your specific market.

Catalyst has worked with companies ranging from pre-revenue startups to Series D businesses on AI search strategy. The citation patterns are consistent across industries. The gap between brands that show up and brands that do not comes down almost entirely to whether they have earned coverage on the right third-party pages.

What Are the 4 Types of Pages AI Cites Most Often?

Not all third-party pages carry equal weight. Analysis of AI citation patterns across the major models points to 4 types that account for the majority of B2B citations:

| Page Type | Why AI Cites It | Examples | Share of Citations | |---|---|---|---| | Best-of lists and comparisons | Pre-curated comparisons match how buyers search; engines lift the rankings verbatim | "Best CRM tools 2026," "Top AI writing tools for B2B" | ~43.8% of ChatGPT citations (Ahrefs) | | Review platforms | Structured data, verified buyers, and category taxonomy that LLMs can parse cleanly | G2, Capterra, TrustRadius | High — matched to purchase-intent queries | | Community threads | Real buyer language, stated pain points, named alternatives; Perplexity weights Reddit heavily | Reddit, niche forums, LinkedIn comments | ~40% of citations across engines (Profound, 30M citations) | | Broad reference sources | High authority, structured data, encyclopedic coverage; ChatGPT was trained heavily on Wikipedia | Wikipedia, Wikidata, Crunchbase | Brands with a Wikipedia entry get cited 4.1x more often (WinWithSEO, 3,200-query study) |

Will described this framework during the workshop: "We borrow the credibility of a source that AI already trusts and lend it to how that source names us." Getting mentioned on a G2 profile or a Wirecutter-style comparison article carries more citation weight than adding another page to your own site.

Why Does Each AI Model Cite Different Sources?

The 4 page types above are consistent, but which specific sources each model prefers is not. Ahrefs' cross-platform study of 76.7M AI Overview responses, 957K ChatGPT citations, and 953K Perplexity citations found that 86% of top citation sources are not shared across ChatGPT, Perplexity, and Google. Only 7 of the top 50 citation sites appear in all 3 engines.

The divergence follows each model's training emphasis:

  • ChatGPT was trained heavily on Wikipedia and structured web content. Formal, wiki-style information architecture and high-authority editorial sources perform best.
  • Perplexity weights community platforms heavily. Reddit is one of its most-cited domains and accounts for roughly 47% of Perplexity citations according to Profound's 30-million-citation panel.
  • Google AI Overview extends Google's E-E-A-T standards and prefers pages already ranking in organic search. Domain authority matters more here than on other models.
  • Gemini and Claude blend editorial and community sources, but both update their citation behavior faster than ChatGPT as new content enters their retrieval index.

Will walked through a live audit of a CRM company in the workshop and found they were showing up consistently on Gemini but nearly invisible on ChatGPT and Perplexity. The fix for each gap pointed to a different type of third-party page: a Wikipedia stub for ChatGPT, a Reddit presence for Perplexity, and comparison articles on category-specific blogs for both.

A single content strategy built around one model's preferences fails the other 3. The audit reveals which models you are missing and which page types to pursue first. See How to Audit Your AI Search Ranking in 20 Minutes for a step-by-step breakdown of the audit process.

What Makes On-Site Content Quotable Enough for AI to Cite?

Third-party pages drive most citations, but your own site content determines whether AI can lift a specific claim, statistic, or recommendation and attribute it to you. Will described the distinction clearly in the workshop: "LLMs have foundation knowledge. If you ask broad, high-level educational questions, they answer without referencing a specific company. To get cited, we need to provide information they do not inherently already know."

2 types of on-site content achieve this:

Proprietary data. Numbers, findings, or structured datasets that only your company can produce. Ramp.com builds published indexes from its own customer spend data. That content is uncopyable and citable because no other source has the same numbers. Research reports built on internal data follow the same principle.

Original expert perspective. A specific, defensible stance on a topic from a credible named author. Generic category overviews are not citable because the model already contains a version of the same information. A specific argument supported by a real case example is citable because it is new.

The research backs this up. The Princeton/Georgia Tech GEO study (Aggarwal et al., KDD 2024) found that adding cited sources, statistics, and direct quotations to content produces up to a 40% visibility lift in AI citation frequency. That is the largest controllable lever in the study. Writing more content at the same level of generality produces no lift.

How Do You Find Which Sources AI Uses in Your Category Right Now?

Knowing the 4 page types is the starting point. The specific sources AI is currently using to answer questions in your market are what determine your target list. These differ by category, by query type, and by model.

The process Will ran live in the workshop:

  1. Run fresh LLM queries using buyer language across ChatGPT, Perplexity, Claude, Gemini, and Google AI Overview. Use bottom-of-funnel queries, not broad category terms. "Best CRM with AI deal automation for mid-market teams" surfaces different sources than "best CRM."
  2. Extract every URL cited across all 5 responses. Each URL is a third-party source the model trusts enough to cite for that query.
  3. Classify each URL by type (listicle, review platform, Reddit thread, editorial blog) and by model.
  4. Build a ranked target list. The sources appearing most often across models are your priority placements. The ones appearing only on Perplexity point to Reddit and community gaps. The ones appearing only on ChatGPT point to editorial and Wikipedia gaps.
  5. Check your competitors against the same list. If a competitor appears on 8 of your 10 target sources and you appear on 2, the citation gap is structural, not a content quality problem.

For the CRM company audited in the workshop, the target list included specific blog articles from monday.com, LinkedIn posts from Cybill, and a set of G2 comparison pages. Getting named in future editions of those pieces was a more direct path to citation improvement than anything they could do on their own site.

Use the Catalyst AEO audit tool to run this process automatically across 5 models and get a ranked map of the sources driving competitor citations in your category.

What Signals Tell AI Your Company Is a Credible Entity?

Citations come from third-party pages, but AI decides whether to include your brand name in an answer partly based on how many credible external sources have independently established your existence and reputation. Will called this "entity authority" in the workshop.

The key signals that build entity authority for B2B companies:

  • G2 or Capterra profile with verified reviews and complete category tagging
  • Crunchbase listing with accurate funding stage, description, and founding team
  • LinkedIn company page with consistent activity and a clear description of what the company does for whom
  • Wikipedia entry with a Wikidata cross-reference (4.1x citation lift, per WinWithSEO's 3,200-query study)
  • Press mentions from industry publications that reference the company in the context of its category

These signals are not a replacement for earning placement on the 4 page types above. They function as verification infrastructure. When an LLM is deciding whether to include your brand in an answer about your category, it cross-references these sources to confirm you are a real, credible company. A brand with strong entity signals and good third-party coverage ranks consistently. A brand with strong entity signals but weak third-party coverage still gets skipped on category queries.

For more on how entity signals interact with AI citation frequency, see How B2B Marketing Teams Get Named in AI Search.

What Does a 30-Day Plan to Close a Citation Gap Look Like?

The CRM company audited in the workshop had a specific gap: they were showing up well on Gemini but nearly absent on ChatGPT and Perplexity. The 30-day plan that came out of the audit had 3 priorities:

Week 1 to 2: Target the highest-traffic comparison articles. Identify 3 to 5 articles in your category (like "best CRM tools 2026") that are already generating ChatGPT and Perplexity citations. Reach out directly to the authors and ask to be included in their next update. Many of these are maintained by independent bloggers or small content teams who respond well to direct outreach. Will noted in the workshop: "You can literally go to these websites and reach out to the owners directly to see if they'll include you on future lists. People do respond to that pretty well."

Week 2 to 3: Fill the Reddit and G2 gaps. If Perplexity is your weakest model, the fastest fix is answering real buyer questions in relevant Reddit threads (genuinely, not as promotion) and ensuring your G2 profile is complete, current, and category-tagged correctly for the queries you want to show up on.

Week 3 to 4: Build or claim one high-authority reference source. If you do not have a Wikipedia entry, create one. If you have one but it lacks a Wikidata sameAs link, add it. If you have original data sitting in internal dashboards, publish a structured summary as a market report. These take longer to index but produce persistent citation lift across all models.

Run the audit again after 4 weeks. AEO gives faster feedback than SEO. Will's observation from running this with clients: "You can actually see major improvements in just a couple of weeks. Running this on a weekly basis allows you to track how you're improving and test what's working." Cite your own improvement numbers when they are real. That kind of proprietary performance data is itself citable.

The Takeaway

AI builds its answers from pages it already trusts. 82% of the time, those pages are not yours. The 4 page types doing most of the work are best-of lists, review platforms, community threads, and broad reference sources. Each major model weights them differently, so the path to consistent cross-model visibility requires a map of which sources each model is using in your specific category today.

The one action to take this week: run fresh queries across 5 models using your buyers' actual language, extract every cited URL, and identify whether your gap is in editorial listicles, review platforms, or community sources. That 30-minute exercise tells you exactly where to spend the next month. Use the Catalyst AEO audit tool to run it automatically and get the source map with competitor coverage already included.

The Content Engineer

Enjoyed this article?

Frameworks like this, weekly. No fluff — just original research and actionable insight.

Ready to turn insight into pipeline?

We work with B2B companies that know content is the moat. Let's talk.