Executive summary and key findings#
TL;DR
We captured ChatGPT's retrieval traffic again, ten months after the original study, and re-measured Reddit's visible citation rate on 291,607 citations. The engine works differently now, and Reddit's role in it has changed with it: it is retrieved on a narrower rule, and where it is retrieved it is now shown and credited. Reddit remains a key surface for marketers; what has changed is which questions it answers and how to win them. The other shift is not about Reddit at all: ChatGPT now favours a company's own website over third-party pages at every step of retrieval, from the search plan to the citation, and describes an instruction to that effect.
Key findings#
- Reddit retrieval and citations both fell: Reddit was retrieved on most queries in 2025 and on about a quarter of search turns now, on opinion-shaped questions. Visible citations fell with it, from 0.35% to 0.12% of ChatGPT citations on 291,607 measured (clustered 95% interval 0.05% to 0.22%, the same in both model eras).
- Reddit is still a major source at citation time: it was retrieved on one in four search turns in our chats, and once read it is cited at the same per-page rate as any other source (29% against 31%), where in 2025 almost none of it was; it was cited on 73% of the turns it was read on. The model's instructions name it as the place for community opinion.
- Which questions still bring Reddit in: Reddit is retrieved when the question asks for people's views, "what do people think about X", "what are people saying about Y", "honest opinions on Z", "is it worth switching, what do people say", for B2B tools and consumer products alike, and when the question names Reddit or a subreddit. It is now retrieved less on comparative prompts: none of the "best X for Y" or "A vs B" software comparisons read it, nor did any pricing, news, docs or brand-research question, nor a prompt that named another community such as Hacker News, Stack Overflow or G2. This is one of the drivers behind visible Reddit citations falling from 0.35% to 0.12%. The dedicated Reddit source fired on the opinion prompts and the subreddit prompt; naming Reddit itself sent the question down a
site:reddit.comrewrite instead. The same subject flips with the phrasing: "is Apollo.io better than ZoomInfo" read no Reddit, "what do people actually think about ZoomInfo data accuracy" read a full Reddit group and Reddit took over half of the citations. Reddit is the validation channel; the comparison question goes to vendor pages. - Reddit has a structural position no other third-party source has: the dedicated Reddit allocation described in the 2025 study is still in place. Reddit is retrieved as its own group, up to twelve threads, and fed into the model's reasoning for citation selection alongside the ranked web results. Review sites, forums and editorial pages compete for places in one list; Reddit arrives separately, with precedence built into the retrieval step.
Retrieval engine mechanisms
- The retrieval operators have been extended, and they are used more: query rewriting is known; what is new is the control the model now attaches to each line. Every line carries a tier (
sloworfast), most carry a freshness window in days, and lines can be scoped to a single host.site:appears on a quarter of lines, often path-scoped, and separate forms exist for product search, local search, page opens and data widgets. Exclusions andintitle:appear when the task is to enumerate a site. Retrieval is being steered line by line, which is where the vendor and Reddit routing described below happens. - Search results come from an index; on ~50% of turns some pages are also fetched live before the answer is written: on ~50% of the turns about our own site nothing was requested from it and the answer was built from index records; on the other ~50%
ChatGPT-Userfetched the pricing, about and service pages seconds after the search plan. About a third of what the model reads per turn is fetched live, and citations draw on both fetched and index-only pages. A page published today can be read today, and blocking the fetcher removes you from that part of retrieval.
Content formats
- Vendor pages gain share between what ChatGPT reads and what it cites, 28% to 36%, and pricing questions cite vendor pages only: the search plan targets a product's own site first, and the live fetcher spends its budget on pricing, about and product pages.
- Docs replaced listicles in ChatGPT's citations: on a B2B SaaS prompt set, docs rose from 36.9% to 55.1% of citations while listicles fell from 6.0% to 1.0%, and the fall was the same for high- and low-authority listicles and on prompts that ask for a list. It is the same official-site policy, observed in the citation data rather than on the wire.
What changed since the last study#
The 2025 study described a fixed pool of 60 search slots per query, split Bing Search 32, Reddit 16, Bing News 12, with Reddit fetched on almost every query and cited on almost none. That was the "hidden influence gap": 27% of retrieval, 0.35% of visible citations.
Four things are different in September 2026.
- Reddit retrieval narrowed to opinion-shaped questions: the dedicated Reddit source fired only on prompts of the "what do people think", "honest opinions", "what's on r/…" kind, and returned a block of up to twelve threads placed after the web and news results. It did not fire on any of the B2B software comparison prompts we ran; on those the engine read vendor pages, docs, review pages and comparison articles. Where in 2025 Reddit was fetched on most queries, it is now fetched on about a quarter of search turns.
- Citations fell with retrieval: Reddit's share of visible ChatGPT citations is 0.121% on 291,607 citations, against the 0.35% we published for Nov–Dec 2025. The clustered interval excludes the old figure, the result holds under the most conservative cut, and it is the same in the gpt-5.4-nano and gpt-5.6 eras, so the model version is not what moved it. The retrieval gate is.
- When Reddit is read it is now cited: on the turns where Reddit entered the model's reading list it was cited visibly on 73% of them, and per page at the same rate as any other source, 29% against 31%. The 2025 pattern of reading Reddit and citing almost none of it no longer describes the engine. The routing itself is built in: the model's instructions name Reddit for community opinion, it writes "Reddit" into its own search lines on those questions, and it has a dedicated Reddit source.
- The retrieval interface itself keeps changing, and differently per engine: ChatGPT's search lines now carry a tier, a freshness window and a host scope, with separate forms for product search, local search, page opens and data widgets, and its source policy moved in March 2026 without a new model. Google AI Mode, Gemini and Claude did not make the same move on listicles. What holds for one engine's retrieval in one quarter does not hold for the others, or for the next quarter.
The visible citation rate moved with the retrieval gate, not with any change in how Reddit is treated once read.

- The drop is a property of the retrieval policy, not of any one model release: the rate is 0.116% in the gpt-5.4-nano era and 0.129% in the gpt-5.6 Luna era, a difference whose interval includes zero, so the July model change did nothing to Reddit; the March change to how ChatGPT searches did.
- That matters for planning: a shift that lives in the search policy can move again with a policy update, without a new model, and it will show up first in retrieval logs rather than in citation counts.
How ChatGPT's retrieval engine works now#
We captured every byte between the ChatGPT web app and OpenAI's servers across a range of prompt types over two days, reconstructed the streams, and tested each claim below against the full set rather than the examples that fit. The full protocol document is in the technical appendix.
2.1 The decision to search is server-side#
Every request the browser sends is identical in shape: model auto, no hints, no search flag. The server classifies the turn (instant search, search, shopping, local, instant answers, text) and decides whether to invoke its search tool. Factual, medical and code-troubleshooting prompts never searched. Nothing the user does in the composer influences this beyond the words of the prompt.
2.2 The search plan: more operators, used more often#
That the model rewrites the prompt before searching is well known. What has changed is how much control each rewritten line now carries. Before it searches, the model writes a query plan: several lines, one per sub-query, in a fixed grammar with a tier, a freshness window and an optional host scope on every line.
The format, one line per sub-query:
tier | query | window | scope
- Tier:
fast(53% of lines) is the default lookup.slow(38%) is the path the model picks for opinion, news and pricing, and it always carries a window. ChatGPT itself describesslowas "intended for harder-to-find or higher-confidence searches" andfastas "suitable for broad searches". Separate line types exist for product search, local business search, page opens and data widgets such as stock charts and weather. - Query: The rewritten search text. It carries the brands the model expects to find, the year, and words like "official", "pricing" and "reviews", and it can hold
site:restrictions and quoted phrases. - Window: The trailing number is a look-back in days, set from cues in the prompt: "today" gives 1, "this week" gives 7, the default is 30, "recent papers" gives 365, and history, evidence and Reddit-oriented lines get 3650. Absent means no constraint. It is strict for news results and a preference for web results.
- Scope: An optional hostname that restricts the line to one site. We saw it for pricing pages, release notes, arXiv, PubMed, NEJM, JAMA and Reddit.
Examples from the capture:
slow|ZoomInfo data accuracy reviews inaccurate contact data Reddit G2 2026|30
fast|ZoomInfo data accuracy reviews contact accuracy G2 Reddit|3650
slow|Notion pricing team 20 members 2026|30|notion.com
fast|site:stripe.com/pricing Europe pricing cards European cards Stripe fees|30|stripe.com
Across every search turn we captured, not one line was the user's prompt. The model adds the brands it expects to find, the year, and words like "official", "pricing" and "reviews". It uses site: on a quarter of lines, often path-scoped (openai.com/news, g2.com/products/notion/reviews, reddit.com/r/BitcoinEurope), and in enumeration tasks it reaches for exclusions and intitle: as well. The operators that matter for marketers are the ones that decide which sources a line can return at all: the tier, the window and the scope.

2.3 One ranked list, Reddit handled differently#
The search tool returns a single ranked list per call, and Reddit is the one source that does not blend into it.
- Web, news, academic and YouTube results interleave in rank order in one list.
- Reddit is retrieved as its own dedicated group and fed into citation selection: when the Reddit source fires, its threads arrive together, up to twelve of them, after the web and news range rather than mixed into it, and the model reasons over that group alongside the ranked list when it picks what to cite. No other third-party source is handled this way, and on those turns Reddit carries from a fifth to all of the citations.
2.4 Index results, with some pages fetched live#
The wire shows which pages the model read, but not whether each one came from OpenAI's stored index or was fetched from the site at the time of the question. That matters because it decides how quickly a change on your site can reach an answer. To find out, we asked ChatGPT about our own domain and compared the capture with Cloudflare's logs for the site, where OpenAI's fetcher identifies itself by user agent.
- Most search results come from the index: on ~50% of the turns about our site, nothing was requested from it at all.
- Some pages are fetched during the turn, before the answer is written: on the other ~50%,
ChatGPT-Userrequested a small number of pages a few seconds after the search plan, mostly the pricing, about and service pages. Fetched pages were usually cited, but pages that were never fetched were cited too, so fetching is not a precondition for being quoted. - The snippet field shows which: result entries that carried text had been fetched in that minute, and entries without text had not. Across the general prompts, roughly a third of what the model read per turn was fetched live.
- Explicit page opens fetch live unless cached: "read this URL" and "browse the site" prompts fetched the named pages, while pages the crawler
OAI-SearchBothad recently visited were served from cache.
Two agents do this work. OAI-SearchBot crawls continuously and builds the index; ChatGPT-User fetches the pages the model is about to quote, as ordinary requests from OpenAI's cloud that respect robots rules. Both appear in CDN logs by user agent, and a rule that blocks either one takes the site out of that part of retrieval.
2.5 Reading versus citing#
The engine reads far more than it cites, and what it drops between the two steps isn't random. We grouped every page the model read, and every page it cited, by who owns it: the vendor whose product the question is about, an editorial site or list post, a forum or social platform, a review site, or a news outlet. Comparing the two stages shows which source types survive the cut.
Vendor-owned pages are 28% of what the model reads and 36% of what it cites. On comparison prompts the step is larger, 24% to 39%. Editorial pages and lists lose share at the same step. Forums hold their share overall because opinion prompts cite them heavily. Review sites are 2% of reading and 6% of citing: rarely fetched, but kept when they are.

Reddit today#
3.1 Where it enters the answer#
Two mechanisms bring Reddit into an answer, and both are triggered by the shape of the question rather than by the topic.
- The dedicated Reddit source: fired on every prompt that asked for people's views or named a subreddit, and on no other prompt. It returned its threads as a group of up to twelve after the web and news results.
- Ordinary web search with a
site:reddit.comrewrite: fired when the user named Reddit in the question, returning Reddit threads through the same path as any other page. - Incidental results: on a few consumer shopping prompts ("best budget standing desk") Reddit threads arrived through ordinary web search among the other results. None were cited.
Retrieved on
- Opinion questions about a product or company: "what do people actually think about ZoomInfo data accuracy", "what do developers think about Cursor vs Claude Code", "what are people saying about the new iPhone battery life", "honest opinions on the Tesla Model Y". A full Reddit group each time, cited every time, with Reddit taking from a fifth to over two thirds of the citations.
- Switching and validation questions that ask for views: "is it worth switching from Slack to Teams, what do people say". A smaller Reddit group, cited.
- Questions that name Reddit or a subreddit: "what does reddit say about the best crypto exchange", "best subreddits for B2B SaaS founders", "what's on r/sysadmin about patch tuesday". A
site:reddit.comsearch each time, and the subreddit prompt also fired the dedicated source. Cited every time.
Not retrieved on
- B2B software comparisons and best-of lists: "best incident management tools for SRE teams", "best CRM for a 10 person startup", "is Apollo.io better than ZoomInfo", "is instantly.ai worth it compared to lemlist", "top observability platforms for kubernetes", "best AI SEO agencies for B2B SaaS". The engine read vendor blogs, docs, review pages and comparison articles instead.
- Pricing: "how much does Notion cost for a team of 20", "how much does incident.io cost per month". Vendor pages only.
- News and live questions: "what happened in AI news this week", "latest news on EU AI Act enforcement", "Nvidia stock price right now".
- Consumer comparisons that do not ask for views: "air fryer vs convection oven", "compare the Sony A7 IV and Canon R6 II for video", "best noise cancelling headphones under 200 euros".
- Questions that name another community or a review site: "what do people on hacker news think about the new MacBook", "what does stack overflow say about python circular imports", "summarize the reviews of Notion AI on G2 and Capterra". The engine went to the named site and not to Reddit.
- Docs, how-to, evidence, local and brand research questions.
The same subject flips with the phrasing. "Is Apollo.io better than ZoomInfo" read no Reddit; "what do people actually think about ZoomInfo data accuracy" read a full Reddit group and Reddit took over half of the citations. "Is instantly.ai worth it compared to lemlist" read no Reddit; "is it worth switching from Slack to Teams, what do people say" did. The ask for people's views is the trigger, and "worth it" on its own was not.

The trigger is the shape of the prompt, not the words on the search line. The word "Reddit" appeared in the query text on every Reddit-bearing turn, including those where no block fired, and the scope host, the site: operator and the ten-year window were each absent on some block turns. The switch sits upstream, in the classifier.
3.2 Why Reddit is still a key surface#
Less Reddit is retrieved, and more of what is retrieved is cited. That leaves Reddit a narrower but more visible channel, concentrated on the questions buyers ask when they are deciding.
- It is cited when it is read: on the turns where Reddit entered the reading list it was cited on 73% of them, at the same per-page rate as any other source (29% against 31%). There is no set of Reddit pages that shape the answer without appearing in it.
- It carries the answer on opinion questions: on the turns where the dedicated source fired, Reddit took from a fifth to all of the citations, shown with the "Reddit +1" chip. Our internal data agrees: Reddit is cited zero times when it is not in the retrieval pool and appears in 16% of visible citations when it is.
- Those are the validation questions: "what do people actually think about X", "is it worth switching" are the questions from the 2025 study's B2B buyer pattern. The comparison question goes to vendor pages; the validation question goes to Reddit.
- The routing is built in: the model's instructions name Reddit for community reactions, reviews and recommendations, it writes "Reddit" into its own search lines on those questions, and it has a dedicated Reddit source. None of that changed between the two model versions we measured.
If a Reddit thread influenced an answer, it is now in the citations.
Review sites: absent from the weights, used at retrieval#
A separation of concerns is developing between what the model remembers and what it looks up. Is your brand in AI training data? found that review platforms leave almost no trace in the weights: G2 content showed zero detectable imprint across the open models tested. On the retrieval side the picture is the opposite.
- Review sites are reached through the search operators when the question asks for reviews: on "what do people actually think about ZoomInfo data accuracy" the plan read and cited G2 and Trustpilot; on "is it worth switching from Slack to Teams, what do people say" it read and cited G2 repeatedly; on "summarize the reviews of Notion AI on G2 and Capterra" it wrote
site:g2.com/products/notion/reviewsandsite:capterra.com/p/lines. Where the prompt did not ask for opinions, review sites were absent from both the search lines and the reading list. - They are not a named source: asked several ways, ChatGPT said G2, Capterra and Trustpilot are not mentioned in the instructions it can describe, while Reddit is named explicitly for community reactions, reviews and recommendations. Review sites have no dedicated source and no block; they arrive through the same
site:and query-text path as any other page. - The practical reading: a brand's review-site standing does not shape what the model believes about it, but it does shape what the model reads and cites when a buyer asks for opinions. Reviews are a retrieval asset, not a memory asset, and they are pulled in by the same operators that route pricing questions to vendor pages and opinion questions to Reddit.
Retrieval content format changes: ChatGPT, Gemini and Claude#
ChatGPT's retrieval now favours a brand's own pages at three separate steps, and the wire shows each one.
- Query construction: "official" appears on 12% of search lines and "pricing" on 15%, and the scope field points at vendor hosts on pricing questions. When the prompt names a product, the model's first search line targets the product's own site.
- Selection: vendor pages gain eight points between reading and citing across all turns, and fifteen on comparison prompts. Pricing questions cited vendor pages exclusively.
- Live fetching: the pages the fetcher pulled live on our own domain were the pricing, about and service pages, not blog posts.
The citation data shows the same shift at scale, and shows it is ChatGPT-specific.

- ChatGPT alone dropped listicles: on one B2B SaaS prompt set they fell from 6.0% of citations to 1.0% between Oct 2025 and Aug 2026, and from 14.6% to 1.4% on a second. Google AI Mode, Gemini and Claude held at 11 to 22%.
- Authority did not protect them: high-authority listicles lost share at the same rate as low-authority ones (4.96% to 0.84% for domain rank 300+, 0.90% to 0.06% below).
- Intent did not either: on "best X tools" prompts, where a list is the right answer, ChatGPT went from 11.5% to 1.2% while Google stayed at 37 to 41%.
- It is a policy, and it is dated: the change arrived in March 2026 with the gpt-5.4-nano release and did not move again in July. It may change again with a later release.
5.1 What it cites instead#

- Docs took the share: on the same prompt set docs rose from 36.9% of ChatGPT citations to 55.1%, product pages held at about 9%, and pricing pages slipped from 15.9% to 8.2%. Official pages of all types went from 63% to 74%; on a second set docs alone are 46% and official pages 71%.
- Docs are now the largest single citation source on ChatGPT for a B2B software brand, so a documentation gap is a citation gap.
- Pricing pages lost share while pricing questions cite vendor pages only, which means the engine looks for the pricing page and often does not find a citable answer on it.
- The weights reward the same asset: Is your brand in AI training data? found that a brand's own properties, help centre, docs, pricing and integrations pages, are what move category accuracy in the model's memory (help-centre content adds 12 points), and that pricing was the least reliably known fact in the weights (0.53 accuracy). Retrieval and memory are separate pillars that reward the same plain, factual, first-party page.
Actions for marketers#
For ChatGPT, this quarter#
- Ramp up the facts on the official site and in the docs: pricing, plans, integrations, limits and comparisons with named competitors, stated plainly in the first screen of text. Docs went from a third of ChatGPT citations to over half, the search plan targets your own site first, and the fetcher pulls those pages live. A documentation gap is now a citation gap.
- Keep doing Reddit: for the reasons in section 3, Reddit is retrieved as its own group on the questions buyers ask when they are deciding, and it is cited when it is read. Map the "what do people think about X" and "is it worth switching" questions in your category and be in those threads with substance the model can quote.
- Leave the other engines' plan alone: Google AI Mode, Gemini and Claude have not moved. They still cite listicles and comparisons at a third to half of citations on list-shaped prompts, so those formats stay in the plan for them.
- Don't flip-flop on a citation-stage metric: citation counts sit at the end of the pipeline and miss the retrieval mechanisms upstream, which questions triggered a search, what was fetched and by which agent. Re-test on each model release rather than rebuilding around a single reading.
- Watch the fetchers in your CDN logs: read
ChatGPT-UserandOAI-SearchBothits alongside citations, and keep both agents allowed in robots and WAF rules.
