FineQuestions: A Map of What the Internet Asks
We extracted 1.49 billion unique questions from twelve years of public web crawls, then classified them by topic and shopping intent. The result is a map of what people actually ask — and a public stand-in for the query logs no one outside the AI labs can see.


Krishna Srinivasan
Dataset: 1.49 billion unique normalized questions extracted from 110 FineWeb snapshots spanning 2013–2025.
Questions are the atomic unit of human intent.
Every question a person types or publishes encodes something they want: to know, to buy, to fix, to compare, to decide.
The marketing industry, specifically SEO and now generative engine optimization (GEO), puts enormous effort into understanding that intent. Companies spend serious money and time to figure out what people mean from what they type into a search box: guessing at questions, inventing personas in workshops, testing content against a signal they cannot see directly.
Search query logs and even more so the questions and prompts to an LLM are proprietary. The web itself, meanwhile, is full of questions already published in forums, FAQs, articles, blog posts and reviews. The raw material has been sitting in front of us, in the public crawls, waiting to be extracted as a corpus.
Consider how a single query “running shoes” could expand into three questions found on the web:
“How often should I replace my running shoes?”
“Should I invest in a neutral running shoe even though I have flat feet?”
“Why are some running shoes so expensive?”
Each opens a different conversation: replacement, suitability, cost. Together, they begin to describe what an answer system, a product team, or a researcher needs to understand about a seemingly simple topic.
We built FineQuestions to explore those differences at scale: 1.49 billion unique normalized questions extracted from FineWeb with occurrence counts and links to their source documents, which provide the contexts in which they appear.
From web pages to a question corpus
FineQuestions draws on 110 snapshots of FineWeb, the filtered English-language corpus derived from Common Crawl, spanning 2013–2025. Extraction yielded roughly 13.4 billion raw questions filling 22 TB. Normalization and aggregation reduced these to approximately 1.5 billion unique questions.
Each question carries its total occurrence count, its frequency tier, and up to 100 source pointers identifying the FineWeb documents, URLs, hosts, and crawl dates where it appeared. Those pointers let you return to the page text and inspect the question in context. Different phrasings of the same underlying question can remain separate entries.
Most questions occur only once or twice
Occurrences in the corpus | Unique questions |
|---|---|
50 or more | 27,375,443 |
10 to 49 | 302,933,413 |
5 to 9 | 269,411,419 |
3 to 4 | 238,021,805 |
1 to 2 | 656,196,725 |
Total | 1,493,938,805 |
Unique questions grouped by how often they appear in the corpus. The distribution has a pronounced long tail toward low-frequency questions. Source: FineQuestions, derived from FineWeb.
The most repeated question occurs 12 million times. At the other end, 894.2 million questions (roughly 60% of the corpus) occur fewer than five times. That tail leaves an enormous space to explore beyond the most familiar questions about a topic.
Organized by topic and commerce
The questions occurring at least five times (about 600 million) are classified by topic using Google’s content classification taxonomy, with categories ranging from Arts & Entertainment and Finance to Health, Law & Government, and Shopping.
Within that group, a classifier flags 112.5 million questions for shopping intent and maps them into the Google Product Taxonomy. The product labels span three levels: 21 top-level categories, 192 second-level categories, and 1,349 leaves.
The top-level categories include everything from Apparel & Accessories to Sporting Goods and Toys & Games to Vehicles.
What the map makes possible
Explore a field’s question space. Filter by content category and you can explore the question space of a field: law, health, finance, travel.
Examine the concerns around a purchase. Filter the shopping subset by product category and you can examine what surrounds a purchase decision: price, compatibility, suitability, maintenance.
Read the person behind the phrasing. A question about replacing worn-out shoes suggests a different situation from one about choosing a first pair. Those distinctions can inform personas and product research, with source context helping test the interpretation.
Trace questions over time. Dated sources offer a way to investigate how questions appear across twelve years of web captures.
And because FineQuestions draws from a corpus used to train language models, it offers a starting point for sampling training data, examining coverage and bias, and constructing task-specific evaluation resources.
Follow the questions
We are preparing FineQuestions for release alongside the extraction and aggregation code, classification methods, mining recipes, evaluation materials and paper.
The web has rich data. Common Crawl and FineWeb have aggregated them. Now FineQuestions extracts the intents.
If Common Crawl is a memory of the web, FineQuestions is a map of its intentions.
About the author

Krishna Srinivasan
Co-founder
Researcher, Prev. Google Deepmind (Search Architecture)