Every keyword tool on the market sells you the same abstraction: estimated volumes for canonical head terms. "hair oil — 40K searches/month." Useful for sizing markets; nearly useless for writing content that AI answer engines cite.
Search Console holds something no tool can replicate: the verbatim strings your actual audience typed to find your actual pages. Typos, dialects, slang, question shapes, year modifiers, mixed languages — recorded at query level, attached to the page each one triggered. For classic SEO this data was a supporting exhibit. For GEO it's the content brief itself, because answer engines match question shape and register when deciding what to cite.
This is the complete mining method.
Why verbatim beats volume for GEO
Compare what the two layers tell you about the same topic:
The keyword tool says: "hair oil — 40K/mo", "dry scalp treatment — 12K/mo", "argan oil benefits — 30K/mo".
Search Console says:
best hair oil for dry scalp
hair oil for itchy scalp winter
which oil reduce hair fall fast
زيت الشعر للقشرة
best hair oil 2026 for curly hair
oil scalp without making hair greasy
Seven real strings → at least five distinct intents, two languages, one seasonal modifier, one constraint ("without making hair greasy"). The tool layer told you the topic is big. The verbatim layer tells you what to write, in what language, in what shape — because these are the literal strings that will collide with your headings inside an answer engine's matcher.
The rule that follows: GEO questions should be your users' queries — grammar cleaned, phrasing preserved. "Which oil reduce hair fall fast" becomes the FAQ question "Which oil reduces hair fall fast?" — not the brand-safe "Which oils help with hair fall?" The first matches; the second paraphrases, and paraphrase costs citations.
The four extractions
Extraction 1 — Question shapes (decides your format)
Pull a page's queries and classify the shape:
- Full interrogatives: "what is the best hair oil for a dry scalp?"
- Fragments: "best hair oil dry scalp"
- Topic phrases: "hair oils for scalp"
Count the dominant shape per page. Full questions → your content should lead with FAQ-format blocks; fragments → heading-and-answer blocks mirror better; a mix → the page needs both, section by section. This one classification settles the format fork with data instead of taste.
Extraction 2 — Register and dialect (decides your voice)
Note the language mix and formality level in the verbatim layer:
- Egyptian Arabic colloquial vs Modern Standard Arabic
- Casual English ("greasy", "flakey") vs clinical ("seborrheic", "desquamation")
- Direct second person ("you") vs impersonal ("users")
AI engines assemble answers that match the asker's register. A clinical-English passage rarely gets quoted for a slang query when a register-matched competitor exists. Whatever register dominates your queries is the register your answers are written in — for the Scrabio workflow, this detection is automatic; manually, it's a read-through of your top 30 queries.
Extraction 3 — Intent splits (decides your page structure)
Cluster the queries by intent verb:
| Pattern | Intent | The passage it becomes |
|---|---|---|
| "best", "top", "vs" | Comparison | A comparison block — named options + the deciding factor |
| "why", "won't", "problem" | Troubleshooting | Cause → fix, stated first |
| "how to", "how do" | Instructional | Steps in order, no preamble |
| "for X" (constraints) | Use-case | One block per constraint: "for curly hair", "without heat" |
The clusters become the page's section plan — each section shaped by its own query cluster, each phrased in its own register. This is how one page legitimately serves five question shapes without diluting any of them.
Extraction 4 — Modifiers as section titles (the free wins)
Modifiers in real queries are pre-validated section titles nobody has written yet:
- "best hair oil 2026" → a freshness section (and an update trigger every January)
- "…for curly hair" → a use-case section
- "…without making hair greasy" → an objection-handling section — objections are the highest-converting FAQ questions that exist
- "…fast" → a speed/expectations section with timelines ("within 2–3 weeks")
Each modifier row in your query data is demand that searched and found a section that doesn't exist. That's not keyword research; that's a to-do list.
Worked example: one page, one mining cycle
Connect GSC to your AI
Query your Search Console data conversationally in Claude or ChatGPT — free MCP server.
Get the free GSC MCPA real-shaped example from a beauty ecommerce property — the verbatim query set for /best-hair-oil (28 days, sorted by impressions):
best hair oil for dry scalp 3,100 imp · pos 8.2
hair oil for itchy scalp 1,900 imp · pos 11.4
which oil reduce hair fall fast 860 imp · pos 14.1
best hair oil 2026 for curly hair 640 imp · pos 9.8
oil scalp without making hair greasy 410 imp · pos 16.3
زيت الشعر للقشرة 2,200 imp · pos 6.5
Six queries → what they dictate:
- Format: 4 of 6 are full questions or near-questions → FAQ-led page, with heading blocks for the fragment shapes
- Register: mixed AR/EN with casual English → answers written in direct "you" voice, Arabic register for the Arabic cluster
- Sections: comparison (best), seasonal (winter/itchy), speed+cause (reduce fall fast), constraint (without greasy), freshness (2026), and the Arabic cluster as its own passage — not a translated afterthought
- Six sections, each mapped to a real string — zero invented topics
That's one page, ten minutes, a complete local content plan. Multiply by your top 20 pages and the quarter's editorial calendar is done — in the audience's language, not the keyword tool's.
Three mistakes that waste the verbatim layer
- Cleaning the phrasing out of existence. "Which oil reduce hair fall fast" polished into "Which oils help reduce hair loss?" loses the matcher's trail. Clean grammar, keep words.
- Translating instead of writing. Arabic-dialect queries answered in formal Arabic (or machine-translated English) get outranked by native-register competitors. The query language is the writing language.
- Mining at the account level only. Account-wide query lists hide page-specific shapes. A "/pricing" page and a "/blog/guide" page attract different registers — mine per page, per cluster.
The workflow, compressed
- Pull queries per top page, verbatim, sorted by impressions (28-day window)
- Classify shapes → decide the format per cluster
- Cluster intents → build the section plan
- Harvest modifiers → the section-title backlog
- Write dual-format, in the dominant register
- Score for citability, publish the winner
Steps 1–2 are one prompt when your Search Console is connected to your AI: "all queries for /best-hair-oil, verbatim, sorted by impressions" — then "classify these by question shape and intent". Steps 5–6 are the automated half of the GEO platform: every cluster becomes dual-format content in the detected language, scored before you publish.
Your keyword tool tells you what the market asks. Search Console tells you what your audience asked you — in their own words, on your own pages. The second list is the one your answer engines are actually matching against.
