Skip to main content
Generative engine optimization (GEO) research measures how AI answer engines discover, describe, cite, compare, and recommend a brand. A useful study does more than ask a few branded questions: it tests realistic user prompts across multiple platforms, preserves the complete answers and sources, and repeats the test under consistent conditions. The goal is to identify where a brand appears, where it is absent or misrepresented, which sources shape the answers, and which content or authority gaps deserve attention.

Start with a research question

Define the decision the research should support before collecting data. Common questions include:
  • Does the brand appear when buyers discover or compare products in its category?
  • Which competitors receive more mentions or stronger recommendations?
  • How accurately do AI platforms describe the brand and its capabilities?
  • Do answers cite the brand’s website, third-party sources, or no visible source?
  • Which topics combine meaningful demand, commercial value, and a large visibility gap?
Also define the target market, language, AI platforms, competitor set, collection period, and model or interface used. These conditions are part of the result: changing them can change the answer.

Build a prompt universe

A prompt universe is the controlled set of questions used in the study. It should represent how real users express needs, not simply turn a keyword list into questions. Design prompts across several dimensions: The Octoparse AI Visibility Prompt Generator uses a brand website, competitors, target market, and prompt language to generate 30 tailored prompts. It balances the set across four stages:
  1. Discovery — questions people ask before they know which brands to consider.
  2. Use case — problem- and scenario-based questions about completing a task.
  3. Comparison — alternatives, brand comparisons, and competitive shortlists.
  4. Evaluation — questions about fit, reputation, and recommendations.
This balance matters. Branded prompts mostly test what happens after a user already knows the company; non-branded discovery and use-case prompts reveal whether the brand enters the consideration set at all.
Use real customer language where possible. Long-form search queries, support conversations, sales questions, community discussions, and onsite search terms can reveal wording that a generic keyword list misses.

Control the sample size

Start with a broad scan, then deepen the areas that matter most.
  • Baseline scan: cover the main categories, use cases, competitors, journey stages, markets, and platforms.
  • Focused study: add more prompts and local-language variations for topics with high value, strong competitor visibility, or a clear brand gap.
Avoid generating every possible combination of website, field, country, and wording. Use reusable prompt patterns, but keep enough variation to represent distinct user intents. For an initial manual study, five to ten prompts per topic is a practical starting range rather than a statistical guarantee.

Run a controlled cross-platform test

Run the same prompt set across the AI platforms relevant to the target audience. A study might include ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews, or Google AI Mode, depending on market availability and research access. For every run, preserve:
  • the exact prompt;
  • the full answer, not just whether the brand appeared;
  • visible citations, source titles, URLs, and domains;
  • platform, model or product surface, market, and language;
  • run timestamp and relevant account or personalization conditions;
  • the target brand and competitors being evaluated.
Do not treat one answer as a stable ranking. AI responses can vary across runs, model versions, locations, and user contexts. Repeat prompts and compare aggregates before drawing conclusions.

Create a structured analysis table

Use one row per prompt and platform run. A practical schema is: Automated classification can accelerate the analysis, but ambiguous mentions, sentiment, ranking, and brand-name collisions need human review. Keep the original answer so every derived label can be audited.

Measure the visibility funnel

GEO performance can be evaluated as a sequence rather than a single score: Use metric definitions consistently. For example, define whether “mention rate” counts one mention per answer or every occurrence, and whether “share of voice” is based on prompts, mentions, or ranked positions. Publish the denominator with the result.

Diagnose visibility gaps

Review the results by topic, intent, journey stage, platform, market, and competitor. Prioritize gaps in this order:
  1. A commercially important topic where the brand is consistently absent.
  2. A topic where the brand appears but is described inaccurately or with outdated information.
  3. A topic where the brand appears but the answer does not cite the official site or another reliable source.
  4. A topic with unstable results that needs a larger sample before action.
A high competitor share of voice does not by itself explain the cause. Inspect the cited pages, the claims supported by those pages, their format and freshness, and whether multiple platforms depend on the same source.

Turn findings into action

Match the user question and evidence gap to the appropriate response: Content should answer the question early, use descriptive headings, state applicable conditions and limits, and make important claims easy to verify. Third-party authority may also matter: technical communities, videos, GitHub examples, industry publications, review sites, and local-language sources can influence the evidence available to answer engines.

Repeat the study over time

A GEO study is a measurement cycle, not a one-time audit:
  1. Freeze a baseline prompt set and research conditions.
  2. Run the prompts and retain the raw evidence.
  3. Prioritize and implement actions.
  4. Rerun the unchanged baseline to measure movement.
  5. Maintain a separate exploratory set for new topics and user language.
Keeping the baseline and exploratory sets separate prevents prompt changes from being mistaken for visibility improvements.

GEO platform vs controlled research workflow

There are two common ways to operationalize the method: use an off-the-shelf GEO platform or build a controlled collection workflow with a tool such as Octoparse. Neither approach is universally better. The right choice depends on whether the priority is speed and standardized reporting or control and research transparency. The main advantage of a controlled workflow is not simply “more data.” It is a clearer chain of evidence:
This chain makes the process easier to inspect, challenge, and reproduce. It also prevents a visibility score from becoming the conclusion without showing which prompts, answers, and rules produced it. An off-the-shelf platform is often the better choice when a team needs a standard dashboard quickly and its supported scope matches the research question. A controlled Octoparse workflow is more suitable when the team needs custom sampling, access to underlying evidence, explicit quality-control steps, or integration with its own analysis model. Some teams use both: a platform for directional monitoring and a controlled workflow to investigate high-value or disputed findings.

Collecting GEO research data with Octoparse

For a small study, prompts and answers can be recorded manually. As the number of prompts, platforms, markets, and repetitions grows, consistent collection becomes the main operational challenge. Octoparse provides a free prompt generator for building the initial question set and ready-made visibility templates for supported surfaces. For example, the ChatGPT Visibility Tracker accepts prompts, a target brand, and competitors, then returns structured fields for complete responses, brand mentions and positions, competitors, and citations. Running the same inputs again can support change tracking. Template availability, fields, access, and the behavior of third-party AI surfaces can change. Review the template page before use, retain timestamps and raw responses, and manually validate important conclusions. Octoparse supports the collection and structuring stage of the research; business prioritization, causal interpretation, and GEO strategy still require analyst judgment.

Research limitations

  • AI answers are probabilistic and may vary even when the prompt is unchanged.
  • Results may depend on model version, product surface, geography, language, account state, and personalization.
  • Visible citations do not necessarily expose every source used to produce an answer.
  • Mention position is not always equivalent to preference or recommendation.
  • Sentiment and description accuracy require context-aware review.
  • Platform terms, access rules, and applicable laws still govern collection and reuse.
State these limitations alongside the findings. A reproducible methodology and preserved evidence are more useful than a precise-looking score with unclear inputs.