> ## Documentation Index
> Fetch the complete documentation index at: https://www.octoparse.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# GEO and AI visibility research

> Learn how to run GEO research: build prompt sets, test AI platforms, measure brand mentions and citations, and prioritize visibility gaps.

Generative engine optimization (GEO) research measures how AI answer engines discover, describe, cite, compare, and recommend a brand. A useful study does more than ask a few branded questions: it tests realistic user prompts across multiple platforms, preserves the complete answers and sources, and repeats the test under consistent conditions.

The goal is to identify where a brand appears, where it is absent or misrepresented, which sources shape the answers, and which content or authority gaps deserve attention.

## Start with a research question

Define the decision the research should support before collecting data. Common questions include:

* Does the brand appear when buyers discover or compare products in its category?
* Which competitors receive more mentions or stronger recommendations?
* How accurately do AI platforms describe the brand and its capabilities?
* Do answers cite the brand's website, third-party sources, or no visible source?
* Which topics combine meaningful demand, commercial value, and a large visibility gap?

Also define the target market, language, AI platforms, competitor set, collection period, and model or interface used. These conditions are part of the result: changing them can change the answer.

## Build a prompt universe

A prompt universe is the controlled set of questions used in the study. It should represent how real users express needs, not simply turn a keyword list into questions.

Design prompts across several dimensions:

| Dimension           | Examples                                                                          |
| ------------------- | --------------------------------------------------------------------------------- |
| Audience            | Individual buyer, small team, enterprise decision-maker, practitioner, specialist |
| Need                | Learn about a category, solve a problem, compare approaches, choose a provider    |
| Context             | Industry, organization size, location, language, level of experience              |
| Evaluation criteria | Price, quality, ease of use, reliability, support, compatibility, reputation      |
| Constraint          | Limited budget, short timeline, regulatory requirements, existing technology      |
| Journey stage       | Discovery, use case, comparison, evaluation                                       |

The [Octoparse AI Visibility Prompt Generator](https://www.octoparse.com/web-tools/geo-prompt-generator) uses a brand website, competitors, target market, and prompt language to generate 30 tailored prompts. It balances the set across four stages:

1. **Discovery** — questions people ask before they know which brands to consider.
2. **Use case** — problem- and scenario-based questions about completing a task.
3. **Comparison** — alternatives, brand comparisons, and competitive shortlists.
4. **Evaluation** — questions about fit, reputation, and recommendations.

This balance matters. Branded prompts mostly test what happens after a user already knows the company; non-branded discovery and use-case prompts reveal whether the brand enters the consideration set at all.

<Tip>
  Use real customer language where possible. Long-form search queries, support conversations, sales questions, community discussions, and onsite search terms can reveal wording that a generic keyword list misses.
</Tip>

### Control the sample size

Start with a broad scan, then deepen the areas that matter most.

* **Baseline scan:** cover the main categories, use cases, competitors, journey stages, markets, and platforms.
* **Focused study:** add more prompts and local-language variations for topics with high value, strong competitor visibility, or a clear brand gap.

Avoid generating every possible combination of website, field, country, and wording. Use reusable prompt patterns, but keep enough variation to represent distinct user intents. For an initial manual study, five to ten prompts per topic is a practical starting range rather than a statistical guarantee.

## Run a controlled cross-platform test

Run the same prompt set across the AI platforms relevant to the target audience. A study might include ChatGPT, Gemini, Claude, Perplexity, Google AI Overviews, or Google AI Mode, depending on market availability and research access.

For every run, preserve:

* the exact prompt;
* the full answer, not just whether the brand appeared;
* visible citations, source titles, URLs, and domains;
* platform, model or product surface, market, and language;
* run timestamp and relevant account or personalization conditions;
* the target brand and competitors being evaluated.

Do not treat one answer as a stable ranking. AI responses can vary across runs, model versions, locations, and user contexts. Repeat prompts and compare aggregates before drawing conclusions.

## Create a structured analysis table

Use one row per prompt and platform run. A practical schema is:

| Field                   | What to record                                                           |
| ----------------------- | ------------------------------------------------------------------------ |
| Topic / intent / prompt | Research category, journey stage, and exact question                     |
| Test conditions         | Platform, model or surface, market, language, timestamp                  |
| Complete response       | Full answer for later review and reclassification                        |
| Brand visibility        | Mentioned or absent, mention position, recommendation position           |
| Description             | How the brand is characterized and whether the statement is accurate     |
| Competitors             | Competitor mentions, positions, and co-occurrence                        |
| Citations               | Source title, URL, domain, and whether the official website is cited     |
| Sentiment               | Positive, neutral, negative, or mixed, with the supporting passage       |
| Review status           | Human-verified, needs review, or excluded as noise                       |
| Recommended action      | Content, technical, third-party authority, product messaging, or monitor |

Automated classification can accelerate the analysis, but ambiguous mentions, sentiment, ranking, and brand-name collisions need human review. Keep the original answer so every derived label can be audited.

## Measure the visibility funnel

GEO performance can be evaluated as a sequence rather than a single score:

| Level                     | Questions and indicators                                                                     |
| ------------------------- | -------------------------------------------------------------------------------------------- |
| Discoverable              | Can relevant crawlers access the site, and which pages or grounding queries appear?          |
| Cited                     | Is the website used as a source? What is the share of official versus third-party citations? |
| Mentioned and recommended | Mention rate, share of voice, recommendation position, competitor co-occurrence, sentiment   |
| Business outcome          | AI referral traffic, branded search, organic visits, sign-ups, inquiries, and conversions    |

Use metric definitions consistently. For example, define whether "mention rate" counts one mention per answer or every occurrence, and whether "share of voice" is based on prompts, mentions, or ranked positions. Publish the denominator with the result.

## Diagnose visibility gaps

Review the results by topic, intent, journey stage, platform, market, and competitor. Prioritize gaps in this order:

1. A commercially important topic where the brand is consistently absent.
2. A topic where the brand appears but is described inaccurately or with outdated information.
3. A topic where the brand appears but the answer does not cite the official site or another reliable source.
4. A topic with unstable results that needs a larger sample before action.

A high competitor share of voice does not by itself explain the cause. Inspect the cited pages, the claims supported by those pages, their format and freshness, and whether multiple platforms depend on the same source.

## Turn findings into action

Match the user question and evidence gap to the appropriate response:

| Observed gap                                           | Possible action                                                                  |
| ------------------------------------------------------ | -------------------------------------------------------------------------------- |
| Users ask how to choose a tool                         | Publish a category guide with explicit evaluation criteria                       |
| Users ask how to complete a task                       | Create a use-case tutorial, workflow, or reusable template                       |
| Competitors dominate comparison prompts                | Publish a sourced, condition-based comparison or alternative page                |
| Product descriptions are inaccurate                    | Clarify stable positioning and product facts on authoritative pages              |
| AI cites third-party sources but not the official site | Improve the relevant official page and strengthen consistent external references |
| Important claims lack evidence                         | Add examples, test results, methodology, limitations, and update dates           |

Content should answer the question early, use descriptive headings, state applicable conditions and limits, and make important claims easy to verify. Third-party authority may also matter: technical communities, videos, GitHub examples, industry publications, review sites, and local-language sources can influence the evidence available to answer engines.

## Repeat the study over time

A GEO study is a measurement cycle, not a one-time audit:

1. Freeze a baseline prompt set and research conditions.
2. Run the prompts and retain the raw evidence.
3. Prioritize and implement actions.
4. Rerun the unchanged baseline to measure movement.
5. Maintain a separate exploratory set for new topics and user language.

Keeping the baseline and exploratory sets separate prevents prompt changes from being mistaken for visibility improvements.

## GEO platform vs controlled research workflow

There are two common ways to operationalize the method: use an off-the-shelf GEO platform or build a controlled collection workflow with a tool such as Octoparse. Neither approach is universally better. The right choice depends on whether the priority is speed and standardized reporting or control and research transparency.

| Research consideration | Off-the-shelf GEO platform                                                                                         | Controlled workflow with Octoparse                                                                                                       |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- |
| Time to first result   | Usually faster because prompts, metrics, and dashboards are predefined                                             | Requires initial setup, field design, and test runs                                                                                      |
| Research scope         | Works best within the platforms, markets, models, and metrics the product supports                                 | Researchers can define the target surfaces, inputs, fields, and sampling plan supported by the selected templates or custom tasks        |
| Prompt control         | May organize or generate prompts within a fixed methodology                                                        | The team controls the baseline set, exploratory set, topic taxonomy, language variants, and test inputs                                  |
| Raw evidence           | Access varies; some products emphasize scores and summaries                                                        | The workflow can retain complete responses, citations, URLs, timestamps, and other extracted evidence as structured records              |
| Metric logic           | Standardized metrics make reporting and benchmarking easier, but calculation details may be fixed or partly opaque | The team defines denominators, exclusions, ranking rules, and aggregation logic outside the collection task                              |
| Conclusion clarity     | Dashboards provide quick summaries, although a composite score can hide the observations behind it                 | Findings can be traced from a conclusion to a metric, record, complete answer, and cited source when the data model preserves them       |
| Process management     | Collection, analysis, and reporting are usually managed inside one product                                         | Collection tasks, schedules, run history, logs, exports, and the downstream analysis process can be managed as separate, explicit stages |
| Quality control        | Validation follows the checks offered by the platform                                                              | Teams can add required fields, review statuses, exclusion reasons, duplicate checks, and manual verification points                      |
| Adaptability           | Convenient for recurring standard reports; unusual research questions may not fit the available schema             | Fields and workflows can be adjusted when the research question, AI surface, or evidence requirements change                             |
| Maintenance            | The vendor maintains collection and metric implementation                                                          | The research team must test outputs and maintain tasks when interfaces or page structures change                                         |
| Expertise required     | Lower setup burden for standard monitoring                                                                         | Higher research and operational burden; clear methodology and analyst review are required                                                |

The main advantage of a controlled workflow is not simply “more data.” It is a clearer chain of evidence:

```text theme={null}
Research question
→ Prompt and test conditions
→ Complete platform response
→ Extracted mention, position, and citation fields
→ Human-reviewed classification
→ Defined metric
→ Conclusion and recommended action
```

This chain makes the process easier to inspect, challenge, and reproduce. It also prevents a visibility score from becoming the conclusion without showing which prompts, answers, and rules produced it.

An off-the-shelf platform is often the better choice when a team needs a standard dashboard quickly and its supported scope matches the research question. A controlled Octoparse workflow is more suitable when the team needs custom sampling, access to underlying evidence, explicit quality-control steps, or integration with its own analysis model. Some teams use both: a platform for directional monitoring and a controlled workflow to investigate high-value or disputed findings.

## Collecting GEO research data with Octoparse

For a small study, prompts and answers can be recorded manually. As the number of prompts, platforms, markets, and repetitions grows, consistent collection becomes the main operational challenge.

Octoparse provides a [free prompt generator](https://www.octoparse.com/web-tools/geo-prompt-generator) for building the initial question set and ready-made visibility templates for supported surfaces. For example, the [ChatGPT Visibility Tracker](https://www.octoparse.com/template/chatgpt-visibility-tracker) accepts prompts, a target brand, and competitors, then returns structured fields for complete responses, brand mentions and positions, competitors, and citations. Running the same inputs again can support change tracking.

Template availability, fields, access, and the behavior of third-party AI surfaces can change. Review the template page before use, retain timestamps and raw responses, and manually validate important conclusions. Octoparse supports the collection and structuring stage of the research; business prioritization, causal interpretation, and GEO strategy still require analyst judgment.

## Research limitations

* AI answers are probabilistic and may vary even when the prompt is unchanged.
* Results may depend on model version, product surface, geography, language, account state, and personalization.
* Visible citations do not necessarily expose every source used to produce an answer.
* Mention position is not always equivalent to preference or recommendation.
* Sentiment and description accuracy require context-aware review.
* Platform terms, access rules, and applicable laws still govern collection and reuse.

State these limitations alongside the findings. A reproducible methodology and preserved evidence are more useful than a precise-looking score with unclear inputs.

## Related resources

* [GEO and AI visibility monitoring with Octoparse](/docs/en/solutions/geo-ai-visibility-monitoring) — turn the research method into a controlled collection workflow
* [Data collection for market and competitive research](/docs/en/academy/market-research) — design recurring research around comparable web data
* [Data collection for AI and LLMs](/docs/en/academy/ai-data-collection) — prepare traceable, structured web data for AI workflows
* [Octoparse templates](https://www.octoparse.com/template) — browse ready-made collection workflows
* [Is web scraping legal?](/docs/en/academy/is-web-scraping-legal) — understand collection scope and responsibility
