Skip to main content
Use this workflow when you need to study how AI answer engines mention, describe, compare, cite, or recommend a brand—and you need the underlying responses and sources, not only a dashboard score. Octoparse supports the data collection and structuring stages of GEO research. You define the prompt set and competitors, run supported AI visibility templates, retain the complete responses and citations, and deliver structured records to your own review or analysis process. The result is a traceable research dataset. It is not an automatic explanation of why an AI platform produced an answer, and it does not replace analyst judgment.

What this workflow helps you do

  • Generate a balanced starting set of buyer prompts for a brand, category, market, and language.
  • Test prompts with ready-made templates for supported AI platforms and search surfaces.
  • Collect complete answers alongside brand mentions, recommendation positions, competitors, and visible citations when those fields are available.
  • Preserve the prompt, market, run timestamp, response, and source URLs needed to review a finding.
  • Repeat a stable baseline prompt set to compare visibility over time.
  • Export structured records for human review, spreadsheets, BI tools, or a custom scoring workflow.
Supported platforms, inputs, output fields, access levels, and usage conditions vary by template and may change as third-party AI surfaces change. Review the current template page and run a sample before designing a recurring study around it.

Choose the right GEO template

The AI Visibility Prompt Generator connects the prompt-design step to seven ready-made collection templates. They fall into two groups: brand visibility trackers that add brand- and competitor-level observations, and AI search scrapers that collect the displayed answer and its cited sources for downstream analysis. The three visibility trackers are the most direct starting point when the research question is about a named brand and competitor set. The four search-surface scrapers are useful when the study begins with queries or prompts and the analysis team wants to classify brands, sources, and topics using its own rules.
Treat Gemini, Google AI Overviews, and Google AI Mode as separate surfaces. They can return different answers and citation patterns even though they are all Google products. Naver AI Overview and Naver AI Tab should also be stored as separate surfaces in the dataset.

Define the study before running a template

The collection task should follow the research design, not determine it. Define these elements first: Keep a fixed baseline prompt set for measurement. Put new questions and experimental wording in a separate exploratory set so changes in the sample are not mistaken for changes in visibility.

Learn the GEO research method

Design the prompt universe, test conditions, metrics, gap analysis, and action framework before configuring the collection workflow.

Build the workflow

1

Generate or import the prompt set

Use the AI Visibility Prompt Generator to create 30 tailored buyer prompts from a brand website, competitors, target market, and language, or bring a prompt universe developed from your own customer and search research. Review the prompts before use and label each one by topic, intent, and journey stage.
2

Choose a supported visibility template

Select one or more templates from the matrix above for the AI platforms or search surfaces you need to study. Check each template’s current inputs, output fields, access level, usage conditions, and update date. Different surfaces expose different response and citation structures, so do not assume that every template produces an identical schema.
3

Configure brands and research inputs

Add the prompt set, target brand, competitors, market, or other inputs required by the selected template. Use consistent brand names and competitor lists across comparable runs. Start with a small sample to confirm that the intended entities and fields are recognized correctly.
4

Run and validate a sample

Compare selected output rows with the complete platform responses. Check brand-name collisions, mention and rank fields, official-domain detection, citation URLs, empty answers, and unexpected response formats. Record any exclusion or correction rule before expanding the run.
5

Preserve the raw evidence

Keep the prompt, platform, market, run timestamp, complete response, citation titles, citation URLs, citation domains, and extracted brand fields. Do not overwrite the complete response with a summary: the original answer is the evidence used to audit classifications and conclusions.
6

Export for analysis

Export the structured records in a format that fits the downstream workflow. Keep raw Octoparse output separate from calculated metrics, manual corrections, charts, and recommendations so the source evidence remains intact.
7

Repeat and compare

Rerun the unchanged baseline under comparable conditions. Append or archive each snapshot instead of replacing the previous one, then calculate changes using the same definitions and denominators.

Use a traceable data model

The exact fields depend on the selected template. The Gemini, ChatGPT, and Claude visibility trackers accept prompts, a target brand, and competitors and return platform-specific answers, citations, and extracted brand observations. Google and Naver search templates are query- or prompt-led and focus on the displayed generated result and source records. Add a normalized layer after collection when the study compares these different output schemas. Organize available fields into three layers: Add analyst-controlled fields outside the raw output:
  • review status and reviewer;
  • exclusion reason;
  • corrected entity or rank;
  • description accuracy;
  • sentiment with a supporting passage;
  • recommended action;
  • methodology version.
This separation keeps automated observations, human judgments, and business recommendations from being mixed into one opaque score.

Keep conclusions connected to evidence

A controlled workflow should make every conclusion auditable:
For example, a low mention rate should link back to the exact prompt-platform runs counted in its numerator and denominator. A claim that the brand is described inaccurately should retain the relevant answer passage and its review decision. A citation opportunity should show which domains and pages repeatedly appear for the topic.

Quality-control checklist

Check these items before relying on the dataset:
Do not treat a template’s extracted fields as final research conclusions without validating a sample against the complete responses. AI answers are variable, and third-party interfaces can change their structure or availability.

Analyze and prioritize the results

Calculate metrics only after defining their rules. Common measures include:
  • Mention rate: the share of valid prompt runs that mention the target brand.
  • Recommendation presence: the share of valid runs in which the brand is presented as an option or recommendation.
  • Position: the brand’s order under a documented ranking rule.
  • Share of voice: the target brand’s visibility relative to a defined competitor set and denominator.
  • Official citation rate: the share of valid runs that cite the official domain.
  • Source frequency: how often domains or pages appear across relevant answers.
Prioritize findings that combine business value with strong evidence: important topics where the brand is consistently absent, descriptions are inaccurate, competitors dominate, or authoritative brand information is not cited. Keep unstable or low-sample findings in a monitoring queue rather than turning them directly into content actions.

When to use this workflow

This Octoparse workflow is a good fit when:
  • you need the complete answers and citation records behind the metrics;
  • your team has a custom prompt taxonomy or scoring method;
  • analysis must preserve an auditable evidence trail;
  • you want to compare multiple supported surfaces using a common downstream schema;
  • manual review and explicit exclusion rules are part of the process.
An off-the-shelf GEO platform may be a better fit when you need a standard dashboard quickly and its platforms, markets, metrics, and reporting model already match the decision. Some teams use both: a dashboard for directional monitoring and Octoparse for controlled collection and deeper investigation.

Boundaries and maintenance

  • Octoparse collects and structures data exposed through supported templates or configured tasks; it does not reveal an AI model’s private training data or complete internal reasoning.
  • Visible citations may not represent every source involved in generating an answer.
  • Template support and schemas are specific to each third-party surface.
  • Repeated runs may still produce different answers because AI output is probabilistic.
  • Researchers remain responsible for prompt design, metric definitions, quality review, interpretation, and lawful use of the collected data.
  • Test the workflow again when a third-party interface changes or output fields become incomplete.

AI Visibility Prompt Generator

Generate a balanced starting set of buyer prompts for a brand and market.

ChatGPT Visibility Tracker

Collect complete responses, brand observations, competitor fields, and citations for the supplied prompts.

Gemini Visibility Tracker

Structure Gemini answers, brand observations, competitors, and available citations.

Claude Visibility Tracker

Collect Claude responses, brand context, recommendation positions, competitors, and citations.

Google AI Mode Scraper

Capture AI Mode responses and their displayed source-page details.

Google AIO Scraper

Collect Google AI Overview content and cited source records for tracked queries.

Naver AIO Scraper

Collect Naver AI Overview content, source counts, descriptions, and cited URLs.

Naver AI Tab Scraper

Capture Naver AI Tab prompt responses and cited source records.

GEO and AI visibility research

Learn how to design the study, metrics, gap analysis, and action framework.

Export formats

Choose a structured file format for review or downstream analysis.