Why Use This AI Crawl Template?
🚀 One URL In, Full Dataset Out Give AI Crawl a single list page — a homepage, category page, channel page, or search result page — and it will automatically pick the right detail links and extract structured data from every one of them.
🧠 AI-Powered URL Selection No include/exclude path rules, no regex, no crawl depth tuning. Just describe in plain language which links you want ("all article detail links", "all product pages posted in 2025", etc.) and AI will read the page's Markdown and select the matching URLs.
🔎 Flexible Field Definition Define fields one by one manually, or just describe what you want in natural language — AI generates the field list for you and extracts each field from every detail page.
💼 Perfect for Teams & Businesses Content teams, e-commerce analysts, researchers, and marketers rely on AI Crawl to turn any list page into a clean, ready-to-use dataset — without writing a single scraper.
Explore Our AI-Powered Scraping Options
Data Output
{"price":"$6.99","name":"Nee doh Jellyfish Squeeze Toys, Drop Malt Sugar Balls Relieve Stress, Sensory Toy with a Super Solid Squish for Adult, Squishy Ultra Squishy and Moldable Slow Rise","brand":"Generic","status_of_stock":"In Stock"},
{"price":"$12.99","name":"2026 New Dragon Fly Clips, 3D Artificial Dragonfly Hat Clip, Dragonfly Garden Decor (10PCS)","brand":"Generic","status_of_stock":"In Stock"},
{"price":"$9.99","name":"Glitter Dumpling Squishy, 2026 Upgrade Dumpling Squishy Mystery Box, Colorful Dumplings Stress Balls Fidget Sensory Toy, Stocking Stuffers with Food Steamer Stretchy Desk Toys","brand":"wartleves","status_of_stock":"In Stock"},
{"price":"$42.00","name":"Pokémon TCG: Mega Evolution—Perfect Order Booster Bundle","brand":"Pokémon","status_of_stock":"Only 1 left in stock - order soon"},
{"price":"$52.50","name":"Google Nest Mini Bluetooth Speaker, Japan Model, Multi Language with English Compatibility Assistant (2nd Gen) Charcoal (Renewed)","brand":"Google","status_of_stock":"In Stock"},Who Is This For?
📰 News Monitoring & Media Analysts
Point AI Crawl at a news site's section or column (TechCrunch AI, a financial daily's policy vertical), filter semantically ("stories mentioning funding rounds", "regulation deep dives"), and extract companies, amounts, and each article's core thesis from the body — not just the headline.
🛒 Discerning Shoppers & Market Researchers
Filter a category or search page semantically ("under $500, 4+ stars, sustainable"; "pet-friendly rentals near transit"), and get each listing's top 3 selling points and complaints — sentiment no scraper can parse.
📚 Academic & Scientific Researchers
Feed in a conference proceedings page or journal issue, filter papers by topic ("agent benchmarks", "uses RLHF"), and distill methods, datasets, and findings from each abstract — fields that only exist once the text is read.
💼 Single-Company Deep Dives — Sales, Recruiting & Strategy
Drop in one company's Careers, Blog, or Case Studies index and pull prose-level insights: top-5 required skills, each post's core argument, or customer problem + outcome numbers — not DOM fields.
🐙 Why Choose Octoparse for AI Crawling?
🧩 No-Code Required You don't need any programming knowledge — just paste one starting URL, describe the links you want, define fields, and run the task.
🔄 Cross-Layout Compatibility AI interprets both the list page's Markdown and each detail page's structure automatically — even when detail pages within the same site have inconsistent layouts, the output stays uniformly structured.
🔐 Login & Proxy Support Access login-protected pages via cookies, and bypass aggressive anti-bot measures with built-in proxy options.
📁 Easy Export Formats Export structured results directly to Excel, CSV, JSON, XML, HTML, or push straight into Google Sheets.
👩💻 Beginner-Friendly UI Ready to use out of the box — describe your links and fields in plain language and let AI handle the rest.
Start Turning Any List Page Into Structured Data. Join 100,000+ users extracting web data with Octoparse.
⚠️ Important Notes Before You Start
- Use responsibly and comply with each target site's Terms of Service.
- AI Crawl uses a strict two-layer model: Starting URL → Detail Pages. It does not auto-paginate or follow "Next Page" links. If you need to cover multiple list pages, submit them as separate tasks.
- The starting URL must be a page that already contains the detail links you want. If those links are hidden behind pagination, filters, or another click, go one level deeper first.
- AI Crawl reads each page once as initially rendered — it does not simulate clicks, scroll, or other user interactions to load additional content.
- Only one site's cookies are supported per task. Cross-site cookie reuse is not supported yet.
- AI extraction is powerful but not deterministic — both URL selection and field results may vary slightly across runs on pages with complex nested structures.
- Ensure personal data is handled ethically and in compliance with applicable regulations.
Ready to Start Crawling With AI?
Try It Now – Free Trial No coding. No rule configuration. Works immediately.
❓ FAQs
Do I need to write any scraping rules or path filters? No. AI Crawl reads the Markdown of your starting URL and picks detail links based on your natural-language prompt — no XPath, CSS selectors, include/exclude rules, or regex required.
What kind of URL should I use as the starting URL? A list-style page that already contains the detail page links you want to scrape — a homepage, category page, channel page, search results page, or index page. If the links you need aren't visible on that page, it won't work as a starting URL.
Can AI Crawl follow pagination or go multiple layers deep? No. AI Crawl is strictly two-layer: starting URL → detail pages. It does not follow "Next Page" links or crawl deeper than one level. For paginated sources, submit each page as a separate task.
Can I extract the same fields from different websites in one task? Not in the same task — each AI Crawl task starts from one list page and scrapes its detail pages. If you need cross-site extraction from a known URL list, use AI Scrape instead.
How many detail pages can AI Crawl handle? AI Crawl scrapes every detail URL your prompt selects from the list page, up to a cap of 9999 URLs per task. For best results, run a small test (narrow your prompt to ~10 URLs) first to validate both URL selection and field extraction before scaling up.
Does AI Crawl support login-protected pages? Yes. Paste valid cookies into advanced settings to simulate a logged-in session. Make sure cookies are still valid before running the task.
What if a site has strong anti-bot protection? Enable proxy IP in the advanced settings. Basic Proxy is free; Enhanced Proxy costs an additional $0.004 per record.
How much does it cost? Page scrape fee is $0.005 per page crawled — AI Crawl bills by the number of pages visited, not by records successfully returned. Duplicate URLs are billed separately: if the same page appears twice in your list, you'll pay twice, even though the extracted content is identical. Proxy fees apply only if you enable Enhanced Proxy.
Can AI Crawl handle infinite scroll or click-to-expand content? No. AI Crawl reads each page as initially rendered. For pages that require user interaction, use Octoparse's standard task builder (OTD) instead.
When should I use AI Scrape instead of AI Crawl? If you already have the list of detail URLs to scrape, use AI Scrape — it's purpose-built for a known URL list. Use AI Crawl when you only have one list page and want AI to discover the detail URLs for you.
🛠 How to Use: Step-by-Step Guide
1. Start the template
Click "Try it!" to open AI Crawl.
2. Enter your starting URL and prompt
On the input screen, enter one list-page URL (homepage, category, channel, or search results page) and write a natural-language prompt describing which links on that page you want to extract (e.g. "extract all article detail links, ignore category and ad links").
3. Input Fields Explained
Input Field | Requirement | Description | Example |
Starting URL | Required | Enter one list-page URL — a homepage, category page, channel page, search results page, or index page that already contains the detail links you want to scrape. Only one URL per task; pagination is not followed automatically. | https://www.example.com/news/tech |
URL Selection Prompt | Required | Describe in plain language which links on the list page you want to extract. AI will read the page's Markdown and select the matching detail URLs. Be specific — vague prompts pull in nav, footer, and ad links. | "Extract all article detail links under the Tech section, ignore category and tag links." |
Extraction Fields (Option A – Manual) | Required (A or B) | Specify each field directly with a field name (required) and field description (optional). Data type is auto-detected by AI. Supported types: Text, Number, Decimal, Date, URL, Image, List. | Article Title, Author, Publish Date, Body Text |
Extraction Fields (Option B – Natural Language) | Required (A or B) | Describe what you want in plain language. AI generates the field list for you, which you can review, edit, add to, or remove before running. | "I want to extract the title, author, publish date, and body text from each article page." |
Proxy | Optional | Enable proxy IP for sites with strong anti-bot measures. Basic Proxy is free; Enhanced Proxy costs an additional $0.004 per page crawled. | Basic / Enhanced |
4. Define the fields for detail pages
Tell AI exactly what data you need from each detail page — title, body text, author, publish date, price, image URLs, etc. You can define fields manually or describe them in natural language.
5. Run the scraper
Click "Start" and select a run mode. AI Crawl will fetch the starting URL, convert it to Markdown, use the prompt to select the target detail URLs, and then scrape each detail page against your field definitions — returning one structured row per detail page.
6. Monitor & handle interruptions
Scraping duration varies based on the number of detail URLs selected and target site response time. If a sub-task shows status "Stopped" (e.g. due to a network glitch or expired cookies), re-run that sub-task to continue scraping.
7. Export your data
Once scraping completes, open the Data Preview tab to review the results, then export in any of the formats listed above
💡 Tips
- Always run a small test first (narrow your prompt to 5–10 URLs) to validate both URL selection and field extraction before scaling up — this avoids unnecessary credit usage from a vague prompt or a wrong starting URL.
- Be specific in your URL-selection prompt. "Get the links" will pull in nav, footer, and ad links. "Extract all article detail links under the 'Tech' section, ignore category and tag links" gives a clean set.
- Use field descriptions to improve accuracy. When a field name is ambiguous, a short description makes a real difference — for example, distinguishing "original price" from "discounted price."
- Pick the right starting URL. The starting URL must be a page that already contains the detail links you want. If they're buried behind pagination or filters, go one level deeper first.
- If you already have a known list of detail URLs, use AI Scrape instead — no URL-discovery step needed.
- For large projects, split starting URLs into multiple tasks (e.g. one per category page) to keep runs stable and easy to monitor.
