logo
languageENdown
menu
Template GalleryTemplate Details

📌 How can AI extract custom fields from a list of web pages?

This template collects structured fields requested by the user from each supplied public page URL. It is useful for no-code data users, research teams, and data operations teams.

Data is collected from User-supplied public web pages. This template is not tied to one source website. It extracts user-requested fields from the public page URLs supplied for the task.


💰 Pricing

This template has no per-line usage fee. Basic Proxy is free. Enhanced Proxy adds $0.04 per URL. Auto tries Basic Proxy first and automatically switches to Enhanced Proxy if Basic fails; that switch can add the Enhanced Proxy charge.


📦 Output

Output fields are defined by the user's extraction instructions, so this template has no fixed output-field schema.

{}

🎯 Use Cases

  • Extract product attributes such as titles and prices when those fields are explicitly requested.
  • Collect article or listing attributes from heterogeneous page layouts using a shared field specification.
  • Prototype a structured dataset before building a site-specific scraper.
  • Process pages that require the selected proxy or supplied cookies when those settings are provided.

🐙 Why Octoparse

  • Built for User-supplied web pages: The Python automation validates each target URL, submits the page together with the user-defined field specification and selected proxy or cookie settings to the AI extraction service, waits for the result, and returns the dynamically shaped record.
  • Inputs match the workflow: The form uses Target URLs (up to 10,000 entries), Fields to Extract, and Proxy.
  • Output follows your request: AI Scraper creates the fields requested at runtime instead of forcing every website into one fixed schema.
  • Ready for repeat use: After AI Scraper runs in the cloud, schedule eligible tasks and export the rows for spreadsheets, dashboards, APIs, or AI workflows.

📝 Input

Complete the following fields:

  • Target URLs (Required) — Enter the public page URLs to process. Starting with no more than 10 URLs is recommended to avoid excessive credit use. Up to 10,000 entries per run.
  • Fields to Extract (Required)
  • Proxy (Required) — Choose Basic Proxy for free processing, Enhanced Proxy for an additional $0.04 per URL, or Auto to try Basic first and automatically switch to Enhanced if Basic fails. Available options: Basic Proxy - Free, Enhanced Proxy - Extra $0.04 per URL, Auto - Try Basic first and switch to Enhanced if Basic fails.
  • Cookies (Optional) — Optionally paste cookies for pages that require login. Use cookies from only one website in each task.

🚀 How to Use

  1. Open AI Scraper and click Try it!.
  2. Complete Target URLs, Fields to Extract, and Proxy using the formats and limits shown in the Input section.
  3. Run it in the Octoparse cloud and start the task.
  4. Review the dynamically generated columns, then export the rows in the format you need.

💡 Tips

  • Test one representative value in Target URLs before submitting a large batch to AI Scraper.
  • Keep the requested field names consistent when combining or comparing repeated exports.

❓ FAQ

How many entries can AI Scraper accept in Target URLs per run?

Enter up to 10,000 values in Target URLs per run.

Which pages does AI Scraper process?

The workflow is configured for web pages; its output columns are created from the fields requested at runtime.

What can I use data from AI Scraper for?

Extract product attributes such as titles and prices when those fields are explicitly requested.


  • AI Crawl — Use the crawl workflow when pages must first be discovered under a starting website.
  • Universal Content Scraper — Use a fixed content-crawling workflow for structured extraction across a website.
  • Google Search Scraper — Use Google search result URLs as target pages for custom extraction.
Share