Web Search Rerank
Search the web from a natural-language request and return relevance-ranked webpage results with content and source metadata
Overview
Web research often starts with a question rather than a known page. This Data App turns a natural-language request into a relevance-ranked set of webpages, preserving the page text, concise summaries, source identity, and publication context needed for analysis.
It is useful for quickly building a research set around a company, product, topic, or event. Freshness and domain filters help focus the search, while account, organization, and search-role filters can narrow the results when those source attributes are available.
Data notes
Each row represents one webpage returned for the submitted search question. Page URL is used to identify a result and remove duplicates. The result set covers webpage records only; a separate image collection returned by the source is not included here.
Page content may be empty even when a title, snippet, or summary is available. Account name, organization name, site icon, cached URL, language, and family-friendly or navigational flags may also be unavailable for individual pages. Publication and crawl timestamps are preserved as source text without inferring a timezone.
What the results look like
Each row is one relevance-ranked webpage result.
| Page URL | Page title | Site name | Published at | Snippet |
|---|---|---|---|---|
| https://www.example.com/article/123 | Xiaomi car delivery momentum continues to rise | Example Site | 2026-08-05 14:23:10 | Xiaomi car deliveries have continued to grow. |
Additional values include the full page content, summary, display and cached URLs, account and organization names, language, source identifiers, and crawl time.
Use cases
- Build a research set for a company, product, market, or event by combining Page title, Page content, Snippet, and Summary.
- Monitor recent coverage by pairing a Freshness filter with Published at and Last crawled at.
- Compare source visibility by grouping results by Site name, Account name, or Organization name.
- Preserve traceable source links by using Page URL and Display URL when reviewing or exporting results.
Scope and boundaries
Each run accepts one natural-language search request and returns up to 200 relevance-ranked webpage records in one response. Separate image results are not included.
- Good for
- When you need relevance-ranked webpage results for a natural-language research question
- When you need webpage titles, summaries, content, source metadata, and publication timestamps in a structured result set
- When you want to narrow web search by freshness, domains, account, organization, or search role
- Do not use for
- When you need image-only results or a dedicated image search dataset
- When you need private, login-only, or non-public webpage content
Failure handling
Failure and retry behavior declared by the author. We recommend including it in your system prompt when integrating.
- 1If no useful webpages are returned, broaden the query or remove restrictive domain and freshness filters
- 2Retry later after rate limiting, temporary service unavailability, or a timeout
- 3If authentication fails, verify the API Key configured for the execution environment
Input
Parameters required to call this App, generated from the input.schema in manifest.json.
| Field | Business name | Type | Required | Default | Enum / Constraints | Example | Description |
|---|---|---|---|---|---|---|---|
| query | Search question | string | Yes | — | — | Xiaomi car delivery and sales performance | The natural-language topic or question to search. Provide at least one character; clearer wording produces more focused results. |
| count | Result count | integer | No | 10 | 1–200 | — | The number of webpage results to request. The default is 10 and the allowed range is 1 to 200; fewer results may be available for a narrow query. |
| freshness | Freshness filter | string | No | noLimit | — | — | Limits results by recency. Use noLimit, oneDay, oneWeek, oneMonth, oneYear, a date such as 2026-08-05, or a date range such as 2026-08-01..2026-08-31. The default is noLimit. |
| include | Included domains | string | No | — | — | — | Search only the listed domains. Separate up to 100 domain values with | or commas, using domain values supplied by the Data Hub domain list. |
| exclude | Excluded domains | string | No | — | — | — | Exclude the listed domains from the search. Separate up to 100 domain values with | or commas, using domain values supplied by the Data Hub domain list. |
| account_name | Account name | string | No | — | — | — | Filters results to an exact account name when the source identifies one. Leave empty to search without an account filter. |
| office_name | Organization name | string | No | — | — | — | Filters results to an exact organization name when the source identifies one. Leave empty to search without an organization filter. |
| role_code | Search role | string | No | — | — | — | Uses a configured Data Hub search role, such as WEB_SEARCH. Leave empty to use the system default role. |
Output
Field structure of a single record, generated from the output.schema in manifest.json.
| Field | Business name | Type | Example | Description |
|---|---|---|---|---|
| account_name | Account name | string | China News Network | The account name associated with the webpage when available. |
| cached_page_url | Cached page URL | string | — | The cached webpage URL when available. |
| content | Page content | string | Xiaomi car deliveries have continued to grow, with stronger showroom traffic and steady delivery scheduling. | The webpage body content when available. It is independent from the snippet and summary and may be empty. |
| date_last_crawled | Last crawled at | string | 2026-08-05 15:00:00 | The source timestamp of the latest crawl, preserved as returned without timezone conversion. |
| date_published | Published at | string | 2026-08-05 14:23:10 | The webpage publication timestamp when available, preserved as returned without timezone conversion. |
| display_url | Display URL | string | https://www.example.com/article/123 | The URL formatted for displaying the webpage result when available. |
| id | Result ID | string | 0 | The source sorting identifier for this webpage result. |
| is_family_friendly | Family-friendly | boolean | false | Whether the source marks the webpage as suitable for family browsing. Empty source values are represented as false. |
| is_navigational | Navigational result | boolean | false | Whether the source marks the result as navigational. Empty source values are represented as false. |
| language | Content language | string | en | The language code reported for the webpage when available. |
| name | Page title | string | Xiaomi car delivery momentum continues to rise | The title of the webpage result. |
| office_name | Organization name | string | — | The organization associated with the webpage when available. |
| site_icon | Site icon URL | string | — | The site icon URL when available. |
| site_name | Site name | string | Example Site | The name of the site hosting the webpage when available. |
| snippet | Snippet | string | Xiaomi car deliveries have continued to grow. | A short source-provided excerpt for the webpage when available. |
| summary | Summary | string | Xiaomi car deliveries have continued to grow, with stronger showroom traffic and steady delivery scheduling. | A more complete source-provided summary of the webpage when available. |
| url | Page URL | string | https://www.example.com/article/123 | The original webpage URL. This value identifies and de-duplicates a result record. |
Record schema
Output is returned record by record. detail.output.idFieldHint
{
"type": "object",
"properties": {
"account_name": {
"type": "string",
"title": "Account name",
"description": "The account name associated with the webpage when available.",
"prefill": "China News Network"
},
"cached_page_url": {
"type": "string",
"title": "Cached page URL",
"description": "The cached webpage URL when available."
},
"content": {
"type": "string",
"title": "Page content",
"description": "The webpage body content when available. It is independent from the snippet and summary and may be empty.",
"prefill": "Xiaomi car deliveries have continued to grow, with stronger showroom traffic and steady delivery scheduling."
},
"date_last_crawled": {
"type": "string",
"title": "Last crawled at",
"description": "The source timestamp of the latest crawl, preserved as returned without timezone conversion.",
"prefill": "2026-08-05 15:00:00"
},
"date_published": {
"type": "string",
"title": "Published at",
"description": "The webpage publication timestamp when available, preserved as returned without timezone conversion.",
"prefill": "2026-08-05 14:23:10"
},
"display_url": {
"type": "string",
"title": "Display URL",
"description": "The URL formatted for displaying the webpage result when available.",
"prefill": "https://www.example.com/article/123"
},
"id": {
"type": "string",
"title": "Result ID",
"description": "The source sorting identifier for this webpage result.",
"prefill": "0"
},
"is_family_friendly": {
"type": "boolean",
"title": "Family-friendly",
"description": "Whether the source marks the webpage as suitable for family browsing. Empty source values are represented as false.",
"prefill": false
},
"is_navigational": {
"type": "boolean",
"title": "Navigational result",
"description": "Whether the source marks the result as navigational. Empty source values are represented as false.",
"prefill": false
},
"language": {
"type": "string",
"title": "Content language",
"description": "The language code reported for the webpage when available.",
"prefill": "en"
},
"name": {
"type": "string",
"title": "Page title",
"description": "The title of the webpage result.",
"prefill": "Xiaomi car delivery momentum continues to rise"
},
"office_name": {
"type": "string",
"title": "Organization name",
"description": "The organization associated with the webpage when available."
},
"site_icon": {
"type": "string",
"title": "Site icon URL",
"description": "The site icon URL when available."
},
"site_name": {
"type": "string",
"title": "Site name",
"description": "The name of the site hosting the webpage when available.",
"prefill": "Example Site"
},
"snippet": {
"type": "string",
"title": "Snippet",
"description": "A short source-provided excerpt for the webpage when available.",
"prefill": "Xiaomi car deliveries have continued to grow."
},
"summary": {
"type": "string",
"title": "Summary",
"description": "A more complete source-provided summary of the webpage when available.",
"prefill": "Xiaomi car deliveries have continued to grow, with stronger showroom traffic and steady delivery scheduling."
},
"url": {
"type": "string",
"title": "Page URL",
"description": "The original webpage URL. This value identifies and de-duplicates a result record.",
"prefill": "https://www.example.com/article/123"
}
},
"required": [],
"additionalProperties": false
}Integration
This App can be integrated via MCP, API, SDK, or file export. All channels share the same capabilities and pricing. Every request authenticates with the Authorization: Bearer header using an API key (long-lived, created in the Data Hub Console); MCP clients can also sign in with OAuth, no key required. More options such as CLI and Skill are on the way.
With the MCP (Model Context Protocol), you can call this App directly from AI clients like Claude and Cursor. Pick your client and auth mode, then copy the config below.
Client config
Replace the value after Bearer with your long-lived API key. Works in any client, CI, or headless environment.
{
"mcpServers": {
"meme_today__data-hub-web-search-rerank": {
"type": "http",
"url": "https://mcp-v2.octoparse.com?pin=meme_today/data-hub-web-search-rerank",
"headers": { "Authorization": "Bearer <YOUR_API_KEY>" }
}
}
}Let AI set it up for you
Don't want to edit configs by hand? Copy the install prompt and paste it into any AI client. It will complete the setup its own way. (The prompt asks the AI to request your API key from you, so credentials never end up in chat history or shared configs.)
detail.access.mcp.composeHint
Pricing
Charged by the number of records successfully returned. Failed runs are not charged. Billed in units of 50 record; any partial unit is rounded up to 50 record.
Multiple billing events accumulate independently. See each item for details. Failed runs are not charged.
View creditsTry it now
Fill in the parameters and run. Results come from a real call.