Skip to content
AgentSearch

agentsearch-web-extract-v1 · Pocket MainNet

URL to markdown extraction API for AI agents and RAG

AgentSearch Web Extract turns any URL into clean page content: markdown or plain text with the title, up to 100 links and fetch metadata. It is web scraping for LLMs without the boilerplate, built for RAG extraction, citations and long-context prompts.

$0.005 per callUSDC via x402 or MPPNo account, no API keyagentsearch-web-extract-v1

What it does

Clean markdown or text

Navigation, scripts and page chrome are stripped. Choose formats markdown, text or both, capped by max_chars (up to 100,000; default 50,000).

Citation-ready metadata

Get the final URL after redirects, the title, links and meta (status code, content type, fetch time, character count and a truncated flag).

Predictable errors

Target-site failures return HTTP 200 with an error object and a retryable flag. Private, loopback and metadata addresses are refused.

When to use it vs. the other two

AgentSearch has three services. Pick by what your agent has in hand and where the content lives.

Find sources

Web Search

You have a question, not a URL. Get up to 5 ranked results with title, URL, snippet, score, domain and date.

Web Search API →
This page

Web Extract

You have a URL and the content is in the HTML the server sends: docs, blogs, news, most marketing pages. Fast and cheap.

Read a JavaScript page

Web Render

The content only appears after JavaScript runs: SPAs, dashboards, client-rendered lists. Headless Chromium renders it first.

Web Render API →

Request and response

Request (Agentic Portal)

curl -X POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract \
  -H 'content-type: application/json' \
  -d '{"url":"https://example.com"}'
# → 402 Payment Required with the terms.
#   An x402 or MPP client pays and retries.

Response

{
  "portal": {
    "provenance": "third-party-supplier",
    "serviceId": "agentsearch-web-extract-v1",
    "schemaCheck": "unchecked"
  },
  "data": {
    "request_id": "05f5a4ebe2684c79867aa9bf3515c33d",
    "url": "https://example.com/",
    "title": "Example Domain",
    "markdown": "# Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n",
    "text": "",
    "links": [
      { "href": "https://iana.org/domains/example", "text": "Learn more" }
    ],
    "meta": {
      "status_code": 200,
      "content_type": "text/html",
      "fetched_at": "2026-09-25T15:53:07.843243+00:00",
      "chars": 167,
      "truncated": false
    },
    "error": null
  }
}

The real captured example from the Agentic Portal service page. Options and error codes: extract OpenAPI spec. Step-by-step guide: Turn any URL into clean markdown for RAG.

How to connect

HTTP

Agentic Portal (pay per call)

POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/…. An unpaid request returns 402 with the terms; an x402 (Base) or MPP (Tempo) client pays $0.005 in USDC and retries. No account, no API key.

Open the portal page →
MCP (recommended)

Pocket Agentic Portal MCP

Add npx -y @pocket-network/agentic-portal-mcp to Claude Desktop, Cursor or Claude Code. Use describe_service and call_service on agentsearch-web-extract-v1. It pays per call from a local wallet with spend limits.

MCP setup guide →
Self-hosted

Direct mode MCP

Prefer to run your own stack? @agentsearchhq/agentsearch-mcp is a self-hosted MCP server with a direct mode.

Self-hosted MCP on npm →

Use cases

RAG ingestion

Turn pages into markdown, chunk by heading, embed, and keep the source URL and fetch time as metadata.

Agent browsing

Let an agent read a changelog, API reference or pricing page in full after Web Search finds it.

Freshness checks

Re-extract stored documents on a schedule using meta.fetched_at, and update what changed.

FAQ

What does the URL to markdown API return?

The final URL, title, markdown and/or text, up to 100 absolute links, and meta (status code, content type, fetch time, character count, truncated flag). error is null on success.

How long can an extracted page be?

Up to 100,000 characters per call through max_chars. The default is 50,000, and meta.truncated tells you if the cap was hit.

Can it extract PDFs or JavaScript-only pages?

No. Non-HTML content such as PDFs returns UNSUPPORTED_CONTENT. For pages that only render with JavaScript, use Web Render.

How much does it cost, and do I need an API key?

It costs $0.005 per call in USDC over x402 (Base) or MPP (Tempo) on the Pocket Agentic Portal. No account and no API key.

Other AgentSearch services