Tools / pdf.extract
PDF text extractor
pdf.extract
Page-by-page text from a text-based PDF, given as a URL or base64, up to 10 MB.
MediaLaunch set2 units per callnot cachedMCP: pdf_extractv1.0.0
When to use it
Use to read the text of a PDF such as a report, contract or paper. Works on text-based PDFs; for scanned pages with no text layer, use image.ocr on a rendered page image instead.
Input
| Field | Type | Description |
|---|---|---|
url | string (≤ 2048 chars) | Absolute http or https URL of the file. Use instead of base64. |
base64 | string (≤ 14000000 chars) | Base64-encoded file bytes, up to 10 MB once decoded. Use instead of url. |
max_pages | integer (≥ 1, ≤ 200) | Pages to read from the start of the document. Default: 50. |
Output
Returns data with: page_count, pages_extracted, truncated, pages: [{page, text}].
Request
curl -X POST https://agentops.tools/v1/pdf.extract \
-H 'content-type: application/json' \
-d '{"url":"https://arxiv.org/pdf/1706.03762","max_pages":2}'Over MCP the tool is named pdf_extract. Over A2A its skill id is pdf.extract. Same input, same envelope.