apyhub
Back
▣ DATA EXTRACTION · FILE CONVERSION

Image & PDF OCR to Markdown API

What it does

Image & PDF OCR to Markdown reads a single image or a single PDF page and returns its content as Markdown, plain text, or structured JSON. Send a photo, scan, or screenshot by public URL or base64 to the image endpoint, or a scanned or digital PDF by public URL or base64 to the document endpoint along with the page you want, and get clean, layout-aware output back in 100+ languages.

Use it when you need accurate text and structure from receipts, invoices, forms, reports, research papers, or screenshots. The JSON response includes markdown, text, blocks, tables, and page, so you get the reading order, page dimensions, typed blocks with bounding boxes, and every detected table as HTML, Markdown, and cell rows without building your own OCR pipeline. Formulas come back as LaTeX, and bar, line, and pie charts can be turned into data tables.

For PDFs, you choose the page to read (1-based), set dpi between 72 and 300 for rendering resolution, and use page.count from the response to walk through the rest of the document one page at a time. Both endpoints let you switch mode between auto, layout, and plain, pick a format of json, markdown, or text, turn block output on or off with include_blocks, and disable chart-to-table conversion with charts.

Limits

Some limits apply to both endpoints, and some apply to only one.

InputLimitApplies to
PDFsUp to 20 MiB and 1,000 pages; password-protected PDFs are not supported; one page per request.Read a PDF Page
PDF resolutiondpi 72–300 (default 200), at most 25 megapixels per rendered page. If the requested dpi would exceed 25 megapixels, the service automatically uses a lower dpi instead of failing. Always read the actual value from page.dpi in the response.Read a PDF Page
ImagesUp to 20 MiB and 40 megapixels; PNG, JPEG, WebP, GIF, BMP, or TIFF; only the first frame of multi-frame images is read.Read an Image
URLsPublic http(s) only, up to 2,048 characters and 3 redirects; private and internal addresses are refused.Both endpoints
Base64Up to 20 MB decoded. Base64 encoding adds about one third to the request size.Both endpoints
Request timeUp to 55 seconds; very complex inputs may time out.Both endpoints
▣ ENDPOINT 01 / 02
POST
Read a PDF page
https://api.eu.apyhub.com/callable-labs/ocr-pdf-to-markdown/v1/ocr/document

QUICKSTART

GUIDE

Quickstart

Read page 1 of a PDF from a public URL.

curl -X POST "https://api.eu.apyhub.com/callable-labs/ocr-pdf-to-markdown/v1/ocr/document" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "document": {
      "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf"
    },
    "page": 1,
    "dpi": 200,
    "format": "json",
    "mode": "auto",
    "include_blocks": true,
    "charts": true
  }'

Expect JSON with recognized markdown, text, blocks, tables, and page metadata.

What you'll get back

You will receive the page's content as Markdown and plain text, typed blocks with bounding boxes in reading order, any tables found, page information including the total page count and the resolution used, the reading mode used, model, cache, and timing information.

{
  "markdown": "Dummy PDF file\n",
  "text": "Dummy PDF file",
  "blocks": [
    {
      "type": "text",
      "label": "text",
      "order": 0,
      "bbox": [153, 193, 506, 246],
      "content": "Dummy PDF file",
      "html": null,
      "markdown": null,
      "latex": null
    }
  ],
  "tables": [],
  "page": {
    "width": 1653,
    "height": 2339,
    "index": 1,
    "count": 1,
    "dpi": 200
  },
  "mode": "layout",
  "coverage": null,
  "model": {
    "layout": "PP-DocLayoutV3@241f8bd",
    "vlm": "PaddleOCR-VL-1.6@c5630ab"
  },
  "cache": "hit",
  "timing": {
    "fetch_ms": 320,
    "decode_ms": 0,
    "queued_ms": 0,
    "ocr_ms": 0
  }
}
TRY ITLIVE · 400 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
DocumentRequest*
The PDF to read. Provide exactly one of `document.url` or `document.base64`.
PDF URL*
Public http(s) URL of the PDF, up to 2,048 characters. The file must be 20 MiB or smaller. Private and internal addresses are refused.
Resolution used to render the page. Range: `72-300`. Default: `200`. The rendered page is capped at 25 megapixels; if the requested `dpi` would exceed this, a lower `dpi` is used automatically.
Reading mode: `auto` for varied documents, `layout` for structured blocks, or `plain` for a single text block. Default: `auto`.
The 1-based page number to read. Range: `1-1000`. Default: `1`.

About this endpoint

What it does

Reads one page of a PDF, scanned or digital, and returns it as Markdown, plain text, typed blocks with bounding boxes, and tables as HTML, Markdown, and cell rows, all in reading order.

Every response includes the total number of pages (page.count and the X-Page-Count header), so you can read a whole document by requesting page: 1, 2, 3, and so on. Each page is a separate request, and the document must be sent with each request.

Request Body Parameter(s)

AttributeTypeMandatoryDescription
documentObjectYesThe PDF to read. Provide exactly one of document.url or document.base64.
document.urlStringNo*Public http(s) URL of the PDF, up to 2,048 characters. The file must be 20 MiB or smaller. Private and internal addresses are refused.
document.base64StringNo*The PDF as base64 (a data: URI is also accepted), 20 MB or smaller once decoded.
pageIntegerNoThe 1-based page number to read. Range: 1-1000. Default: 1.
dpiIntegerNoResolution used to render the page. Range: 72-300. Default: 200. The rendered page is capped at 25 megapixels; if the requested dpi would exceed this, a lower dpi is used automatically.
formatStringNoResponse format: json (full result), markdown (page as text/markdown), or text (page as text/plain). Default: json.
modeStringNoReading mode: auto for varied documents, layout for structured blocks, or plain for a single text block. Default: auto.
include_blocksBooleanNoWhether to include blocks in the JSON response. Default: true.
chartsBooleanNoWhether to convert bar, line, and pie charts into data tables. Default: true.

* Exactly one of document.url or document.base64 must be provided. Sending both, or neither, returns an error.

Supported documents: PDFs up to 20 MiB and 1,000 pages. Password-protected PDFs are not supported.

Response

When format is json, the response contains all fields listed in Common Response Fields. For PDFs:

When format is markdown or text, the response is the page as text/markdown or text/plain, and the total page count is still available in the X-Page-Count header.

▣ ENDPOINT 02 / 02
POST
Read an image (OCR to Markdown)
https://api.eu.apyhub.com/callable-labs/ocr-pdf-to-markdown/v1/ocr/image

QUICKSTART

GUIDE

Quickstart

Set APY_TOKEN to your ApyHub API key. Replace the image URL with a public image you are authorized to process.

curl -X POST "https://api.eu.apyhub.com/callable-labs/ocr-pdf-to-markdown/v1/ocr/image" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "image": {
      "url": "https://upload.wikimedia.org/wikipedia/commons/0/0b/ReceiptSwiss.jpg"
    },
    "format": "json",
    "mode": "auto",
    "include_blocks": true
  }'

Expect JSON with markdown, text, blocks, tables, and page. Review receipt amounts manually. For a local file, encode it with Python base64 and supply image.base64 instead of image.url.

TRY ITLIVE · 400 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
ImageRequest*
The image to read. Provide exactly one of `image.url` or `image.base64`.
Image URL*
Public http(s) URL of the image, up to 2,048 characters. The file must be 20 MB or smaller. Private and internal addresses are refused.
Reading mode: `auto` for varied documents, `layout` for structured blocks, or `plain` for a single text block. Default: `auto`.
Whether to convert bar, line, and pie charts into data tables. Default: `true`.
Response format: `json` (full result), `markdown` (page as `text/markdown`), or `text` (page as `text/plain`). Default: `json`.

About this endpoint

What it does

Reads one photo, scan, or screenshot and returns its content as Markdown, plain text, typed blocks with bounding boxes, and tables as HTML, Markdown, and cell rows, all in reading order.

Use this for single-page inputs such as receipts, invoices, forms, whiteboards, screenshots, and photographed documents. Send the image as a public URL or as base64-encoded content.

Request Body Parameter(s)

AttributeTypeMandatoryDescription
imageObjectYesThe image to read. Provide exactly one of image.url or image.base64.
image.urlStringNo*Public http(s) URL of the image, up to 2,048 characters. The file must be 20 MB or smaller. Private and internal addresses are refused.
image.base64StringNo*The image as base64 (a data: URI is also accepted), 20 MB or smaller once decoded.
formatStringNoResponse format: json (full result), markdown (page as text/markdown), or text (page as text/plain). Default: json.
modeStringNoReading mode: auto for varied documents, layout for structured blocks, or plain for a single text block. Default: auto.
include_blocksBooleanNoWhether to include blocks in the JSON response. Default: true.
chartsBooleanNoWhether to convert bar, line, and pie charts into data tables. Default: true.

* Exactly one of image.url or image.base64 must be provided. Sending both, or neither, returns an error.

Supported image formats: PNG, JPEG, WebP, GIF, BMP, and TIFF. Images can be up to 40 megapixels. For animated or multi-frame images, only the first frame is read.

Response

When format is json, the response contains all fields listed in Common Response Fields. For images, page.index, page.count, and page.dpi are null.

When format is markdown or text, the response is the page as text/markdown or text/plain.

▣ COMMON ERRORS

Errors any endpoint can return

400bad_request

Required parameter missing or malformed body.

401unauthorized

API key missing, revoked, or not authorized for this service.

429rate_limited

Your plan's per-second rate exceeded. Retry with exponential backoff.

503upstream_busy

Backend temporarily unavailable. Try again in a few seconds.