apyhub
Back
▣ DATA EXTRACTION · SEO

Extract Metadata From URL API

What it does

The URL Metadata API reads an article or blog post and returns its publishing details as JSON. Send one page as url to the GET endpoint, or an array of urls to the POST batch endpoint, and get back title, author, publicationName, publicationDate, platform, wordCount, and readingTime for each page. Some platforms, such as Medium, also return a coverImage.

It works best on article pages such as blog posts and news stories. Known platforms like Medium are named in platform, and other sites come back as Generic. Dates follow the source page's format, so normalize publicationDate before you sort or store it. When a page can't be read, the response falls back to defaults such as Untitled Article, Unknown Author, and a wordCount of 0, so treat those as a miss.

Use the URL metadata API to build reading lists and content catalogs with author and date, show reading time next to saved links, index a blog archive, or audit your own posts for missing authors and dates during an SEO review. The batch endpoint enriches a whole list of URLs in one request.

For the full article text, use the Article Content API. To collect every URL on a site first, run the Sitemap Extraction API. For preview cards with images, descriptions, and favicons, the Link Preview API returns Open Graph data.

▣ ENDPOINT 01 / 02
GET
Extract Metadata From a URL
https://api.eu.apyhub.com/namastesumalya/extract-metadata-from-url-api

QUICKSTART

GUIDE

Quickstart

Fetch metadata for a webpage by passing its url as a query parameter.

curl -X GET "https://api.eu.apyhub.com/namastesumalya/extract-metadata-from-url-api?url=https://apyhub.com" \
  -H "apy-token: $APY_TOKEN"

What you'll get back

Returns a JSON object with these top-level string fields: url, title, author, platform, reading_time, publication_date, and publication_name.

{
  "url": "https://apyhub.com/article",
  "title": "Understanding AI in Modern Applications",
  "author": "John Doe",
  "platform": "Medium",
  "reading_time": "4 min",
  "publication_date": "2025-02-10T00:00:00Z",
  "publication_name": "Medium"
}
TRY ITLIVE · 300 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.

About this endpoint

What it does

Extracts metadata from the URL you provide and returns the discovered page/article fields in a JSON object.

Query Parameter(s)

AttributeTypeDescription
urlStringThe URL to extract metadata from.

Response

Returns a JSON object containing metadata fields about the supplied URL. The response wrapper includes the url string plus optional string fields for title, author, platform, reading_time, publication_date, and publication_name.

ParameterTypeDescription
urlStringThe URL that metadata was extracted from.
titleStringThe page or article title.
authorStringThe content author.
platformStringThe publishing platform.
reading_timeStringThe estimated reading time.
publication_dateStringThe publication date/time as a string.
publication_nameStringThe name of the publication.
▣ ENDPOINT 02 / 02
POST
Batch Extract Metadata From Multiple URLs
https://api.eu.apyhub.com/namastesumalya/extract-metadata-from-url-api

QUICKSTART

GUIDE

Quickstart

Send a list of URLs to extract metadata from multiple pages in one request.

curl -X POST "https://api.eu.apyhub.com/namastesumalya/extract-metadata-from-url-api" \
  -H "apy-token: $APY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://apyhub.com","https://apify.com"]}'

What you'll get back

Returns a JSON object with a data array. Each item is an object that may include url, title, author, platform, reading_time, publication_date, and publication_name.

{
  "data": [
    {
      "url": "https://example.com/article",
      "title": "Understanding AI in Modern Applications",
      "author": "John Doe",
      "platform": "Medium",
      "reading_time": "4 min",
      "publication_date": "2025-02-10T00:00:00Z",
      "publication_name": "Medium"
    }
  ]
}
TRY ITLIVE · 500 ATOMS
Loading your default key…
The full key is used to call the gateway and stays in this tab — never sent to orbit or saved.
body*
urls

About this endpoint

What it does

Extracts metadata for multiple URLs in a single request and returns an array of metadata objects for the submitted URLs.

Request Body

ParameterTypeDescription
urlsString ArrayList of URLs to extract metadata from.

Response

Returns a JSON object with a data array field. Each item in data is an object containing metadata for one URL, with fields for the source URL and extracted page details.

ParameterTypeDescription
dataObject ArrayArray of metadata objects.
data[].urlStringThe URL that was processed.
data[].titleStringThe extracted page title.
data[].authorStringThe extracted author name.
data[].platformStringThe platform or site name.
data[].reading_timeStringThe extracted reading time.
data[].publication_dateStringThe publication date.
data[].publication_nameStringThe publication name.
▣ COMMON ERRORS

Errors any endpoint can return

400bad_request

Required parameter missing or malformed body.

401unauthorized

API key missing, revoked, or not authorized for this service.

429rate_limited

Your plan's per-second rate exceeded. Retry with exponential backoff.

503upstream_busy

Backend temporarily unavailable. Try again in a few seconds.