apyhub
Cover illustration for How to Scrape a Website with an API: A Step-by-Step Guide for Beginners
Engineering · Python

How to Scrape a Website with an API: A Step-by-Step Guide for Beginners

How to Scrape a Website with an API: A Step-by-Step Guide for Beginners

01Introduction

Web scraping means using a program to collect information from websites, such as the text of an article, a list of links or a product table. The easiest way to start is with a scraping API: you send a page address, and the API sends back the content, already cleaned up. You skip writing HTML parsers, and you skip running a headless browser.

Here is the whole idea in five lines of Python:

python

· python
import requests

response = requests.get(
    "https://api.eu.apyhub.com/apyhub/extract-visible-text-from-a-webpage",
    params={"url": "https://example.com"},
    headers={"apy-token": "YOUR_API_KEY"},
)
print(response.json()["data"])

That prints the readable text of the page. The rest of this guide builds on it step by step: getting text, collecting links, pulling structured data for AI tools, and scraping a small site. Every step uses an API from the ApyHub catalog.

If you want the background first (how scraping works, what changed in 2026 and what's legal), read our companion guide: Web Scraping in 2026: What Changed, What's Legal, and What to Use Instead.

02What you need

  • An ApyHub API key. Sign up for free and copy your key from the dashboard. The free plan needs no credit card.
  • Python 3 with the requests library. Install it with pip install requests.

No Python? You can follow along without code by running each request in Voiden, the open-source API client. More on that below.

03Step 1: Get the text from a page

The Extract Text from Webpage API returns the text a visitor sees on a page, with the HTML, scripts and styling removed. Turn on paragraph mode to keep the text readable:

python

· python
import requests

API_KEY = "YOUR_API_KEY"
BASE = "https://api.eu.apyhub.com/apyhub"

def get_text(url):
    response = requests.get(
        f"{BASE}/extract-visible-text-from-a-webpage",
        params={"url": url, "preserve_paragraphs": "true"},
        headers={"apy-token": API_KEY},
    )
    return response.json()["data"]

print(get_text("https://example.com"))

Use this when you need the words on a page: to search them, count keywords or feed them into another tool.

04Step 2: Collect the links

To go beyond one page, you need to know where a page links to. The Extract Links from Webpage API returns every link on a page as a full address:

python

· python
def get_links(url):
    response = requests.post(
        f"{BASE}/extract-links-from-webpage",
        json={"url": url},
        headers={"apy-token": API_KEY},
    )
    return response.json()["data"].get("links", [])

links = get_links("https://example.com")
print(f"Found {len(links)} links")

This API also has a safety check built in. Before fetching a page, it checks the address against a database of known malicious sites. If the page is flagged, you get a warning back instead of the links, so your script doesn't visit it.

05Step 3: Get structured data for AI tools

Plain text works for simple jobs. For AI tools, you usually want the page broken into parts: the title, headings, sections, tables and a clean Markdown version of the content.

The AI-ready Clean Data Extractor API does this in one call:

python

· python
def get_structured(url):
    response = requests.post(
        f"{BASE}/webpage-extractor-api",
        json={"url": url},
        headers={"apy-token": API_KEY},
    )
    return response.json()["data"]

page = get_structured("https://example.com")
print(page["title"])
print(page["headings"])

Along with the title and headings, the response includes the page type (article, product page, documentation and so on), any tables as rows and columns, and a clean Markdown version you can pass straight to a language model.

06Step 4: Scrape a small site

Now combine the steps. This script starts on one page, follows the links that stay on the same site, and saves the text of each page. It stops after 10 pages and waits between requests, so it stays polite to the site you're scraping.

python

· python
import time
from urllib.parse import urlparse

start = "https://example.com"
site = urlparse(start).netloc

to_visit = [start]
seen = set()
results = {}

while to_visit and len(seen) < 10:
    url = to_visit.pop(0)
    if url in seen:
        continue
    seen.add(url)

    results[url] = get_text(url)

    for link in get_links(url):
        if urlparse(link).netloc == site and link not in seen:
            to_visit.append(link)

    time.sleep(1)

print(f"Scraped {len(results)} pages")

That's a working scraper in about 25 lines, and you never had to deal with raw HTML.

07Scrape responsibly

A few habits keep your scraping on the right side of site owners and the law:

  • Check the site's robots.txt and terms. Many sites say what they allow automated tools to access.
  • Go slowly. Keep a pause between requests and limit how many pages you fetch.
  • Stick to public pages. Don't scrape content behind a login.
  • Be careful with personal data. Names, emails and other personal details come with privacy rules, such as GDPR in Europe.
  • Look for an official API first. If a site offers one, it's usually the more reliable and more welcome option.

For the legal detail, see Web Scraping in 2026.

08Test without writing code

You can try every request in this guide before writing a script. Open Voiden, the open-source API client, create a request with the API address and your apy-token header, and run it. It's a quick way to see what each API returns and decide which one fits your project.

You can also test each API directly from its page in the ApyHub catalog.

09Scraping with AI agents

Every API in this guide is also available through ApyHub MCP. An AI agent can discover, evaluate and call them directly, without a hand-written wrapper or tool definition. Ask your agent to "summarize the pricing pages of these three competitors", and it can fetch each page with the Clean Data Extractor and work from the structured result.

Explore the data extraction APIs →

10Conclusion

Scraping with an API turns a job that used to mean HTML parsing and browser automation into a few simple requests. Start with the Extract Text from Webpage API to get readable text, add the Extract Links from Webpage API to move across pages, and use the AI-ready Clean Data Extractor when you need structured content for AI tools.

Scrape slowly, respect each site's rules, and you have a scraper you can build on.

11FAQ

What is web scraping? Web scraping is using a program to collect information from websites automatically, such as text, links, prices or tables. Our 2026 guide covers how it works in more depth.

Is web scraping legal? Scraping public pages is often allowed, but it depends on the site's terms, the country and the type of data. Personal data and content behind a login carry more risk. See the legal section of our 2026 guide for details.

Do I need to know Python to scrape a website? No. The examples use Python because it's beginner-friendly, but the APIs work from any language. You can also run each request in an API client like Voiden without writing code.

How can I test these APIs before building a scraper? Run them from their pages in the ApyHub catalog, or open the requests in Voiden, the open-source API client, and run them with your API key. Both work on the free plan.

Why use a scraping API instead of writing my own scraper? An API handles fetching the page, stripping the HTML and returning clean data, so you don't maintain parsers that break every time a site changes its layout. You write the logic that matters to your project.

Which API should I use for AI projects? The AI-ready Clean Data Extractor. It returns the page as structured data and clean Markdown, which is easier for language models to work with than raw text.

Can my AI agent scrape websites with these APIs? Yes. Every endpoint is available through ApyHub MCP, so an agent can fetch and read pages as part of its work.

What if a page comes back empty? Some pages load their content with JavaScript after the page opens. If the text you expect is missing, check the result with the Clean Data Extractor, or look for an official API or data feed from that site.

12About ApyHub

ApyHub is a curated API catalog and trusted operational layer for developers and AI agents. The catalog offers over 1,500 endpoints and capabilities, and it keeps growing. Every API is verified before listing and carries machine-readable certification for GDPR, SOC 2 and ISO 27001.

One subscription covers the whole catalog. Usage is measured in atoms, a per-call unit that reflects the compute work each request does. Every endpoint is MCP-ready by default, so AI agents can call it through ApyHub MCP without custom wrappers.

ApyHub is headquartered in Amsterdam, with offices in the Netherlands, Greece and India, and serves 65,000+ monthly developer workspaces. The free plan needs no credit card. API providers can publish their APIs to the catalog through the provider program.

Start free with ApyHub →