Irmu
Free tool

HTML heading extractor

The heading structure of a page as JSON, grouped by level — the fastest way to see how a document is actually organised.

HTML input
JSON output
0 B · 0 linesExtracted locally in your browser — nothing is uploaded.
How it works

What this tool does

  • Headings are grouped by tag, so you can see at a glance whether a page has one h1 or seven.
  • Empty headings are skipped and text is whitespace-normalised.
  • Useful for SEO audits, accessibility checks and building tables of contents.
  • Heading structure is also the best natural chunk boundary for RAG — AI-ready mode uses exactly this.
With Irmu

When you also need to fetch the page

A browser cannot fetch arbitrary URLs. Irmu returns the HTML — rendering, proxies and retries included — and you run the same extraction over it.

crawl_and_extract.py
# 1. Fetch the page with Irmu — rendering, proxies and retries handled for you
# 2. Extract the fields you need locally

import os, json, requests
from bs4 import BeautifulSoup

resp = requests.get(
    "https://app.irmu.com/api/crawl",
    headers={"Authorization": f"Bearer {os.environ['IRMU_API_KEY']}"},
    params={"url": "https://store.example.com/products/acme-headphones", "render": "true"},
    timeout=60,
)
resp.raise_for_status()

soup = BeautifulSoup(resp.text, "html.parser")

data = {
    "name": soup.select_one("h1").get_text(strip=True),
    "price": soup.select_one(".price").get_text(strip=True),
    "reviews": [r.get_text(strip=True) for r in soup.select(".review")],
}

print(json.dumps(data, indent=2))
Code examples

Do the same thing in your own pipeline

The equivalent extraction with the standard HTML parser in each language.

# pip install beautifulsoup4

import json
from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")

data = {"outline": [
    {"level": int(h.name[1]), "text": h.get_text(strip=True)}
    for h in soup.select("h1, h2, h3, h4, h5, h6")
]}

print(json.dumps(data, indent=2))
FAQ

Questions about this tool

Turn any URL into clean JSON

Fetch, render and extract with one API. 1,000 free credits every month, no card required.