Free tool
HTML table extractor
Turn every table on a page into JSON. Header cells become keys, so each row arrives as a record you can load straight into a dataframe or a database.
HTML input
JSON output
0 B · 0 linesExtracted locally in your browser — nothing is uploaded.
How it works
What this tool does
- Each table returns its caption, header row, raw rows and — when headers exist — an array of keyed records.
- All tables on the page are extracted at once, in document order.
- Cell text is whitespace-normalised, so line breaks and indentation in the markup do not leak into your data.
- Works on the messy tables real sites ship: no thead, th scattered through the body, nested markup inside cells.
With Irmu
When you also need to fetch the page
A browser cannot fetch arbitrary URLs. Irmu returns the HTML — rendering, proxies and retries included — and you run the same extraction over it.
crawl_and_extract.py
# 1. Fetch the page with Irmu — rendering, proxies and retries handled for you
# 2. Extract the fields you need locally
import os, json, requests
from bs4 import BeautifulSoup
resp = requests.get(
"https://app.irmu.com/api/crawl",
headers={"Authorization": f"Bearer {os.environ['IRMU_API_KEY']}"},
params={"url": "https://store.example.com/products/acme-headphones", "render": "true"},
timeout=60,
)
resp.raise_for_status()
soup = BeautifulSoup(resp.text, "html.parser")
data = {
"name": soup.select_one("h1").get_text(strip=True),
"price": soup.select_one(".price").get_text(strip=True),
"reviews": [r.get_text(strip=True) for r in soup.select(".review")],
}
print(json.dumps(data, indent=2))Code examples
Do the same thing in your own pipeline
The equivalent extraction with the standard HTML parser in each language.
# pip install beautifulsoup4
import json
from bs4 import BeautifulSoup
soup = BeautifulSoup(html, "html.parser")
tables = []
for table in soup.find_all("table"):
rows = [[c.get_text(strip=True) for c in tr.find_all(["th", "td"])] for tr in table.find_all("tr")]
headers = rows.pop(0) if rows and table.find("th") else []
tables.append({
"headers": headers,
"rows": rows,
"records": [dict(zip(headers, r)) for r in rows] if headers else None,
})
data = {"tables": tables}
# Or, for data analysis: pandas.read_html(html) returns a DataFrame per table.
print(json.dumps(data, indent=2))FAQ
Questions about this tool
Turn any URL into clean JSON
Fetch, render and extract with one API. 1,000 free credits every month, no card required.