Proxy-Cheap
Proxies & Business
October 9, 2026
7 min

How to read HTML tables with pandas read_html()

Alex Sadovskij
Alex Sadovskij
CEO Proxy-Cheap
How to read HTML tables with pandas read_html()
Summary
Read HTML tables with pandas read_html(), select the right table, fix "No tables found" and 403 responses, and collect public table data reliably at scale.

pandas.read_html() reads every HTML table on a page or in a string and returns them as a list of DataFrames. Pass a URL, file path, or HTML string, then select the table you want by its list index. It handles most static tables in one line, and pairs with the requests library plus a standard browser User-Agent header when a server returns a 403 response or reports that no tables were found.

Key takeaways

  • read_html() returns a list of DataFrames, one per table. Index the list (for example, tables[0]) to get the one you need.
  • Use the match and attrs parameters to target a specific table by its text or its HTML id/class.
  • If a server returns a 403 response or "No tables found," fetch the page with requests and a standard User-Agent header, then pass the HTML text to read_html() wrapped in StringIO.
  • For large-scale, geo-specific collection of publicly available tables, route those requests through residential or datacenter proxies for reliable access.

How to read HTML tables with pandas read_html()

Install pandas and a parser (pip install pandas lxml), then call pandas.read_html() with a URL, file path, or HTML string. It returns a list of DataFrames, one per table found on the page. Select the table you need by list index, for example, tables[0]. One line reads the tables; from there, you work with standard pandas DataFrames.

bash pip install pandas lxml

Here's a first read, using the Wikipedia list of the largest companies by revenue:

python import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue" tables = pd.read_html(    url,    storage_options={"User-Agent": "table-reader/1.0 ([email protected])"}, ) print(f"Tables found: {len(tables)}") df = tables[0] df.head()

The core of it is pd.read_html(url). The storage_options line is there because Wikipedia's User-Agent policy asks scripts to send a descriptive User-Agent with contact details. Plenty of other sites work with the bare call, and the troubleshooting section below explains the difference.

The return value is always a list, even when the page holds a single table. len(tables) tells you how many pandas DataFrames were found. Every item is a regular DataFrame, so everything you already know about pandas applies.

One thing to keep in mind: read_html() parses the HTML the server sends, so it works on static, server-rendered tables like most reference pages and reports. If a table only appears after JavaScript runs in the browser, pandas won't see it. The full parameter list lives in the pandas read_html() documentation, and this walkthrough from Proxy-Cheap sticks to the options you'll actually use.

Reading a specific table with match, attrs, and header

Busy pages can contain many tables, so guessing the index is fragile. The match parameter keeps tables containing specific text and supports regex. The attrs parameter targets HTML attributes such as id or class.

tables = pd.read_html(url, match="Revenue") df = tables[0] # Or target an attribute: tables = pd.read_html(url, attrs={"id": "constituents"})

For example, the Wikipedia S&P 500 table uses id="constituents". Inspect your target with browser developer tools to find the right attributes.

Headers also need checking. pandas treats <th> cells as headers, but tables using <td> may produce columns named 0, 1, 2. Use header=0 to promote the first row and index_col=0 to use the first column as row labels.

If a title row appears above the header, use skiprows. Since pandas applies it before header, skiprows=1 and header=0 skip the title and use the next row. This is common in market research data collection, where table formats vary between sources.

Fixing "No tables found" and 403 responses

read_html() sends a plain request, so some servers return 403 or HTML without the tables a browser sees, causing ValueError: No tables found. A reliable fix is to fetch the page with requests and a standard User-Agent, then pass response.text to read_html() via StringIO. For large-scale or geo-specific public table collection, proxies can provide more reliable access.

When given a URL, pandas uses urllib, which identifies itself as Python-urllib/3.x. Some servers reject this generic value. A 403 appears as HTTPError, while missing table markup produces ValueError.

Pandas documents the issue, and storage_options was added in pandas 2.1 (see the pandas 2.1.0 release notes). For more control, use this pattern:

from io import StringIO import requests import pandas as pd url = "https://example.com/report" headers = {    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) "                  "AppleWebKit/537.36 (KHTML, like Gecko) "                  "Chrome/124.0 Safari/537.36" } resp = requests.get(url, headers=headers, timeout=30) resp.raise_for_status() tables = pd.read_html(StringIO(resp.text)) print(f"Tables found: {len(tables)}")

raise_for_status() catches HTTP errors before parsing. Use StringIO because current pandas treats a bare string as a file path. Follow any site-specific User-Agent guidance.

SymptomLikely causeFix
ValueError: No tables foundMissing User-Agent or no server-rendered tableUse requests + User-Agent + StringIO
HTTP Error 403Server rejects the requestSend a standard User-Agent; use proxies at scale
No tables matching patternmatch text doesn't appearCheck spelling, case, or regex
FileNotFoundErrorHTML passed as a file pathWrap it in StringIO
Stops during a large jobToo many requests from one IPAdd pacing and rotate proxies
Region-specific data missingPage varies by marketUse an IP in the target market

For larger jobs, such as public data collection, rotating residential proxies can spread requests and improve consistency. For fast, uniform sources, datacenter proxies provide higher throughput. Geo-specific pages can use rotating residential proxies in the target market.

proxy = "http://USERNAME:PASSWORD@PROXY_HOST:PORT" resp = requests.get(    url,    headers=headers,    proxies={"http": proxy, "https": proxy},    timeout=30, )Cleaning and shaping the parsed table

Tables built for human readers rarely arrive ready for analysis, with stacked headers and numbers stored as text. Most fixes are one-liners.

On the Wikipedia revenue table, "Revenue" and "Profit" sit above a shared "USD (in billions)" row, so pandas builds a two-level MultiIndex and names the column ("Revenue", "USD (in billions)"). You can ask for that shape explicitly with header=[0, 1]. To flatten it, drop the level you don't need:

python df = tables[0] df.columns = df.columns.droplevel(1)          # keep the top header row df["Revenue"] = (df["Revenue"].astype(str)                 .str.replace(r"[^\d.]", "", regex=True)                 .astype(float)) df.to_csv("companies.csv", index=False)

Which level to drop depends on the table. Here the useful names live in level 0, so level 1 goes, but other tables work the other way round. Print df.columns first.

For numbers, thousands="," is the default, and decimal="." sets the decimal point. European sources often flip both, and thousands=".", decimal="," reads "1.234,50" as 1234.5. Use converters={"CIK": str} to keep leading zeros in IDs.

Dates often carry notes like "(est.)". Strip them with df["Date"].str.replace(r"\(.*\)", "", regex=True) and convert with pd.to_datetime(). You'll repeat this cleanup constantly in price and catalog monitoring, where the same table gets pulled every day.

Choosing a parser: lxml, html5lib, and bs4 (the flavor argument)

The flavor argument picks the engine that parses the HTML. By default, pandas tries lxml and falls back to BeautifulSoup with html5lib if lxml fails. You can also choose one yourself:

python tables = pd.read_html(url, flavor="bs4")   # tolerant parser for messy markup

  • flavor="lxml": fast and strict. Install with pip install lxml.
  • flavor="bs4" or flavor="html5lib": two names for the same engine. It parses the way a browser does, so it's slower but far more forgiving. Install with pip install beautifulsoup4 html5lib.

Stick with lxml for clean pages and switch to bs4 when a table comes out wrong, for example with shifted columns or cells under the wrong header. pandas understands colspan and rowspan, but broken markup around those attributes can trip a strict parser. html5lib repairs the markup the way a browser would, which usually fixes the shape.

Speed starts to matter at volume: if you're parsing thousands of static documentation tables through datacenter IPv4 proxies, lxml keeps the parsing step quick.

The literal-string change: wrap HTML in StringIO

This one trips up a lot of people. From pandas 2.1 through the 2.x releases, passing a literal HTML string to read_html() raises a FutureWarning. In pandas 3.0 the change took full effect: a bare string is now treated as a file path, so you get a FileNotFoundError that prints your HTML back at you.

Wrap the string in io.StringIO instead:

python from io import StringIO import pandas as pd html = "<table><tr><th>A</th></tr><tr><td>1</td></tr></table>" tables = pd.read_html(StringIO(html))

Use io.BytesIO for raw bytes, while URLs and file paths work as before.

If you've searched "read_html deprecated," here's the short answer: the function isn't. Only the literal-string input form changed. The rule covers HTML from any source, whether that's resp.text, a cached page, or a service you use to automate data collection with an API. Not sure which rules apply to you? Run pd.__version__.

What pandas read_html() is, and when to reach for a full scraper instead

pandas.read_html() parses HTML <table> elements into DataFrames and handles merged cells automatically. It’s the simplest option when your data is in static, server-rendered tables.

For more complex jobs:

  • Pagination: Loop through page URLs with requests and pass each response to read_html().
  • JavaScript tables: Render the page with a headless browser first.
  • Non-table content: Use requests with BeautifulSoup or a crawling framework.

At scale, choose proxies based on the workload. Rotating residential proxies suit broad collection across sites and regions. Static residential proxies, also known as ISP proxies, keep one IP for consistent sessions. Datacenter proxies suit high-throughput public content. See proxy type differences.

For multi-page or geo-specific public data, Proxy-Cheap rotating residential proxies provide reliable access with pay-as-you-go pricing and no monthly commitment.

Frequently Asked Questions

Yes, for static, server-rendered HTML tables, where it needs only one line. For JavaScript-rendered tables, pagination, or content outside tables, pair it with a full scraper.

The HTML pandas received has no elements. Often the server sent a trimmed page because the request lacked a standard User-Agent header, so fetch with requests and a header, then pass StringIO(resp.text). If the error continues, the table is probably built by JavaScript.
Use match="some text" to filter by content, or attrs={"id": "..."} to target a table by its HTML attribute. Combining both works best on busy pages.

Yes. Pass the file path directly, for example, pd.read_html("report.html"), or wrap an HTML string in io.StringIO.

No, it reads the HTML as served. If JavaScript builds the table in the browser, render the page first with a headless browser, then pass that HTML to read_html().

The server declined the request, and pandas reports it as HTTPError: HTTP Error 403: Forbidden. Send a standard User-Agent header; for large or geo-specific jobs, route requests through residential or datacenter proxies.

No, the function is current. Only the literal-string input form changed in pandas 2.1, and pandas 3.0 now reads bare strings as file paths. Wrap raw HTML in io.StringIO.

pandas uses
cells as headers automatically. If the header row uses plain cells, set header=0, or header=[0, 1] for two header rows. Flatten multi-level headers later with df.columns.droplevel().
read_html() parses HTML tables into DataFrames. to_html() does the reverse: it renders a DataFrame as an HTML table.