

pandas.read_html() reads every HTML table on a page or in a string and returns them as a list of DataFrames. Pass a URL, file path, or HTML string, then select the table you want by its list index. It handles most static tables in one line, and pairs with the requests library plus a standard browser User-Agent header when a server returns a 403 response or reports that no tables were found.
Key takeaways
Install pandas and a parser (pip install pandas lxml), then call pandas.read_html() with a URL, file path, or HTML string. It returns a list of DataFrames, one per table found on the page. Select the table you need by list index, for example, tables[0]. One line reads the tables; from there, you work with standard pandas DataFrames.
bash pip install pandas lxml
Here's a first read, using the Wikipedia list of the largest companies by revenue:
python import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue" tables = pd.read_html( url, storage_options={"User-Agent": "table-reader/1.0 ([email protected])"}, ) print(f"Tables found: {len(tables)}") df = tables[0] df.head()
The core of it is pd.read_html(url). The storage_options line is there because Wikipedia's User-Agent policy asks scripts to send a descriptive User-Agent with contact details. Plenty of other sites work with the bare call, and the troubleshooting section below explains the difference.
The return value is always a list, even when the page holds a single table. len(tables) tells you how many pandas DataFrames were found. Every item is a regular DataFrame, so everything you already know about pandas applies.
One thing to keep in mind: read_html() parses the HTML the server sends, so it works on static, server-rendered tables like most reference pages and reports. If a table only appears after JavaScript runs in the browser, pandas won't see it. The full parameter list lives in the pandas read_html() documentation, and this walkthrough from Proxy-Cheap sticks to the options you'll actually use.
Busy pages can contain many tables, so guessing the index is fragile. The match parameter keeps tables containing specific text and supports regex. The attrs parameter targets HTML attributes such as id or class.
tables = pd.read_html(url, match="Revenue") df = tables[0] # Or target an attribute: tables = pd.read_html(url, attrs={"id": "constituents"})
For example, the Wikipedia S&P 500 table uses id="constituents". Inspect your target with browser developer tools to find the right attributes.
Headers also need checking. pandas treats <th> cells as headers, but tables using <td> may produce columns named 0, 1, 2. Use header=0 to promote the first row and index_col=0 to use the first column as row labels.
If a title row appears above the header, use skiprows. Since pandas applies it before header, skiprows=1 and header=0 skip the title and use the next row. This is common in market research data collection, where table formats vary between sources.
read_html() sends a plain request, so some servers return 403 or HTML without the tables a browser sees, causing ValueError: No tables found. A reliable fix is to fetch the page with requests and a standard User-Agent, then pass response.text to read_html() via StringIO. For large-scale or geo-specific public table collection, proxies can provide more reliable access.
When given a URL, pandas uses urllib, which identifies itself as Python-urllib/3.x. Some servers reject this generic value. A 403 appears as HTTPError, while missing table markup produces ValueError.
Pandas documents the issue, and storage_options was added in pandas 2.1 (see the pandas 2.1.0 release notes). For more control, use this pattern:
from io import StringIO import requests import pandas as pd url = "https://example.com/report" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) " "AppleWebKit/537.36 (KHTML, like Gecko) " "Chrome/124.0 Safari/537.36" } resp = requests.get(url, headers=headers, timeout=30) resp.raise_for_status() tables = pd.read_html(StringIO(resp.text)) print(f"Tables found: {len(tables)}")
raise_for_status() catches HTTP errors before parsing. Use StringIO because current pandas treats a bare string as a file path. Follow any site-specific User-Agent guidance.
| Symptom | Likely cause | Fix |
|---|---|---|
| ValueError: No tables found | Missing User-Agent or no server-rendered table | Use requests + User-Agent + StringIO |
| HTTP Error 403 | Server rejects the request | Send a standard User-Agent; use proxies at scale |
| No tables matching pattern | match text doesn't appear | Check spelling, case, or regex |
| FileNotFoundError | HTML passed as a file path | Wrap it in StringIO |
| Stops during a large job | Too many requests from one IP | Add pacing and rotate proxies |
| Region-specific data missing | Page varies by market | Use an IP in the target market |
For larger jobs, such as public data collection, rotating residential proxies can spread requests and improve consistency. For fast, uniform sources, datacenter proxies provide higher throughput. Geo-specific pages can use rotating residential proxies in the target market.
proxy = "http://USERNAME:PASSWORD@PROXY_HOST:PORT" resp = requests.get( url, headers=headers, proxies={"http": proxy, "https": proxy}, timeout=30, )Cleaning and shaping the parsed table
Tables built for human readers rarely arrive ready for analysis, with stacked headers and numbers stored as text. Most fixes are one-liners.
On the Wikipedia revenue table, "Revenue" and "Profit" sit above a shared "USD (in billions)" row, so pandas builds a two-level MultiIndex and names the column ("Revenue", "USD (in billions)"). You can ask for that shape explicitly with header=[0, 1]. To flatten it, drop the level you don't need:
python df = tables[0] df.columns = df.columns.droplevel(1) # keep the top header row df["Revenue"] = (df["Revenue"].astype(str) .str.replace(r"[^\d.]", "", regex=True) .astype(float)) df.to_csv("companies.csv", index=False)
Which level to drop depends on the table. Here the useful names live in level 0, so level 1 goes, but other tables work the other way round. Print df.columns first.
For numbers, thousands="," is the default, and decimal="." sets the decimal point. European sources often flip both, and thousands=".", decimal="," reads "1.234,50" as 1234.5. Use converters={"CIK": str} to keep leading zeros in IDs.
Dates often carry notes like "(est.)". Strip them with df["Date"].str.replace(r"\(.*\)", "", regex=True) and convert with pd.to_datetime(). You'll repeat this cleanup constantly in price and catalog monitoring, where the same table gets pulled every day.
The flavor argument picks the engine that parses the HTML. By default, pandas tries lxml and falls back to BeautifulSoup with html5lib if lxml fails. You can also choose one yourself:
python tables = pd.read_html(url, flavor="bs4") # tolerant parser for messy markup
Stick with lxml for clean pages and switch to bs4 when a table comes out wrong, for example with shifted columns or cells under the wrong header. pandas understands colspan and rowspan, but broken markup around those attributes can trip a strict parser. html5lib repairs the markup the way a browser would, which usually fixes the shape.
Speed starts to matter at volume: if you're parsing thousands of static documentation tables through datacenter IPv4 proxies, lxml keeps the parsing step quick.
This one trips up a lot of people. From pandas 2.1 through the 2.x releases, passing a literal HTML string to read_html() raises a FutureWarning. In pandas 3.0 the change took full effect: a bare string is now treated as a file path, so you get a FileNotFoundError that prints your HTML back at you.
Wrap the string in io.StringIO instead:
python from io import StringIO import pandas as pd html = "<table><tr><th>A</th></tr><tr><td>1</td></tr></table>" tables = pd.read_html(StringIO(html))
Use io.BytesIO for raw bytes, while URLs and file paths work as before.
If you've searched "read_html deprecated," here's the short answer: the function isn't. Only the literal-string input form changed. The rule covers HTML from any source, whether that's resp.text, a cached page, or a service you use to automate data collection with an API. Not sure which rules apply to you? Run pd.__version__.
pandas.read_html() parses HTML <table> elements into DataFrames and handles merged cells automatically. It’s the simplest option when your data is in static, server-rendered tables.
For more complex jobs:
At scale, choose proxies based on the workload. Rotating residential proxies suit broad collection across sites and regions. Static residential proxies, also known as ISP proxies, keep one IP for consistent sessions. Datacenter proxies suit high-throughput public content. See proxy type differences.
For multi-page or geo-specific public data, Proxy-Cheap rotating residential proxies provide reliable access with pay-as-you-go pricing and no monthly commitment.
| cells as headers automatically. If the header row uses plain | cells, set header=0, or header=[0, 1] for two header rows. Flatten multi-level headers later with df.columns.droplevel(). read_html() parses HTML tables into DataFrames. to_html() does the reverse: it renders a DataFrame as an HTML table. |
|---|