

Key takeaways
A parse error in Python means the interpreter or a library could not read your input as valid. There are two families. Python raises SyntaxError or IndentationError when it cannot parse your source code before it runs. Libraries raise their own ParseError or ParserError when they cannot parse the data you gave them, such as a CSV, XML, or JSON payload.
Start by reading the exception name and where it was raised. If Python stops before running your program, you have a SyntaxError or IndentationError in your source. If your program starts and then fails while reading a file or a response, a library raised the error: pandas, xml.etree, json, lxml, or configparser. The table below maps each one to its trigger and fix.
| You see | Raised by | Most common trigger | Jump to |
|---|---|---|---|
| SyntaxError: invalid syntax | Python interpreter | missing colon, = vs ==, unclosed quote | Fix your source code |
| SyntaxError: unexpected EOF while parsing / '(' was never closed | Python interpreter | unclosed bracket, parenthesis, or quote | Fix your source code |
| IndentationError | Python interpreter | wrong indent, mixed tabs and spaces | Fix your source code |
| xml.etree.ElementTree.ParseError | xml.etree.ElementTree | malformed or truncated XML | Fix library ParseErrors |
| pandas.errors.ParserError | pandas read_csv / read_html | wrong delimiter, ragged rows, HTML in a CSV | Fix library ParseErrors |
| json.JSONDecodeError | json | non-JSON body (often HTML), empty response | Fix library ParseErrors |
| lxml.etree.XMLSyntaxError | lxml | invalid markup, wrong encoding | Fix library ParseErrors |
| configparser.ParsingError | configparser | malformed .ini, stray indentation | Fix library ParseErrors |
At Proxy-Cheap, we work with developers whose data-collection runs stall on the library rows. If your input comes from the web, choosing the right proxy type for the job is part of the fix, covered at the end.
Python's official errors and exceptions guide says syntax errors are also known as parsing errors. Python reads your whole file before it runs a single line. If the structure breaks the grammar, nothing runs.
These fixes apply whether you're writing a quick script or working with a Python API client. We tested each example on Python 3.9 and on 3.10+, since the wording changed a lot in 3.10.
Every statement that opens an indented suite needs a colon at the end of its first line.
python def greet(name) print("Hello", name)
text File "app.py", line 1 def greet(name) ^ SyntaxError: expected ':'
Python 3.9 just says invalid syntax, with the caret in the same spot. Add the colon: def greet(name):.
SyntaxError: unexpected EOF while parsing means Python reached the end of the file while still expecting more code, almost always an unclosed bracket, parenthesis, or quote earlier in the file. Scan upward from the reported line, not just the line itself. On Python 3.10 and newer, the message is often more precise; for example, '(' was never closed.
python total = sum([1, 2, 3] print(total)
text File "app.py", line 1 total = sum([1, 2, 3] ^ SyntaxError: '(' was never closed
Python 3.10+ prints the output above. Python 3.9 blames line 2 instead, with plain invalid syntax. For quotes, 3.9 says EOL while scanning string literal, and 3.10+ says unterminated string literal.
A single = assigns a value. A double == compares two values.
python if status = 200: print("OK")
Python 3.10+ even suggests the fix: SyntaxError: invalid syntax. Maybe you meant '==' or ':=' instead of '='?. Change it to if status == 200:.
Without a comma, Python can't tell where one item ends and the next begins.
python config = { "timeout": 10 "retries": 3 }
Python 3.10+ flags line 2 with Perhaps you forgot a comma?. Python 3.9 flags line 3 with invalid syntax. Add a comma after 10.
A misspelled keyword breaks the grammar. fro item in [1, 2, 3]: raises SyntaxError: invalid syntax, because Python has no idea what fro is.
A misspelled built-in is different. prnt("done") parses fine and fails only when it runs, with NameError: name 'prnt' is not defined. If you see NameError, you're already past the parsing stage.
Python uses indentation to define structure, so a wrong indent is a parse error. It raises IndentationError, a subclass of SyntaxError.
python def check(value): return value > 0text File "app.py", line 2 return value > 0 ^ IndentationError: expected an indented block after function definition on line 1
You may also see unexpected indent or unindent does not match any outer indentation level. Mixed tabs and spaces raise TabError, a subclass of IndentationError. PEP 8 recommends 4 spaces per level. Never mix tabs and spaces in one file.
Code pasted from a web page or PDF often carries characters Python can't read. Curly quotes raise SyntaxError: invalid character '“' (U+201C). A non-breaking space is worse, because the line looks perfect: invalid non-printable character U+00A0. Delete the line and retype it by hand.
Here, your code is valid. It fails on a specific input, so the fix lies in the data or in how you read it. We tested these examples on Python 3.11 with pandas 3.0.2, lxml 6.1.0, and requests 2.33.1.
pandas.errors.ParserError is raised when read_csv or read_html cannot tokenize the input, for example, "Error tokenizing data. C error: Expected 5 fields in line 3, saw 7." The usual causes are the wrong delimiter, quoted fields with stray commas, ragged rows, or an HTML error page saved with a .csv name. Fix it by setting the correct sep, using quoting/on_bad_lines, or checking that the file is really CSV.
The pandas ParserError reference calls it a generic error for problems parsing file contents. Here's a handling pattern:
python import pandas as pd path = "orders.csv" # Look at the first bytes: is this really CSV, or an HTML page? with open(path, "rb") as f: print(f.read(80)) try: df = pd.read_csv(path) except pd.errors.ParserError as exc: print(f"Strict read failed: {exc}") df = pd.read_csv(path, on_bad_lines="skip")
on_bad_lines="skip" needs pandas 1.3 or newer, and it drops rows silently, so log what you lose. A wrong delimiter may not raise an error at all: a semicolon file often loads as one wide column, so pass sep=";". engine="python" is slower but handles regex separators. If the first bytes show <!DOCTYPE html>, you saved an error page, and no setting will fix that.
xml.etree.ElementTree.ParseError is raised on malformed XML, with messages like "not well-formed (invalid token)", "mismatched tag", or "no element found". "no element found" usually means the response was empty or truncated. Fix it by validating the source, handling truncation, and wrapping ET.fromstring() or ET.parse() in try/except ET.ParseError.
The exception has a position attribute with the line and column, as the xml.etree.ElementTree docs explain:
python import xml.etree.ElementTree as ET def parse_xml(xml_text): if not xml_text.strip(): raise ValueError("Empty response, nothing to parse") try: return ET.fromstring(xml_text) except ET.ParseError as exc: line, column = exc.position print(f"Bad XML at line {line}, column {column}: {exc}") print(xml_text[:200]) raise
A bare & in text is a common cause of "invalid token". Write it as &.
json.loads() raises JSONDecodeError, a subclass of ValueError, when text isn't valid JSON. The classic message is Expecting value: line 1 column 1 (char 0). The parser failed on the first character, so you got HTML or an empty body. Expecting ',' delimiter often means a truncated payload.
Check the response before parsing it:
python import json import requests def fetch_json(url): response = requests.get(url, timeout=10) content_type = response.headers.get("Content-Type", "") if response.status_code != 200 or "json" not in content_type: print(f"Unexpected response: {response.status_code} {content_type}") print(response.text[:200]) return None try: return json.loads(response.text) except json.JSONDecodeError as exc: print(f"Invalid JSON at line {exc.lineno}, column {exc.colno}: {exc.msg}") print(response.text[:200]) return None
With response.json(), requests raises its own JSONDecodeError. In our tests, except json.JSONDecodeError still caught it.
lxml raises XMLSyntaxError on invalid markup, such as Opening and ending tag mismatch. Wrong encoding is the other big cause. A file that declares UTF-8 but holds Latin-1 bytes raises Invalid bytes in character encoding. Pass bytes (response.content) so lxml can read the encoding declaration.
For messy markup, try the recovering parser:
python from lxml import etree raw = b"<item><name>Mug</item>" try: root = etree.fromstring(raw) except etree.XMLSyntaxError as exc: print(f"Strict parse failed: {exc}") root = etree.fromstring(raw, etree.XMLParser(recover=True)) print(etree.tostring(root)) # b'<item><name>Mug</name></item>'
recover=True guesses at the structure, so check the output. It can't rescue an empty document either.
configparser raises ParsingError when an .ini line doesn't fit the format, such as a line with no = or :. A missing [section] header raises MissingSectionHeaderError, a subclass. Duplicate keys raise DuplicateOptionError, which strict=False turns off. Stray indentation often raises nothing and silently glues the line onto the value above it.
python import configparser config = configparser.ConfigParser(strict=False) # last duplicate key wins try: if not config.read("settings.ini"): raise FileNotFoundError("settings.ini was not found") except configparser.Error as exc: # parent of ParsingError and friends print(f"Bad config file: {exc}") raise
The file check matters because config.read() skips missing files silently.
Catch the specific exception, not a bare except. For XML use except ET.ParseError, for JSON use except json.JSONDecodeError, for pandas use except pandas.errors.ParserError. Log the raw input that failed so you can see whether the problem is your parser or the data itself. A bare except hides the real cause and makes ParseErrors harder to fix.
| Library | Import | Exception to catch |
|---|---|---|
| xml.etree | import xml.etree.ElementTree as ET | ET.ParseError |
| json | import json | json.JSONDecodeError |
| pandas | import pandas as pd | pd.errors.ParserError |
| lxml | from lxml import etree | etree.XMLSyntaxError |
| configparser | import configparser | configparser.Error |
If your handler never fires, check the module. except ET.ParseError won't catch an lxml error, and xml.parsers.expat.ExpatError won't catch an ElementTree one. Also make sure the parse call sits inside your try clause. A SyntaxError in your own file can't be caught there at all, because the file never starts.
A source-code parse error prints a short report with four parts.
The file name and line number come first. Treat the line number as a starting point, because the real mistake is often one line earlier, especially on Python 3.9. Next is the offending line, then the caret (^). It points at, or just after, the token that confused the parser, and 3.10+ can underline a whole span with ^^^^.
The last line is the error type and message. Notice there's no Traceback (most recent call last) header, since your code never started. Still stuck? Copy the lines into a new file and cut them down until the error disappears. The last thing you removed is the problem.
When a scraper raises a ParseError, the problem is often the input, not your code. The response you parsed was not the document you expected: a truncated page, an empty body, or an error page returned with an HTTP 200 status. Feeding that to json.loads() or ET.fromstring() raises a parse error. First, log the raw response. Then make your requests return complete, consistent content by routing data collection through reliable IPs.
Logging answers the real question: is your parser wrong, or is the data wrong? Check the status code and Content-Type header, then print the first 200 characters of the body. The fetch_json() function above does both.
In web data collection, inconsistent or incomplete responses are a common source of malformed-input parse errors. A server under load may send a partial page. A busy endpoint may answer with an HTML notice instead of JSON. A dropped connection leaves you with half an XML feed.
Reliable residential, ISP, or datacenter IPs return complete, well-formed responses that parse cleanly. Match the type to the workload:
Rotating residential is pay-as-you-go, with no monthly commitment, so you pay only for the traffic you use. Start with a small top-up and compare your parse-error logs before and after.
Parsing means checking input against grammar rules. For code, that's the Python language itself. For a data file, it's the format's specification.
It happens in two stages. In lexical analysis, a tokenizer splits raw text into tokens, such as names and operators. In syntax analysis, the parser checks whether those tokens form a valid structure. A parse error happens at this second stage, where every token is fine, but the structure never completes.
Data parsing follows the same idea. json.loads() tokenizes text and validates it against JSON rules, raising a library error instead of a SyntaxError. Our blog covers other Python and networking basics too.
Parse errors stop your program before it runs, because the structure is invalid. Runtime exceptions like TypeError or IndexError happen while valid code is running. Library ParseErrors technically belong here, since your code works until it meets bad data.
Logic bugs run without an error but give the wrong result, so only tests catch them.
So how do you fix a TypeError? It isn't a parse error. You used the wrong type, like adding a string to a number, so check the types on the flagged line.
Follow PEP 8, especially 4-space indentation. Use a modern IDE like VS Code or PyCharm, which underlines syntax errors as you type. Add a linter and formatter such as ruff and black, and run them in a pre-commit hook along with python -m py_compile yourfile.py.
Write and run code in small increments. For data, check status codes and headers early and test against real sample files. When you check how pages load from different locations, browser testing tools let you switch proxies without touching system settings.