

Key takeaways
A headless browser is a real browser engine, such as Chrome, Firefox, or WebKit, that runs without a visible window. In Python, libraries like Selenium, Playwright, and Pyppeteer let you load pages, run JavaScript, and access the fully rendered HTML that a basic HTTP request may miss.
This matters for single-page apps where content only appears after JavaScript executes. A headless browser builds the same DOM as a regular browser, but without displaying it, making it ideal for servers, CI pipelines, and scheduled jobs.
Python’s main options are Selenium, Playwright, and Pyppeteer. All three support proxy connections, which we’ll cover later.
To run headless Chrome with Selenium in Python, create a ChromeOptions object, add the --headless=new argument, and pass it to webdriver.Chrome. Selenium Manager, built into Selenium 4.6 and later, fetches the matching driver automatically. Call driver.get(url), wait for an element with WebDriverWait, then read driver.page_source or extract elements by CSS selector.
We tested everything below with Selenium 4.49.0, Python 3.11, and Chrome 148.
bash pip install selenium
Older tutorials add the webdriver-manager package, but you don't need it anymore. Selenium Manager finds your Chrome version and downloads the right ChromeDriver.
This script opens the JavaScript version of the Quotes to Scrape sandbox, waits for the quotes to render, and prints them:
python from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument("--window-size=1920,1080") driver = webdriver.Chrome(options=options) # Selenium Manager resolves the driver try: driver.get("https://quotes.toscrape.com/js/") WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote")) ) for quote in driver.find_elements(By.CSS_SELECTOR, "div.quote"): text = quote.find_element(By.CSS_SELECTOR, "span.text").text author = quote.find_element(By.CSS_SELECTOR, "small.author").text print(f"{author}: {text}") finally: driver.quit()
A few details matter here:
To parse with BeautifulSoup instead, pass it driver.page_source. For the networking side, our explainer on the proxy API covers how proxies expose endpoints to code. The full reference is in the Selenium documentation.
Install Playwright with pip install playwright, then playwright install chromium. Launch with p.chromium.launch(headless=True), open a page, and call page.goto(url). Playwright waits for elements automatically, so you can call page.locator(...) without manual sleeps. It drives Chromium over the Chrome DevTools Protocol, which makes it faster than Selenium's WebDriver path.
bash pip install playwright playwright install chromium # On a fresh Linux server, add system libraries too: # playwright install --with-deps chromium
Playwright downloads its own browser builds, so there's no driver to match. Current releases need Python 3.10 or later, and we tested with Playwright 1.63.0.
Here's the same quotes job with the sync API:
python from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto("https://quotes.toscrape.com/js/") quotes = page.locator("div.quote") quotes.first.wait_for() # explicit wait for JS-rendered content for i in range(quotes.count()): text = quotes.nth(i).locator("span.text").inner_text() author = quotes.nth(i).locator("small.author").inner_text() print(f"{author}: {text}") page.screenshot(path="quotes.png", full_page=True) browser.close()
Locator actions like inner_text() and click() wait for the element on their own. count() doesn't, which is why the script waits on quotes.first before counting. The screenshot line is a one-liner, and page.pdf(path="quotes.pdf") does the same for PDFs in Chromium.
The async API is where Playwright pulls ahead, because one browser can load several pages at once:
python import asyncio from playwright.async_api import async_playwright URLS = [f"https://quotes.toscrape.com/js/page/{n}/" for n in range(1, 6)] async def scrape(browser, url): page = await browser.new_page() await page.goto(url) await page.locator("div.quote").first.wait_for() count = await page.locator("div.quote").count() await page.close() return url, count async def main(): async with async_playwright() as p: browser = await p.chromium.launch(headless=True) results = await asyncio.gather(*(scrape(browser, u) for u in URLS)) for url, count in results: print(url, count) await browser.close() asyncio.run(main())
By default, headless=True uses a lightweight Chromium headless shell. Pass channel="chromium" to launch() if you want a full, new headless Chrome instead. Pick Playwright for new projects, async jobs, or when you need Chromium, Firefox, and WebKit from one API. It also pairs well with rotating residential proxies for multi-page jobs. More options are in the Playwright Python documentation.
Pyppeteer is a Python port of Puppeteer that drives Chromium through the DevTools Protocol and is async by default. In 2026, use pyppeteer==2.0.0 in its own virtualenv. It requires websockets 10.x and can conflict with Selenium 4’s urllib3 2.x dependency.
Older tutorials often pin websockets==8.1 for Pyppeteer 0.x, so follow current version guidance. Our install tests found:
Use a separate environment and pin the version:
bash python -m venv pyppeteer-env source pyppeteer-env/bin/activate pip install pyppeteer==2.0.0python import asyncio from pyppeteer import launch async def main(): browser = await launch(headless=True, args=["--no-sandbox"]) page = await browser.newPage() await page.goto("https://quotes.toscrape.com/js/") await page.waitForSelector("div.quote") quotes = await page.querySelectorAllEval( "div.quote", "nodes => nodes.map(n => n.querySelector('small.author').innerText" " + ': ' + n.querySelector('span.text').innerText)", ) for line in quotes: print(line) await browser.close() asyncio.run(main())
On first launch, Pyppeteer downloads a Chromium build of about 100MB, or you can use an existing browser with executablePath. Recent Chrome versions may print a harmless “Future exception was never retrieved” message on close.
Pyppeteer is Chromium-only and no longer actively maintained. Its README recommends Playwright, and the latest release is 2.0.0 from February 2024. It still works well for porting Puppeteer scripts, but the Pyppeteer documentation shows the older 0.0.25 version, so check PyPI for current pins.
Each library accepts a proxy at launch. Selenium uses --proxy-server=host:port with ChromeOptions and a BiDi auth handler. Playwright accepts proxy={"server", "username", "password"} at launch. Pyppeteer uses --proxy-server and page.authenticate.
Residential or ISP proxies in your target market help avoid per-IP rate limits and ensure region-specific pages match what local visitors see. This is useful for localization testing, price checks, and QA on public content.
The snippets use placeholder credentials. Get your real values from the Proxy-Cheap dashboard, where you can select a country and choose rotating IPs or sticky sessions of about 30 minutes. We tested all three over HTTPS through a local password-protected HTTP proxy.
Playwright is the simplest. The proxy is a dict passed to launch():
python from playwright.sync_api import sync_playwright PROXY = { "server": "http://your-proxy-host:your-proxy-port", "username": "your-username", "password": "your-password", } with sync_playwright() as p: browser = p.chromium.launch(headless=True, proxy=PROXY) page = browser.new_page() page.goto("https://httpbin.org/ip") print(page.inner_text("body")) browser.close()
Selenium takes the proxy address as a Chrome flag, but Chrome won't accept a username and password there. Selenium 4 handles the login through WebDriver BiDi:
python from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC PROXY_HOST = "your-proxy-host" # copy these from the Proxy-Cheap dashboard PROXY_PORT = "your-proxy-port" PROXY_USER = "your-username" PROXY_PASS = "your-password" options = webdriver.ChromeOptions() options.add_argument("--headless=new") options.add_argument(f"--proxy-server=http://{PROXY_HOST}:{PROXY_PORT}") options.enable_bidi = True # WebDriver BiDi handles the proxy login driver = webdriver.Chrome(options=options) driver.network.add_auth_handler(PROXY_USER, PROXY_PASS) def open_page(url): # Navigate over BiDi so the auth handler can answer the proxy challenge driver.browsing_context.navigate( context=driver.current_window_handle, url=url, wait="complete" ) try: open_page("https://httpbin.org/ip") print(driver.find_element(By.TAG_NAME, "body").text) # exit IP open_page("https://quotes.toscrape.com/js/") WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "div.quote")) ) print(len(driver.find_elements(By.CSS_SELECTOR, "div.quote")), "quotes") finally: driver.quit()
One gotcha from our testing: with the auth handler active, a plain driver.get() stalled until the page load timeout on Selenium 4.49 and Chrome 148. Classic navigation waits for the page, while the proxy login waits on a BiDi reply. Navigating with driver.browsing_context.navigate() keeps both on BiDi, and the page loaded in under a second. On IP whitelist plans, skip the handler and keep only the --proxy-server flag.
Pyppeteer takes the flag in args and the login through page.authenticate():
python import asyncio from pyppeteer import launch PROXY_HOST = "your-proxy-host" PROXY_PORT = "your-proxy-port" PROXY_USER = "your-username" PROXY_PASS = "your-password" async def main(): browser = await launch( headless=True, args=["--no-sandbox", f"--proxy-server=http://{PROXY_HOST}:{PROXY_PORT}"], ) page = await browser.newPage() await page.authenticate({"username": PROXY_USER, "password": PROXY_PASS}) await page.goto("https://httpbin.org/ip") print(await page.evaluate("document.body.innerText")) await browser.close() asyncio.run(main())
Rotating residential proxies use username and password, built for high-volume rotation. To keep credentials out of code, static residential (ISP) proxies and datacenter plans also support IP whitelist authentication. See our web data collection page for related workloads, and the Proxy-Cheap API docs for managing proxies from code.
Match the proxy to the job. Use rotating residential for high-volume, distributed collection where each request can come from a fresh IP. Use static residential (ISP) for account-bound sessions that must keep one identity. Use datacenter IPv4 for high-throughput work on publicly available, unprotected pages where speed and cost matter most.
| Workload | Recommended Proxy-Cheap product | Auth | Why it fits |
|---|---|---|---|
| High-volume collection across many pages | Rotating residential | Username/password | New IP per request or a sticky session of about 30 minutes, 180+ locations, pay-as-you-go per GB |
| Account-bound or logged-in sessions | Static residential (ISP) | Username/password or IP whitelist | One fixed IP for the whole subscription, country and ISP targeting, unlimited bandwidth |
| High-throughput runs on public pages | Datacenter IPv4 (IPv6 also available) | Username/password or IP whitelist | Per-IP pricing, unlimited bandwidth, fixed IPs for repeatable daily crawls |
| Mobile-first sites and mobile QA | Rotating mobile | IP whitelist | Carrier IPs from 5G, 4G, and LTE networks in 100+ countries |
Rotating residential has no limit on concurrent sessions, which suits async Playwright jobs. Static residential and datacenter plans are sized for about 100 concurrent connections per proxy. With datacenter proxies, per-IP pricing keeps costs predictable for a crawl you run every day.
4G and 5G mobile proxies show you what visitors on mobile networks see, which pairs well with Playwright's device emulation. For background on each type, read our guide to residential vs datacenter proxies.
Ready to test your setup? Proxy-Cheap runs on pay-as-you-go billing with no monthly commitment, so you can try a small plan against your own script before scaling. For session-heavy headless work, start with static and rotating ISP proxies and drop the credentials into the snippets above.
Playwright is the fastest and most modern, with automatic waiting and native async support. Selenium supports the widest browser range and has the largest community. Pyppeteer is a lightweight, Chromium-only choice for porting Puppeteer code. For most new Python projects, start with Playwright; use Selenium when browser coverage matters.
| Library | Browsers | Speed (20 URLs) | Async | Best for |
|---|---|---|---|---|
| Playwright 1.63 | Chromium, Firefox, WebKit | 7.4s sync, 5.4s async | Yes | New projects |
| Selenium 4.49 | Chrome, Firefox, Edge, Safari | 10.2s | No native async | Cross-browser testing |
| Pyppeteer 2.0.0 | Chromium | Not benchmarked | Async only | Porting Puppeteer |
Benchmark: Selenium and Playwright loaded 20 JavaScript-rendered pages and extracted 10 quotes from each on a 2-vCPU, 8GB Linux sandbox. Figures are the median of three runs, including browser startup. Playwright was about 1.4x faster sequentially and 1.9x faster with five concurrent pages.
Real sites will take longer, but the ranking should hold. Speed matters most for recurring jobs such as SEO research workflows.
Memory matters too. One headless Chrome instance used about 310MB in our sandbox, within the reported 200 to 500MB range.
Yes. Sites can detect automation through signals like navigator.webdriver, missing browser features, unusual request patterns, and the "HeadlessChrome" user agent. Hundreds of requests per minute from one IP can also stand out.
For reliable, location-accurate sessions, use residential or ISP IPs that match your target market and collect only public data. Be a considerate client:
Headless browsers are heavier than plain HTTP requests. Each instance uses roughly 200-500MB of RAM and adds time to every page, so a single server handles only a handful of concurrent instances. They also need browser dependencies that are easy to miss on a fresh server, which is why scripts that work locally fail there.
The same Firecrawl analysis estimates a 4GB server handles about 5 to 8 concurrent browser instances. It puts a self-hosted headless setup at $190 to $780 a month, covering compute, proxies, storage, and monitoring.
Server fragility is the second pain point. These fixes come up most often:
Sometimes the best fix is no browser at all. If the data is in the initial HTML, or loads from a JSON endpoint you can call directly, requests plus BeautifulSoup is faster and cheaper. Check the page source and network tab first.
Yes. Selenium remains widely used in 2026 because it supports Chrome, Firefox, Safari, and Edge, has the largest community, and integrates with most test frameworks. Playwright is faster for new scraping projects, but Selenium is still the pragmatic pick when you need broad browser coverage or existing test infrastructure.
It keeps improving, too. Selenium Manager removed most driver headaches, and WebDriver BiDi now handles proxy logins that used to need extra packages.