

Where HttpProxyMiddleware sits in the Scrapy request lifecycle.
Scrapy is an open-source Python framework for crawling and structured data extraction. It is asynchronous, built on Twisted, and ships with item pipelines, selectors, an autothrottle, and a downloader middleware stack that lets you inject behavior at the request and response level.
A spider running with default settings can issue 16 concurrent requests to a single domain in seconds. At that volume, a single origin IP triggers rate limits, soft bans, and CAPTCHA pages. Proxies route Scrapy requests through pools of residential, ISP, datacenter, or mobile IPs so a crawl distributes load, collects publicly available content from multiple geographic locations for compliance and QA, and maintains stable session quality across long runs.
The brief version: if your spider hits the same target more than a few hundred times an hour, proxies are not optional infrastructure.
The built-in HttpProxyMiddleware accepts proxy URLs with the http:// and https:// schemes. The scheme refers to how Scrapy talks to the proxy, not to the target. For HTTPS destinations, Scrapy opens an HTTP CONNECT tunnel through the proxy and then negotiates TLS with the target inside that tunnel.
SOCKS5 is not supported by Scrapy core. The Twisted-based HTTP/1.1 download handler only speaks HTTP CONNECT. If you need SOCKS, the practical options are to run a local HTTP-to-SOCKS bridge process and point Scrapy at the local HTTP listener, or to swap in a custom download handler that wraps a SOCKS-aware agent. Many teams skip this complexity by using a SOCKS5 proxy endpoint with HTTP fallback for tools that need it and HTTP proxies for Scrapy itself.
The simplest pattern sets the proxy on a single request through the meta dictionary. HttpProxyMiddleware reads meta['proxy'] on every request, parses any embedded credentials, and writes the corresponding Proxy-Authorization header before the request leaves the downloader.
import scrapy
class IpCheckSpider(scrapy.Spider):
name = "ip_check"
def start_requests(self):
proxy = "http://user:[email protected]:8080"
yield scrapy.Request(
url="https://httpbin.org/ip",
meta={"proxy": proxy},
callback=self.parse,
dont_filter=True,
)
def parse(self, response):
self.logger.info("status=%s body=%s", response.status, response.text)Use this approach for one-off scripts, debugging, or spiders where a small set of URLs needs a specific exit location. The meta['proxy'] value takes precedence over any environment variable and is the only way to force a request through a proxy when no_proxy would otherwise apply.
To opt a single request out of an otherwise proxied crawl, set request.meta["proxy"] = None explicitly.
If every request in a project should go through the same proxy, environment variables are the lowest-friction option. HttpProxyMiddleware calls urllib.request.getproxies() at startup, which picks up http_proxy, https_proxy, and no_proxy in either lowercase or uppercase.
import os
os.environ["http_proxy"] = "http://user:[email protected]:8080"
os.environ["https_proxy"] = "http://user:[email protected]:8080"
os.environ["no_proxy"] = "localhost,127.0.0.1,.internal.example"
# After this point, every Scrapy request that does not set meta["proxy"]
# explicitly will route through the configured proxy. HttpProxyMiddleware
# is enabled by default at order 750 when HTTPPROXY_ENABLED is True.The environment-variable path respects no_proxy, which is useful when a spider needs to reach internal endpoints alongside public ones. Note that meta['proxy'] ignores no_proxy entirely, so the two methods are not strictly equivalent.
Production projects almost always end up with a custom downloader middleware. It centralizes proxy logic, supports rotation, and keeps the spider code clean. The pattern below assigns a single configured proxy URL to every request that does not already have one.
# myproject/middlewares.py
from scrapy.exceptions import NotConfigured
class ForceProxyMiddleware:
"""Routes every outgoing request through the configured proxy.
Order it below 750 so HttpProxyMiddleware can still extract
credentials from the URL and convert them into a header.
"""
def __init__(self, proxy_url):
if not proxy_url:
raise NotConfigured("PROXY_URL is not configured")
self.proxy_url = proxy_url
@classmethod
def from_crawler(cls, crawler):
return cls(crawler.settings.get("PROXY_URL"))
def process_request(self, request, spider):
request.meta.setdefault("proxy", self.proxy_url)
return NoneEnable it in settings.py:
PROXY_URL = "http://user:[email protected]:8080"
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.ForceProxyMiddleware": 610,
}Middleware ordering matters. Place your custom proxy middleware below 750 so the built-in HttpProxyMiddleware still parses embedded credentials and converts them to a Proxy-Authorization header. Place it above 750 only if you want the final word on what meta['proxy'] contains.
HttpProxyMiddleware handles basic auth automatically when credentials are embedded in the URL. The middleware calls urllib.request._parse_proxy, strips the user and password from the URL, base64-encodes them using HTTPPROXY_AUTH_ENCODING (default latin-1), and writes the result to request.headers[b"Proxy-Authorization"].
Two practical rules follow from that behavior.
First, percent-encode special characters in your credentials. Passwords containing @, #, %, or $ will mangle the URL parse and produce silent failures. Wrap them with urllib.parse.quote(password, safe="") before assembling the URL.
Second, when the URL approach fails, set the header yourself. This is the canonical fix for HTTPS targets where credentials need to survive a CONNECT tunnel:
import base64
import scrapy
def build_proxy_auth_header(username, password):
"""Return a Basic auth value for the Proxy-Authorization header."""
token = base64.b64encode(f"{username}:{password}".encode("latin-1"))
return b"Basic " + token
class AuthHeaderSpider(scrapy.Spider):
name = "auth_header"
def start_requests(self):
yield scrapy.Request(
url="https://httpbin.org/ip",
meta={"proxy": "http://proxy.example.net:8080"},
headers={
"Proxy-Authorization": build_proxy_auth_header("user", "pass"),
},
callback=self.parse,
dont_filter=True,
)
def parse(self, response):
self.logger.info("ok status=%s", response.status)w3lib.http.basic_auth_header(user, password) returns the same value and ships as a transitive dependency of Scrapy. Either approach is fine.
If credentials contain non-latin-1 characters, set HTTPPROXY_AUTH_ENCODING = "utf-8" in settings.py.
Rotation prevents any single IP from absorbing the full request rate. Four patterns cover the field.
Per-request rotation via custom middleware. A list of proxies, picked at random per request, with a process_exception hook that swaps the proxy on connection failures.
import random
from scrapy.exceptions import NotConfigured
class RotatingProxyMiddleware:
"""Picks a random proxy from PROXY_LIST for each request.
Order below 750 so HttpProxyMiddleware parses user:pass.
"""
def __init__(self, proxies):
if not proxies:
raise NotConfigured("PROXY_LIST is empty")
self.proxies = list(proxies)
@classmethod
def from_crawler(cls, crawler):
return cls(crawler.settings.getlist("PROXY_LIST"))
def process_request(self, request, spider):
if "proxy" in request.meta:
return None
request.meta["proxy"] = random.choice(self.proxies)
return None
def process_exception(self, request, exception, spider):
current = request.meta.get("proxy")
candidates = [p for p in self.proxies if p != current] or self.proxies
new_request = request.copy()
new_request.meta["proxy"] = random.choice(candidates)
new_request.dont_filter = True
return new_requestSticky session rotation. Same exit IP for the duration of a session, keyed by meta['cookiejar']. Required for login flows, multi-step checkouts, and any target that fingerprints the IP and cookie pair together. Read how sticky and rotating sessions compare for the trade-offs.
Time-based rotation. A new exit IP every N minutes. Useful for crawls that want some stability without locking to a single IP for the whole run.
Provider gateway rotation. A single endpoint that rotates server-side. This is the simplest pattern from the Scrapy side: you set one proxy URL, the provider returns a different exit IP per request. The Proxy-Cheap single hub gateway is typically used this way, with port 8080 for HTTP and SOCKS5. Background on how IP rotation reduces detection covers the underlying mechanics.

Which Proxy-Cheap product fits which Scrapy workload.
Default Scrapy settings are tuned for direct connections. Proxies add latency and shift the failure modes, so a handful of settings deserve attention.
CONCURRENT_REQUESTS defaults to 16 and CONCURRENT_REQUESTS_PER_DOMAIN to 8. With a rotating pool, concurrency becomes per-proxy when rotation is active, so a 50-proxy pool with the defaults can fire 50 sessions in parallel. Match your concurrency to pool size deliberately.
DOWNLOAD_DELAY defaults to 0. RANDOMIZE_DOWNLOAD_DELAY (default True) jitters the actual wait to a value between 0.5x and 1.5x of the delay you set. A small delay of 0.25 to 1 second per request often outperforms aggressive concurrency on protected targets.
AUTOTHROTTLE_ENABLED defaults to False. Turn it on for any production crawl. It adjusts the per-slot delay based on observed response latency, so slow proxies do not overwhelm the reactor. Set AUTOTHROTTLE_TARGET_CONCURRENCY to a small fraction of your pool size (1 to 4 is typical) and let AUTOTHROTTLE_MAX_DELAY cap the worst case at 30 to 60 seconds.
DOWNLOAD_TIMEOUT defaults to 180 seconds. That is too generous for rotating proxies. Lower it to 45 to 60 seconds so dead exit nodes are dropped quickly and RetryMiddleware can reissue the request onto a fresh proxy.
RETRY_TIMES defaults to 2. RETRY_HTTP_CODES defaults to [500, 502, 503, 504, 522, 524, 408, 429]. Add 407 to the list while debugging proxy auth so you can observe failures instead of looping. RetryMiddleware also catches TunnelError automatically, which covers the HTTPS CONNECT failure path.
The right product depends on the target site and the shape of the crawl.
| Crawl profile | Fit | Why |
|---|---|---|
| High-volume crawl across many domains, rotating IPs preferred | Rotating residential proxies | Large pool across 195+ countries with sticky and rotating sessions |
| Session-bound crawl (login, cart, paginated dashboards) | Static residential ISP proxies | Fixed IP per session, datacenter-grade speed with ISP origin |
| Stable identity, fast crawls on lightly protected targets | Dedicated ISP proxies | Static IP with high concurrency and predictable latency |
| Documentation crawls, sitemaps, unprotected public content | Dedicated datacenter proxies | High throughput at low cost per IP |
| Targets that filter aggressively on network origin | Mobile 4G and 5G proxies | Real carrier IPs from 3G, 4G, and 5G networks |
For projects that need a quick map between provider choice and crawl pattern, the proxies for data scraping overview walks through the same decisions from a use-case angle.
HTTP 407 Proxy Authentication Required. The proxy received no Proxy-Authorization header or rejected the one it got. Confirm credentials are URL-encoded, set HTTPPROXY_AUTH_ENCODING = "utf-8" if the password contains non-latin characters, and add 407 to HTTPERROR_ALLOWED_CODES so the response surfaces in your callback instead of being filtered.
TunnelError: Could not open CONNECT tunnel. Raised when the HTTPS CONNECT handshake to the proxy returns non-200. Most often this is a 407 on the CONNECT line. Switch to the Proxy-Authorization header pattern shown above, and verify that the proxy listens HTTP on the port you are using (not TLS to the proxy itself).
SSL certificate verification errors. Some proxies intercept TLS with their own CA. To run through a known-trustworthy intercepting proxy, set DOWNLOADER_CLIENT_TLS_VERIFY = False for that crawl, or supply a custom DOWNLOADER_CLIENTCONTEXTFACTORY that trusts the intercepting CA. Do this with care and only when the proxy is yours.
DOWNLOAD_TIMEOUT exceeded. A residential proxy with a slow exit node holds the socket open until the 180-second default fires. Lower DOWNLOAD_TIMEOUT to 45 to 60 seconds, raise RETRY_TIMES, and add twisted.internet.error.TCPTimedOutError to RETRY_EXCEPTIONS if you see TCP-level stalls.
Mixed http:// and https:// in the proxy URL. The scheme of meta['proxy'] describes the connection to the proxy, not the target. Use http://host:port unless the provider explicitly documents a TLS-terminated proxy endpoint. The destination URL can still be https://, and Scrapy will do CONNECT plus TLS to the target correctly.
Rotating middleware overwritten by another component. If rotation appears to work but every request still uses the same IP, check middleware order. A rotating middleware at order 610 runs before HttpProxyMiddleware at 750. A naive middleware at 800 may overwrite the assigned proxy without re-syncing the auth header.
Soft bans behind a 200 response. A 200 status code with CAPTCHA HTML in the body means concurrency outran rotation. Lower CONCURRENT_REQUESTS_PER_DOMAIN, raise DOWNLOAD_DELAY, and rotate the User-Agent header alongside the IP.
A handful of habits separate stable production crawls from spiders that need constant attention.
Rotate User-Agent headers in lockstep with IPs. A fresh exit IP paired with the same UA fingerprint is trivially correlatable, and many targets cluster on the combination rather than either alone.
Log the proxy used per request. A two-line process_request hook that emits spider.logger.debug("proxy=%s url=%s", request.meta.get("proxy"), request.url) removes a category of "proxy not working" bugs that are really "proxy not assigned".
Set ROBOTSTXT_OBEY explicitly. The default in new projects is True, and the first request a spider makes is to /robots.txt. That request goes through the proxy too. Confirm it is what you want.
Match concurrency to pool size. With a list-based rotator, the working rule is CONCURRENT_REQUESTS close to pool_size multiplied by CONCURRENT_REQUESTS_PER_DOMAIN. With a provider gateway that rotates server-side, concurrency is bounded by the gateway’s documented thread limit instead.
Track proxy stats. crawler.stats.inc_value("proxies/good") and crawler.stats.inc_value("proxies/dead") from your middleware turn rotation into observable data. The seven proxy types best matched to web scraping covers the upstream decisions that make those stats look healthy in the first place. Pay-as-you-go billing on residential and datacenter pools means you can scale a Scrapy crawl up and down without provisioning friction, so start with a small pool, watch the stats, and grow the pool as your concurrency settings warrant.