Proxy-Cheap
Integrações
Scrapy proxy integration: a complete setup and rotation guide

Scrapy Proxy Integration

Scrapy ships proxy support out of the box through HttpProxyMiddleware, but production crawls need more: rotation, authentication headers, and settings tuned for proxy latency. This guide covers all three integration methods, plus fixes for the errors that come up most.
Obter proxies para Scrapy proxy integration: a complete setup and rotation guide
Scrapy proxy integration: a complete setup and rotation guide
What is Scrapy?
Scrapy is an open-source Python framework for crawling and structured data extraction, built on Twisted with a middleware stack for request and response handling. It's used for high-volume, asynchronous crawls across many domains. Proxies get assigned through a meta key, environment variables, or a custom middleware, and pairing them with the right settings keeps a crawl from tripping rate limits or soft bans.

Key takeaways

  • Scrapy ships with HttpProxyMiddleware enabled by default at order 750, which reads the proxy meta key and the http_proxy / https_proxy environment variables.
  • Three canonical integration methods cover almost every project: per-request meta['proxy'], environment variables, and a custom downloader middleware that injects a proxy on every request.
  • For HTTPS targets, embedded credentials are parsed into a Proxy-Authorization header automatically, and you can set that header manually when the URL approach fails.
  • Pair proxy rotation with AUTOTHROTTLE_ENABLED, a tuned DOWNLOAD_TIMEOUT, and per-domain concurrency limits to avoid 407 loops, tunnel errors, and slow responses.

Scrapy request lifecycle with HttpProxyMiddleware

Where HttpProxyMiddleware sits in the Scrapy request lifecycle.

What Scrapy is and why proxies matter at scale

Scrapy is an open-source Python framework for crawling and structured data extraction. It is asynchronous, built on Twisted, and ships with item pipelines, selectors, an autothrottle, and a downloader middleware stack that lets you inject behavior at the request and response level.

A spider running with default settings can issue 16 concurrent requests to a single domain in seconds. At that volume, a single origin IP triggers rate limits, soft bans, and CAPTCHA pages. Proxies route Scrapy requests through pools of residential, ISP, datacenter, or mobile IPs so a crawl distributes load, collects publicly available content from multiple geographic locations for compliance and QA, and maintains stable session quality across long runs.

The brief version: if your spider hits the same target more than a few hundred times an hour, proxies are not optional infrastructure.

Proxy schemes Scrapy understands natively

The built-in HttpProxyMiddleware accepts proxy URLs with the http:// and https:// schemes. The scheme refers to how Scrapy talks to the proxy, not to the target. For HTTPS destinations, Scrapy opens an HTTP CONNECT tunnel through the proxy and then negotiates TLS with the target inside that tunnel.

SOCKS5 is not supported by Scrapy core. The Twisted-based HTTP/1.1 download handler only speaks HTTP CONNECT. If you need SOCKS, the practical options are to run a local HTTP-to-SOCKS bridge process and point Scrapy at the local HTTP listener, or to swap in a custom download handler that wraps a SOCKS-aware agent. Many teams skip this complexity by using a SOCKS5 proxy endpoint with HTTP fallback for tools that need it and HTTP proxies for Scrapy itself.

Method 1: per-request proxy via meta

The simplest pattern sets the proxy on a single request through the meta dictionary. HttpProxyMiddleware reads meta['proxy'] on every request, parses any embedded credentials, and writes the corresponding Proxy-Authorization header before the request leaves the downloader.

import scrapy
 
 
class IpCheckSpider(scrapy.Spider):
    name = "ip_check"
 
    def start_requests(self):
        proxy = "http://user:[email protected]:8080"
        yield scrapy.Request(
            url="https://httpbin.org/ip",
            meta={"proxy": proxy},
            callback=self.parse,
            dont_filter=True,
        )
 
    def parse(self, response):
        self.logger.info("status=%s body=%s", response.status, response.text)

Use this approach for one-off scripts, debugging, or spiders where a small set of URLs needs a specific exit location. The meta['proxy'] value takes precedence over any environment variable and is the only way to force a request through a proxy when no_proxy would otherwise apply.

To opt a single request out of an otherwise proxied crawl, set request.meta["proxy"] = None explicitly.

Method 2: environment variables read by HttpProxyMiddleware

If every request in a project should go through the same proxy, environment variables are the lowest-friction option. HttpProxyMiddleware calls urllib.request.getproxies() at startup, which picks up http_proxy, https_proxy, and no_proxy in either lowercase or uppercase.

import os
 
os.environ["http_proxy"] = "http://user:[email protected]:8080"
os.environ["https_proxy"] = "http://user:[email protected]:8080"
os.environ["no_proxy"] = "localhost,127.0.0.1,.internal.example"
 
# After this point, every Scrapy request that does not set meta["proxy"]
# explicitly will route through the configured proxy. HttpProxyMiddleware
# is enabled by default at order 750 when HTTPPROXY_ENABLED is True.

The environment-variable path respects no_proxy, which is useful when a spider needs to reach internal endpoints alongside public ones. Note that meta['proxy'] ignores no_proxy entirely, so the two methods are not strictly equivalent.

Method 3: custom downloader middleware

Production projects almost always end up with a custom downloader middleware. It centralizes proxy logic, supports rotation, and keeps the spider code clean. The pattern below assigns a single configured proxy URL to every request that does not already have one.

# myproject/middlewares.py
from scrapy.exceptions import NotConfigured
 
 
class ForceProxyMiddleware:
    """Routes every outgoing request through the configured proxy.
 
    Order it below 750 so HttpProxyMiddleware can still extract
    credentials from the URL and convert them into a header.
    """
 
    def __init__(self, proxy_url):
        if not proxy_url:
            raise NotConfigured("PROXY_URL is not configured")
        self.proxy_url = proxy_url
 
    @classmethod
    def from_crawler(cls, crawler):
        return cls(crawler.settings.get("PROXY_URL"))
 
    def process_request(self, request, spider):
        request.meta.setdefault("proxy", self.proxy_url)
        return None

Enable it in settings.py:

PROXY_URL = "http://user:[email protected]:8080"
 
DOWNLOADER_MIDDLEWARES = {
    "myproject.middlewares.ForceProxyMiddleware": 610,
}

Middleware ordering matters. Place your custom proxy middleware below 750 so the built-in HttpProxyMiddleware still parses embedded credentials and converts them to a Proxy-Authorization header. Place it above 750 only if you want the final word on what meta['proxy'] contains.

Proxy authentication patterns

HttpProxyMiddleware handles basic auth automatically when credentials are embedded in the URL. The middleware calls urllib.request._parse_proxy, strips the user and password from the URL, base64-encodes them using HTTPPROXY_AUTH_ENCODING (default latin-1), and writes the result to request.headers[b"Proxy-Authorization"].

Two practical rules follow from that behavior.

First, percent-encode special characters in your credentials. Passwords containing @, #, %, or $ will mangle the URL parse and produce silent failures. Wrap them with urllib.parse.quote(password, safe="") before assembling the URL.

Second, when the URL approach fails, set the header yourself. This is the canonical fix for HTTPS targets where credentials need to survive a CONNECT tunnel:

import base64
import scrapy
 
 
def build_proxy_auth_header(username, password):
    """Return a Basic auth value for the Proxy-Authorization header."""
    token = base64.b64encode(f"{username}:{password}".encode("latin-1"))
    return b"Basic " + token
 
 
class AuthHeaderSpider(scrapy.Spider):
    name = "auth_header"
 
    def start_requests(self):
        yield scrapy.Request(
            url="https://httpbin.org/ip",
            meta={"proxy": "http://proxy.example.net:8080"},
            headers={
                "Proxy-Authorization": build_proxy_auth_header("user", "pass"),
            },
            callback=self.parse,
            dont_filter=True,
        )
 
    def parse(self, response):
        self.logger.info("ok status=%s", response.status)

w3lib.http.basic_auth_header(user, password) returns the same value and ships as a transitive dependency of Scrapy. Either approach is fine.

If credentials contain non-latin-1 characters, set HTTPPROXY_AUTH_ENCODING = "utf-8" in settings.py.

Rotation strategies that work in production

Rotation prevents any single IP from absorbing the full request rate. Four patterns cover the field.

Per-request rotation via custom middleware. A list of proxies, picked at random per request, with a process_exception hook that swaps the proxy on connection failures.

import random
from scrapy.exceptions import NotConfigured
 
 
class RotatingProxyMiddleware:
    """Picks a random proxy from PROXY_LIST for each request.
 
    Order below 750 so HttpProxyMiddleware parses user:pass.
    """
 
    def __init__(self, proxies):
        if not proxies:
            raise NotConfigured("PROXY_LIST is empty")
        self.proxies = list(proxies)
 
    @classmethod
    def from_crawler(cls, crawler):
        return cls(crawler.settings.getlist("PROXY_LIST"))
 
    def process_request(self, request, spider):
        if "proxy" in request.meta:
            return None
        request.meta["proxy"] = random.choice(self.proxies)
        return None
 
    def process_exception(self, request, exception, spider):
        current = request.meta.get("proxy")
        candidates = [p for p in self.proxies if p != current] or self.proxies
        new_request = request.copy()
        new_request.meta["proxy"] = random.choice(candidates)
        new_request.dont_filter = True
        return new_request

Sticky session rotation. Same exit IP for the duration of a session, keyed by meta['cookiejar']. Required for login flows, multi-step checkouts, and any target that fingerprints the IP and cookie pair together. Read how sticky and rotating sessions compare for the trade-offs.

Time-based rotation. A new exit IP every N minutes. Useful for crawls that want some stability without locking to a single IP for the whole run.

Provider gateway rotation. A single endpoint that rotates server-side. This is the simplest pattern from the Scrapy side: you set one proxy URL, the provider returns a different exit IP per request. The Proxy-Cheap single hub gateway is typically used this way, with port 8080 for HTTP and SOCKS5. Background on how IP rotation reduces detection covers the underlying mechanics.

Scrapy crawl profile to Proxy-Cheap product fit matrix

Which Proxy-Cheap product fits which Scrapy workload.

Settings that pair with proxy use

Default Scrapy settings are tuned for direct connections. Proxies add latency and shift the failure modes, so a handful of settings deserve attention.

CONCURRENT_REQUESTS defaults to 16 and CONCURRENT_REQUESTS_PER_DOMAIN to 8. With a rotating pool, concurrency becomes per-proxy when rotation is active, so a 50-proxy pool with the defaults can fire 50 sessions in parallel. Match your concurrency to pool size deliberately.

DOWNLOAD_DELAY defaults to 0. RANDOMIZE_DOWNLOAD_DELAY (default True) jitters the actual wait to a value between 0.5x and 1.5x of the delay you set. A small delay of 0.25 to 1 second per request often outperforms aggressive concurrency on protected targets.

AUTOTHROTTLE_ENABLED defaults to False. Turn it on for any production crawl. It adjusts the per-slot delay based on observed response latency, so slow proxies do not overwhelm the reactor. Set AUTOTHROTTLE_TARGET_CONCURRENCY to a small fraction of your pool size (1 to 4 is typical) and let AUTOTHROTTLE_MAX_DELAY cap the worst case at 30 to 60 seconds.

DOWNLOAD_TIMEOUT defaults to 180 seconds. That is too generous for rotating proxies. Lower it to 45 to 60 seconds so dead exit nodes are dropped quickly and RetryMiddleware can reissue the request onto a fresh proxy.

RETRY_TIMES defaults to 2. RETRY_HTTP_CODES defaults to [500, 502, 503, 504, 522, 524, 408, 429]. Add 407 to the list while debugging proxy auth so you can observe failures instead of looping. RetryMiddleware also catches TunnelError automatically, which covers the HTTPS CONNECT failure path.

Which Proxy-Cheap product fits which Scrapy use case

The right product depends on the target site and the shape of the crawl.

Crawl profileFitWhy
High-volume crawl across many domains, rotating IPs preferredRotating residential proxiesLarge pool across 195+ countries with sticky and rotating sessions
Session-bound crawl (login, cart, paginated dashboards)Static residential ISP proxiesFixed IP per session, datacenter-grade speed with ISP origin
Stable identity, fast crawls on lightly protected targetsDedicated ISP proxiesStatic IP with high concurrency and predictable latency
Documentation crawls, sitemaps, unprotected public contentDedicated datacenter proxiesHigh throughput at low cost per IP
Targets that filter aggressively on network originMobile 4G and 5G proxiesReal carrier IPs from 3G, 4G, and 5G networks

For projects that need a quick map between provider choice and crawl pattern, the proxies for data scraping overview walks through the same decisions from a use-case angle.

Common Scrapy proxy errors and fixes

HTTP 407 Proxy Authentication Required. The proxy received no Proxy-Authorization header or rejected the one it got. Confirm credentials are URL-encoded, set HTTPPROXY_AUTH_ENCODING = "utf-8" if the password contains non-latin characters, and add 407 to HTTPERROR_ALLOWED_CODES so the response surfaces in your callback instead of being filtered.

TunnelError: Could not open CONNECT tunnel. Raised when the HTTPS CONNECT handshake to the proxy returns non-200. Most often this is a 407 on the CONNECT line. Switch to the Proxy-Authorization header pattern shown above, and verify that the proxy listens HTTP on the port you are using (not TLS to the proxy itself).

SSL certificate verification errors. Some proxies intercept TLS with their own CA. To run through a known-trustworthy intercepting proxy, set DOWNLOADER_CLIENT_TLS_VERIFY = False for that crawl, or supply a custom DOWNLOADER_CLIENTCONTEXTFACTORY that trusts the intercepting CA. Do this with care and only when the proxy is yours.

DOWNLOAD_TIMEOUT exceeded. A residential proxy with a slow exit node holds the socket open until the 180-second default fires. Lower DOWNLOAD_TIMEOUT to 45 to 60 seconds, raise RETRY_TIMES, and add twisted.internet.error.TCPTimedOutError to RETRY_EXCEPTIONS if you see TCP-level stalls.

Mixed http:// and https:// in the proxy URL. The scheme of meta['proxy'] describes the connection to the proxy, not the target. Use http://host:port unless the provider explicitly documents a TLS-terminated proxy endpoint. The destination URL can still be https://, and Scrapy will do CONNECT plus TLS to the target correctly.

Rotating middleware overwritten by another component. If rotation appears to work but every request still uses the same IP, check middleware order. A rotating middleware at order 610 runs before HttpProxyMiddleware at 750. A naive middleware at 800 may overwrite the assigned proxy without re-syncing the auth header.

Soft bans behind a 200 response. A 200 status code with CAPTCHA HTML in the body means concurrency outran rotation. Lower CONCURRENT_REQUESTS_PER_DOMAIN, raise DOWNLOAD_DELAY, and rotate the User-Agent header alongside the IP.

Best practices for Scrapy crawls at scale

A handful of habits separate stable production crawls from spiders that need constant attention.

Rotate User-Agent headers in lockstep with IPs. A fresh exit IP paired with the same UA fingerprint is trivially correlatable, and many targets cluster on the combination rather than either alone.

Log the proxy used per request. A two-line process_request hook that emits spider.logger.debug("proxy=%s url=%s", request.meta.get("proxy"), request.url) removes a category of "proxy not working" bugs that are really "proxy not assigned".

Set ROBOTSTXT_OBEY explicitly. The default in new projects is True, and the first request a spider makes is to /robots.txt. That request goes through the proxy too. Confirm it is what you want.

Match concurrency to pool size. With a list-based rotator, the working rule is CONCURRENT_REQUESTS close to pool_size multiplied by CONCURRENT_REQUESTS_PER_DOMAIN. With a provider gateway that rotates server-side, concurrency is bounded by the gateway’s documented thread limit instead.

Track proxy stats. crawler.stats.inc_value("proxies/good") and crawler.stats.inc_value("proxies/dead") from your middleware turn rotation into observable data. The seven proxy types best matched to web scraping covers the upstream decisions that make those stats look healthy in the first place. Pay-as-you-go billing on residential and datacenter pools means you can scale a Scrapy crawl up and down without provisioning friction, so start with a small pool, watch the stats, and grow the pool as your concurrency settings warrant.

Perguntas frequentes

The working rule is one healthy proxy for every 8 concurrent requests against a single domain, scaled by the target’s tolerance. A 16-thread crawl against a moderately protected site usually runs cleanly on a 20 to 50 proxy pool. Protected targets need 200 plus, or a rotating gateway that handles the math for you.

No. The built-in HttpProxyMiddleware and the HTTP/1.1 download handler only speak HTTP CONNECT. To use SOCKS5 with Scrapy, run a local HTTP-to-SOCKS bridge process and point Scrapy at the local HTTP listener, or install a custom download handler that wraps a SOCKS-aware agent.

Implement process_exception on your rotating middleware to swap the proxy and reissue the request, and add 403, 407, and 429 to RETRY_HTTP_CODES so soft bans trigger the retry path. For response-body bans (CAPTCHA HTML behind a 200), add a custom process_response hook that inspects the body and raises IgnoreRequest or returns a fresh request with a new proxy.

Run scrapy shell -s HTTPPROXY_ENABLED=True "https://httpbin.org/ip" after setting http_proxy and https_proxy in the shell. The response body returns the exit IP, which is the simplest sanity check for credential format, scheme, and reachability.

Yes. Use spider-level custom_settings to override PROXY_LIST, PROXY_URL, or the middleware configuration on a per-spider basis. Each spider then loads its own pool when its crawl starts, and the rotating middleware reads the spider-specific setting through from_crawler.

A proxy list is a static collection of host:port:user:pass entries that you assign yourself in code. A rotating gateway is a single endpoint that returns a different exit IP per request, with the rotation logic handled by the provider. Lists give you fine control; gateways trade that control for a much simpler client integration.