Skip to content
ProxyForge

Scrapy proxy setup: rotation, retries and 429 handling

ProxyForge engineeringUpdated 7 min read

A Scrapy proxy is set per request with request.meta['proxy'], and the built-in HttpProxyMiddleware does the rest: it strips the credentials out of the proxy URL, sends them as a Proxy-Authorization header, and tunnels HTTPS requests through CONNECT. Rotation needs no extra package. Point requests at a rotating gateway, or write a short downloader middleware that assigns a proxy per request or per domain and picks a new one on retry.

This guide builds that middleware, then covers the settings that decide whether a proxied crawl is fast, polite and cheap: concurrency, AutoThrottle, timeouts and retries, including what Scrapy does and does not do with a 429 Too Many Requests. Code targets Scrapy 2.13 and later. Connection strings are read from the environment and look like http://USERNAME:[email protected]:PORT; your dashboard generates the exact string for the country and session mode you choose.

Setting a Scrapy proxy per request with meta['proxy']

The smallest working setup is a spider that sets meta['proxy'] on each request it creates:

import os

import scrapy


class CatalogSpider(scrapy.Spider):
    name = "catalog"

    async def start(self):
        yield scrapy.Request(
            "https://httpbin.org/ip",
            meta={"proxy": os.environ["PROXY_URL"]},
        )

    def parse(self, response):
        yield {"url": response.url, "body": response.text}

HttpProxyMiddleware is enabled by default. When it sees credentials in the proxy URL, it moves them into a Proxy-Authorization header and rewrites meta['proxy'] without them, so they do not appear in logs of the request meta. Credentials must be percent-encoded if they contain characters such as @ or :; the dashboard's connection string already is.

Two things to know about the default behavior:

  • Environment variables apply when meta does not. The middleware reads http_proxy and https_proxy at start-up and uses them for any request without meta['proxy']. That is convenient for a quick run and surprising in a container image that happens to set them.
  • Follow-up requests do not inherit meta. response.follow() creates a new request with fresh meta. Set the proxy on every request, which is exactly what a middleware is for.

The async def start() method is the Scrapy 2.13+ replacement for start_requests(); on older versions, use start_requests() with the same body.

A downloader middleware for rotation and sticky sessions

A middleware keeps spiders free of Scrapy proxy logic. This one takes a list of proxy URLs from PROXY_URLS and assigns one to every request, keyed by domain so that all requests to one site share an exit, and moves to a different entry when Scrapy retries a request.

import hashlib
import os
from urllib.parse import urlsplit


class ProxyPoolMiddleware:
    def __init__(self, proxy_urls):
        if not proxy_urls:
            raise ValueError("No proxy URLs configured")
        self.proxy_urls = proxy_urls

    @classmethod
    def from_crawler(cls, crawler):
        return cls(os.environ.get("PROXY_URLS", "").split())

    def _pick(self, key):
        digest = hashlib.sha256(key.encode()).digest()
        return self.proxy_urls[int.from_bytes(digest[:4], "big") % len(self.proxy_urls)]

    def process_request(self, request, spider=None):
        if "proxy" in request.meta and "proxy_session" not in request.meta:
            return None
        session = request.meta.setdefault("proxy_session", urlsplit(request.url).hostname or "")
        attempt = request.meta.get("retry_times", 0)
        request.meta["proxy"] = self._pick(f"{session}:{attempt}")
        return None

How it behaves:

  1. A request with its own meta['proxy'] is left alone. Spiders can still override the pool for a specific request.
  2. Everything else gets a proxy chosen by proxy_session, which defaults to the hostname. A stable hash means every request to shop.example.com uses the same entry, which is what you want when each entry is a sticky session or a dedicated address.
  3. A spider can set its own proxy_session. For a login flow, put an account identifier in meta['proxy_session'] and pass it to follow-up requests; every step of that account's session then shares an exit.
  4. Retries move. RetryMiddleware copies the request and increments retry_times, so the retried copy hashes to a different entry. A retry after a 429 or a timeout goes out through a different session instead of the one that just failed.

What goes into PROXY_URLS depends on the line. For residential or mobile, it can be a single rotating gateway string (the gateway rotates for you, and the middleware still gives you the retry and override hooks), or several sticky-session strings generated in the dashboard. For ISP or datacenter, it is one entry per dedicated address. Whether a target needs rotation or stickiness is the subject of rotating vs sticky proxies.

Register it before HttpProxyMiddleware, which sits at priority 750, so the proxy is set before credentials are processed. The spider=None signature works on current Scrapy, which no longer passes spider to middleware methods that do not require it, and on older releases that always pass it.

Settings that matter behind a proxy

DOWNLOADER_MIDDLEWARES = {
    "myproject.middlewares.ProxyPoolMiddleware": 350,
    "myproject.middlewares.SlowDownOn429Middleware": 560,
}

CONCURRENT_REQUESTS = 32
CONCURRENT_REQUESTS_PER_DOMAIN = 4
DOWNLOAD_TIMEOUT = 30

AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1.0
AUTOTHROTTLE_MAX_DELAY = 30.0
AUTOTHROTTLE_TARGET_CONCURRENCY = 2.0

RETRY_ENABLED = True
RETRY_TIMES = 3
RETRY_HTTP_CODES = [429, 500, 502, 503, 504, 522, 524, 408]

The reasoning behind each group:

  • CONCURRENT_REQUESTS_PER_DOMAIN is the politeness limit; CONCURRENT_REQUESTS is the throughput limit. A proxy pool does not change what a single site will tolerate. Spreading requests across more exits hides your concurrency from per-IP limits, not from per-account, per-session or site-wide ones.
  • DOWNLOAD_TIMEOUT defaults to 180 seconds. Behind a proxy that is far too long: a stalled exit holds a concurrency slot for three minutes. Thirty seconds is a better starting point for most HTML pages.
  • AutoThrottle adjusts the per-domain delay toward a target concurrency based on observed latency. It is a good default for crawls where you do not know the site's tolerance. The AutoThrottle documentation describes the algorithm.
  • RETRY_HTTP_CODES above matches Scrapy's default list. 403 is deliberately absent: a 403 is usually a block that a quick retry will not fix. 407 is also absent, because a proxy authentication failure is a configuration error, not a transient one.
  • Cookies follow the cookiejar, not the proxy. If each proxy_session represents an identity, set meta['cookiejar'] to the same value so cookies and exit IP change together.

Handling 429 responses without hammering the site

Scrapy retries 429 by default, but it does so without waiting: the retried request goes back into the scheduler with a lower priority and is downloaded as soon as a slot is free. RetryMiddleware does not read Retry-After, and AutoThrottle does not slow down on a 429; it only refuses to speed up after non-200 responses. With a single exit, that means retrying straight into the same rate limit.

Two things fix most of it. First, the pool middleware above moves each retry to a different entry, which helps when the limit is per IP. Second, a small middleware can raise the delay of the download slot that received the 429, which helps when the limit is per site:

class SlowDownOn429Middleware:
    def __init__(self, crawler, max_delay):
        self.crawler = crawler
        self.max_delay = max_delay

    @classmethod
    def from_crawler(cls, crawler):
        return cls(crawler, crawler.settings.getfloat("AUTOTHROTTLE_MAX_DELAY", 60.0))

    def process_response(self, request, response, spider=None):
        if response.status == 429:
            slot = self.crawler.engine.downloader.slots.get(request.meta.get("download_slot"))
            if slot is not None:
                slot.delay = min(max(slot.delay * 2, 1.0), self.max_delay)
        return response

It doubles the slot's delay on each 429, up to the AutoThrottle maximum, and AutoThrottle brings the delay back down as healthy responses return. Its priority, 560, places it just above RetryMiddleware at 550, so it sees the 429 before the response is turned into a retry. The download slot is the same internal structure AutoThrottle adjusts; it is stable in practice but not a documented public API, so pin your Scrapy version and re-test on upgrades.

If 429s persist after this, the limit is probably tied to something the proxy does not change: an account, an API key, or a fingerprint. Proxy 403 and 429 errors covers how to tell these cases apart.

Verifying the exit IP and measuring bandwidth

Before a long crawl, confirm the Scrapy proxy pool actually works: request an IP echo endpoint with a few different proxy_session values and check that the addresses differ where you expect them to and match where you expect stickiness.

import scrapy


class ExitIpSpider(scrapy.Spider):
    name = "exit_ip"

    async def start(self):
        for session in ("a", "b", "c"):
            yield scrapy.Request(
                "https://httpbin.org/ip",
                meta={"proxy_session": session},
                dont_filter=True,
                cb_kwargs={"session": session},
            )

    def parse(self, response, session):
        yield {"session": session, "exit": response.json()["origin"]}

dont_filter=True is needed because the duplicate filter would otherwise drop the second and third request to the same URL.

For per-GB lines, the crawl's stats already contain what you need: downloader/response_bytes at the end of every run is the number to compare against your usage. Scrapy fetches only the URLs you request, with no images or scripts, so per-page bandwidth is low compared with browser automation.

When pages need a browser: scrapy-playwright

If a target renders its content with JavaScript, scrapy-playwright hands those requests to a real browser. It does not use meta['proxy']; its documentation lists that as unsupported. Configure the proxy in PLAYWRIGHT_LAUNCH_OPTIONS for the whole browser, or per context in PLAYWRIGHT_CONTEXTS, which maps well onto one sticky session per context. The browser-side details are in the Playwright proxy guide.

Which proxy line fits a Scrapy crawl

Crawl Line Notes
Broad crawls of consumer sites that score IP reputation Residential Rotating per request, or sticky up to 60 minutes per proxy_session
Repeated monitoring of the same pages, logged-in crawls ISP One static, dedicated address per pool entry, billed per address
High-volume crawls of tolerant targets and public datasets Datacenter Dedicated addresses from our own ASN; cost does not grow with page weight
Mobile-specific pages and carrier-dependent content Mobile Rotating, or sticky up to 30 minutes

For price and catalog monitoring specifically, the price monitoring use case describes a typical setup.

Running Scrapy on ProxyForge

Every ProxyForge line accepts username and password or IP allowlist authentication, so HttpProxyMiddleware works with the connection strings from the dashboard as they are, and the pool middleware above works across lines without changes. Because billing draws on a prepaid wallet, a runaway spider can only spend what the wallet holds.

Billing is pay-as-you-go from a prepaid wallet, from 1 GB or 1 IP, with published rates on the pricing page. If you are replacing a provider under an existing Scrapy deployment, the migration process mirrors your current endpoint and session format so your spiders run unchanged against both during the comparison.

FAQ

Related questions

Do I need a third-party package for rotating proxies in Scrapy?

No. Scrapy's built-in HttpProxyMiddleware handles the proxy and its credentials, and a downloader middleware of about twenty lines is enough to assign proxies per request or per domain. A rotating gateway endpoint does the per-request rotation for you.

Does Scrapy read the HTTP_PROXY and HTTPS_PROXY environment variables?

Yes. HttpProxyMiddleware reads them when the crawler starts and applies them to requests that do not set meta['proxy']. A proxy set in request meta always takes precedence over the environment.

Why does my spider still get 429s with AutoThrottle enabled?

AutoThrottle adjusts delay based on response latency and never lowers the delay after a non-200 response, but it does not treat a 429 as a signal to slow down. Lower CONCURRENT_REQUESTS_PER_DOMAIN, or add a middleware that raises the download slot's delay when a 429 arrives.

Can I use a SOCKS5 proxy with Scrapy?

Not with the default download handler, which supports HTTP and HTTPS proxies only. Recent releases include an optional httpx-based handler that can use SOCKS, but Scrapy marks it as not yet recommended for production, so use your provider's HTTP endpoint, which carries HTTPS through a CONNECT tunnel.

Run it on a network you can account for

Order from 1 GB or 1 IP with no monthly minimum, or talk to an engineer about your workload first.

One business day, from a named engineer.