Skip to content
ProxyForge

aiohttp proxy and httpx proxy setup for async Python

ProxyForge engineeringUpdated 6 min read

An aiohttp proxy is set per request with proxy="http://USERNAME:PASSWORD@host:port", or once for the whole session with ClientSession(proxy=...); httpx takes the same URL as AsyncClient(proxy=...). Both send HTTPS through an HTTP proxy as a CONNECT tunnel and read credentials from the URL. For SOCKS5, aiohttp needs the aiohttp-socks package and httpx needs the httpx[socks] extra.

This guide covers both libraries in turn, then the parts that matter once a crawler runs at volume: connection limits, timeouts, retries with backoff and a bounded-concurrency crawler. Every example was run against a local authenticating HTTP proxy and a SOCKS5 proxy, using Python 3.12, aiohttp 3.14.3, aiohttp-socks 0.12.0 and httpx 0.28.1. Every example reads the connection string from PROXY_URL, shaped like http://USERNAME:[email protected]:PORT. The dashboard generates the exact string for the country and session mode you choose.

aiohttp proxy configuration

Per request and per session

The smallest working aiohttp proxy setup passes the URL on the request:

import asyncio
import os

import aiohttp

PROXY_URL = os.environ["PROXY_URL"]


async def main() -> None:
    timeout = aiohttp.ClientTimeout(total=30, sock_connect=10)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get("https://api.ipify.org?format=json", proxy=PROXY_URL) as response:
            response.raise_for_status()
            print(await response.json())


asyncio.run(main())

Current aiohttp releases also accept proxy= on ClientSession, which becomes the default for every request made through it. That is usually the better shape: one session per proxy identity, created once and reused.

async with aiohttp.ClientSession(proxy=os.environ["PROXY_URL"]) as session:
    async with session.get("https://api.ipify.org?format=json") as response:
        print(await response.json())

aiohttp decodes percent-encoded credentials in the proxy URL, so a password containing @ or : must be written as %40 or %3A. When the proxy rejects the credentials, aiohttp raises ClientHttpProxyError with status 407. It does not retry that, and neither should your code.

proxy_auth and its replacement

Older examples pass the credentials separately with proxy_auth=aiohttp.BasicAuth(user, password). That still works, but aiohttp 3.14 emits a DeprecationWarning for both BasicAuth and proxy_auth, and states they will be removed in aiohttp 4.0. If your credentials arrive as separate values, the forward-compatible form is a Proxy-Authorization header built with encode_basic_auth:

import os

import aiohttp

proxy = os.environ["PROXY_ENDPOINT"]
proxy_headers = {
    "Proxy-Authorization": aiohttp.encode_basic_auth(
        os.environ["PROXY_USERNAME"], os.environ["PROXY_PASSWORD"]
    )
}


async def fetch_ip(session: aiohttp.ClientSession) -> dict:
    async with session.get(
        "https://api.ipify.org?format=json", proxy=proxy, proxy_headers=proxy_headers
    ) as response:
        return await response.json()

aiohttp sends proxy_headers on the CONNECT request for HTTPS targets, which is where the proxy checks them. Credentials embedded in the URL produce no warning, so if you can build a single PROXY_URL, that remains the simplest route.

Environment variables with trust_env

With trust_env=True, aiohttp reads HTTP_PROXY, HTTPS_PROXY and NO_PROXY from the environment:

async with aiohttp.ClientSession(trust_env=True) as session:
    async with session.get("https://api.ipify.org?format=json") as response:
        print(await response.json())

It is off by default in aiohttp, which is the opposite of httpx. That difference catches people who move code between the two: a container that sets HTTPS_PROXY routes httpx traffic through the proxy and leaves aiohttp traffic direct unless you opt in.

SOCKS5 with aiohttp-socks

aiohttp has no SOCKS support of its own. aiohttp-socks supplies a connector:

import asyncio
import os

import aiohttp
from aiohttp_socks import ProxyConnector


async def main() -> None:
    connector = ProxyConnector.from_url(os.environ["SOCKS_PROXY_URL"], rdns=True)
    async with aiohttp.ClientSession(connector=connector) as session:
        async with session.get("https://api.ipify.org?format=json") as response:
            print(await response.json())


asyncio.run(main())

Two details. The connector belongs to the session, so every request in it uses that SOCKS endpoint; use one session per endpoint. And from_url accepts socks5:// but rejects socks5h:// with ValueError: Invalid scheme component. Remote DNS resolution is controlled by rdns=True instead. In our tests the proxy received the hostname rather than an IP address, which is what you want: the exit resolves the name, so geo-routed CDNs answer for the exit's location, not your server's. SOCKS5 vs HTTP proxy covers when SOCKS is worth using at all; for HTTPS scraping it usually is not.

httpx proxy configuration

proxy= and the removal of proxies=

httpx takes a single proxy URL with proxy=:

import asyncio
import os

import httpx


async def main() -> None:
    async with httpx.AsyncClient(proxy=os.environ["PROXY_URL"], timeout=30.0) as client:
        response = await client.get("https://api.ipify.org?format=json")
        response.raise_for_status()
        print(response.json())


asyncio.run(main())

proxy= arrived in httpx 0.26, and httpx 0.28 removed the older proxies= dictionary. Code written against earlier releases fails with TypeError: AsyncClient.__init__() got an unexpected keyword argument 'proxies'. A rejected proxy login raises httpx.ProxyError: 407 Proxy Authentication Required.

Note timeout=30.0. httpx's default timeout is five seconds for each phase, which is short for a proxied request that opens a tunnel and then a TLS session through it, especially on residential exits. Set it explicitly.

httpx also reads HTTP_PROXY, HTTPS_PROXY, ALL_PROXY and NO_PROXY by default, because trust_env defaults to True. If a job must only ever use the proxy you configured in code, pass trust_env=False.

Per-scheme and per-host routing with mounts

mounts maps URL patterns to transports. It replaces the old proxies dictionary and does more: a pattern can match a scheme, a host, or a wildcard domain, and mapping a pattern to None sends matching requests direct.

import os

import httpx

proxy = os.environ["PROXY_URL"]
mounts = {
    "http://": httpx.AsyncHTTPTransport(proxy=proxy),
    "https://": httpx.AsyncHTTPTransport(proxy=proxy),
    "all://*.internal.example.com": None,
}

client = httpx.AsyncClient(mounts=mounts, timeout=30.0)

The most specific pattern wins. Keeping internal APIs, health checks and metadata endpoints off the proxy matters when the proxy is billed per GB. Reducing proxy bandwidth covers the other savings.

SOCKS5 with httpx[socks]

Install the extra with pip install "httpx[socks]", then pass a SOCKS URL as the proxy:

async with httpx.AsyncClient(proxy=os.environ["SOCKS_PROXY_URL"], timeout=30.0) as client:
    print((await client.get("https://api.ipify.org?format=json")).json())

httpx 0.28.1 accepted both socks5:// and socks5h://, and in both cases our proxy received the hostname, so resolution happened at the exit.

Connection limits, timeouts and keep-alive

Both libraries pool connections, and the pool is the concurrency control you already have.

Setting aiohttp httpx
Total connections TCPConnector(limit=100) (default 100) Limits(max_connections=100) (default 100)
Per target host TCPConnector(limit_per_host=0) (0 means unlimited) none; bound it yourself
Idle keep-alive TCPConnector(keepalive_timeout=...) (default 15 s) Limits(max_keepalive_connections=20, keepalive_expiry=5.0) (defaults)
Overall timeout ClientTimeout(total=...) (default 300 s) Timeout(5.0) (default, per phase)
Connect timeout ClientTimeout(sock_connect=...) (default 30 s) Timeout(..., connect=...)
Waiting for a pooled connection counted in total Timeout(..., pool=...)
New connection per request TCPConnector(force_close=True) Limits(max_keepalive_connections=0)

A configured httpx client for proxy work looks like this:

import os

import httpx

limits = httpx.Limits(max_connections=20, max_keepalive_connections=10)
timeout = httpx.Timeout(30.0, connect=5.0, pool=10.0)

client = httpx.AsyncClient(proxy=os.environ["PROXY_URL"], limits=limits, timeout=timeout)

Keep-alive has a consequence specific to proxies. A pooled connection to an HTTPS site is a tunnel through the gateway, and most gateways fix the exit IP when the tunnel opens. In our tests, two requests on one httpx client produced a single CONNECT: both requests left from the same exit. That is what you want on a sticky session, where a login and the pages behind it should share an address. On a rotating endpoint where each request should get a new exit, disable reuse with the last row of the table; both settings produced one CONNECT per request. Rotating vs sticky proxies covers choosing between the modes.

Retries with backoff

Neither library retries on status codes. httpx's transport has a retries argument, but it only retries failed connection attempts (ConnectError and ConnectTimeout), not 429s, 5xx responses or read timeouts. A small helper covers the rest:

import asyncio
import random

import httpx

RETRY_STATUSES = {429, 500, 502, 503, 504}


async def get_with_retry(client: httpx.AsyncClient, url: str, attempts: int = 4) -> httpx.Response:
    for attempt in range(1, attempts + 1):
        try:
            response = await client.get(url)
        except (httpx.ConnectError, httpx.ReadTimeout, httpx.RemoteProtocolError):
            if attempt == attempts:
                raise
        else:
            if response.status_code not in RETRY_STATUSES or attempt == attempts:
                return response
            retry_after = response.headers.get("Retry-After", "")
            if retry_after.isdigit():
                await asyncio.sleep(int(retry_after))
                continue
        await asyncio.sleep(0.5 * 2 ** attempt + random.uniform(0, 0.5))
    raise RuntimeError("unreachable")

It honors Retry-After when the server sends a number of seconds and otherwise backs off exponentially with jitter, so parallel workers do not retry in lockstep. It deliberately leaves out httpx.ProxyError, which is what a 407 raises: wrong credentials do not become right on the next attempt, and a retry loop around them only adds load to the gateway. The same goes for 403 from the target, which is usually a block rather than a transient fault. Proxy 403 and 429 errors covers how to tell them apart. Only wrap idempotent requests.

A bounded-concurrency crawler

The crawler below combines the pieces for aiohttp. A semaphore bounds how many URLs are in flight, the connector bounds sockets, and the retry logic lives inside the semaphore so a page waiting out a backoff still counts against the limit. That keeps a burst of 429s from turning into a burst of new requests.

import asyncio
import os
import random

import aiohttp

PROXY_URL = os.environ["PROXY_URL"]
CONCURRENCY = 10
RETRY_STATUSES = {429, 500, 502, 503, 504}


async def fetch(session: aiohttp.ClientSession, url: str, attempts: int = 4) -> str:
    for attempt in range(1, attempts + 1):
        try:
            async with session.get(url) as response:
                if response.status in RETRY_STATUSES and attempt < attempts:
                    retry_after = response.headers.get("Retry-After", "")
                    delay = float(retry_after) if retry_after.isdigit() else 0.5 * 2 ** attempt
                    await asyncio.sleep(delay + random.uniform(0, 0.5))
                    continue
                response.raise_for_status()
                return await response.text()
        except (aiohttp.ClientConnectionError, asyncio.TimeoutError):
            if attempt == attempts:
                raise
            await asyncio.sleep(0.5 * 2 ** attempt + random.uniform(0, 0.5))
    raise RuntimeError("unreachable")


async def crawl(urls: list[str]) -> dict[str, str | BaseException]:
    semaphore = asyncio.Semaphore(CONCURRENCY)
    connector = aiohttp.TCPConnector(limit=CONCURRENCY)
    timeout = aiohttp.ClientTimeout(total=60, sock_connect=10, sock_read=30)

    async with aiohttp.ClientSession(
        connector=connector, timeout=timeout, proxy=PROXY_URL
    ) as session:

        async def worker(url: str) -> str:
            async with semaphore:
                return await fetch(session, url)

        results = await asyncio.gather(*(worker(u) for u in urls), return_exceptions=True)
    return dict(zip(urls, results))


if __name__ == "__main__":
    pages = [f"https://example.com/page/{n}" for n in range(1, 51)]
    for url, result in asyncio.run(crawl(pages)).items():
        if isinstance(result, BaseException):
            print("error", url, repr(result))
        else:
            print("ok", url, len(result))

Against a local test server that answered some pages with 503 and 429 before succeeding, all 50 pages completed, with the extra attempts visible in the proxy's log. A 407 from the proxy surfaces as ClientHttpProxyError, which is a ClientResponseError rather than a connection error, so it is not retried.

return_exceptions=True keeps one failed page from cancelling the rest; inspect the results rather than assuming success. For large URL lists, feed a fixed pool of worker tasks from an asyncio.Queue instead of creating one task per URL, so memory stays flat. If you outgrow a hand-written crawler, the Scrapy rotating proxies guide covers the framework route, and the Python Requests proxy guide covers the synchronous equivalent.

Set CONCURRENCY per target site, not per process, and start low. The proxy will carry far more than most sites tolerate from one client.

Running async Python on ProxyForge

Every ProxyForge line accepts HTTP, HTTPS and SOCKS5 with username and password or IP allowlist authentication, so each pattern above works unchanged whichever line you use; only PROXY_URL differs. For consumer sites that score IP reputation, the residential line rotates per request or holds a sticky session for up to 60 minutes. For high-volume calls to tolerant endpoints, datacenter addresses are dedicated and announced from our own ASN.

Keep the credentials out of the code: load PROXY_URL from your secrets store at start-up, as the proxy credentials guide describes. Billing is pay-as-you-go from a prepaid wallet with published rates on the pricing page.

FAQ

Related questions

Why does httpx say AsyncClient got an unexpected keyword argument 'proxies'?

httpx 0.28 removed the proxies argument. Pass a single URL with proxy=, or use mounts with a transport per URL pattern when different schemes or hosts need different proxies.

Does aiohttp support SOCKS5 proxies?

Not natively. Install aiohttp-socks and pass its ProxyConnector to the ClientSession. The connector applies to the whole session, so use one session per SOCKS endpoint.

Why do all my requests through a rotating proxy get the same IP in aiohttp?

The session keeps connections alive, and a reused connection is a reused CONNECT tunnel, which usually keeps its exit. Create the connector with force_close=True, or use one short-lived session per request, when each request needs a fresh exit.

How many concurrent requests should I send through a proxy with asyncio?

Start low, around 5 to 20 per target site, and raise it while watching 429 and 403 rates. The target's tolerance is usually the constraint, not the proxy or the event loop.

Run it on a network you can account for

Order from 1 GB or 1 IP with no monthly minimum, or talk to an engineer about your workload first.

One business day, from a named engineer.