Skip to content
ProxyForge

Airflow proxy configuration for data pipelines

ProxyForge engineeringUpdated 7 min read

Airflow proxy configuration comes down to three choices: where the proxy credential is stored, which tasks use the proxy, and how retries and concurrency are limited so the pipeline does not overwhelm the proxy or the target. Keep the credential in an Airflow Connection, backed by a secrets backend in production, apply the proxy per task rather than to the whole worker, and use a pool to cap how many tasks send traffic through the proxy at once. Retries should be aligned with the backoff your HTTP client already does, not stacked on top of it.

The examples target Airflow 3 and the Task SDK imports, with notes where Airflow 2 differs. The proxy connection string has the shape http://USERNAME:[email protected]:PORT; the provider's dashboard generates the exact value for your country and session mode.

Where to store proxy credentials in Airflow

Airflow gives you three places to put a secret. They are not equivalent:

Option Masked in task logs Visible in the UI Suitable for a proxy credential
Variable Only if the Variable's name contains a sensitive keyword Yes, to anyone who can view Variables No
Connection in the metadata database Password field always; extra keys only with sensitive names Password hidden Yes
Connection from a secrets backend (Vault, AWS Secrets Manager, and others) Same as a Connection Not listed in the UI Yes, preferred in production

A Connection is the right container because it has a password field that Airflow's secrets masker always redacts from task logs. With a secrets backend configured, Airflow looks up Connections in your external store before the metadata database, so the credential lives in the same place as every other production secret and rotating it needs no Airflow change. Managing proxy credentials in Vault and AWS Secrets Manager covers the store side, including access policies and rotation.

For development, add the Connection from the CLI, reading values from the environment so they do not land in shell history:

airflow connections add scraping_proxy \
  --conn-type http \
  --conn-host gateway.proxyforge.io \
  --conn-port "$PROXY_PORT" \
  --conn-login "$PROXY_USERNAME" \
  --conn-password "$PROXY_PASSWORD"

Avoid a worker-wide HTTP_PROXY and HTTPS_PROXY. Every library in the worker that honors them will use the proxy: remote logging to object storage, cloud SDK calls from other operators, and, in Airflow 3, the Task SDK's own HTTP calls from the task to the API server. Each of those then needs a NO_PROXY exemption, and forgetting one produces failures that look unrelated to the proxy.

Setting the Airflow proxy per task with @task

The most predictable pattern is a TaskFlow task that reads the Connection at run time and builds a Requests session with the proxy set explicitly. Nothing is read at DAG parse time, and only this task's traffic uses the proxy:

from datetime import datetime, timedelta
from urllib.parse import quote

import requests
from airflow.exceptions import AirflowFailException
from airflow.sdk import BaseHook, dag, task
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

PROXY_CONN_ID = "scraping_proxy"


def proxy_session() -> requests.Session:
    conn = BaseHook.get_connection(PROXY_CONN_ID)
    user = quote(conn.login, safe="")
    password = quote(conn.password, safe="")
    proxy_url = f"http://{user}:{password}@{conn.host}:{conn.port}"

    retry = Retry(
        total=3,
        backoff_factor=1.0,
        status_forcelist=(429, 502, 503, 504),
        allowed_methods=frozenset({"GET", "HEAD"}),
        respect_retry_after_header=True,
    )
    session = requests.Session()
    session.trust_env = False
    session.proxies = {"http": proxy_url, "https": proxy_url}
    session.mount("https://", HTTPAdapter(max_retries=retry))
    session.mount("http://", HTTPAdapter(max_retries=retry))
    return session


@dag(
    schedule="@daily",
    start_date=datetime(2026, 9, 1),
    catchup=False,
    default_args={
        "retries": 3,
        "retry_delay": timedelta(minutes=5),
        "retry_exponential_backoff": 2.0,
        "max_retry_delay": timedelta(minutes=30),
    },
    tags=["scraping"],
)
def competitor_prices():
    @task
    def list_batches() -> list[list[str]]:
        urls = [f"https://www.example.com/products?page={page}" for page in range(1, 201)]
        return [urls[i : i + 25] for i in range(0, len(urls), 25)]

    @task(pool="scraping_proxy", execution_timeout=timedelta(minutes=20))
    def fetch_batch(urls: list[str]) -> list[dict]:
        session = proxy_session()
        results = []
        for url in urls:
            response = session.get(url, timeout=(5, 30))
            if response.status_code == 407:
                raise AirflowFailException("Proxy rejected the credential; check the scraping_proxy connection")
            if response.status_code >= 500 or response.status_code == 429:
                response.raise_for_status()
            results.append({"url": url, "status": response.status_code, "bytes": len(response.content)})
        return results

    fetch_batch.expand(urls=list_batches())


competitor_prices()

A few details are deliberate:

  • Batches, not one task per URL. Each task instance carries scheduling overhead. Mapping over batches of 25 keeps the number of task instances manageable, while the pool still bounds concurrency.
  • The URL is built from separate fields and percent-encoded. Passwords containing @, : or / otherwise break the URL.
  • trust_env = False so that a proxy variable set on the worker later cannot override the explicit setting.
  • The assembled URL is never logged. Airflow masks the stored password, but the percent-encoded form inside a URL is a different string and may not be recognized.

On Airflow 2, import dag and task from airflow.decorators and BaseHook from airflow.hooks.base; the rest is the same. The schedule argument has been available since Airflow 2.4.

For background on the Requests settings themselves, including why both the http and https keys are needed, see setting up a Python Requests proxy.

HttpOperator and HttpHook: proxies in the Connection extra

If you call a single HTTP API rather than crawl pages, HttpOperator from apache-airflow-providers-http may be simpler. Here the Connection points at the target API, and the proxy goes into that Connection's extra field as a Requests-style proxies dictionary:

{
  "proxies": {
    "http": "http://USERNAME:[email protected]:PORT",
    "https": "http://USERNAME:[email protected]:PORT"
  }
}

HttpHook, which the operator uses, reads proxies (or proxy) from the extra and applies it to its session, as described in the provider's HTTP connection documentation. Three cautions before relying on it:

  1. It needs provider 4.9.0 or later. Support for request parameters such as proxies in the extra was added in 4.9.0. Before that, every extra key was sent as an HTTP request header, which would send your proxy credential to the target. Check the installed version with pip show apache-airflow-providers-http.
  2. Unrecognized keys still become headers. The hook removes the keys it knows (proxies, proxy, verify, stream, cert and a few others) and applies whatever remains as request headers. A misspelled proxys would be sent to the target along with its credential.
  3. The extra is not masked by default. Airflow redacts extra keys only when their names contain a sensitive keyword, and proxies does not. Add it to AIRFLOW__CORE__SENSITIVE_VAR_CONN_NAMES, then confirm in a test task's log that the value is redacted.

With that in place, the operator itself needs no proxy arguments:

from datetime import timedelta

from airflow.providers.http.operators.http import HttpOperator

fetch_prices = HttpOperator(
    task_id="fetch_prices",
    http_conn_id="pricing_api_via_proxy",
    endpoint="v1/prices",
    method="GET",
    pool="scraping_proxy",
    retries=3,
    retry_delay=timedelta(minutes=5),
)

Define the operator inside a DAG as usual. Note that HttpOperator defaults to POST, and that SimpleHttpOperator was removed in provider 5.0.0.

Aligning Airflow retries with proxy backoff

A pipeline that fetches through an Airflow proxy usually has two retry layers, and they should do different jobs:

Layer Timescale Handles Example setting
HTTP client (urllib3 Retry) Seconds A single 429 with Retry-After, a dropped connection, a brief 502 3 attempts, backoff factor 1.0
Airflow task retries Minutes A target rate-limiting the whole run, a gateway incident, a network outage 3 retries, 5 minutes doubling to a 30-minute cap

The layers multiply. A request that always fails is attempted up to four times by urllib3 in each of four task attempts: sixteen requests through the proxy, all billed. Keep both layers small, and make sure execution_timeout covers the worst case of the inner layer so the task is not killed mid-backoff.

Match the outer delay to what you are waiting out. A target that returns 429 for a whole run needs minutes to recover, not seconds, so a retry_delay of five minutes with exponential growth is a sensible start. On Airflow 3's Task SDK, retry_exponential_backoff is a multiplier (2.0 doubles the delay each time); on Airflow 2 it is a boolean, and 2.0 is treated as true, so the setting above works on both.

Do not retry what cannot succeed. A 407 means the credential is wrong or revoked; a 404 means the page does not exist. The example raises AirflowFailException for the 407 so the task fails at once and the alert names the actual problem. Proxy 403 and 429 errors explains how to tell a target-side block from a proxy-side failure, which decides which layer should handle it.

If the pipeline uses sticky sessions, remember that a task retry minutes later will usually get a different exit address. Design multi-step flows so a retry restarts the whole sequence, not the middle of it. Rotating vs sticky proxies covers when each mode fits.

Capping concurrency with pools

Dynamic task mapping makes it easy to launch hundreds of fetch tasks at once, and Airflow will run as many as your executor has capacity for. A pool is the control that stops that from turning into hundreds of simultaneous connections through the proxy to one site. Create it once:

airflow pools set scraping_proxy 8 "Concurrent tasks sending traffic through the scraping proxy"

Every task with pool="scraping_proxy" then waits for a free slot, across all DAGs. That global scope is the point: two DAGs scraping on the same schedule share one budget, rather than each assuming it has the proxy to itself. The pools documentation covers the scheduling details.

Size the pool from the constraints, not from worker capacity:

  • The number of concurrent connections your proxy account allows, if the provider sets a limit. Ask.
  • The request rate each target tolerates before returning 429s.
  • How many requests each task makes at a time. A task that fetches with four threads should take pool_slots=4, so the pool counts connections rather than tasks.

Use max_active_tis_per_dag on a task to limit one DAG further without affecting others. And since per-GB traffic is usually the largest cost of a scraping pipeline, pair concurrency limits with the techniques in reducing proxy bandwidth.

Using ProxyForge in Airflow

ProxyForge fits either pattern above: store the connection string the dashboard generates, or its username, password, host and port, in an Airflow Connection or your secrets backend. Every line accepts HTTP, HTTPS and SOCKS5 with username and password or IP allowlist authentication, so workers leaving from fixed NAT addresses can drop the credential entirely.

For recurring collection such as price monitoring, per-GB residential proxies suit pages that score IP reputation, and per-address datacenter proxies suit high-volume APIs. Billing is pay-as-you-go from a prepaid wallet with no monthly minimum; see the pricing page for current rates.

FAQ

Related questions

Can I set HTTP_PROXY for the whole Airflow deployment?

You can, but it applies to everything the worker process does, including remote logging, cloud SDK calls and, in Airflow 3, the task's calls to the API server. If you do it, maintain a NO_PROXY list for all internal and cloud endpoints. Scoping the proxy to the tasks that fetch external pages is easier to reason about.

Are Airflow Variables safe for storing a proxy password?

Only partly. Variables are masked in logs only when their name contains a sensitive keyword, and their values are visible to anyone who can view Variables in the UI. A Connection keeps the password in a field that is always masked, and a secrets backend keeps it out of the metadata database altogether.

Why is my Airflow task retrying a request that can never succeed?

Airflow retries any exception unless you signal otherwise. For failures a retry cannot fix, such as a 407 from the proxy or a 404 from the target, raise AirflowFailException so the task fails immediately without using up its retries.

How many pool slots should a scraping pool have?

Start from the lower of two numbers: the concurrency your proxy account and budget allow, and the request rate the target tolerates without returning 429s. Set the pool to that, divided by the number of concurrent requests each task makes, and adjust from observed error rates.

Run it on a network you can account for

Order from 1 GB or 1 IP with no monthly minimum, or talk to an engineer about your workload first.

One business day, from a named engineer.