Airflow proxy configuration comes down to three choices: where the proxy credential is stored, which tasks use the proxy, and how retries and concurrency are limited so the pipeline does not overwhelm the proxy or the target. Keep the credential in an Airflow Connection, backed by a secrets backend in production, apply the proxy per task rather than to the whole worker, and use a pool to cap how many tasks send traffic through the proxy at once. Retries should be aligned with the backoff your HTTP client already does, not stacked on top of it.
The examples target Airflow 3 and the Task SDK imports, with notes where Airflow 2 differs. The proxy connection string has the shape http://USERNAME:[email protected]:PORT; the provider's dashboard generates the exact value for your country and session mode.
Where to store proxy credentials in Airflow
Airflow gives you three places to put a secret. They are not equivalent:
| Option | Masked in task logs | Visible in the UI | Suitable for a proxy credential |
|---|---|---|---|
| Variable | Only if the Variable's name contains a sensitive keyword | Yes, to anyone who can view Variables | No |
| Connection in the metadata database | Password field always; extra keys only with sensitive names | Password hidden | Yes |
| Connection from a secrets backend (Vault, AWS Secrets Manager, and others) | Same as a Connection | Not listed in the UI | Yes, preferred in production |
A Connection is the right container because it has a password field that Airflow's secrets masker always redacts from task logs. With a secrets backend configured, Airflow looks up Connections in your external store before the metadata database, so the credential lives in the same place as every other production secret and rotating it needs no Airflow change. Managing proxy credentials in Vault and AWS Secrets Manager covers the store side, including access policies and rotation.
For development, add the Connection from the CLI, reading values from the environment so they do not land in shell history:
airflow connections add scraping_proxy \
--conn-type http \
--conn-host gateway.proxyforge.io \
--conn-port "$PROXY_PORT" \
--conn-login "$PROXY_USERNAME" \
--conn-password "$PROXY_PASSWORD"
Avoid a worker-wide HTTP_PROXY and HTTPS_PROXY. Every library in the worker that honors them will use the proxy: remote logging to object storage, cloud SDK calls from other operators, and, in Airflow 3, the Task SDK's own HTTP calls from the task to the API server. Each of those then needs a NO_PROXY exemption, and forgetting one produces failures that look unrelated to the proxy.
Setting the Airflow proxy per task with @task
The most predictable pattern is a TaskFlow task that reads the Connection at run time and builds a Requests session with the proxy set explicitly. Nothing is read at DAG parse time, and only this task's traffic uses the proxy:
from datetime import datetime, timedelta
from urllib.parse import quote
import requests
from airflow.exceptions import AirflowFailException
from airflow.sdk import BaseHook, dag, task
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
PROXY_CONN_ID = "scraping_proxy"
def proxy_session() -> requests.Session:
conn = BaseHook.get_connection(PROXY_CONN_ID)
user = quote(conn.login, safe="")
password = quote(conn.password, safe="")
proxy_url = f"http://{user}:{password}@{conn.host}:{conn.port}"
retry = Retry(
total=3,
backoff_factor=1.0,
status_forcelist=(429, 502, 503, 504),
allowed_methods=frozenset({"GET", "HEAD"}),
respect_retry_after_header=True,
)
session = requests.Session()
session.trust_env = False
session.proxies = {"http": proxy_url, "https": proxy_url}
session.mount("https://", HTTPAdapter(max_retries=retry))
session.mount("http://", HTTPAdapter(max_retries=retry))
return session
@dag(
schedule="@daily",
start_date=datetime(2026, 9, 1),
catchup=False,
default_args={
"retries": 3,
"retry_delay": timedelta(minutes=5),
"retry_exponential_backoff": 2.0,
"max_retry_delay": timedelta(minutes=30),
},
tags=["scraping"],
)
def competitor_prices():
@task
def list_batches() -> list[list[str]]:
urls = [f"https://www.example.com/products?page={page}" for page in range(1, 201)]
return [urls[i : i + 25] for i in range(0, len(urls), 25)]
@task(pool="scraping_proxy", execution_timeout=timedelta(minutes=20))
def fetch_batch(urls: list[str]) -> list[dict]:
session = proxy_session()
results = []
for url in urls:
response = session.get(url, timeout=(5, 30))
if response.status_code == 407:
raise AirflowFailException("Proxy rejected the credential; check the scraping_proxy connection")
if response.status_code >= 500 or response.status_code == 429:
response.raise_for_status()
results.append({"url": url, "status": response.status_code, "bytes": len(response.content)})
return results
fetch_batch.expand(urls=list_batches())
competitor_prices()
A few details are deliberate:
- Batches, not one task per URL. Each task instance carries scheduling overhead. Mapping over batches of 25 keeps the number of task instances manageable, while the pool still bounds concurrency.
- The URL is built from separate fields and percent-encoded. Passwords containing
@,:or/otherwise break the URL. trust_env = Falseso that a proxy variable set on the worker later cannot override the explicit setting.- The assembled URL is never logged. Airflow masks the stored password, but the percent-encoded form inside a URL is a different string and may not be recognized.
On Airflow 2, import dag and task from airflow.decorators and BaseHook from airflow.hooks.base; the rest is the same. The schedule argument has been available since Airflow 2.4.
For background on the Requests settings themselves, including why both the http and https keys are needed, see setting up a Python Requests proxy.
HttpOperator and HttpHook: proxies in the Connection extra
If you call a single HTTP API rather than crawl pages, HttpOperator from apache-airflow-providers-http may be simpler. Here the Connection points at the target API, and the proxy goes into that Connection's extra field as a Requests-style proxies dictionary:
{
"proxies": {
"http": "http://USERNAME:[email protected]:PORT",
"https": "http://USERNAME:[email protected]:PORT"
}
}
HttpHook, which the operator uses, reads proxies (or proxy) from the extra and applies it to its session, as described in the provider's HTTP connection documentation. Three cautions before relying on it:
- It needs provider 4.9.0 or later. Support for request parameters such as
proxiesin the extra was added in 4.9.0. Before that, every extra key was sent as an HTTP request header, which would send your proxy credential to the target. Check the installed version withpip show apache-airflow-providers-http. - Unrecognized keys still become headers. The hook removes the keys it knows (
proxies,proxy,verify,stream,certand a few others) and applies whatever remains as request headers. A misspelledproxyswould be sent to the target along with its credential. - The extra is not masked by default. Airflow redacts extra keys only when their names contain a sensitive keyword, and
proxiesdoes not. Add it toAIRFLOW__CORE__SENSITIVE_VAR_CONN_NAMES, then confirm in a test task's log that the value is redacted.
With that in place, the operator itself needs no proxy arguments:
from datetime import timedelta
from airflow.providers.http.operators.http import HttpOperator
fetch_prices = HttpOperator(
task_id="fetch_prices",
http_conn_id="pricing_api_via_proxy",
endpoint="v1/prices",
method="GET",
pool="scraping_proxy",
retries=3,
retry_delay=timedelta(minutes=5),
)
Define the operator inside a DAG as usual. Note that HttpOperator defaults to POST, and that SimpleHttpOperator was removed in provider 5.0.0.
Aligning Airflow retries with proxy backoff
A pipeline that fetches through an Airflow proxy usually has two retry layers, and they should do different jobs:
| Layer | Timescale | Handles | Example setting |
|---|---|---|---|
HTTP client (urllib3 Retry) |
Seconds | A single 429 with Retry-After, a dropped connection, a brief 502 |
3 attempts, backoff factor 1.0 |
| Airflow task retries | Minutes | A target rate-limiting the whole run, a gateway incident, a network outage | 3 retries, 5 minutes doubling to a 30-minute cap |
The layers multiply. A request that always fails is attempted up to four times by urllib3 in each of four task attempts: sixteen requests through the proxy, all billed. Keep both layers small, and make sure execution_timeout covers the worst case of the inner layer so the task is not killed mid-backoff.
Match the outer delay to what you are waiting out. A target that returns 429 for a whole run needs minutes to recover, not seconds, so a retry_delay of five minutes with exponential growth is a sensible start. On Airflow 3's Task SDK, retry_exponential_backoff is a multiplier (2.0 doubles the delay each time); on Airflow 2 it is a boolean, and 2.0 is treated as true, so the setting above works on both.
Do not retry what cannot succeed. A 407 means the credential is wrong or revoked; a 404 means the page does not exist. The example raises AirflowFailException for the 407 so the task fails at once and the alert names the actual problem. Proxy 403 and 429 errors explains how to tell a target-side block from a proxy-side failure, which decides which layer should handle it.
If the pipeline uses sticky sessions, remember that a task retry minutes later will usually get a different exit address. Design multi-step flows so a retry restarts the whole sequence, not the middle of it. Rotating vs sticky proxies covers when each mode fits.
Capping concurrency with pools
Dynamic task mapping makes it easy to launch hundreds of fetch tasks at once, and Airflow will run as many as your executor has capacity for. A pool is the control that stops that from turning into hundreds of simultaneous connections through the proxy to one site. Create it once:
airflow pools set scraping_proxy 8 "Concurrent tasks sending traffic through the scraping proxy"
Every task with pool="scraping_proxy" then waits for a free slot, across all DAGs. That global scope is the point: two DAGs scraping on the same schedule share one budget, rather than each assuming it has the proxy to itself. The pools documentation covers the scheduling details.
Size the pool from the constraints, not from worker capacity:
- The number of concurrent connections your proxy account allows, if the provider sets a limit. Ask.
- The request rate each target tolerates before returning 429s.
- How many requests each task makes at a time. A task that fetches with four threads should take
pool_slots=4, so the pool counts connections rather than tasks.
Use max_active_tis_per_dag on a task to limit one DAG further without affecting others. And since per-GB traffic is usually the largest cost of a scraping pipeline, pair concurrency limits with the techniques in reducing proxy bandwidth.
Using ProxyForge in Airflow
ProxyForge fits either pattern above: store the connection string the dashboard generates, or its username, password, host and port, in an Airflow Connection or your secrets backend. Every line accepts HTTP, HTTPS and SOCKS5 with username and password or IP allowlist authentication, so workers leaving from fixed NAT addresses can drop the credential entirely.
For recurring collection such as price monitoring, per-GB residential proxies suit pages that score IP reputation, and per-address datacenter proxies suit high-volume APIs. Billing is pay-as-you-go from a prepaid wallet with no monthly minimum; see the pricing page for current rates.