Proxy redundancy means holding at least two independent proxy providers so that losing one, whether to an outage, a block wave or a permanent shutdown, degrades your data collection instead of stopping it. In practice it means a thin routing layer that can send any workload to either provider, health checks that detect failure within minutes, and a second vendor held warm at low cost. It is cheap to set up while nothing is broken and very expensive to improvise during an incident.
This guide covers why a single vendor is a concentration risk, the two basic topologies, how to route and fail over, how to keep a standby warm without paying for idle capacity, and the due diligence the second vendor must pass.
Why is a single proxy vendor a concentration risk?
Most teams treat their proxy provider as a utility. It is closer to a critical supplier with an unusual failure profile:
- Outages are correlated. When a provider's gateway fails, every job on it fails at the same time, across every team that uses it.
- Block waves are provider-specific. A target that blocks a vendor's ranges or detects its fingerprint blocks all of your traffic through that vendor together.
- Shutdowns happen with no notice. A top-tier residential provider was seized by federal authorities in July 2026 after its network was found to be built on compromised devices. Its customers lost service immediately, and those with no second vendor spent days rebuilding while their data went stale. Our playbook for when a proxy provider shuts down covers that recovery.
- Commercial changes land all at once. A price change, a new minimum commitment or a policy change on permitted targets affects every workload on the vendor at once, and you have little room to negotiate if you cannot move.
Proxy redundancy does not remove these risks. It caps the damage at the share of traffic that was on the failed provider, for the few minutes it takes to reroute.
Active-active or active-passive?
There are two basic topologies for proxy redundancy, and most mature setups end up with a mix.
| Active-active | Active-passive (warm standby) | |
|---|---|---|
| Traffic in normal operation | Split between both providers | All on the primary; a small canary on the standby |
| Failover | Shift the failed provider's share to the other | Promote the standby to carry everything |
| Readiness of the second path | Proven continuously by real traffic | Proven only by canaries and drills |
| Cost overhead | Little, if both are per-GB | The canary traffic, plus any per-address rentals |
| Operational effort | Two vendors to watch at full scale | One vendor at scale, one at low volume |
| Capacity risk on failover | Lower; the survivor is already carrying load | Higher; the standby must absorb a sudden jump |
Active-active gives you the strongest guarantee that the second path works, because it is always carrying real traffic. It also gives you a live comparison of the two vendors on the same targets. The cost is running two full integrations and watching both.
Active-passive is cheaper and simpler. The risk is that the passive path decays: credentials expire, a configuration drifts, a country you depend on is not provisioned on the standby, and nobody notices until the day it is needed. The canary traffic and drills described below exist to prevent that.
Whichever you choose, confirm that the survivor can absorb the full load. Ask the standby vendor what notice it needs for a sudden increase, and whether per-address products can be provisioned within the time you can tolerate being degraded.
Routing by workload or by percentage
There are two ways to split traffic between providers, the same two used during a migration and described in our guide to switching proxy providers without downtime.
By workload. Each job or target is assigned to a provider. This is easier to reason about: one job's metrics come from one vendor, and a target that works better on one provider can stay there. Failover moves whole workloads.
By percentage. Each request, or each session, is assigned to a provider by weight, for example 80/20. This spreads every target across both vendors, so a provider-specific block wave only affects a fraction of each job. Failover changes the weights.
Workload routing suits active-passive setups and estates with a few large jobs. Percentage routing suits active-active setups and estates with many targets. A common combination is percentage routing for broad scraping and workload pinning for anything with sessions, since a logged-in flow should never switch provider mid-session.
Keeping a warm standby cheaply
A standby that has not carried traffic recently should be treated as broken. Keeping it warm costs less than most teams assume.
- Run a canary workload on it permanently. Pick a small, representative slice of real targets and send it through the standby every few minutes. Validate the content, not just the status code.
- Exercise every configuration you would fail over. If production uses three countries and sticky sessions, the canary must too. An unused country is the most common thing to find broken on the day.
- Prefer usage-based billing for the standby. Per-GB products cost only what the canary uses. Per-address products cost the full monthly rate for each address whether or not it carries traffic, so hold the smallest allocation that proves the path and know how quickly you can add more.
- Hold enough prepaid balance or credit to absorb a failover. A standby with an empty wallet fails at the worst moment. Size the balance for a few days of full production volume and alert on it.
- Drill it. At least once a quarter, take the primary out of rotation during working hours and let the standby carry production for an hour.
Contract structure decides whether this is affordable. A vendor with a monthly minimum or an annual commitment makes a standby expensive, because you pay for capacity you hope never to use. A vendor with no minimum and pay-as-you-go billing costs you only the canary traffic.
Health checks and automatic failover
Failover is only as fast as detection. Two kinds of signal are needed:
- Passive signals from real traffic. Connection failures, proxy authentication errors (407), gateway errors and timeouts on production requests, tracked per provider.
- Active probes. A request every minute or so through each provider to a stable endpoint you control, so the standby is checked even when it carries no traffic.
The hard part is telling a provider failure from a target block. A 403 or a captcha page is usually the target reacting to you, and failing over will not fix it; a refused connection, a failed CONNECT tunnel or a 407 is the provider. Count only provider-side failures toward failover, and track target-side blocks separately; our guide to proxy 403 and 429 errors covers the distinction, and proxy monitoring covers the metrics worth alerting on.
A minimal failover selector in Python, using Requests, looks like this. Each provider's URL comes from an environment variable shaped like http://USERNAME:[email protected]:PORT, with the other vendor's gateway in the other variable.
import os
import threading
import time
from dataclasses import dataclass
import requests
PROVIDER_ERRORS = (requests.exceptions.ProxyError, requests.exceptions.ConnectTimeout)
PROVIDER_STATUS = {407, 502, 503, 504}
@dataclass
class Upstream:
name: str
proxy_url: str
failures: int = 0
down_until: float = 0.0
@property
def proxies(self):
return {"http": self.proxy_url, "https": self.proxy_url}
class FailoverSelector:
def __init__(self, upstreams, max_failures=5, cooldown=120.0):
self.upstreams = upstreams
self.max_failures = max_failures
self.cooldown = cooldown
self._lock = threading.Lock()
def select(self):
now = time.monotonic()
with self._lock:
for upstream in self.upstreams:
if upstream.down_until <= now:
return upstream
return min(self.upstreams, key=lambda u: u.down_until)
def record(self, upstream, ok):
with self._lock:
if ok:
upstream.failures = 0
return
upstream.failures += 1
if upstream.failures >= self.max_failures:
upstream.down_until = time.monotonic() + self.cooldown
upstream.failures = 0
selector = FailoverSelector([
Upstream("primary", os.environ["PROXY_URL_PRIMARY"]),
Upstream("secondary", os.environ["PROXY_URL_SECONDARY"]),
])
def fetch(url, **kwargs):
upstream = selector.select()
try:
response = requests.get(url, proxies=upstream.proxies, timeout=30, **kwargs)
except PROVIDER_ERRORS:
selector.record(upstream, ok=False)
raise
selector.record(upstream, ok=response.status_code not in PROVIDER_STATUS)
return response
def probe(upstream, health_url):
try:
response = requests.get(health_url, proxies=upstream.proxies, timeout=10)
ok = response.status_code == 200
except requests.RequestException:
ok = False
selector.record(upstream, ok)
return ok
Upstreams are listed in order of preference. After a run of provider-side failures, an upstream is taken out of rotation for the cooldown period, then becomes eligible again, so traffic fails back to the primary without manual work. If every upstream is down, the one due back soonest is used rather than failing outright. Run probe against each upstream on a timer so that the standby's state is known before it is needed.
This is a sketch, not a finished component. A production version would keep state across processes, emit metrics on every transition, and treat gateway status codes according to each vendor's documentation, since a 502 can come from either the gateway or the target depending on the vendor.
Abstract credentials and configuration
Proxy redundancy only works if job code does not know which vendor it is using. Keep three things out of it:
- Credentials. Store each provider's credentials in a secrets manager and inject them at runtime; our guide to managing proxy credentials covers the patterns.
- Connection syntax. Vendors encode country, session and rotation options differently. Translate a neutral request, such as "Germany, sticky, 10 minutes", into each vendor's format in one place.
- Vendor-specific error handling. Map each vendor's error responses onto your own small set of outcomes (provider failure, target block, success) in the routing layer.
The thin abstraction described in switching proxy providers is the same layer; building it for redundancy also makes any future migration a configuration change.
The second vendor needs the same due diligence
Redundancy adds a supplier, and the second supplier's risk becomes yours the moment traffic flows through it. A standby chosen in a hurry, or chosen because it was cheap to hold, can reintroduce exactly the risk it was meant to reduce. A backup that turns out to run on compromised devices does not protect you from the next seizure; it makes you a customer of it.
Apply the same review to both vendors before either carries production traffic:
- sourcing evidence proportionate to the proxy type, including a sample trace (see verifying a provider's IP sourcing);
- a DPA, sub-processor list and retention terms your privacy team accepts;
- security questionnaire answers and audit status;
- customer screening (KYC) and an acceptable use policy that keeps abusive co-customers off the network;
- the commercial terms that decide whether the standby is affordable: minimums, commitments, and how quickly capacity can grow.
The proxy provider due diligence checklist lists the questions in a form you can send to both vendors at once.
Adding ProxyForge as a second provider
ProxyForge bills pay-as-you-go from a prepaid wallet with no monthly minimum, and you can order from 1 GB or a single address, so holding us as a warm standby costs the canary traffic and little else. ISP and datacenter addresses have a paid trial. Our migration process supports the same pattern from the other direction: we mirror your current endpoint structure, session syntax and auth format on our gateway, and you run both providers in parallel on published rates without cancelling anything.
Our sourcing page sets out the supply-chain policy and twice-yearly independent audit, so the second vendor in your estate can pass the same review as the first.