Skip to content
ProxyForge

How to benchmark proxy providers before you buy

ProxyForge engineeringUpdated 6 min read

To benchmark proxy providers fairly, send the same representative set of production targets through each provider at the same concurrency, in the same time windows, with requests randomly interleaved. Count a request as successful only when the response contains the content you needed, and compare validated success rate, block rate, p50 and p95 latency, geo accuracy and cost per successful request per target. Collect enough requests that the differences you see are larger than the noise.

Most published comparisons fail at least one of those conditions, which is why two teams testing the same vendors can reach opposite conclusions. The sections below cover each condition, then give a small Python harness you can adapt.

Design a test that is fair to every provider

A benchmark is only as good as its controls. When you benchmark proxy providers, the goal is that the only variable is the provider.

  • Representative targets. Use the sites and page types you will actually fetch in production, in roughly the same proportions. A test on easy public pages tells you nothing about the login-walled marketplace that is half your volume. Include your hardest target even if it drags every provider's numbers down.
  • Same proxy type and targeting. Compare residential with residential, in the same countries and session modes. A rotating pool against a sticky one is two different experiments.
  • Same concurrency and time windows. Give each provider the same number of simultaneous connections and run them at the same time. Targets change their defenses by hour and day of week; a provider tested on Sunday night against one tested on Monday morning is a comparison of Sundays and Mondays.
  • Randomized interleaving. Shuffle requests across providers and targets rather than running provider A's batch and then provider B's. That spreads any time-dependent effect evenly.
  • Identical clients. Same HTTP client, headers, TLS behavior and timeouts. If production uses a headless browser, benchmark with one.
  • Fresh connections for rotating pools. An HTTP client that reuses a CONNECT tunnel keeps the same exit address for every request on that connection, which quietly turns a rotating test into a sticky one.

Why HTTP 200 is not success

A 200 response can carry a challenge page, a login wall, an empty search template, a consent interstitial or a price in the wrong currency. If success is measured by status code, the provider that returns the most polite-looking block pages wins.

Define success per target as the presence of the content you need: a price element, a result list, a JSON field. Check for a marker that only appears on a real page, and optionally for a minimum body size. Record the result as a separate field from the status code so you can see both.

Detecting blocks and captchas

Classify blocks explicitly, because a block and a timeout have different causes and different fixes:

  • Status codes 403 and 429, and 503 on some targets.
  • Challenge and captcha pages, which often return 200. Match on markers in the body.
  • Redirects to a challenge or login path.
  • Soft blocks: pages that load but show no results, reduced content or a different region.

Our guide to handling proxy 403 and 429 errors covers what each pattern usually means.

Measure latency as percentiles, not averages

Report p50 and p95 latency per provider and per target, measured to the complete response. The average hides the tail, and the tail decides whether a job finishes inside its window. Measure only successful requests for the headline figures, and report time-to-failure separately; a provider that fails fast looks quick on average while returning nothing.

Check geo accuracy and session stickiness

Geo accuracy

For a sample of requests, fetch an IP-echo endpoint through the proxy and record the exit address. Look each address up in an independent geolocation database, not one supplied by the provider, and compare the result with the country or city you requested. Then check what the target itself shows: currency, language, local results or shipping options. The target's view is the one that matters, and geolocation databases disagree with each other at city level.

Session stickiness

For sticky sessions, open a session and request the IP-echo endpoint at intervals, recording the exit address each time. Measure how long the address actually holds against the duration you asked for, and what happens when it changes: an error, or a silent switch. Repeat across many sessions, since residential and mobile sessions end early whenever a peer device drops. The explainer on rotating vs sticky proxies covers why that matters for multi-step flows.

Compute cost per successful request

List prices are not comparable across billing models. Normalize to cost per validated success:

  • Per-GB pools: billed gigabytes for the test, from the provider's own usage report, times the rate, divided by validated successes. Use the provider's figure rather than response sizes you measured, because metering may include request bytes, headers or failed requests.
  • Per-address pools: the monthly cost of the addresses needed to carry the production rate without blocks, divided by the validated successes they would carry in a month.

A provider with a higher rate and higher success can cost less per usable result. The guide to proxy pricing per GB and per IP shows how to estimate the inputs.

How many requests are enough?

Success rate is a proportion, so its uncertainty shrinks with the square root of the sample. At a true success rate of 90 percent, 400 requests give a 95 percent confidence interval of roughly plus or minus three percentage points; narrowing that to one point takes about 3,500. If two providers differ by two points and each has a few hundred samples, you have not measured a difference.

Noise also comes from time. Run the test across several days and at the hours your production jobs run, and look at each day separately before pooling. A result that flips from day to day is not yet a result.

A practical procedure:

  1. Pick targets and success markers, and write the pass criteria down before the first request.
  2. Run a small pilot to check that the markers and block detection classify responses correctly.
  3. Run the full test across several days and time windows, interleaved.
  4. Summarize per provider and per target, with sample sizes shown next to every rate.
  5. Pull usage reports from each provider and compute cost per successful request.
  6. Spot-check geo accuracy and stickiness separately.

Common mistakes when teams benchmark proxy providers

The same errors recur in internal comparisons:

  • Testing against an IP-echo service only. It proves the proxy connects and nothing about how your targets treat it.
  • Ignoring retries. If the production client retries three times, a provider whose first attempt fails but whose second succeeds costs more and takes longer. Record attempts, not just outcomes.
  • Letting one provider warm up. Running a provider for a week before testing the other lets sessions, cookies and target reputation differ. Start both together from a clean state.
  • Mixing pool types. A shared pool and a dedicated one behave differently; record which one each provider was tested on.
  • Reporting one aggregate number. A single success rate across all targets hides the target where a provider fails completely.

A small Python harness

The harness below reads proxy URLs from environment variables, interleaves requests across providers and targets at equal concurrency, and writes status, latency, size, block and content-check results to CSV. It uses httpx with asyncio and disables keep-alive so each request opens a new tunnel. For SOCKS5 proxy URLs, install httpx[socks].

pip install httpx
export PROVIDERS="a,b"
export PROXY_URL_A="http://USERNAME:[email protected]:PORT"
export PROXY_URL_B="http://USERNAME:[email protected]:PORT"
python benchmark.py

Replace TARGETS with your own pages and a marker string that appears only on a successful response.

import asyncio
import csv
import os
import random
import time
from datetime import datetime, timezone

import httpx

PROVIDERS = {
    name: os.environ[f"PROXY_URL_{name.upper()}"]
    for name in os.environ.get("PROVIDERS", "a,b").split(",")
}
TARGETS = [
    ("https://www.example.com/product/12345", 'itemprop="price"'),
    ("https://www.example.org/search?q=widgets", 'class="result"'),
]
BLOCK_MARKERS = ("captcha", "access denied", "unusual traffic")
REQUESTS_PER_PAIR = int(os.environ.get("REQUESTS_PER_PAIR", "200"))
CONCURRENCY = int(os.environ.get("CONCURRENCY", "10"))
OUT = os.environ.get("OUT", "benchmark.csv")


async def fetch(client, semaphore, provider, url, marker):
    async with semaphore:
        row = {
            "ts": datetime.now(timezone.utc).isoformat(),
            "provider": provider,
            "url": url,
            "status": "",
            "latency_ms": "",
            "bytes": "",
            "blocked": False,
            "content_ok": False,
            "error": "",
        }
        started = time.perf_counter()
        try:
            response = await client.get(url)
            body = response.text
            row["status"] = response.status_code
            row["bytes"] = len(response.content)
            row["blocked"] = response.status_code in (403, 429) or any(
                m in body.lower() for m in BLOCK_MARKERS
            )
            row["content_ok"] = (
                response.status_code == 200 and marker in body and not row["blocked"]
            )
        except httpx.HTTPError as exc:
            row["error"] = type(exc).__name__
        row["latency_ms"] = round((time.perf_counter() - started) * 1000)
        return row


async def main():
    jobs = [
        (provider, url, marker)
        for provider in PROVIDERS
        for url, marker in TARGETS
        for _ in range(REQUESTS_PER_PAIR)
    ]
    random.shuffle(jobs)

    clients = {
        name: httpx.AsyncClient(
            proxy=proxy_url,
            timeout=httpx.Timeout(30.0),
            limits=httpx.Limits(max_keepalive_connections=0),
            follow_redirects=True,
        )
        for name, proxy_url in PROVIDERS.items()
    }
    semaphores = {name: asyncio.Semaphore(CONCURRENCY) for name in PROVIDERS}
    try:
        rows = await asyncio.gather(
            *(fetch(clients[p], semaphores[p], p, url, marker) for p, url, marker in jobs)
        )
    finally:
        for client in clients.values():
            await client.aclose()

    with open(OUT, "w", newline="") as fh:
        writer = csv.DictWriter(fh, fieldnames=list(rows[0]))
        writer.writeheader()
        writer.writerows(rows)


if __name__ == "__main__":
    asyncio.run(main())

A short summary per provider and target:

import csv
import os
import statistics
from collections import defaultdict

groups = defaultdict(list)
with open(os.environ.get("OUT", "benchmark.csv"), newline="") as fh:
    for row in csv.DictReader(fh):
        groups[(row["provider"], row["url"])].append(row)

for (provider, url), rows in sorted(groups.items()):
    ok = [r for r in rows if r["content_ok"] == "True"]
    blocked = sum(r["blocked"] == "True" for r in rows)
    latencies = [int(r["latency_ms"]) for r in ok]
    cuts = statistics.quantiles(latencies, n=100) if len(latencies) >= 2 else [float("nan")] * 99
    print(
        f"{provider:<8} {url[:60]:<60} n={len(rows):<6} "
        f"success={len(ok) / len(rows):.1%} blocked={blocked / len(rows):.1%} "
        f"p50={cuts[49]:.0f}ms p95={cuts[94]:.0f}ms"
    )

Keep the raw CSV. When a result surprises you, the individual rows, their timestamps and error classes are what explain it. If you then move production traffic, the same measurements carry over to the parallel run described in how to switch proxy providers.

Benchmarking ProxyForge

ProxyForge accounts are pay-as-you-go from a prepaid wallet with no monthly minimum, so a benchmark can start from a single gigabyte of residential or mobile traffic or a single ISP or datacenter address, and ISP and datacenter lines have a paid trial. Test traffic is billed at the published rates on our pricing page. The product overview describes each proxy line and its session options.

If you are benchmarking because you are considering a move, our migration process includes a parallel run on your own dashboards, and a named engineer can help set up targeting and sessions so the comparison is fair.

FAQ

Related questions

How many requests do I need to compare two proxy providers?

Enough per target and per provider that the confidence interval on success rate is narrower than the difference you care about. A few hundred requests per pair detects large gaps; detecting a difference of one or two percentage points takes thousands.

Should I trust a proxy provider's published success rate?

Treat it as a claim about someone else's targets. Success rates vary widely by site, time and request pattern, so only a test against your own targets tells you what you will get.

Can I benchmark proxies with a free trial?

A trial is enough for a first pass if it allows the concurrency and volume your test needs. Check whether trial traffic uses the same pools and gateways as paid traffic, since a separate trial pool measures the wrong thing.

What is a good block rate for a proxy benchmark?

There is no universal figure, because block rates depend mostly on the target and on how the request is made. Compare providers against each other on the same targets rather than against an absolute number.

Start with the evidence

Ask us to trace an address, send you the sourcing attestation, or price your current volume at our published rates. A named engineer will help with your technical and procurement review.

One business day, from a named engineer.