To reduce proxy bandwidth, measure first, then remove bytes you never use. For most scraping and browser automation, the biggest savings come from blocking images, fonts, media and third-party analytics in the browser, asking for compressed responses, re-fetching only what has changed with conditional requests, fetching the JSON a page loads instead of the whole page where that is permitted, and making sure retries do not download the same large response repeatedly. Traffic that does not need a per-GB line at all can often move to addresses billed per IP.
The sections below take these in order of typical impact, with code tested against a local proxy.
Measure bytes per page before changing anything
Before you try to reduce proxy bandwidth, get a number: bytes transferred per page, or better, per successful record. Without it you cannot tell which change mattered, and you will spend effort on assets that were never the problem.
In a browser, Playwright reports the size of each request and response. The script below loads a page twice through the proxy, once as-is and once with heavy assets blocked, and prints the transfer broken down by resource type. It reads PROXY_URL in the shape http://USERNAME:[email protected]:PORT.
import asyncio
import os
import sys
from collections import Counter
from urllib.parse import unquote, urlsplit
from playwright.async_api import async_playwright
proxy_url = urlsplit(os.environ["PROXY_URL"])
PROXY = {
"server": f"{proxy_url.scheme}://{proxy_url.hostname}:{proxy_url.port}",
"username": unquote(proxy_url.username or ""),
"password": unquote(proxy_url.password or ""),
}
BLOCKED_TYPES = {"image", "media", "font"}
BLOCKED_HOSTS = ("google-analytics.com", "googletagmanager.com", "doubleclick.net")
def is_blocked_host(host):
return any(host == h or host.endswith(f".{h}") for h in BLOCKED_HOSTS)
async def block_heavy(route):
request = route.request
host = urlsplit(request.url).hostname or ""
if request.resource_type in BLOCKED_TYPES or is_blocked_host(host):
await route.abort()
else:
await route.continue_()
async def measure(browser, url, block):
context = await browser.new_context(service_workers="block")
if block:
await context.route("**/*", block_heavy)
page = await context.new_page()
by_type = Counter()
async def on_finished(request):
sizes = await request.sizes()
by_type[request.resource_type] += (
sizes["requestHeadersSize"] + sizes["requestBodySize"]
+ sizes["responseHeadersSize"] + sizes["responseBodySize"]
)
page.on("requestfinished", on_finished)
await page.goto(url, wait_until="networkidle")
await context.close()
return by_type
async def main(url):
async with async_playwright() as p:
browser = await p.chromium.launch(proxy=PROXY)
for block in (False, True):
by_type = await measure(browser, url, block)
label = "blocked" if block else "full"
total_kib = sum(by_type.values()) / 1024
print(f"{label:>8}: {total_kib:,.0f} KiB", dict(by_type.most_common()))
await browser.close()
asyncio.run(main(sys.argv[1]))
Two runs illustrate why measuring comes first. On a public scraping sandbox with a product grid, the page transferred about 350 KiB in full and 113 KiB with images, fonts and media blocked, a two-thirds reduction. On a large news homepage, the same blocking cut about 2.1 MiB to 1.5 MiB, because scripts accounted for most of the weight and were left alone. The same technique saves very different amounts on different targets.
These are client-side counts. A provider's meter usually also includes TLS overhead and traffic in both directions, so expect its figure to be somewhat higher. Compare the two over a day of traffic and track the ratio; if it is stable, your own measurements are good enough to optimize against. For ongoing tracking of bytes per successful request, see proxy monitoring in production.
Block images, fonts, media and analytics
If your job extracts text, prices or structured data, the browser does not need to download pictures, video, web fonts or analytics beacons. Blocking them is the single largest saving for most browser workloads.
In Playwright, register a route handler on the context, as in block_heavy above, and abort requests by resource_type and by host. Two details from the Playwright documentation matter here. First, requests made by a service worker bypass route handlers, so create the context with service_workers="block". Second, enabling routing disables the browser's HTTP cache for that context. For short-lived contexts, where the cache rarely helps anyway, that is a good trade; for a long session that revisits the same assets, measure both ways. The Playwright proxy guide covers the proxy setup itself, including per-context proxies.
In Puppeteer, the equivalent is request interception. Chrome's --proxy-server flag does not accept credentials, so they are supplied with page.authenticate:
import puppeteer from 'puppeteer';
const proxy = new URL(process.env.PROXY_URL);
const BLOCKED_TYPES = new Set(['image', 'media', 'font']);
const BLOCKED_HOSTS = ['google-analytics.com', 'googletagmanager.com', 'doubleclick.net'];
const browser = await puppeteer.launch({
args: [`--proxy-server=${proxy.protocol}//${proxy.host}`],
});
const page = await browser.newPage();
await page.authenticate({
username: decodeURIComponent(proxy.username),
password: decodeURIComponent(proxy.password),
});
await page.setRequestInterception(true);
page.on('request', (request) => {
if (request.isInterceptResolutionHandled()) return;
const host = new URL(request.url()).hostname;
const blocked =
BLOCKED_TYPES.has(request.resourceType()) ||
BLOCKED_HOSTS.some((h) => host === h || host.endsWith(`.${h}`));
if (blocked) request.abort();
else request.continue();
});
await page.goto(process.argv[2], { waitUntil: 'networkidle2' });
console.log(await page.title());
await browser.close();
Some cautions apply to both:
- Do not block what the data depends on. Blocking stylesheets breaks layout-dependent selectors and visibility checks. Blocking first-party scripts breaks pages that render content client-side. Start with images, media and fonts, verify your extraction still works, then consider more.
- Block third-party hosts by name, not by pattern guessing. Use your measurement to list the hosts that carry the most bytes, and check each one before adding it.
- Validate results after any change. A page that silently renders an empty list costs less bandwidth and is worth nothing.
Ask for compressed responses
Text compresses well. An HTML page that is 300 KB uncompressed is often 60 KB or less with gzip or Brotli. Over HTTPS, the proxy relays an encrypted tunnel and cannot compress anything for you; compression happens only if your client asks the target for it with an Accept-Encoding header.
Browsers always ask. HTTP clients vary:
- curl does not ask by default. Fetching a Wikipedia article in our test transferred 317 KB without flags and 60 KB with
--compressed. - Python Requests sends
Accept-Encoding: gzip, deflateby default, and addsbrwhen thebrotlipackage is installed. - httpx behaves the same way: gzip and deflate always, Brotli when
brotliis installed.
curl --compressed -x "$PROXY_URL" -o page.html https://example.com/
If you set headers manually, make sure you have not overwritten Accept-Encoding with identity or dropped it. And measure the wire size rather than the decoded size: in httpx, response.num_bytes_downloaded gives bytes as received, before decompression.
Use HTTP caching and conditional requests
Many jobs re-fetch pages that have not changed since the last run. HTTP has a built-in mechanism for this, defined in RFC 9110: store the ETag or Last-Modified value from a response, send it back as If-None-Match or If-Modified-Since next time, and the server answers 304 Not Modified with no body when nothing has changed.
import json
import os
from pathlib import Path
import requests
PROXY_URL = os.environ["PROXY_URL"]
PROXIES = {"http": PROXY_URL, "https": PROXY_URL}
CACHE = Path("http-cache.json")
def fetch_if_changed(session, url):
cache = json.loads(CACHE.read_text()) if CACHE.exists() else {}
entry = cache.get(url, {})
headers = {}
if "etag" in entry:
headers["If-None-Match"] = entry["etag"]
if "last_modified" in entry:
headers["If-Modified-Since"] = entry["last_modified"]
response = session.get(url, headers=headers, proxies=PROXIES, timeout=30)
if response.status_code == 304:
return None
response.raise_for_status()
validators = {}
if "ETag" in response.headers:
validators["etag"] = response.headers["ETag"]
if "Last-Modified" in response.headers:
validators["last_modified"] = response.headers["Last-Modified"]
cache[url] = validators
CACHE.write_text(json.dumps(cache))
return response.content
Run twice against the plain-text copy of RFC 9110 through a proxy, this downloaded the full 500 KB document the first time (about 140 KB on the wire, compressed) and received a 304 with headers only the second. A None return means "unchanged, use what you stored". In production, keep the validators in the same store as the content, not in a JSON file.
Support varies by site. Static files, feeds, sitemaps and APIs usually honor validators; dynamic HTML pages often do not. Test your targets: if every conditional request comes back 200 with a full body, drop the logic for that target rather than carrying it.
Fetch less: JSON endpoints, sitemaps and HEAD requests
The most effective way to reduce proxy bandwidth is to not download the page at all.
JSON endpoints. Many pages load their data from a JSON endpoint and render it in the browser. Open the page with developer tools, filter the network panel to Fetch/XHR, and you will often find the exact data you want in a response a small fraction of the page's size, with no browser needed. Use these only where the site's terms permit it, and prefer documented public APIs where they exist.
Sitemaps and feeds. An XML sitemap with lastmod dates or an RSS feed tells you which pages changed, so you can re-crawl only those instead of the whole site.
HEAD requests. A HEAD request returns the headers without the body: enough to check whether a URL exists, its size, its type and when it last changed.
curl -sI --suppress-connect-headers -x "$PROXY_URL" https://example.com/catalog.pdf
Some servers mishandle HEAD and return 405 or different headers than for GET. Fall back to a conditional GET for those.
URL hygiene. Deduplicate URLs before fetching, normalize query parameters, and stay out of faceted navigation that produces thousands of URLs for the same listings in different sort orders. A crawler that fetches each product once instead of five times saves more than any header trick.
Keep other traffic off the proxy. Health checks, geolocation lookups, your own APIs and assets from your own CDN do not need to go through a metered line. Playwright's proxy settings accept a bypass list, and HTTP clients honor NO_PROXY when they read proxy settings from the environment.
Stop retries from re-downloading
Retries are necessary, and badly designed ones are a common cause of bandwidth that nobody can explain. Each retry of a full page downloads the full page again.
- Store the raw response before parsing. If your parser fails, fix it and re-parse from storage. Re-fetching because of your own bug costs bandwidth and adds load on the target.
- Do not retry blocks at full speed. A challenge or block page still costs bytes. Treat a rising block rate as a signal to slow down, as described in diagnosing proxy 429 and 403 errors, and use a circuit breaker so a blocked target stops consuming bandwidth.
- Cap response size. A misrouted link to a video or a large archive can cost more than a day of normal pages. Stream the response and stop reading past a limit:
import os
import httpx
MAX_BYTES = 2 * 1024 * 1024
def get_capped(client, url):
with client.stream("GET", url) as response:
response.raise_for_status()
if int(response.headers.get("Content-Length", 0)) > MAX_BYTES:
return None
chunks = []
for chunk in response.iter_bytes():
chunks.append(chunk)
if response.num_bytes_downloaded > MAX_BYTES:
return None
return b"".join(chunks)
with httpx.Client(proxy=os.environ["PROXY_URL"], timeout=30, follow_redirects=True) as client:
body = get_capped(client, "https://www.rfc-editor.org/rfc/rfc9110.txt")
The check on Content-Length avoids starting a download that is declared too large; the check on num_bytes_downloaded stops responses that do not declare a length, such as chunked ones, once they pass the limit. We tested both paths through a local proxy.
Move bulk fetches to per-IP lines
Some traffic does not need a rotating residential or mobile address: public datasets, sitemaps, documentation, APIs that do not score address types, and your own monitoring. Running that traffic on a per-GB line pays bandwidth rates for work that a static address would handle just as well.
Datacenter and ISP proxies are billed per address per month rather than per gigabyte, so moving high-volume, low-sensitivity fetches onto them changes the cost model entirely. The crossover point depends on your volume and the targets involved; per-GB vs per-IP proxy pricing walks through the calculation. Keep residential and mobile for the targets that need them, and measure success rate on both lines before moving anything permanently.
A bandwidth checklist
Use this list to reduce proxy bandwidth on an existing pipeline:
- Bytes per page, or per record, measured on each major target
- Images, media and fonts blocked in browser jobs, with extraction still validated
- Third-party analytics and ad hosts blocked by name
- Service workers blocked in Playwright contexts that use routing
- Compression requested by every HTTP client, including curl
- Conditional requests on targets that honor
ETagorLast-Modified - JSON endpoints, feeds or sitemaps used where permitted and available
- URLs deduplicated and faceted navigation excluded
- Non-target traffic bypassing the proxy
- Raw responses stored before parsing; retries capped; response size capped
- Bulk, low-sensitivity traffic moved to per-IP lines
How ProxyForge bills bandwidth
On ProxyForge, residential and mobile proxies are billed per GB from a prepaid wallet, with no monthly minimum, and ISP and datacenter proxies are billed per address per month. Current rates are on the pricing page. Because the same account and gateway serve all four lines, moving a workload from a per-GB line to a per-address line is a change of connection string rather than a new vendor.
If your bandwidth per record is higher than you expect and you cannot see why, a named engineer can review a sample of your traffic with you; the contact page is the quickest route.