Proxy monitoring in production means measuring outcomes, not uptime: the share of requests that returned valid data, per target; how often you were blocked or challenged; latency at the 50th, 95th and 99th percentiles; bytes transferred per successful request; and the cost of each record you actually kept. Break errors down by class so you can tell a proxy failure (407, gateway 5xx, timeouts) from a target refusing you (403, 429, a challenge page). Instrument the client that makes the requests, keep label cardinality low, and alert on sustained changes against your own baseline.
This guide covers each metric, a Python instrumentation example with prometheus_client, alert rules validated with promtool, the dashboards worth building, and how to use the data when a vendor's service falls short.
Which metrics matter for proxy monitoring?
A proxy sits between your code and a target you do not control, so a single "is it up" check says very little. These are the metrics that describe whether the pipeline is working and what it costs.
| Metric | Definition | Why it matters | Common mistake |
|---|---|---|---|
| Validated success rate | Share of requests whose response passed content validation, per target | The only rate that reflects usable data | Counting any 200 as success, including challenge pages and empty results |
| Block and challenge rate | Share of 403, 429 and detected challenge pages |
Early signal of rate or fingerprint problems | Folding it into a generic error rate |
| Latency p50, p95, p99 | Time from request start to full body, completed responses only | Sizing timeouts, concurrency and throughput | Averaging, which hides the tail |
| Bytes per successful request | Response bytes divided by validated successes | Drives the bill on per-GB lines | Measuring decompressed size instead of wire size |
| Cost per successful record | Spend divided by records kept | The number finance and procurement care about | Dividing by requests instead of records |
| Error class breakdown | Counts by outcome: proxy auth, proxy error, unreachable, timeout, blocked, challenge, invalid | Tells you whose problem it is | One errors_total counter with no class |
The error classes deserve their own list, because they point to different owners:
proxy_auth: the gateway returned407. Credentials or the IP allowlist are wrong. This is a configuration fault that retries will not fix, and it should page someone.proxy_error: the gateway answered theCONNECTwith a5xx, meaning it could not reach the target through the chosen exit. Transient at low rates; a vendor problem at high rates.proxy_unreachable: your client could not connect to the gateway at all. Check your own network and DNS first, then the vendor's status.timeout: no response within your timeout. Could be either hop.blocked,captcha,invalid: the target answered, and the answer was a refusal, a challenge or unusable content. These are about rate, client consistency or proxy type, not about gateway health. Diagnosing proxy 429 and 403 errors covers how to act on each.
Instrumenting a Python client with prometheus_client
The example below wraps an httpx client with four metrics: a request counter labeled by outcome, a latency histogram, a counter of response bytes as received on the wire, and a counter of extracted records. A gauge carries your contracted price per GB so that cost can be computed in queries. It reads PROXY_URL from the environment, in the shape http://USERNAME:[email protected]:PORT, and needs pip install httpx prometheus-client.
import os
import sys
import time
import httpx
from prometheus_client import Counter, Gauge, Histogram, start_http_server
PROXY_URL = os.environ["PROXY_URL"]
PROVIDER = os.environ.get("PROXY_PROVIDER", "primary")
LINE = os.environ.get("PROXY_LINE", "residential")
LABELS = ["provider", "line", "target"]
REQUESTS = Counter(
"proxy_requests_total",
"Proxied requests by outcome.",
LABELS + ["outcome"],
)
LATENCY = Histogram(
"proxy_request_duration_seconds",
"Time from request start to full response body, completed responses only.",
LABELS,
buckets=(0.25, 0.5, 1, 2, 4, 8, 16, 32, 64),
)
RESPONSE_BYTES = Counter(
"proxy_response_bytes_total",
"Response body bytes as received on the wire, before decompression.",
LABELS,
)
RECORDS = Counter(
"proxy_records_total",
"Records extracted from validated responses.",
LABELS,
)
PRICE_PER_GB = Gauge(
"proxy_price_per_gb",
"Contracted price per GB for this provider and line.",
["provider", "line"],
)
PRICE_PER_GB.labels(provider=PROVIDER, line=LINE).set(
float(os.environ.get("PROXY_PRICE_PER_GB", "0"))
)
BLOCK_STATUSES = {403, 429}
def fetch(client, target, url, validate):
labels = {"provider": PROVIDER, "line": LINE, "target": target}
started = time.perf_counter()
try:
response = client.get(url)
except httpx.ProxyError as exc:
outcome = "proxy_auth" if str(exc).startswith("407") else "proxy_error"
REQUESTS.labels(**labels, outcome=outcome).inc()
return None
except httpx.TimeoutException:
REQUESTS.labels(**labels, outcome="timeout").inc()
return None
except httpx.ConnectError:
REQUESTS.labels(**labels, outcome="proxy_unreachable").inc()
return None
except httpx.TransportError:
REQUESTS.labels(**labels, outcome="connection_error").inc()
return None
LATENCY.labels(**labels).observe(time.perf_counter() - started)
downloaded = sum(r.num_bytes_downloaded for r in [*response.history, response])
RESPONSE_BYTES.labels(**labels).inc(downloaded)
if response.status_code == 407:
outcome, records = "proxy_auth", 0
elif response.status_code in BLOCK_STATUSES:
outcome, records = "blocked", 0
elif response.status_code >= 500:
outcome, records = "target_error", 0
elif response.status_code != 200:
outcome, records = "other_status", 0
else:
outcome, records = validate(response.content)
REQUESTS.labels(**labels, outcome=outcome).inc()
if records:
RECORDS.labels(**labels).inc(records)
return response.content if outcome == "success" else None
def validate_product_page(body):
text = body.decode("utf-8", errors="replace")
if "captcha" in text.lower():
return "captcha", 0
if 'itemprop="price"' not in text:
return "invalid", 0
return "success", 1
if __name__ == "__main__":
start_http_server(int(os.environ.get("METRICS_PORT", "9108")))
timeout = httpx.Timeout(60, connect=10)
with httpx.Client(proxy=PROXY_URL, timeout=timeout, follow_redirects=True) as client:
for line in sys.stdin:
fetch(client, "example-shop", line.strip(), validate_product_page)
We ran this against a local proxy and a stub gateway that answers CONNECT with 407 or 502, and each case landed in the expected outcome. Some details are worth explaining:
- httpx rather than Requests, for the byte count. httpx reports
num_bytes_downloaded: body bytes as they arrived, before gzip or Brotli decoding. Requests has no reliable equivalent.response.raw.tell()looks like one, but in our tests with urllib3 2.8 it returned0for responses sent with chunked transfer encoding, which is common for dynamically generated pages. Measuring the decoded size instead overstates transfer several times over: a page that decompresses to 300 KB may have cost 60 KB. Redirect hops are downloaded too, so the code adds them fromresponse.history. - For HTTPS targets, a failed
CONNECTis an exception, not a response. httpx raisesProxyErrorwith the gateway's status line as its message, which is how407is told apart from a gateway5xx. With a proxy configured, every connection goes to the gateway first, so aConnectErrormeans the gateway itself was unreachable. - The byte count is still a lower bound. It excludes response headers, request bytes and TLS overhead, all of which a provider meters. Reconcile against the provider's usage figures and track the ratio, which should be stable.
- Latency is observed only for completed responses. Timeouts would otherwise pile up at the timeout value and distort every percentile. They are counted as an outcome instead. Make the top histogram bucket at least as large as your read timeout.
- Validation is per target. The
validatefunction is where your domain knowledge goes: a selector that must exist, a minimum number of items, the absence of a challenge marker. Without it, "success rate" measures the target's willingness to answer, not your pipeline's output.
If your workers run as several processes on one host, for example under a process manager, read the prometheus_client documentation on multiprocess mode; each process otherwise exports its own counters on its own port.
Label cardinality: what never to label by
Every unique combination of label values creates a separate time series in Prometheus. Four labels with a handful of values each is a few hundred series. One label with an unbounded set of values is millions, and it will slow or break your monitoring before it tells you anything.
Never label by exit IP. A rotating residential pool can hand you a different address on every request. Labeling by it creates a new series per request, and none of them has enough samples to be meaningful. The same applies to:
- full URLs, paths or query strings
- session IDs, request IDs or trace IDs
- user or customer IDs
- raw hostnames from crawled URLs, when you crawl open-ended sets of sites
Do label by a small, fixed set of values you define: provider, line (residential, mobile, ISP, datacenter), target as a named group such as retailer-a rather than a hostname, and outcome from a fixed list.
Per-address analysis still has a place, for instance when you investigate whether particular exits fail more often. Put that in logs or a trace store, sampled, with the exit IP as a field, and query it when you need it. Metrics are for aggregates; logs are for individual events.
Recording rules for dashboards
Recording rules precompute the expressions your dashboards and alerts use. These were checked with promtool check rules and unit tested with promtool test rules:
groups:
- name: proxy-recording
rules:
- record: proxy:outcome_share:rate5m
expr: |
sum by (provider, line, target, outcome) (rate(proxy_requests_total[5m]))
/ ignoring (outcome) group_left
sum by (provider, line, target) (rate(proxy_requests_total[5m]))
- record: proxy:latency_seconds:p50_5m
expr: histogram_quantile(0.50, sum by (le, provider, line) (rate(proxy_request_duration_seconds_bucket[5m])))
- record: proxy:latency_seconds:p95_5m
expr: histogram_quantile(0.95, sum by (le, provider, line) (rate(proxy_request_duration_seconds_bucket[5m])))
- record: proxy:latency_seconds:p99_5m
expr: histogram_quantile(0.99, sum by (le, provider, line) (rate(proxy_request_duration_seconds_bucket[5m])))
- record: proxy:cost_per_record:1d
expr: |
sum by (provider, line, target) (increase(proxy_response_bytes_total[1d])) / 1e9
* on (provider, line) group_left max by (provider, line) (proxy_price_per_gb)
/
sum by (provider, line, target) (increase(proxy_records_total[1d]))
The max by in the cost rule matters: every worker exports the price gauge with its own instance label, and without collapsing those, the join fails with a many-to-many matching error. Cost per record is computed from response bytes only, so it is a lower bound; scale it by your measured ratio to the provider's metered usage if you need it to match the invoice.
Alert rules that catch real problems
Good proxy monitoring alerts only on conditions that need a human, sustained long enough to rule out noise, with a minimum traffic guard so a quiet target does not alert on three failed requests. The thresholds below are starting points; replace them with values derived from your own baseline.
groups:
- name: proxy-alerts
rules:
- record: proxy:success_ratio:rate15m
expr: |
sum by (provider, line, target) (rate(proxy_requests_total{outcome="success"}[15m]))
/
sum by (provider, line, target) (rate(proxy_requests_total[15m]))
- record: proxy:bytes_per_success:rate1h
expr: |
sum by (provider, line, target) (rate(proxy_response_bytes_total[1h]))
/
sum by (provider, line, target) (rate(proxy_requests_total{outcome="success"}[1h]))
- alert: ProxyAuthFailing
expr: sum by (provider, line) (rate(proxy_requests_total{outcome="proxy_auth"}[5m])) > 0
for: 5m
labels:
severity: page
annotations:
summary: "{{ $labels.provider }}/{{ $labels.line }}: gateway is rejecting credentials"
- alert: ProxySuccessRateLow
expr: |
proxy:success_ratio:rate15m < 0.85
and
sum by (provider, line, target) (rate(proxy_requests_total[15m])) > 0.2
for: 15m
labels:
severity: ticket
annotations:
summary: "{{ $labels.target }} via {{ $labels.provider }}: validated success {{ $value | humanizePercentage }}"
- alert: ProxyGatewayErrorsHigh
expr: |
sum by (provider, line) (rate(proxy_requests_total{outcome=~"proxy_error|proxy_unreachable|timeout"}[10m]))
/
sum by (provider, line) (rate(proxy_requests_total[10m]))
> 0.05
for: 10m
labels:
severity: page
annotations:
summary: "{{ $labels.provider }}/{{ $labels.line }}: {{ $value | humanizePercentage }} of requests failing at the proxy hop"
- alert: ProxyLatencyP95High
expr: |
histogram_quantile(0.95,
sum by (le, provider, line) (rate(proxy_request_duration_seconds_bucket[10m]))
) > 16
for: 15m
labels:
severity: ticket
annotations:
summary: "{{ $labels.provider }}/{{ $labels.line }}: p95 latency {{ $value | humanizeDuration }}"
- alert: ProxyBytesPerSuccessRising
expr: |
proxy:bytes_per_success:rate1h
/
proxy:bytes_per_success:rate1h offset 1d
> 1.5
for: 1h
labels:
severity: ticket
annotations:
summary: "{{ $labels.target }}: bytes per successful request up {{ $value | printf \"%.1f\" }}x on yesterday"
The split in severity is deliberate. Authentication failures and proxy-hop errors are paged, because they stop every job on that line and are fixable only by a person. A falling success rate on one target and a latency drift are tickets: they need investigation, usually into rate, fingerprint or target changes, not a 3 a.m. response. The bytes-per-success alert catches a quieter problem, such as a target that started serving a heavier page or a change that disabled compression, which otherwise shows up only on the invoice. How to reduce proxy bandwidth covers the fixes.
A unit test file for promtool test rules is worth keeping next to these rules. Feeding it a series where half of the requests are blocked and confirming that ProxySuccessRateLow fires is how you know the expression does what its name says.
Dashboards that answer questions
A useful proxy monitoring dashboard answers a small number of questions quickly. One row per question:
- Is data coming in? Validated success rate by target, with the week-ago line for comparison, and records per minute.
- If not, whose problem is it? A stacked chart of
proxy:outcome_share:rate5mby outcome. A band ofproxy_authorproxy_errorpoints at the proxy hop; a band ofblockedorcaptchapoints at the target. - How fast is it? p50, p95 and p99 latency per provider and line, on a logarithmic axis.
- What does it cost? Bytes per successful request and cost per record per target, plus daily spend.
- Is the gateway itself healthy? A synthetic check, described below, per provider and line.
For the fifth row, run a canary: a small job that requests a static page you host yourself, through each provider and line, every 30 to 60 seconds. Because you control the target, a failure can only come from the proxy hop or the network path. It is the cleanest availability signal you can have, and it keeps measuring when production traffic is idle.
Using your data to hold a vendor to its SLA
Most proxy SLAs cover gateway availability and sometimes support response times, not success rates against arbitrary targets. That is reasonable, since no vendor controls a target's bot management. It also means you need your own evidence to make a claim, and the metrics above give you exactly that. What to look for in a proxy SLA covers the contract side.
When you raise a breach:
- Separate the hops. Report
proxy_auth,proxy_error,proxy_unreachableand canary failures, which are about the gateway, apart fromblockedandcaptcha, which are about the target. - Give timestamps and rates, not impressions: the start and end of the incident, the error share per minute, and the volume affected.
- Show a control. The same period through a second provider or line (see proxy redundancy for running one), or the canary against your own endpoint, shows whether the problem was specific to the vendor.
- Keep the raw evidence. Sampled request logs with timestamps, exit IPs and gateway responses let the vendor trace the incident on their side.
The same data drives larger decisions. Per-provider labels let you run two vendors side by side on your own traffic, which is the basis of a fair proxy provider benchmark and of any migration you decide to make.
Monitoring ProxyForge traffic
The classification above relies only on standard HTTP proxy behavior, 407 for authentication and 5xx on CONNECT when an exit cannot reach the target, so it works with ProxyForge as with any other provider. Every ProxyForge line, residential, mobile, ISP and datacenter, is available over HTTP, HTTPS and SOCKS5, and using a provider and line label per connection string gives you the per-line view.
If you are evaluating us against an incumbent, the migration process is built around exactly this kind of measurement: you run both providers in parallel at published rates, compare them on your own dashboards, and shift traffic when the numbers justify it.