Skip to content
ProxyForge

HAProxy and Nginx in front of a proxy gateway

ProxyForge engineeringUpdated 8 min read

An HAProxy forward proxy setup is really HAProxy in front of a forward proxy: HAProxy accepts your scrapers' proxy connections and hands them to one of several provider gateway endpoints, adding failover, health checks and an access boundary. Nginx can do the relay part with its stream module. What separates them is the credential. In TCP mode, HAProxy and Nginx both pass bytes through untouched, so each client still sends its own Proxy-Authorization. Only HAProxy in HTTP mode can add the credential for the client, and only HAProxy can fail over between two providers whose credentials differ.

The configurations below were run on HAProxy 3.2 and Nginx 1.29 (official haproxy:3.2-alpine and nginx:1.29-alpine images) in front of two local stand-in gateways that require Basic proxy authentication. Replace gateway.proxyforge.io:PORT, USERNAME and PASSWORD with your provider's values. A second endpoint is shown as gateway-b.example.net:PORT; use whatever second path you actually have.

What a relay can and cannot do

A proxy client opens a TCP connection to the gateway and speaks the proxy protocol over it: CONNECT host:443 with a Proxy-Authorization header for HTTPS targets, an absolute-form GET http://... for plain HTTP, or a SOCKS5 handshake. What sits in the middle decides which parts of that it can see and change:

Capability HAProxy, TCP mode HAProxy, HTTP mode Nginx stream
Relays HTTP proxy traffic, including CONNECT Yes Yes Yes
Relays SOCKS5 Yes No Yes
Adds the credential so clients hold none No Yes No
Logs the target host of each tunnel No Yes No
Active health checks Yes Yes No, passive only
Fails over between providers with different credentials No Yes No
Re-resolves the gateway hostname while running Yes, with resolvers Yes, with resolvers Yes, with resolve on recent versions

Neither tool parses the proxy authentication for you in TCP mode. Whatever the client sends reaches the gateway unchanged, and a gateway that rejects it answers 407 straight back to the client. Keeping credentials off application hosts therefore takes one of three routes: HTTP mode with header injection, IP allowlisting of the relay's own egress address, or a forward proxy such as Squid holding the credential. Running Squid as an upstream proxy covers the third.

HAProxy in TCP mode: failover with client-held credentials

This configuration listens on an internal address, accepts only one network, and sends every connection to the primary endpoint unless it is down:

global
    log stdout format raw local0
    maxconn 20000

defaults
    log global
    timeout connect 5s
    timeout client 10m
    timeout server 10m
    timeout tunnel 10m
    timeout check 5s

resolvers system
    parse-resolv-conf
    hold valid 30s

frontend proxy_tcp
    bind 10.20.0.6:3128
    mode tcp
    option tcplog
    tcp-request connection reject unless { src 10.20.0.0/24 }
    default_backend gateway_tcp

backend gateway_tcp
    mode tcp
    option tcp-check
    tcp-check connect
    tcp-check send "CONNECT status.example.com:443 HTTP/1.1\r\nHost: status.example.com:443\r\n\r\n"
    tcp-check expect rstring "^HTTP/1\.[01] 407"
    default-server check inter 5s fall 2 rise 3 resolvers system init-addr last,libc,none
    server primary gateway.proxyforge.io:PORT
    server secondary gateway-b.example.net:PORT backup

Clients point at HAProxy with their own credential, exactly as they would at the gateway: http://USERNAME:[email protected]:3128. A client outside 10.20.0.0/24 has its connection reset before a byte is read. Long idle tunnels are common in scraping, since HTTP keep-alive holds a CONNECT open between requests, so set timeout client, timeout server and timeout tunnel well above the short values in most example configurations.

When we stopped the primary gateway, HAProxy logged Server gateway_tcp/primary is DOWN, reason: Layer4 connection problem followed by Running on backup, and new connections went to the secondary. Be precise about the gap: a connection that arrived after the gateway failed but before the check marked it down was aborted (curl reported Proxy CONNECT aborted), even with retries and option redispatch set, because a backup server is not eligible while the active one is still marked up. inter multiplied by fall is the window your clients must cover with their own retry.

The TCP log line names the backend and server but nothing about the destination:

10.20.0.50:40088 [08/Oct/2026:18:05:00.687] proxy_tcp gateway_tcp/secondary 1/0/6 3279 -- 1/1/0/0/0 0/0

Health checks that speak the proxy protocol

A plain TCP check only proves the port accepts connections. The tcp-check sequence above sends a real CONNECT and matches the status line, and there are two useful things to expect:

  • Expect 407 with no credential, as above. It proves the gateway is answering at the proxy layer, costs no metered traffic, and keeps the credential out of HAProxy's configuration. It cannot see a gateway that authenticates you but can no longer reach anything: in our test, a stand-in gateway that answered every authenticated request with 500 stayed up under this check.
  • Expect 200 with a credential. Add Proxy-Authorization: Basic ${PROXY_AUTH_B64} to the sent request and expect rstring "^HTTP/1\.[01] 200". The same broken gateway was marked down within two intervals (TCPCHK did not match content (regex) at step 3). The cost is a credential in HAProxy's environment and a tunnel opened through the provider on every check, which can count toward metered usage.

Point the check at a host you control and that is always up. Checking against a target you collect from adds load to it and turns its outages into yours. HAProxy expands ${VAR} inside double-quoted strings from its environment, which is how the credential stays out of the file itself.

HAProxy in HTTP mode: adding the credential

In mode http, HAProxy parses each request. It forwarded both CONNECT and absolute-form plain-HTTP requests to the gateway in our tests, switched to a tunnel after the 200, and let us set the authentication header on the way:

frontend proxy_http
    bind 10.20.0.6:3129
    mode http
    option httplog
    http-request deny unless { src 10.20.0.0/24 }
    default_backend provider_a

backend provider_a
    mode http
    http-request set-header Proxy-Authorization "Basic ${PROVIDER_A_AUTH}"
    option tcp-check
    tcp-check connect
    tcp-check send "CONNECT status.example.com:443 HTTP/1.1\r\nHost: status.example.com:443\r\nProxy-Authorization: Basic ${PROVIDER_A_AUTH}\r\n\r\n"
    tcp-check expect rstring "^HTTP/1\.[01] 200"
    server gateway gateway.proxyforge.io:PORT check inter 5s fall 2 rise 3 resolvers system init-addr last,libc,none

Clients now use http://10.20.0.6:3129 with no credential at all, and set-header replaces any Proxy-Authorization a client does send. Produce the value with printf '%s' 'USERNAME:PASSWORD' | base64 -w0 (GNU base64 wraps long output at 76 characters without -w0) and pass it to the HAProxy process from your secrets store as PROVIDER_A_AUTH.

HTTP mode also logs the tunnel's destination, which TCP mode cannot:

10.20.0.50:43608 [08/Oct/2026:18:05:00.288] proxy_http provider_b/gateway 0/0/0/0/6 200 3294 - - ---- 1/1/0/0/0 0/0 "CONNECT www.example.com:443 HTTP/1.1"

What HTTP mode cannot do is carry SOCKS5, so a client that only speaks SOCKS needs a TCP-mode frontend. And because anyone allowed by the src rule can now spend the credential, the network boundary in front of HAProxy is the access control; treat it that way in review.

Failing over between two providers

TCP mode can only fail over between endpoints that accept the same credential, because the client's header goes to whichever server is chosen. Two providers issue different credentials and usually different session syntax. HTTP mode handles it with one backend per provider, each setting its own header, and a frontend rule that switches when the first has no live servers:

frontend proxy_http
    bind 10.20.0.6:3129
    mode http
    option httplog
    http-request deny unless { src 10.20.0.0/24 }
    use_backend provider_b if { nbsrv(provider_a) eq 0 }
    default_backend provider_a

provider_b is a copy of provider_a with its own server line and PROVIDER_B_AUTH. With the first gateway stopped, our second stand-in gateway logged the second credential's username on the next request, and traffic returned to provider_a once its checks passed again. This only works if job code does not encode provider-specific options in the request; proxy redundancy covers how to keep the vendor syntax out of the jobs, and the two-vendor topologies this enables.

Nginx stream: the same relay with passive checks

Nginx's stream block goes at the top level of nginx.conf, beside http, not inside it. Check that your build includes the module with nginx -V; distribution packages sometimes ship it as a separate dynamic module that needs a load_module line.

worker_processes auto;
worker_rlimit_nofile 16384;

events {
    worker_connections 8192;
}

stream {
    log_format egress '$remote_addr [$time_local] upstream=$upstream_addr '
                      'status=$status sent=$bytes_sent received=$bytes_received '
                      'time=$session_time';
    access_log /var/log/nginx/egress.log egress;

    resolver 10.20.0.2 valid=30s;

    upstream provider_gateway {
        zone provider_gateway 64k;
        server gateway.proxyforge.io:PORT max_fails=2 fail_timeout=30s resolve;
        server gateway-b.example.net:PORT backup resolve;
    }

    server {
        listen 10.20.0.7:3128;
        allow 10.20.0.0/24;
        deny all;

        proxy_pass provider_gateway;
        proxy_connect_timeout 5s;
        proxy_timeout 10m;
        proxy_next_upstream on;
        proxy_next_upstream_tries 2;
    }
}

Each relayed session holds two connections, one to the client and one to the gateway, so size worker_connections for twice your concurrency and raise the file limit with it. resolver should be a resolver the host can reach; 10.20.0.2 is a stand-in.

Failover behaves differently from HAProxy. With proxy_next_upstream on, a refused connection to the primary is retried on the next server inside the same client session, backup included. When we stopped the primary, no client saw an error, and the log showed both attempts:

10.20.0.50 [08/Oct/2026:18:02:48 +0000] upstream=198.51.100.21:3128, 198.51.100.22:3128 status=200 sent=3279 received=1930 time=0.006

After max_fails failures the error log reads upstream server temporarily disabled, and the primary is skipped for fail_timeout. The limit is what Nginx counts as a failure: only connection errors and timeouts. Pointed at the stand-in gateway that answered every request with 500, Nginx logged status=200 for each session while the client printed CONNECT tunnel failed, response 500, and it never moved to the backup. The open-source build has no active checks to catch that: adding health_check fails with unknown directive "health_check".

The relay is protocol-agnostic. The same configuration carried a SOCKS5 session with socks5h:// credentials when we pointed it at a SOCKS server, and a TCP-mode HAProxy frontend did the same.

DNS: names that change and names that do not resolve

Both tools resolve hostnames in their configuration at startup by default, and both refuse to start if a name does not resolve. Nginx without resolve stops with host not found in upstream; HAProxy with default settings stops with could not resolve address and Failed to initialize server(s) addr. A gateway hostname that is briefly unresolvable during a deploy should not take your egress down, and one whose addresses change should be followed.

  • HAProxy: a resolvers section with parse-resolv-conf, plus resolvers system init-addr last,libc,none on each server. With that, an unresolvable name logs could not resolve address ..., disabling server and HAProxy starts with that server disabled until DNS answers.
  • Nginx: a resolver directive, a shared-memory zone in the upstream, and resolve on each server. Nginx 1.28 and 1.29 accepted this in our tests; Nginx 1.26 rejected it with invalid parameter "resolve", so on older versions the name is resolved once at startup.

Never add send-proxy or send-proxy-v2 to a server line pointing at a provider. The PROXY protocol header arrives where the gateway expects an HTTP request or SOCKS greeting; our test gateway answered 400 and logged error:invalid-request.

HAProxy or Nginx?

You need Use
Clients hold no credential, and the relay must add it HAProxy, HTTP mode
Failover between two providers HAProxy, HTTP mode, one backend per provider
Failover between endpoints that accept the same credential Either; Nginx retries inside the client session, HAProxy fails over after its check
Detection of a gateway that accepts connections but fails requests HAProxy, with a credentialed check
A relay for SOCKS5 clients HAProxy TCP mode or Nginx stream
One more listener on an Nginx fleet you already run Nginx stream, accepting passive checks

Whichever you choose, run two instances. A single relay in front of two gateways is still a single point of failure; put the instances behind a virtual address or a cloud load balancer, and alert on the log lines above. Proxy monitoring covers which signals separate a gateway problem from a target blocking you, and the proxy checker shows what a target sees through a connection, including the exit address and headers, which is a quick way to confirm the relay changed nothing you did not intend.

Putting this in front of ProxyForge

ProxyForge gateways accept username and password or an IP allowlist on every proxy line. The gateway endpoint and credentials for an order are shown on that order in the dashboard. For the HTTP-mode setup, base64-encode that credential into the relay's environment; for TCP mode or Nginx, either leave the credential with the clients or allowlist the relay's public address in the dashboard so clients need none.

If ProxyForge is the second provider behind HAProxy rather than the first, the migration page describes how we run in parallel with your current vendor, and pricing is pay-as-you-go with no monthly minimum, which keeps a standby backend cheap to hold.

FAQ

Related questions

Can HAProxy act as a forward proxy?

Not on its own: it does not resolve and fetch arbitrary destinations. In HTTP mode it can sit in front of a forward proxy, pass CONNECT and absolute-form requests through to it, and add a Proxy-Authorization header on the way. In TCP mode it relays bytes and never reads the request.

Can Nginx add proxy credentials to CONNECT requests?

No. The http module does not proxy CONNECT, and the stream module relays TCP without parsing it, so the client must send its own credential or the gateway must authenticate the Nginx host by IP allowlist.

Does open-source Nginx have active health checks for stream upstreams?

No. The health_check directive is rejected as unknown in the open-source build. Stream upstreams are checked passively: max_fails connection errors or timeouts within fail_timeout take a server out of rotation for fail_timeout.

Should I enable the PROXY protocol towards the provider gateway?

No. A proxy gateway expects the first bytes of a connection to be an HTTP request or a SOCKS greeting. With send-proxy on the server line, our test gateway answered 400 Bad Request to every connection.

Run it on a network you can account for

Order from 1 GB or 1 IP with no monthly minimum, or talk to an engineer about your workload first.

One business day, from a named engineer.