A Squid upstream proxy setup puts one Squid server between your scrapers and a proxy provider's gateway. Clients send their traffic to Squid with no credentials; Squid checks who they are and where they are going, then forwards every request to the gateway with a cache_peer line that carries the provider credential, and never_direct allow all makes sure nothing leaves any other way. The credential lives on one host, policy lives in one file, and one log records what every team sent.
This guide builds that configuration step by step and explains what each line does. Every configuration below was run on Squid 6.13 (the ubuntu/squid container image) against a local stand-in gateway that requires Basic proxy authentication, and the outputs and log lines shown come from those runs, with lab hostnames replaced by example domains. Replace gateway.proxyforge.io:PORT, USERNAME and PASSWORD with the values your provider issues.
Why put Squid in front of a proxy gateway?
Pointing every scraper at the provider directly works until there are several teams, several hosts and a security review. A central forward proxy gives you four things the direct setup does not:
- One copy of the credential. Application hosts, CI runners and notebooks hold the address of your Squid server, not the provider's password. Rotating the credential is a change on one machine.
- One place for policy. Which networks may use the egress, which ports they may tunnel to, and which teams may reach which domains are ACLs in
squid.conf, reviewed like any other configuration. - One log. Squid's access log records the client address, the method, the destination and the result of every request, including HTTPS tunnels. That is the record a reviewer asks for, and the first thing you read during an incident.
- One egress address. If the provider authenticates by IP allowlist instead of a password, you register the Squid host's public address once, rather than every NAT gateway your workloads might leave from. IP allowlist vs username and password compares the two modes.
The cost is one more hop and one more component to keep alive. Run two Squid instances behind a load balancer, or the HAProxy setup in HAProxy and Nginx in front of a proxy gateway, if the egress must survive a host failure.
A minimal squid.conf for an upstream gateway
This is a complete configuration, not a fragment to merge into the distribution's default file. Replace /etc/squid/squid.conf with it, or point Squid at it with squid -f:
http_port 10.20.0.5:3128
acl scrapers src 10.20.0.0/24
acl SSL_ports port 443
acl Safe_ports port 80 443
acl CONNECT method CONNECT
http_access deny !Safe_ports
http_access deny CONNECT !SSL_ports
http_access allow scrapers
http_access deny all
cache_peer gateway.proxyforge.io parent PORT 0 no-query no-digest name=provider login=USERNAME:PASSWORD connect-timeout=10 connect-fail-limit=5
never_direct allow all
cache deny all
forwarded_for delete
via off
logformat egress %ts.%03tu %>a %rm %ru %>Hs %<st %Ss/%Sh/%<A %tr
access_log daemon:/var/log/squid/access.log egress
What each part does:
| Directive | Effect |
|---|---|
http_port 10.20.0.5:3128 |
Listens on the internal address only, so the proxy is not reachable from outside |
acl scrapers src ... and http_access |
Only that network may use the proxy; everyone else gets 403 |
deny CONNECT !SSL_ports |
Tunnels are allowed to port 443 only, which stops the proxy being used to reach arbitrary TCP services |
cache_peer ... parent PORT 0 |
The provider gateway is a parent; 0 disables ICP, a cache-to-cache protocol with no role here |
no-query no-digest |
No ICP queries and no cache-digest fetches to the gateway |
login=USERNAME:PASSWORD |
Squid sends this as Basic Proxy-Authorization on every request to the peer |
connect-timeout=10 connect-fail-limit=5 |
How long a connection to the gateway may take, and how many failures mark it dead |
never_direct allow all |
Squid never connects to a destination itself; if no peer is available the request fails instead of leaking out |
forwarded_for delete, via off |
Squid stops announcing your internal client addresses and its own hostname upstream |
The last two lines matter more than they look. With Squid's defaults, our test gateway received X-Forwarded-For: 10.20.0.50 and Via: 1.1 <squid hostname> (squid/6.13) on every request, including the CONNECT for an HTTPS target, and plain-HTTP targets received both headers too. That is your internal addressing and host naming sent to the provider and to the sites you collect from. With via off, Squid logs WARNING: HTTP requires the use of Via at startup; the warning is expected.
Check the file before you load it with squid -k parse, then reload a running instance with squid -k reconfigure.
How Squid handles HTTPS: CONNECT, not interception
Almost every target is HTTPS, and this is where people expect to need certificates. You do not. The client sends CONNECT www.example.com:443 to Squid. Squid applies its ACLs, sends its own CONNECT www.example.com:443 to the gateway with the Proxy-Authorization header attached, and once the gateway answers 200 it relays bytes in both directions. The TLS session runs end to end between the client and the target; Squid and the gateway see only the hostname and port.
That has three consequences worth stating to a reviewer:
- No TLS interception, no CA to distribute.
ssl_bumpand certificate generation are separate features for inspecting traffic. An egress to a proxy provider does not need them, and leaving them off keeps Squid out of the trust path. - ACLs work on hostnames, not URLs. For HTTPS,
dstdomainand port ACLs are all the policy Squid can enforce. URL-path rules only apply to plain-HTTP requests. - Squid does not resolve your targets. With
never_directand hostname ACLs, Squid hands the name to the gateway. In our test, a hostname that Squid's host could not resolve at all was fetched successfully, because only the gateway needed to resolve it. That suits locked-down networks with no external DNS, and it means lookups for target names happen at the provider. Adst(IP address) ACL would force Squid to resolve every destination, so preferdstdomain.
Credentials: escaping, storage and rotation
Squid URL-unescapes the login= value. A colon or @ in the password needs no escaping; our test password contained both. A literal % must be written as %%; we confirmed that a password containing % authenticated only when written that way.
Keep the credential out of the main file so the policy can be reviewed and versioned without exposing it. Put the peer lines in their own file, render it from your secrets manager at deploy time, and make it readable only by root and the Squid user (proxy on Debian and Ubuntu):
include /etc/squid/provider-peers.conf
Rotating the credential is then: write the new file, run squid -k reconfigure, and watch the log. We changed the password in the included file and reloaded; the next requests used it immediately. Managing proxy credentials in Vault and AWS Secrets Manager covers the store side and a rotation order that avoids an outage.
One design question comes with a shared credential. If the connection string your provider generated pins a sticky session, every client behind Squid now shares that session, and with it one exit address. For a shared egress, a per-request rotating credential is the safer default, and teams that need sticky behavior get their own peer, as below. If clients must choose their own session parameters, login=PASSTHRU forwards whatever credential each client sends; it worked in our test for both CONNECT and plain HTTP, but it puts credentials back on the clients and gives up the main reason for the central proxy. Rotating vs sticky proxies explains when each mode fits.
Per-team ACLs and separate credentials
Two teams, two source networks, two provider credentials, and one team limited to the domains it is approved to collect from. These lines replace the scrapers ACL, the http_access rules and the peer of the minimal file; the port ACLs and the rest stay as they were:
acl pricing_team src 10.20.0.0/24
acl research_team src 10.20.1.0/24
acl research_sites dstdomain .example.com .example.org
http_access deny !Safe_ports
http_access deny CONNECT !SSL_ports
http_access allow pricing_team
http_access allow research_team research_sites
http_access deny all
cache_peer gateway.proxyforge.io parent PORT 0 no-query no-digest name=pricing login=USERNAME:PASSWORD connect-timeout=10 connect-fail-limit=5
cache_peer gateway.proxyforge.io parent PORT 0 no-query no-digest name=research login=USERNAME_2:PASSWORD_2 connect-timeout=10 connect-fail-limit=5
cache_peer_access pricing allow pricing_team
cache_peer_access pricing deny all
cache_peer_access research allow research_team
cache_peer_access research deny all
never_direct allow all
Both peers point at the same host and port, which Squid accepts only because each has a distinct name=. cache_peer_access decides which peer a request may use, so each team's traffic arrives at the gateway under its own credential, and the provider's usage reporting splits along the same line. A research request to a domain outside research_sites is refused with 403 before it reaches the gateway.
A dstdomain entry with a leading dot matches the domain and all of its subdomains. Do not list both example.com and .example.com: Squid 6 treats that as a fatal configuration error ('.example.com' is a subdomain of 'example.com') and refuses to start.
Why caching is mostly off
Squid is a caching proxy first, and its defaults reflect that. For an egress to a proxy gateway, cache deny all is the right starting point:
- HTTPS cannot be cached here. Tunneled responses are opaque to Squid. Since nearly every target is HTTPS, a cache would hold almost nothing.
- Collection wants change. A price check or a ranking measurement exists to see the current page. A cached copy is a silent wrong answer.
- Responses vary by exit. The same URL fetched through different exit addresses or countries can return different content. A shared cache would serve one team's geo-specific page to another.
- A cache hides blocks. If a cached challenge page or error is served, the client never learns that the target started refusing it.
The exception is a large, static, plain-HTTP asset that many jobs fetch, which is rare enough that it is better handled in the application than by turning a cache on for everything.
Logging what the egress carries
The egress log format above records the epoch time, client address, method, URL (or host:port for a tunnel), status, bytes, Squid's result and hierarchy code, the peer name and the response time. Lines from our test run:
1791481810.868 10.20.0.50 GET http://www.example.com/ 200 287 TCP_MISS/FIRSTUP_PARENT/pricing 0
1791481811.363 10.20.1.50 CONNECT www.example.org:443 200 3279 TCP_TUNNEL/FIRSTUP_PARENT/research 8
1791481811.833 10.20.0.50 CONNECT www.example.com:8443 403 3348 TCP_DENIED/HIER_NONE/- 0
Squid strips query strings from logged URLs by default (strip_query_terms on), which keeps tokens in query strings out of the log. Leave it on unless you need the full URL for debugging. The egress format logs no headers, so the credential Squid adds never appears in the access log. Keep it that way: do not add %>h to the format or turn on log_mime_hdrs without filtering.
The results that tell you where a failure is:
| What the client sees | Squid's log | Where the problem is |
|---|---|---|
403 |
TCP_DENIED/HIER_NONE/- |
Your own ACLs refused the request; it never reached the gateway |
407, or an aborted CONNECT |
407 against the peer name |
The gateway rejected the credential in login=, since this Squid does not authenticate clients at all |
503 with X-Squid-Error: ERR_CONNECT_FAIL 111 |
503 against the peer name |
Squid could not connect to the gateway |
200 |
TCP_TUNNEL/FIRSTUP_PARENT/<peer> |
The tunnel opened; anything wrong after this is between the client and the target |
Ship the log to the same place as your other metrics and alert on the second and third rows; proxy monitoring covers which rates are worth paging on.
Failover to a second gateway
List a second parent after the first, with the same options and its own name= and credential. Squid uses the first live parent in configuration order. When we stopped the first gateway, Squid logged Detected DEAD Parent: primary, sent requests to the second, logged with the hierarchy code ANY_OLD_PARENT or FIRSTUP_PARENT and the second peer's name, without the client seeing an error, and moved back on its own after Detected REVIVED Parent: primary once the first returned.
That covers a gateway that refuses connections. It does not cover a gateway that accepts the connection and then answers with an error, which Squid passes back to the client like any other response. Health checks that speak the proxy protocol, and failover between two providers with different credentials, are easier in HAProxy; the HAProxy and Nginx guide shows both.
Replacing a legacy forward-proxy appliance
Many organizations already run a forward proxy for staff browsing: a Windows proxy suite such as WinGate, a hardware appliance, or an older Squid that nobody has touched in years. These products bundle an HTTP proxy, often a SOCKS server, a user database, a cache, access rules and a way to send traffic on to another proxy. When scraping egress has to move onto a maintained, reviewable configuration, map each function before you switch anything:
| Function in the old product | In Squid | Notes |
|---|---|---|
| HTTP and HTTPS proxy on a LAN port | http_port |
Keep the old address and port, or update the PAC file or client policy that points at it |
| Sending traffic on to another proxy | cache_peer with never_direct |
One peer per credential or provider |
| Rules by network, site, time or user | acl with src, dstdomain, time, proxy_auth, then http_access |
First match wins, as in most appliances |
| Staff logins | auth_param with a helper; the Ubuntu package ships negotiate_kerberos_auth, basic_ldap_auth and others |
For a scraping egress, source-network ACLs are usually enough |
| Cache | cache_mem, cache_dir |
Off for scraping traffic, as above |
| SOCKS server | Not Squid | A SOCKS daemon such as Dante; see running your own SOCKS5 proxy with Dante |
| Usage reports | access_log with a logformat |
Build reports in your log pipeline |
| NAT, DHCP, mail relay, VPN | Not a proxy function | Move them to the firewall or to dedicated services first |
Run Squid on a new address alongside the old product, move one team's clients at a time, and compare both logs for the same jobs before retiring anything. Rules that the old product applied implicitly, such as which ports it would tunnel to, are the usual surprise; write them down as explicit ACLs.
Chaining through a corporate proxy
Sometimes the only way out of the network is an existing corporate proxy, and the provider gateway has to sit behind it. There are two cases.
You administer the corporate proxy. Add the provider as a parent there, used only by the scraping hosts, while everything else keeps leaving directly:
acl scraping_hosts src 10.20.0.0/24
cache_peer gateway.proxyforge.io parent PORT 0 no-query no-digest name=provider login=USERNAME:PASSWORD
cache_peer_access provider allow scraping_hosts
cache_peer_access provider deny all
never_direct allow scraping_hosts
In our test, a request from the scraping network was logged as FIRSTUP_PARENT/provider while a request from another office network went HIER_DIRECT as before.
Someone else administers it. Squid cannot open a tunnel to its parent through another proxy, so put a small tunnel in front of it. socat can open a CONNECT tunnel through the corporate proxy to the gateway and expose it on a local port:
socat TCP-LISTEN:13128,bind=127.0.0.1,fork,reuseaddr \
PROXY:corp-proxy.internal.example:gateway.proxyforge.io:PORT,proxyport=8080
Your egress Squid then peers with 127.0.0.1 port 13128 instead of the gateway, with the same login=. Add proxyauth=user:password to the socat address if the corporate proxy requires a login. Start the tunnel before Squid; when Squid started first in our test, it marked the peer dead and refused tunnels until its next probe revived it.
Three things to settle with the network and security teams before relying on this:
- The corporate proxy must permit CONNECT to the gateway's host and port. Many permit CONNECT to port 443 only, and the gateway port is often different.
- The corporate proxy's own log will show only
CONNECT gateway.proxyforge.io:PORTand a byte count, never the target sites, because the request to the gateway runs inside the tunnel. Your egress Squid's log becomes the record of what was collected, so say so in the review. - The gateway sees connections from the corporate proxy's public address. That is the address to allowlist if you use IP authentication.
Testing the setup
Test from a client host on an allowed network, with no credential in the proxy URL:
curl -sS -x http://10.20.0.5:3128 https://api.ipify.org; echo
curl -sS -o /dev/null -w '%{http_code}\n' -x http://10.20.0.5:3128 http://www.example.com/
The first prints the exit address the provider assigned, not your own. Then confirm the controls hold: the same command from a host outside the allowed network should return 403 (curl reports CONNECT tunnel failed, response 403), and https://www.example.com:8443/ should be refused the same way. Read /var/log/squid/access.log after each test; the hierarchy field should name your peer for every allowed request.
Running Squid in front of ProxyForge
ProxyForge gateways accept username and password or an IP allowlist on every proxy line. Squid peers with the HTTP proxy port; the gateway endpoint and credentials are shown on each order in the dashboard, so copy the values from there into provider-peers.conf instead of assembling them by hand. If you would rather allowlist your Squid host's public address, drop login= from the peer line and register the address in the dashboard.
Billing is pay-as-you-go from a prepaid wallet with no monthly minimum, so the egress can start with one team's traffic and grow from there; current rates are on the pricing page. If the egress is part of a move away from another provider, the migration page explains how we run both in parallel.