Skip to content
ProxyForge

Kubernetes egress proxy setup for scraping workloads

ProxyForge engineeringUpdated 10 min read

A Kubernetes egress proxy setup for scraping workloads has four parts: the proxy URL injected into the pod from a Secret, a NO_PROXY list that keeps cluster-internal traffic off the proxy, a stable egress IP if the provider authenticates you by allowlist, and a restart procedure for when the credential changes. Kubernetes has no built-in "proxy" setting for pods; each of these is ordinary configuration you apply per workload, which is also what lets you scope the proxy to the scrapers and nothing else.

This guide assumes a scraping service that reads its proxy from PROXY_URL, shaped like http://USERNAME:[email protected]:PORT. The provider's dashboard generates the exact string for the country and session mode you choose. All manifests below are standard v1 and apps/v1 objects and apply to any conformant cluster.

Store the proxy URL in a Secret

The credential belongs in a Secret in the same namespace as the workload. Pods can only reference Secrets in their own namespace, so that is the natural boundary for who can use which credential.

Create it from your shell or pipeline rather than from a file in the repository:

kubectl create secret generic proxy-credentials \
  --namespace pricing-scrapers \
  --from-literal=PROXY_URL="$PROXY_URL"

The equivalent manifest, for reference or for a templating tool that renders it at deploy time, looks like this. Do not commit it with a real value in it:

apiVersion: v1
kind: Secret
metadata:
  name: proxy-credentials
  namespace: pricing-scrapers
type: Opaque
stringData:
  PROXY_URL: "http://USERNAME:[email protected]:PORT"

If your organization already keeps credentials in Vault or AWS Secrets Manager, sync them into the cluster with a controller such as the External Secrets Operator instead of creating the Secret by hand. Managing proxy credentials in Vault and AWS Secrets Manager covers the store side: layout, access policy and rotation.

Two properties of Secrets are worth stating plainly to whoever reviews the setup. They are only base64-encoded in the API; encryption at rest in etcd is a cluster setting you have to enable (managed services usually offer it with a KMS key). And anyone who can create pods in the namespace can mount any Secret in that namespace, so RBAC on the Secret object alone does not fully contain it. The Kubernetes Secrets documentation lists the remaining hardening options.

Inject the Kubernetes egress proxy into a Deployment

Reference the Secret with valueFrom.secretKeyRef. The Deployment below exposes PROXY_URL to application code, and also sets the conventional HTTP_PROXY and HTTPS_PROXY variables for tools that only read those, using Kubernetes' $(VAR) expansion so the credential is still defined in exactly one place:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: price-scraper
  namespace: pricing-scrapers
spec:
  replicas: 3
  selector:
    matchLabels:
      app: price-scraper
  template:
    metadata:
      labels:
        app: price-scraper
    spec:
      containers:
        - name: scraper
          image: registry.example.com/price-scraper:1.4.2
          env:
            - name: PROXY_URL
              valueFrom:
                secretKeyRef:
                  name: proxy-credentials
                  key: PROXY_URL
            - name: HTTP_PROXY
              value: "$(PROXY_URL)"
            - name: HTTPS_PROXY
              value: "$(PROXY_URL)"
            - name: NO_PROXY
              value: "localhost,127.0.0.1,kubernetes.default.svc,.svc,.cluster.local,10.96.0.0/12,10.244.0.0/16,169.254.169.254"
          resources:
            requests:
              cpu: 250m
              memory: 256Mi
            limits:
              memory: 512Mi

$(PROXY_URL) only expands if PROXY_URL is defined earlier in the same env list, so keep the order. The pod spec stores the literal string $(PROXY_URL), not the credential, so kubectl get deployment -o yaml does not reveal it.

Whether to set HTTP_PROXY and HTTPS_PROXY at all is a real decision. They are convenient, because most HTTP clients honor them without code changes. They are also indiscriminate: every library in the process that respects them, including cloud SDKs, telemetry exporters and your secrets client, will route through the proxy unless NO_PROXY exempts the destination. If your scraper code passes the proxy explicitly (a proxies dict in Requests, proxy= in httpx), prefer exposing only PROXY_URL and leave the global variables unset. Setting up a Python Requests proxy shows the explicit pattern and how environment variables interact with it.

For CronJobs and Jobs, the same env block goes under spec.jobTemplate.spec.template.spec.containers.

Keep cluster traffic off the proxy with NO_PROXY

With HTTPS_PROXY set, a pod that calls another Service, the Kubernetes API or a cloud metadata endpoint will try to send that request to the proxy gateway. At best it fails; at worst it succeeds and you pay per-GB rates to reach your own database. NO_PROXY is the exemption list, and in a cluster it needs:

Entry Covers
localhost,127.0.0.1 Sidecars and local agents
kubernetes.default.svc The API server by its in-cluster name
.svc,.cluster.local Every Service reached by DNS name
Service CIDR (for example 10.96.0.0/12) Clients that connect by ClusterIP, including the API server via KUBERNETES_SERVICE_HOST
Pod CIDR (for example 10.244.0.0/16) Direct pod-to-pod calls
169.254.169.254 Cloud instance metadata, used for workload identity on several providers

The CIDRs in the example are common defaults, not universal ones. Read your own: the Service range is a control-plane flag (or a field in your managed cluster's configuration), and the pod range depends on the CNI. On clusters where pods take addresses from the VPC, the pod range is your VPC subnets.

Two client behaviors make this list less reliable than it looks:

  • CIDR support varies. Go's standard library, Python Requests and curl from 7.86 onward match addresses against CIDR entries. Other clients match only hostnames and suffixes. Because in-cluster clients often connect by IP (the Kubernetes client libraries use KUBERNETES_SERVICE_HOST, which is an IP), keep both the names and the ranges in the list and test the clients you actually ship.
  • Case varies. curl reads only the lowercase http_proxy for plain HTTP targets; Go and Python read both cases. If an image mixes tools, set the lowercase variants too, with the same values.

Why IP allowlisting needs a stable egress IP

Username and password authentication works from anywhere, which suits a cluster. IP allowlist authentication, where the proxy gateway accepts connections from registered source addresses without a credential, only works if the cluster leaves the internet from addresses that do not change. By default, it does not.

Pod traffic leaving the cluster is normally source-NATed to the node's address. Nodes come and go: the cluster autoscaler adds and removes them, upgrades replace them, spot or preemptible capacity is reclaimed, and a replaced node rarely gets the same public IP. If you allowlisted node IPs, scraping breaks the first time the node pool rolls. IP allowlist vs username and password compares the two modes in general; the Kubernetes-specific options for a stable egress are:

Option What stays stable Scope Notes
Private nodes behind a NAT gateway with a static IP The NAT's public address Everything in those subnets AWS NAT gateway uses an Elastic IP per gateway, so one address per availability zone. On Google Cloud, reserve addresses and use manual allocation on Cloud NAT; auto-allocated addresses can change. Azure NAT Gateway takes a static public IP or prefix.
Dedicated node pool in its own subnet and NAT A separate address set Workloads scheduled onto that pool Use a taint and toleration so only scrapers land there. Keeps scraping egress separate from the rest of the cluster.
Egress gateway (service mesh or CNI feature) Addresses on designated gateway nodes Selected pods or namespaces Per-workload control without a separate subnet, at the cost of another component to operate.

Whichever you choose, allowlist every address the path can use, one per zone for zonal NAT gateways, and confirm from inside a pod that the address the proxy sees is the one you registered:

kubectl exec -n pricing-scrapers deploy/price-scraper -- \
  sh -c 'curl -sS --max-time 15 https://api.ipify.org; echo'

Run that with the proxy variables unset in a debug pod to see the cluster's raw egress address, and with -x "$PROXY_URL" to see the exit address the target will see.

Per-namespace credentials for teams

When several teams run scrapers on one cluster, give each team its own namespace and its own Secret, and grant each team's service accounts and humans access only to their namespace. The result is three things you want in an incident: a credential leak is scoped to one team, you can rotate one team's credential without touching the others, and proxy usage can be attributed to a namespace in your own metrics.

A Role that lets a team manage its workloads but not read Secrets looks like the usual pattern: grant create, update and patch on Deployments, and grant nothing on secrets to humans. Remember that list on Secrets returns their contents, not just their names, so a "read-only" role that includes it is not read-only in practice.

How far you can separate credentials on the provider side depends on the vendor's account model. Ask how credentials are issued and whether they can be revoked independently. On ProxyForge, credentials and the gateway address are issued per account in the dashboard, and dashboard access for a team is managed through an organization with Owner, User and Read only roles; proxy team access walks through how that maps onto a platform team and its consumers.

Can NetworkPolicy force traffic through a Kubernetes egress proxy?

A common request is "make the scraper pods use the proxy and nothing else". NetworkPolicy can help, but it is important to be precise about what it does. It does not route traffic and it cannot force a connection through a proxy. It is an allow list of destinations (pods, namespaces and IP blocks, by port) enforced by the CNI plugin, and only if the plugin implements it. What you can do is remove every other path, so that a pod that ignores its proxy settings fails instead of leaking out directly.

Standard NetworkPolicy has no concept of hostnames, and a provider's gateway hostname can resolve to addresses that change. The dependable pattern is therefore an in-cluster forwarding proxy: a small Deployment (Squid and similar forwarders can chain to an authenticated upstream) that holds the provider credential and is the only thing allowed to reach the internet. Scraper pods get PROXY_URL pointing at the forwarder's Service and a policy that permits only DNS and the forwarder:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: price-scraper-egress
  namespace: pricing-scrapers
spec:
  podSelector:
    matchLabels:
      app: price-scraper
  policyTypes:
    - Egress
  egress:
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: kube-system
          podSelector:
            matchLabels:
              k8s-app: kube-dns
      ports:
        - protocol: UDP
          port: 53
        - protocol: TCP
          port: 53
    - to:
        - podSelector:
            matchLabels:
              app: egress-forwarder
      ports:
        - protocol: TCP
          port: 3128

Details that decide whether this works:

  1. The policy applies only to pods it selects. Other workloads in the namespace are unaffected, which is what you want when only some workloads should be proxied.
  2. Ports are the pod's ports, not the Service's. Policies are evaluated after Service address translation, so use the forwarder container's port.
  3. DNS must be allowed explicitly. Check your DNS pods' labels; if the cluster runs NodeLocal DNSCache, pods query a link-local address instead and need an ipBlock rule for it.
  4. The forwarder needs its own egress rule if the namespace has a default-deny policy.
  5. Confirm enforcement. Some CNIs accept NetworkPolicy objects without enforcing them. Test by running a pod with the scraper's labels and checking that a direct request to a public site fails.

Some CNIs extend NetworkPolicy with hostname-based egress rules (Cilium's FQDN policies, for example), which lets you allow the provider gateway directly without a forwarder. That is a CNI-specific resource, not portable Kubernetes. The NetworkPolicy reference documents the portable behavior.

Rotating the Secret and restarting pods

Environment variables from a Secret are resolved when the container starts. Updating the Secret changes nothing in running pods; they keep the old credential until they are replaced. Rotating a Kubernetes egress proxy credential therefore has a fixed shape:

  1. Obtain the new credential and confirm it works from a debug pod with curl -x.
  2. Update the Secret, by kubectl apply, by your pipeline, or by letting the secrets controller sync the new version from the store.
  3. Restart the workloads that consume it: kubectl rollout restart deployment/price-scraper -n pricing-scrapers. A rolling restart replaces pods gradually, so capacity is maintained.
  4. Watch for 407 Proxy Authentication Required in logs and metrics during the rollout. A 407 after the restart means a pod picked up a stale or malformed value.
  5. Revoke the old credential only when no pod still uses it.

If your templates are rendered by Helm or a similar tool, add a checksum of the Secret data as a pod template annotation so that a changed credential changes the template and triggers a rollout on the next deploy. A controller that watches Secrets and restarts their consumers does the same thing continuously.

The alternative is to mount the Secret as a file and have the application re-read it, for example whenever the gateway returns a 407. Mounted Secrets are refreshed by the kubelet after a sync delay, not instantly, and never when mounted with subPath. It avoids restarts, at the cost of reload logic in every service. For most scraping fleets, a rolling restart is simpler and easier to audit. Either way, alert on 407 rates as part of your proxy monitoring: it is the signal that distinguishes a credential problem from a target blocking you.

Running this on ProxyForge

Every ProxyForge line, residential, mobile, ISP and datacenter, accepts both username and password and IP allowlist authentication over HTTP, HTTPS and SOCKS5, so you can start with a credential in a Secret and move to allowlisting once your cluster has a stable NAT address. The dashboard generates the connection string for the country and session mode you pick; paste it into the Secret rather than assembling it by hand.

If your scrapers need fixed addresses of their own on the target side as well, ISP proxies and datacenter proxies are billed per address and dedicated to one customer. Pricing is pay-as-you-go from a prepaid wallet, with current rates on the pricing page.

FAQ

Related questions

Can I set a proxy for every pod in a Kubernetes cluster at once?

Not through a native Kubernetes setting. Proxy environment variables are set per container, so a cluster-wide default needs a mutating admission webhook, a shared Helm values block, or a base image convention. Most teams scope it per workload instead, because a cluster-wide proxy also captures traffic that should never leave the cluster.

Does NO_PROXY support CIDR ranges?

It depends on the client. Go's standard library, Python Requests and recent curl releases match IP ranges in CIDR notation, while other clients only match hostnames and domain suffixes. List both the ranges and the internal domain suffixes so every client in the image behaves the same way.

Do pods see an updated Secret without a restart?

Environment variables are read once when the container starts, so a pod keeps the old value until it is recreated. A Secret mounted as a volume is refreshed by the kubelet after a delay, except when it is mounted with subPath, but your application still has to re-read the file.

Should the proxy credential live in a ConfigMap or a Secret?

A Secret. ConfigMaps are readable by anyone with broad read access to the namespace and are routinely printed by tooling, while Secrets can be restricted separately with RBAC and encrypted at rest in etcd.

Run it on a network you can account for

Order from 1 GB or 1 IP with no monthly minimum, or talk to an engineer about your workload first.

One business day, from a named engineer.