Skip to content
ProxyForge

AWS Lambda proxy setup: credentials, egress IPs and timeouts

ProxyForge engineeringUpdated 8 min read

An AWS Lambda proxy setup has three decisions: how the function authenticates to the proxy, which address its traffic leaves from, and how proxy timeouts fit inside the function's own timeout. Lambda functions have no stable egress IP unless they run in a VPC behind a NAT gateway with an Elastic IP, so username and password authentication is the usual default. Configure the proxy explicitly in your HTTP client rather than through HTTPS_PROXY, keep AWS service calls off the proxy, and create the HTTP session outside the handler so warm invocations reuse it.

The examples use Python with Requests, but every point applies to Node.js, Go and Java functions with their respective clients. The proxy URL has the shape http://USERNAME:[email protected]:PORT; the provider dashboard generates the exact string for your country and session mode.

Where a Lambda function's traffic comes from

A function's outbound address depends on its network configuration, and that determines which proxy authentication modes are available to you:

Configuration Egress address Proxy auth that works Cost to consider
No VPC (default) AWS-owned public addresses, shared and changing Username and password only None beyond Lambda
VPC, private subnets, NAT gateway with Elastic IP The Elastic IP of the NAT in the function's availability zone Username and password, or IP allowlist NAT gateway hourly charge plus per-GB processing on all traffic through it
VPC, no NAT gateway None; the function cannot reach the internet Neither Not usable for scraping

A function outside a VPC runs on infrastructure whose outbound addresses come from AWS's large public pools. You do not control them, other customers use the same ranges, and a new execution environment can leave from a different address than the last one. Allowlisting them is not possible in any meaningful sense.

Attaching the function to a VPC changes the picture, but only with the right routing. Lambda functions in a VPC never receive public IP addresses, even in a public subnet. The documented pattern is to place the function in private subnets whose default route points at a NAT gateway in a public subnet; the NAT gateway's Elastic IP is then the function's stable egress address. AWS's guide to internet access for VPC-connected functions walks through the subnets and route tables.

For availability you will usually run one NAT gateway per availability zone, which means one Elastic IP per zone. Register all of them with the proxy provider. IP allowlist vs username and password covers the trade-offs of each mode in more depth; for Lambda, the practical rule is that credentials are simpler unless you already run functions in a VPC for other reasons.

Handling the AWS Lambda proxy credential

Lambda environment variables are encrypted at rest, but anyone with permission to read the function's configuration sees them in plaintext in the console and API. A proxy URL with an embedded password is a credential, so store it in AWS Secrets Manager and read it during initialization, putting only the secret's name in the environment.

The execution role needs secretsmanager:GetSecretValue on that one secret. Managing proxy credentials in Vault and AWS Secrets Manager has a least-privilege policy, a caching loader, and a rotation workflow that also applies here. Two Lambda-specific details:

  • Read the secret in the init phase, not per invocation. Code at module level runs once per execution environment, and its result is reused by every warm invocation.
  • Rotation reaches warm environments only when they refresh. An environment that read the old value at init keeps it until it is recycled. Either refresh on a timer and on a 407, as the loader in that guide does, or publish a new function version after rotating so new environments start with the new value.

HTTPS_PROXY versus an explicit proxy in code

There are two ways to point a function at a proxy, and they behave differently.

Environment variables. Setting HTTP_PROXY and HTTPS_PROXY on the function configuration makes Requests, urllib3, and most other clients use the proxy without code changes. It also makes the AWS SDK use it. botocore honors the same variables, so every S3 upload, SQS message and Secrets Manager read now goes through the scraping proxy: slower, billed as proxy traffic, and broken outright for VPC interface endpoints, whose private addresses are unreachable from the proxy's exit. If you use this approach, exclude AWS and local endpoints:

NO_PROXY=localhost,127.0.0.1,.amazonaws.com

localhost and 127.0.0.1 cover the Lambda runtime API and any extensions listening locally, such as the AWS Parameters and Secrets extension. .amazonaws.com covers AWS service endpoints; with private DNS enabled on an interface endpoint, the SDK still addresses the service by its public name, so the same entry exempts it.

Explicit configuration. The more predictable option is to leave the environment variables unset and give the proxy only to the HTTP client that fetches target pages. AWS SDK calls then go directly (or through the NAT and VPC endpoints), and there is no exemption list to get wrong. The example below does this, and sets trust_env = False so that a proxy variable added to the function later cannot change the behavior. Setting up a Python Requests proxy explains the precedence rules between the two approaches.

If the function runs in a VPC, add a gateway endpoint for S3 and interface endpoints for the other services it calls. That keeps AWS traffic off the NAT gateway's per-GB processing charge, and makes the explicit-proxy approach cleaner still.

A Python handler with connection reuse and a time budget

import json
import os

import boto3
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

CONNECT_TIMEOUT_S = 5.0
MAX_READ_TIMEOUT_S = 20.0
SAFETY_MARGIN_S = 2.0

_secrets = boto3.client("secretsmanager")
_secret = _secrets.get_secret_value(SecretId=os.environ["PROXY_SECRET_ID"])
_proxy_url = json.loads(_secret["SecretString"])["proxy_url"]

_retry = Retry(
    total=2,
    connect=2,
    read=1,
    status=2,
    backoff_factor=0.5,
    status_forcelist=(429, 502, 503, 504),
    allowed_methods=frozenset({"GET", "HEAD"}),
    respect_retry_after_header=True,
)
_adapter = HTTPAdapter(max_retries=_retry, pool_connections=4, pool_maxsize=4)

session = requests.Session()
session.trust_env = False
session.proxies = {"http": _proxy_url, "https": _proxy_url}
session.mount("http://", _adapter)
session.mount("https://", _adapter)


def handler(event, context):
    url = event["url"]
    remaining_s = context.get_remaining_time_in_millis() / 1000 - SAFETY_MARGIN_S
    read_timeout_s = min(MAX_READ_TIMEOUT_S, remaining_s - CONNECT_TIMEOUT_S)
    if read_timeout_s < 1:
        raise TimeoutError(f"only {remaining_s:.1f}s left, not starting a proxied request")

    response = session.get(url, timeout=(CONNECT_TIMEOUT_S, read_timeout_s))
    return {
        "url": url,
        "status": response.status_code,
        "bytes": len(response.content),
    }

Everything above handler runs once per execution environment: the secret read, the retry policy and the Session with its connection pool. Warm invocations skip all of it and reuse open connections, which removes the TCP, CONNECT and TLS handshakes from every request after the first. Package requests with the function or in a layer; the Python runtimes include boto3 but not Requests.

Connection reuse has a side effect worth knowing. For HTTPS targets, a pooled connection is an open tunnel through the proxy, and most gateways choose the exit address when the tunnel opens. A warm environment making repeated requests to the same host will usually keep the same exit IP, even on a rotating endpoint. That is what you want for a sticky workflow and not what you want for per-request rotation. Rotating vs sticky proxies covers how to choose, and the Requests guide shows how to force fresh connections when rotation matters.

A frozen environment can also hold connections the gateway has since closed. urllib3 discards most of these when it takes them from the pool, and the connect retries handle the rest.

Fitting AWS Lambda proxy timeouts inside the function timeout

A new Lambda function has a default timeout of three seconds. A proxied request to a slow target routinely takes longer than that on its own, so the first change to any scraping function is its timeout, which can go up to 15 minutes.

Work out the worst case explicitly. With the settings above, a single attempt can take up to 5 seconds to connect and 20 to read. Three attempts plus backoff is roughly 80 seconds in the worst case, so a 90-second function timeout leaves room for the handler to return a result instead of being killed. The handler also reads get_remaining_time_in_millis() and shrinks the read timeout when the invocation is already late, so a request never starts that cannot finish.

The distinction matters because the two failures look different. A request timeout raises inside your code, where you can log the URL, return a structured failure and let the caller decide. A Lambda timeout ends the invocation with no return value and, for asynchronous invocations and queue triggers, leads to redelivery, which means the same page is fetched again through the proxy and billed again.

Concurrency: Lambda's versus the proxy's and the target's

Lambda scales by running more execution environments, one invocation at a time each. A burst of 500 queued URLs can mean 500 simultaneous connections to the proxy gateway and, more importantly, 500 simultaneous requests to one target site. Three limits should be set deliberately:

  1. Function concurrency. Reserved concurrency caps how many environments the function can run at once. For SQS triggers, the event source mapping's maximum concurrency setting caps the pollers without reserving capacity.
  2. Per-environment connections. With one request per invocation, pool_maxsize of a few is enough. If a handler fans out with threads, the pool size bounds its parallelism.
  3. Target tolerance. A target that starts returning 429 or 403 under load is telling you the concurrency is too high for it, whatever the proxy can carry. Proxy 403 and 429 errors explains how to tell target blocks from proxy failures.

Ask your provider whether the account has a limit on concurrent connections or requests, and set function concurrency below it, so that throttling happens in Lambda, where it is visible, rather than as connection errors from the gateway.

AWS Lambda proxy checklist before you deploy

  • The proxy credential is in Secrets Manager, and the execution role can read only that secret.
  • HTTP_PROXY and HTTPS_PROXY are unset on the function, or NO_PROXY includes localhost,127.0.0.1,.amazonaws.com.
  • The Session and the secret read are at module level, not inside the handler.
  • The function timeout exceeds the worst-case proxied request, retries included.
  • Reserved concurrency or the event source's maximum concurrency is set.
  • If you use IP allowlisting, every NAT gateway Elastic IP is registered with the provider.
  • A test invocation logs the exit IP seen through the proxy, so you know traffic is not going direct.

Using ProxyForge from Lambda

Every ProxyForge line supports both username and password and IP allowlist authentication, so you can run functions outside a VPC with a credential from Secrets Manager and move to allowlisting your NAT gateway addresses later without changing providers. The dashboard generates the connection string for the country and session mode you choose.

For bursty serverless workloads fetching public pages, per-GB residential proxies usually fit best; for high-volume API collection, per-address datacenter proxies give a predictable cost. Billing is pay-as-you-go from a prepaid wallet with no monthly minimum, and current rates are on the pricing page.

FAQ

Related questions

Does a Lambda function have a fixed public IP address?

No. A function outside a VPC sends traffic from addresses in AWS's shared ranges that change over time and between execution environments. A fixed address requires attaching the function to private subnets whose default route goes through a NAT gateway with an Elastic IP.

Is the requests library available in the Lambda Python runtime?

No. The Python runtimes include the AWS SDK but not requests, so package it with your function or in a layer, and pin the version rather than relying on whatever a runtime update ships.

Why do my boto3 calls fail after I set HTTPS_PROXY on the function?

The AWS SDK honors the standard proxy environment variables, so setting them on the function also routes S3, SQS and Secrets Manager calls through the scraping proxy. Add .amazonaws.com and localhost to NO_PROXY, or pass the proxy explicitly in your HTTP client and leave the environment variables unset.

Can Lambda keep a sticky proxy session between invocations?

Only within one execution environment, and only while it stays warm. Consecutive invocations can land on different environments, so if a workflow needs one exit address for several steps, run those steps inside a single invocation or carry the session identifier in the event.

Run it on a network you can account for

Order from 1 GB or 1 IP with no monthly minimum, or talk to an engineer about your workload first.

One business day, from a named engineer.