Skip to content
ProxyForge

Scraper API vs proxies: what you hand over and what you keep

ProxyForge engineeringUpdated 8 min read

Scraper API vs proxies is a choice about which unit you buy and which work you keep. A scraper API takes a URL and returns the page, or structured data parsed from it, and does the rendering, retries, rotation and unblocking on its side, billed per request. Proxies give you exit addresses, billed per gigabyte or per address, and leave everything above the network to your code. Buy the API when the hard part is unblocking and nobody on the team should own it. Build on proxies when you need control over requests and sessions, a cost you can reason about per byte, or page content that never leaves your infrastructure. Many mature teams end up with both.

This guide sets out what each side takes over and gives up, a decision table by team profile, and a method for running a bounded comparison on your own targets, so the decision rests on your data rather than a vendor's.

What a scraper API takes over

A scraper API bundles several jobs that a proxy-based pipeline does in your code:

  • Rendering. Pages that build their content in JavaScript need a headless browser. The API runs the browser fleet, its memory and its crashes.
  • Retries and rotation. A failed or challenged request is retried through other exits, with backoff, before anything comes back to you.
  • Unblocking. Fingerprint consistency, header ordering, cookies and the handling of challenge pages are the vendor's problem and its specialty.
  • Parsing. For common page types, some APIs return structured fields instead of HTML.
  • Targeting by parameter. Country, and sometimes device or language, is a field on the call rather than a property of the exit you chose.

Each of these is real engineering. Rendering alone means operating browsers at scale, and unblocking is an arms race with sites that invest in bot defense. For a team without someone to own that, the API turns weeks of tuning into a parameter.

What you give up

The same bundling removes things a proxy pipeline gives you by default:

  • Control of the request. Headers, cookies, request order, POST bodies and session identity are whatever the API exposes. Logged-in flows and multi-step sessions are where the gaps show first.
  • Cost visibility. You see a price per request, often with multipliers for rendering or for targets the vendor classes as difficult. Whether retries, timeouts and failed requests are billed varies, and the attempts behind each billed request are invisible. The cost of a job becomes a property of the targets rather than of your traffic, and it can change when a target changes its defenses.
  • Debuggability. When a request fails, you get the API's error, not the exit address, the number of attempts or the response the target actually sent. Diagnosing a drop in data quality means opening a ticket.
  • Data residency. The vendor fetches the page and hands it to you, so it handles the full content, may log or store it, and processes it wherever its infrastructure runs. If pages contain personal data, the vendor is a processor, and it needs a DPA and a place on your sub-processor list. GDPR and web scraping with proxies covers the obligations, and what a proxy DPA should cover applies to an API vendor equally.
  • Method. The vendor's unblocking methods become your methods. If it solves challenges or rotates through residential exits, that is how your organization is collecting, whether or not your policy allows it.
  • Portability. Request parameters, parsed-output schemas and vendor-specific options end up in job code. Leaving means rewriting the jobs and rebuilding the parsers.

Proxies keep these with you. With HTTPS targets, a proxy provider relays an encrypted tunnel and sees the destination hostname, not the page, so content stays between your code and the target. The exit, the request, the retries and the billing unit are all yours to observe, at the price of owning everything the API would have done.

What building on proxies costs

The proxy route is cheaper per unit only if you count all of it. You are signing up to run:

  • a headless-browser tier for targets that need one, sized for peak and patched like any other fleet;
  • retry, backoff and session policy per target, and the monitoring that tells you when it stops working;
  • parsers, and their maintenance every time a target changes its markup;
  • block handling: telling a target's refusal from a network failure, which proxy 403 and 429 errors covers;
  • an egress design for credentials, policy and logs, such as the central Squid egress;
  • someone on call for all of it.

Teams that already run a crawler fleet have paid most of this. Teams that have not should price it honestly, in engineer time, before comparing it with an API invoice.

Comparing cost: per valid record, not per request

The two models bill different units, per request on one side and per gigabyte or per address on the other, so neither price list can be compared with the other directly. Normalize both to the unit you actually want, a valid record:

  • For the API, take the billed usage for the test window from the vendor's own reporting, including any billed retries or failures, not your count of calls.
  • For proxies, take the metered gigabytes or the address rental for the same window from the provider's reporting. Client-side byte counts miss TLS and protocol overhead.
  • Divide each by the records that passed validation, then add the engineering hours each arm consumed during the trial and is likely to consume in operation.

Per-request pricing pays for rendering and unblocking whether or not a page needed them, so it tends to look best on hard targets and worst on simple ones. Per-gigabyte pricing punishes heavy pages and rewards lean fetching; reducing proxy bandwidth and per-GB vs per-IP pricing cover the levers on that side.

Which fits your team?

Team profile Lean towards Why
Analysts or a small team with no scraping engineer, a few difficult targets Scraper API Buys the expertise; the volume keeps per-request pricing tolerable
Data-engineering team with an existing crawler fleet and monitoring Proxies The hard parts are built; control and unit cost matter more
Platform team serving many internal consumers Proxies behind a central egress, an API for exceptions One policy, one log, one credential store; the API covers targets nobody wants to own
Workflows with logins, carts or multi-step sessions Proxies with sticky sessions Session control varies across APIs
Pages with personal data, strict residency or a demanding DPA review Proxies, unless the API vendor passes the same review Content stays in your infrastructure
Many heavily defended targets at modest volume Scraper API Unblocking is the most expensive part to build
High volume of simple, static pages Proxies Per-request pricing pays for work these pages do not require
An unproven data need Scraper API for the trial, decide afterwards Fastest route to a first dataset

The split rarely holds for a whole estate. A common shape is proxies for the core targets that carry most of the volume, and an API for the long tail of hard or rarely visited targets that would not repay the tuning.

How to run a bounded comparison

A fair comparison is small, fixed in advance and run on your own URLs. Agree the following before the trial starts:

  1. The question. Which targets, which record, what freshness and what volume the production job needs.
  2. The sample. A URL list per target, stratified so that easy and hard pages, static and rendered pages, and each country you need are all represented.
  3. Validity. A per-target check that the record is usable, such as required fields present and parseable. An HTTP 200 is not success; challenge and empty pages often return it. Benchmarking proxy providers covers detection and gives a harness for the proxy arm.
  4. The window. Both arms run on the same URLs over the same period, interleaved rather than one after the other, with the same concurrency ceiling, so that time-of-day and target changes affect both.
  5. What to record. Per request: arm, URL, outcome (valid, blocked, empty, timeout, error) and latency.
  6. What to collect afterwards. Each vendor's billed usage for the window, and the engineering hours each arm took.
  7. The exit criteria. Thresholds for valid rate, latency and cost per valid record, and the data-handling review, agreed before anyone sees results.

A small scorer turns the per-request records into the comparison. It reads a CSV with the columns arm, url, outcome and latency_ms:

import csv
import sys
from collections import Counter, defaultdict


def percentile(values, pct):
    if not values:
        return float("nan")
    ordered = sorted(values)
    rank = max(1, round(pct / 100 * len(ordered)))
    return ordered[rank - 1]


def score(path):
    arms = defaultdict(list)
    with open(path, newline="") as handle:
        for row in csv.DictReader(handle):
            arms[row["arm"]].append(row)

    for arm, rows in sorted(arms.items()):
        outcomes = Counter(row["outcome"] for row in rows)
        valid = [row for row in rows if row["outcome"] == "valid"]
        latency = [float(row["latency_ms"]) for row in valid]
        print(f"{arm}: {len(valid)}/{len(rows)} valid ({len(valid) / len(rows):.1%})")
        print(f"  latency of valid responses: p50 {percentile(latency, 50):.0f} ms, p95 {percentile(latency, 95):.0f} ms")
        print("  outcomes: " + ", ".join(f"{name} {count}" for name, count in outcomes.most_common()))


if __name__ == "__main__":
    score(sys.argv[1])

Run it per target as well as overall. An arm that wins on aggregate can lose badly on the one target the business cares about, and the outcome breakdown shows whether failures were blocks, which more tuning might fix, or timeouts, which point at capacity.

Questions to ask a scraper API vendor

Send these before the trial, so the answers can be checked against what you observe:

  • What is billed: attempts or successes? Are rendering, retries, timeouts and failed requests billed, and at what multiplier?
  • How does the API decide a request succeeded: status code, or content?
  • Where are requests sent from, which exit types does the pool use, and how were those addresses sourced? A scraper API still runs on proxies, so verifying IP sourcing applies to its pool as much as to any proxy provider's.
  • Is page content logged or stored, for how long and in which jurisdiction? Who are the sub-processors, and will you sign a DPA?
  • What does the API do against bot defenses, and is any of it something your acceptable use policy, or your targets' terms, would prohibit?
  • Which parameters control headers, cookies, sessions, geography and request method?
  • What are the concurrency limits, timeouts and maximum response sizes?
  • If you leave, can you export or reproduce the parsed output formats your jobs depend on?

Where ProxyForge fits

ProxyForge sells proxies today: residential, mobile, ISP and datacenter, from one account, billed pay-as-you-go from a prepaid wallet. That is the build side of this comparison, and the products overview and pricing pages carry the current lines and rates. Scraper, SERP and Unblocker APIs are coming soon; they are not on sale yet, so nothing in this guide describes them. Talk to sales to be told when they open.

There is no monthly minimum, so the proxy arm of a trial costs only the traffic or addresses it uses, and the sourcing page gives the reviewer the evidence for where those addresses come from, the same question you will be asking the API vendor.

FAQ

Related questions

Does a scraper API remove the need to check where IP addresses come from?

No. A scraper API still sends its requests through a pool of exit addresses, so the sourcing question applies to the vendor's pool exactly as it would to a proxy provider. The difference is that the pool is hidden behind the API, so you have to ask for the evidence rather than test it yourself.

Is a scraper API more expensive than proxies?

It depends on the targets. Per-request pricing pays for rendering, retries and unblocking whether or not a target needs them, so it compares well on hard targets and poorly on simple ones. Compare cost per valid record from a trial on your own URLs, including the engineering time the proxy route needs.

Can I use a scraper API for logged-in or multi-step sessions?

Sometimes. Session control, custom headers, cookies and POST bodies vary a great deal between APIs, and some are designed for one-shot fetches only. If a workflow depends on a sequence of requests from one identity, confirm the API supports it on a real flow before committing.

Who can see the scraped content with each approach?

With proxies and HTTPS targets, the provider relays encrypted traffic and sees destination hostnames, not page content. A scraper API fetches the page and returns it to you, so the vendor necessarily handles the full content, which makes it a processor of any personal data the pages contain.

Start with the evidence

Ask us to trace an address, send you the sourcing attestation, or price your current volume at our published rates. A named engineer will help with your technical and procurement review.

One business day, from a named engineer.