Skip to content
ProxyForge

GDPR web scraping: where the proxy provider fits

ProxyForge engineeringUpdated 8 min read

GDPR web scraping obligations sit with the organization doing the collecting, not with the network it collects through. If the pages you fetch contain personal data about people in the EU or UK, you are the controller of what you collect: you need a lawful basis, usually legitimate interests, you must collect no more than the purpose requires, and at scale you will likely need a data protection impact assessment. A proxy provider is a processor for the traffic metadata it handles on your behalf, and using one does not make the underlying collection lawful or unlawful.

This article walks through the framework as it applies to proxy-backed collection, sets out who holds which role, and lists what to ask a provider. It describes the framework, not your specific obligations; for those, talk to your data protection officer or counsel.

Does the GDPR apply to web scraping at all?

It applies whenever the data you collect is personal data and your processing falls within the regulation's territorial scope. Both conditions are broader than many teams assume.

Personal data is any information relating to an identified or identifiable natural person (GDPR Article 4(1)). Names and email addresses are obvious. Less obvious: usernames, profile photos, review text that identifies its author, a sole trader's business listing, a seller's name on a marketplace, and any field that becomes identifying once combined with others. Collection is itself processing under Article 4(2), so the obligation starts at fetch time, not when you first query the data.

Publicly available personal data is still personal data. Nothing in the regulation exempts information because it was visible without a login. Its public character is relevant to the balancing test below, but it does not remove the need for a lawful basis.

Territorial scope reaches organizations established in the EU, and also organizations outside it whose processing relates to monitoring the behavior of people in the EU (Article 3(2)(b)). Systematic, repeated collection of profiles or activity can fall within that. The UK GDPR mirrors these provisions for people in the UK.

Many scraping workloads involve no personal data: product prices, stock levels, search rankings, published tariffs. For those, the GDPR analysis may end here. Mixed pages are the common case, which is why minimization, covered below, carries so much of the practical weight.

Lawful basis: legitimate interests and the balancing test

Of the six lawful bases in Article 6(1), consent is impractical when you have no relationship with the people concerned, and contract does not apply. Most collectors rely on legitimate interests under Article 6(1)(f). Both the EDPB's Guidelines 1/2024 on legitimate interests and the ICO's legitimate interests guidance describe the same three cumulative tests:

  1. Purpose. Is the interest you pursue lawful, clearly articulated and real rather than speculative? The Court of Justice confirmed in 2024 (Case C-621/22) that a commercial interest can qualify, provided it is lawful.
  2. Necessity. Is collecting this personal data necessary for that purpose, or could you achieve it with less, or with none?
  3. Balancing. Do the interests, rights and freedoms of the people concerned override yours? Relevant factors include what they would reasonably expect when they published the information, the nature of the data, the scale of collection, whether data is combined with other sources, and the safeguards you apply.

Record the outcome in a legitimate interests assessment before collection starts, and revisit it when the purpose or scope changes. Two further points tend to decide the balance in scraping cases.

Special category data. Health, political opinion, religion, sexual orientation and the other categories in Article 9 need a separate condition, and scraped text can contain them without anyone intending it. Filtering them out at the edge is usually simpler than justifying their retention.

Transparency and objection. Where data is not obtained from the person, Article 14 requires you to tell them about the processing. Article 14(5)(b) relaxes this where informing each person would involve disproportionate effort, but still requires appropriate measures, including making the information publicly available. People also keep the right to object under Article 21, so you need a way to receive and honor objections.

Data minimization in a scraping pipeline

Article 5(1)(c) requires personal data to be adequate, relevant and limited to what is necessary for the purpose. In a collection pipeline that principle translates into engineering decisions:

  • Define the fields you need before the crawler is written, and parse only those.
  • Discard personal fields at the parsing stage rather than storing raw pages "in case".
  • Where you need to count or deduplicate people rather than identify them, pseudonymize at ingestion.
  • Set retention periods per dataset, and delete on schedule (Article 5(1)(e), storage limitation).
  • Keep raw HTML only as long as debugging requires, and treat it as the most sensitive copy you hold.

A pipeline that stores full pages indefinitely will struggle to pass the necessity test, however legitimate its purpose.

When a DPIA is required

Article 35 requires a data protection impact assessment before processing that is likely to result in a high risk to individuals. Large-scale collection, systematic monitoring, combining datasets from different sources, and processing people are unaware of are all indicators, and scraping projects often meet several at once. The ICO's list of processing likely to be high risk includes "invisible processing", where you rely on the Article 14 disproportionate effort exemption, which describes much web collection directly.

A GDPR web scraping DPIA should describe the sources, the fields collected, the purpose and legitimate interests assessment, the minimization and retention controls, the processors involved, including your proxy provider, and the residual risks. The ICO's DPIA guidance includes a template.

Who is controller, processor and third party?

Roles follow from who decides the purposes and means of processing, not from what a contract calls each party. In a typical proxy-backed collection setup they fall out as follows.

Party Typical role What it processes What governs it
You, the collector Controller The personal data you scrape, and your own logs Your lawful basis, LIA, DPIA, records of processing
Proxy provider Processor for your traffic; controller for its own account, billing and KYC data Traffic metadata: timestamps, destination hosts, bytes, your source address or credentials An Article 28 data processing agreement
Residential or mobile peer Data subject in the provider's supply chain Their own connection, which carries your requests The peer's consent to the provider or its SDK partner
Target website operator Independent controller of its own data The requests you send it, including the exit address Its own obligations; its terms are a separate, non-GDPR question

The peer relationship deserves attention. When you use residential or mobile proxies, your requests leave the internet from a real person's connection, and that person's address appears in the target's logs next to your request. You have no contract with them. What protects them, and you, is the provider's evidence that they knowingly agreed to carry third-party traffic, can withdraw, and were compensated. That is a sourcing question, covered in what ethically sourced residential proxies should mean, but it belongs in your DPIA as well.

On confidentiality: for HTTPS requests sent through an HTTP proxy, the client opens a tunnel with the CONNECT method, so the provider sees the destination host and port but not the page content. Plain HTTP requests are visible to anyone on the path, which is one more reason to prefer HTTPS targets where you have the choice.

What to ask your proxy provider

Your provider is a processor, so Article 28 requires a written contract with specific terms, and your records of processing must list it. These questions cover what your DPIA and vendor file will need:

  • Will you sign a data processing agreement before we commit, and does it incorporate the EU standard contractual clauses and the UK addendum for transfers?
  • Which categories of traffic metadata do you process on our behalf, for what purposes, and for how long?
  • Where is that metadata stored, and under which transfer mechanism if outside the EEA or UK?
  • Is your sub-processor list published, and how much notice do you give before it changes?
  • How quickly will you notify us of a personal data breach affecting our traffic?
  • How will you assist with data subject requests and with our DPIA, as Article 28(3) requires?
  • For residential and mobile pools, can you produce the consent record for an exit address we choose?

The article on what a proxy provider DPA should cover goes clause by clause, and the proxy vendor security questionnaire covers the security side.

What a proxy does not change

Because proxies change where requests appear to come from, it is easy to assume they change the GDPR web scraping analysis. They do not.

  • A proxy does not make collection lawful. Lawful basis, necessity and minimization are properties of what you collect and why, not of the route it takes.
  • A proxy does not make you anonymous to the law. You remain the controller and remain accountable for the data you hold.
  • A proxy does not remove transparency or objection duties. Articles 14 and 21 apply however the data was fetched.
  • A proxy adds a processor. It creates Article 28 obligations and belongs in your records and your DPIA.
  • A proxy does not settle non-GDPR questions. Site terms, database rights, copyright and computer misuse law are separate analyses.

The last point cuts both ways. Collection that is lawful does not become unlawful because it uses a proxy, and a well-run provider with clear records is a smaller risk in your DPIA than an undocumented one.

How ProxyForge fits into your GDPR file

ProxyForge provides a data processing addendum with EU standard contractual clauses and the UK addendum before you sign, publishes a versioned sub-processor list with 30 days' notice of change, and retains traffic metadata for 30 days. Residential peers join through a disclosed, compensated and reversible opt-in via partner apps, and we can trace any address in your pools to its consent record, so the peer relationship in your DPIA rests on evidence rather than assurance.

The DPA and the sourcing page contain the documents most DPIAs ask for. If your use case is large-scale collection, the web scraping use case covers the product side.

FAQ

Related questions

Is scraping publicly available data legal under the GDPR?

Public availability does not take data outside the GDPR. If the pages contain information about identifiable people, collecting it is processing, and you need a lawful basis, a defined purpose and proportionate safeguards like any other processing.

Do I need consent to scrape personal data?

Rarely in practice, because you usually have no way to ask. Most collectors rely on legitimate interests instead, which requires a documented assessment showing the purpose is legitimate, the collection is necessary, and individuals' interests do not override it.

Does the UK GDPR treat web scraping differently from the EU GDPR?

The core principles, lawful bases and controller and processor roles are the same, and the ICO's guidance follows the same three-part legitimate interests test. Differences arise mainly in transfer mechanisms, in amendments the UK has made since leaving the EU, and in each regulator's priorities, so check the current UK text.

Is my proxy provider a joint controller of the data I scrape?

Normally not. A provider that only carries your requests on your instructions does not decide why or how the data is collected, which is what makes a controller. It is a processor for the traffic metadata it handles for you and a controller of its own account and billing data.

Start with the evidence

Ask us to trace an address, send you the sourcing attestation, or price your current volume at our published rates. A named engineer will help with your technical and procurement review.

One business day, from a named engineer.