GDPR-compliant web scraping: what in-house teams must verify

19 September 2026
14 minutes read
Summary generated by AI:

GDPR web scraping now draws some of the heaviest data protection enforcement across the EU. The Dutch AP treats commercial scraping of personal data as almost always unlawful, and the CNIL’s 19 June 2025 focus sheet sets out what you have to document before collection. The EDPB followed in July 2026 with its first scraping-specific guidance for AI training.

Supply-chain liability is also widening: the EDPB’s Opinion 28/2024 stated that an AI model built on unlawfully collected data may taint whoever deploys it, unless it’s genuinely anonymised. The same reasoning puts bought datasets or rented IP pools on your risk register.

Most in-house pipelines, though, were built for coverage and uptime, leaving compliance questions unowned: which lawful basis applies, how the vendor sourced its IPs, and how long data is kept. The scope of personal information you touch also depends on the collection method, the distinction our web scraping vs web crawling guide lays out. A defensible pipeline answers all those questions before deployment, with a named owner and dated documentation across legal, engineering, procurement, and security.

Why GDPR compliance matters for enterprise web scraping projects

Compliance failures in large-scale scraping rarely surface as warnings; enforcement usually shows up as a fine when the data already feeds production. Two cases set the current baseline:

  • Clearview AI drew a €30.5 million fine from the Dutch AP in September 2024 for amassing billions of facial images without a lawful basis, the largest scraping-specific penalty by an EU authority to date. The French, Italian, and Greek regulators each imposed another €20 million fine, and none of the four has been paid.
  • Three months later, in December 2024, the CNIL fined KASPR €240,000 for collecting LinkedIn contact details that users had restricted, reinforcing that limited-visibility data remains off-limits.

A public profile is still personal data. Visibility doesn’t equal consent, and “it was already online” isn’t a lawful basis for processing.

Regulators can also order deletion of an unlawfully collected dataset, which can cost more than the fine, while a scraping case invites scrutiny of everything the organization collects next.

When does GDPR apply to my web scraping project?

GDPR web scraping obligations begin when two things are true: you’re processing personal data, and it relates to individuals in the EU or EEA. GDPR Article 3 ties territorial scope to those people, not your server location. A US team running GDPR web scraping against EU marketplace sellers is still in scope, and moving the crawler offshore changes nothing.

Product prices, stock levels, aggregate listings, and other non-personal fields fall outside the regulation entirely.

What counts as personal data in a scraping context?

Personal data is any information that identifies a person, directly or indirectly. In web scraping, this covers names, emails, phone numbers, photographs, and profile URLs that single out an individual without additional context, plus online identifiers like IP addresses, cookie IDs, and device fingerprints.

Context also matters: a company name, job title, city, and department can identify a specific employee even without collecting their name. Proxy and request logs also hold personal data, and hashing or pseudonymizing identifiers doesn’t remove them from GDPR scope if they still point to one person.

Special-category data, including biometrics, health information, race, religious beliefs, or political views, is subject to stricter GDPR requirements and needs a separate Article 9(2) exception.

When your scraping falls outside GDPR scope

GDPR doesn’t apply in two cases: datasets with no personal fields, or data subjects entirely outside the EU and EEA. However, other regimes still reach personal data: CCPA in California, LGPD in Brazil, PIPEDA in Canada. So outside GDPR rarely means outside privacy law.

GDPR web scraping requirements under Article 5 principles

Article 5 puts the seven principles at the core of GDPR requirements for web scraping. Each maps to a specific owner or team and comes with a practical way to verify project compliance.

Article 5 principle

What to verify

Owner

Lawfulness, fairness and transparency

Documented lawful basis and public privacy notice

Legal

Purpose limitation

One named scraping purpose with no scope creep 

Product

Data minimization

Only required fields are collected 

Engineering

Accuracy

Defined refresh cadence and correction process 

Data engineering

Storage limitation

Retention windows enforced automatically 

Data engineering

Integrity and confidentiality

Encryption, access control, and audit logging in place 

Security

Accountability

Written LIA and DPIA on file where required

Legal / DPO

GDPR accountability underpins the other six principles, demanding proof of how compliance decisions were made rather than just stating the right controls exist.

During audit, the regulator asks for evidence behind each row: lawful basis assessments, records of processing activities, retention schedules, and deletion logs. A GDPR web scraping pipeline with no artifact for any of those principles fails on accountability alone.

Use clean-sourced residential IPs, ready for production. Proxy-Seller provides consent-based sourcing and chain-of-custody documentation across a 47M+ residential pool, 220+ locations, and a 99.5% success rate.

Establishing a lawful basis for data collection

GDPR web scraping compliance starts with selecting one of the six lawful bases for processing personal data under Article 6. This choice shapes the entire lifecycle, from what data you can collect and how long to keep it to what documentation you’ll need to prove compliance.

The CNIL frames legitimate interest as the practical default for commercial B2B scraping and the only basis that scales for public-data collection, since consent can’t realistically be gathered from every person. Whether that basis holds comes down to the safeguards. Other bases, like contract, legal obligation, vital interest, and public task almost never cover commercial scraping and shouldn’t be used without legal sign-off.

The EDPB’s Guidelines 03/2026, written in July 2026 for scraping that feeds AI training, confirm that reading. Commercial purpose isn’t the obstacle here: in C-621/22 the Court of Justice ruled it can’t disqualify you on its own, provided the scrape is necessary and passes the balancing test.

Legitimate interest and the three-part balancing test

Regulators apply a three-part test during an audit, and each answer needs a written record before GDPR web scraping begins:

  1. Purpose: state a specific, lawful business interest with no vague labels.
  2. Necessity: confirm the scrape is genuinely required for that purpose.
  3. Balancing: check whether people can reasonably expect this use of their data.

The third question is where most scraping fails. A person who posts a public profile doesn’t expect a vendor to harvest, enrich, and resell their contact details. If you conclude the use is still justified, capture your reasoning in a legitimate interest assessment (LIA).

The CNIL’s June 2025 focus sheet (English version published on 5 Jan 2026) ties legitimate interest to specific safeguards your LIA should reflect, including precise collection criteria, exclusion of opt-out sites, and prompt deletion of irrelevant data. 

Proxy and IP infrastructure compliance for GDPR web scraping

The proxy layer is one of the most overlooked compliance surfaces in GDPR web scraping. An IP address can constitute personal data when it relates to an identifiable individual, so how a vendor builds its pools joins your compliance chain alongside your technical stack. Teams that buy proxy SOCKS5 for compliant data collection at scale should treat vendor provenance as a procurement gate.

GDPR liability when choosing a proxy provider

Before asking procurement questions, it’s worth understanding how GDPR classifies the relationship between your organization and the proxy vendor.

It splits every data processing operation into two roles:

  • Data controller decides why and how personal data gets processed.
  • Data processor acts on the controller’s instructions.

When you scrape, you’re the controller, and your proxy vendor typically operates as the processor.  

A common claim is that using a vendor with unlawful sourcing makes you a joint controller with it. The EDPB’s Guidelines 07/2020 reject that as automatic: sharing a provider’s infrastructure alone isn’t enough. Article 26 defines joint controllership explicitly: both parties must determine the purpose and means of the same processing. Because a proxy vendor never decides what you scrape or why, this role doesn’t apply.

Two other rules create the real exposure: Article 28 requires controllers to use only processors that provide sufficient guarantees of GDPR compliance, and Article 5(2) keeps you accountable for the personal data you process. So vendors whose IPs lack consent become your own due-diligence failure, and you answer for every request routed through those IPs, however clean your scraper is.

How to verify a GDPR-compliant proxy provider

Before the first request, require written evidence for each of the following from every prospective vendor:

  • consent-based sourcing for every residential and mobile IP, with opt-in and opt-out records
  • a signed DPA and a current sub-processor list
  • security certifications such as SOC 2 and ISO 27001
  • transparency reporting and a documented chain of custody

A missing or hedged answer disqualifies the vendor. Keep every response on file, as this is the due-diligence documentation a regulator asks for first. As a proxy infrastructure provider, Proxy-Seller supplies this documentation on request, including a DPA, SCCs, and ISO/IEC 27001:2022 certification.

Of those, the DPA carries the most weight because it turns the guarantees into an enforceable contract. Under Article 28(3), it should:

  • name the processing scope and duration,
  • commit the vendor to security measures such as uptime guarantees and encryption,
  • list sub-processors with a change-notification duty,
  • grant you audit rights,
  • specify data deletion or return when the relationship ends.

Residential and mobile proxy sourcing considerations

Residential and mobile proxies route through real user devices, putting user consent at the center of sourcing. Unlike provider-owned datacenter ranges, the compliance profile of residential and mobile pools depends entirely on how the vendor obtained device access.

The sourcing model is what separates a secure proxy network from a risky one. Explicit opt-in (usually through an SDK that trades bandwidth for an app feature), backed by working opt-out and supporting records, puts the vendor at the safe end. Consent buried inside bundled app terms is much harder to justify because users never knowingly agree to it. No consent at all makes the pool unlawful, and the exposure lands on you under Article 28.

Two takedowns in 2026 showed how that plays out. In January, Google disrupted IPIDEA, a backend network that secretly powered 13 proxy and VPN brands and enrolled millions of devices through a trojanized SDK in 600+ Android apps. In July, Google, the FBI, and Lumen dismantled NetNut, a mainstream provider with an enterprise client base and a publicly traded parent, on the same grounds. Google also warned that many well-known brands resell that infrastructure under their own name, so the label on your invoice proves nothing about the exit nodes; only the audit trail does.

Price tier is beside the point: for a cheap mobile proxy or a premium one, IP pool quality comes down to the same evidence, documented consent and per-IP provenance.

Cross-border data transfers under Article 44

Any scraped personal data that leaves the EEA triggers Article 44. That encompasses every part of the scraping infrastructure: proxy gateways, cloud environments, logging, analytics, backups, and support systems.

Before data exits the EEA, verify the transfer relies on a Chapter V mechanism: Standard Contractual Clauses (SCCs) or an adequacy decision. The EU–US Data Privacy Framework covers certified US vendors. As of August 2026, the Framework remains in force, but its adequacy decision is under appeal before the Court of Justice. The US Supreme Court’s June 2026 ruling in Trump v. Slaughter on FTC independence weakens another of its pillars, so re-verify mechanisms quarterly.

Run a Transfer Impact Assessment (TIA) if personal data reaches jurisdictions with broad government-access powers. Confirm the exit country of every proxy route, since a residential IP abroad is itself a cross-border transfer. If an adequacy decision later fails, SCCs supported by a TIA can still hold. 

Pin every request to a known exit country. Proxy-Seller offers mobile proxies from $10 per week with country-level routing and real carrier IPs.

Processing and storage obligations for scraped data

Enforcement in GDPR web scraping usually begins after data has already been collected. Regulators examine how organizations process, retain, and govern scraped datasets rather than focusing solely on extraction. In the KASPR case mentioned earlier, the scrape itself proved survivable, but retention practices and delayed notification ultimately triggered the CNIL fine.

Post-extraction governance revolves around three separate GDPR obligations: data minimization, storage limitation, and transparency to data subjects, each typically owned by a different team.

Data minimization at the selector level

Owner: Engineering.

Article 5.1.c (data minimization) limits collection to the fields a project actually needs, since processing begins the moment data is captured. Collecting everything for later filtering already breaches the principle, making parser configuration one of the first GDPR web scraping controls.

Field-level allowlists put that principle into practice. For example, a price intelligence workflow typically needs product identifiers, prices, availability, currency, and timestamps. Seller phone numbers, personal email addresses, customer names, and profile URLs rarely serve that purpose but immediately expand the compliance scope.

Retention windows and automated deletion

Owner: Data engineering / DBA.

Article 5.1.e (storage limitation) requires personal data to be kept only for as long as the processing purpose justifies. This requirement applies to production databases, backups, exports, temporary storage, and replicated environments.

Assign short TTLs (hours to a few days) to raw personal data, give longer retention to aggregated output, and enforce both through TTL rules and automated deletion jobs. A retention policy documented as intent with no job behind it fails the accountability principle.

Deletion has to be comprehensive: cascade it across backups, search indexes, exports, and proxy access logs. Missing even one replica or log copy is often what an audit finds first.

Article 14 notification and the disproportionate-effort trap

Owner: Compliance / Legal.

Article 14 (transparency obligations) requires organizations to notify people within one month when their personal data comes from a source other than them, unless a GDPR exemption applies.

One of the most cited exemptions is disproportionate effort, but European regulators interpret it narrowly. In the Polish Bisnode case, the company pulled data from public business registers and informed only the 90,000 people whose email addresses it held. It called postal notice to six million more a disproportionate effort. UODO rejected that and fined it around €220,000. On appeal, the courts upheld the duty to notify but reopened the amount, still unsettled seven years later.

A Data Protection Impact Assessment (DPIA) is sometimes cited alongside Article 14, but it doesn’t replace the notification requirement as a general rule. It may support an exemption where affected individuals can’t be contacted, particularly when the scraped dataset contains no contact details.

Technical safeguards and governance for in-house teams

Policies alone don’t demonstrate GDPR compliance and need technical controls behind them. Encryption, access control, and a documented audit trail make compliant web scraping provable. Teams implementing Python web scraping tactics or similar collection frameworks should incorporate these controls into the pipeline before deployment.

Encryption, RBAC, and MFA at the storage layer

Any dataset holding personal data needs baseline protection: encryption at rest and in transit, role-based access control, and multi-factor authentication on every account. A regulator treats their absence as negligence after a breach.

Higher-sensitivity data demands more: keep encryption keys outside the systems they protect, pseudonymize direct identifiers at rest so leaked tables reveal nothing useful, and log every read. Absence of access logs reads as no access control at all.

DPIA and legitimate interest assessment as audit artifacts

During a GDPR web scraping audit, regulators ask for evidence that compliance decisions were written down before processing began. For a scraping pipeline that usually means:

  • a legitimate interest assessment (LIA)
  • a data protection impact assessment (DPIA) for higher-risk processing
  • a record of processing activities (RoPA) under Article 30

Each should state why the data is processed, what safeguards apply, and who signed off. An assessment you only talked through counts for nothing.

DSAR workflow for scraped data

Individuals retain their GDPR rights even when their data came from scraping rather than a direct customer relationship. Organizations should therefore be able to respond to Data Subject Access Requests (DSARs) covering access, rectification, deletion, restriction, and other applicable rights.

A workable DSAR workflow finds every record tied to one person across raw and processed stores and responds within one month. Article 12 allows three months for genuinely complex requests, if the person hears why within the first month. A canonical subject key and a data map of every personal-data store prevent such requests from becoming weeks of manual search. Our web scraping guide covers the extraction layer that feeds those stores.

Technical signals to respect: robots.txt, CAPTCHA, and rate limits

Ignoring a site’s technical signals weakens the legitimate-interest case under GDPR web scraping review. The CNIL’s 2025 focus sheet is explicit: scraping sites that opt out through robots.txt or CAPTCHA falls outside data subjects’ reasonable expectations.

A defensible setup covers the following legal web scraping practices:

  • robots.txt and ai.txt fetched once per domain and cached
  • crawl-delay and disallowed paths respected
  • repeated CAPTCHAs or anti-bot challenges treated as a stop sign, not as an obstacle to bypass
  • backoff on 429 and 403 before any retry

None of this binds legally on its own. What weighs in an audit is the log: proof the crawler fetched, parsed, and honored each site’s signals demonstrates the measures the CNIL expects. Logged proxy performance metrics, like 429 and 403 rates, back that up.

GDPR compliance checklist for in-house web scraping teams

Before deploying a GDPR web scraping project to production, verify that each area has completed the controls relevant to its responsibilities.

Owner

Verify before launch

Legal

  • personal-data scope confirmed
  • lawful basis documented in an LIA
  • Article 14 notice plan or documented exemption prepared
  • DPIA completed for higher-risk processing
  • records of processing activities (RoPA) maintained per pipeline

Engineering

  • selectors scoped to required fields
  • robots.txt and ai.txt fetched, cached, and respected per domain
  • backoff logic implemented for 429 and 403 responses
  • CAPTCHA and anti-bot challenges treated as a stop signal

Data engineering

  • retention enforced by deletion jobs covering backups, exports, and logs
  • DSAR workflow built around a subject key and data map

Procurement

  • vendor DPA and sub-processor list secured
  • consent-based lawful IP sourcing confirmed
  • SOC 2 or ISO 27001 evidence on file
  • cross-border transfers covered by SCCs or adequacy, TIA where needed, proxy exit country verified

Security

  • encryption at rest and in transit
  • encryption keys stored in a dedicated secrets manager
  • RBAC and MFA enabled for every store holding personal data
  • audit logging enabled on admin access and sensitive reads

How to audit existing web scraping pipelines

Pipelines that predate the checklist above should be assessed against current legal web scraping practices to ensure GDPR compliance: 

  1. Inventory every scraped dataset and its fields.
  2. Identify records relating to EU or EEA data subjects.
  3. Verify a documented lawful basis for each dataset.
  4. Check documented retention policies against actual deletion jobs.
  5. Confirm the DSAR and Article 14 workflows are operational.
  6. Remediate or anonymize non-compliant data.

Compliance starts with the right proxy network. Enterprise-grade datacenter and ISP proxies with endpoint-level logs fit audit-ready scraping. Wholesale pricing from 1,000 units at 25–30% off.

Frequently asked questions

Is GDPR web scraping legal for public business data?

Yes, within limits. GDPR web scraping of public business data is legal provided the applicable data protection requirements are met. Public availability doesn’t remove GDPR obligations where personal data is involved, including sole-trader names and direct contact details. A documented lawful basis, Article 14 notification where required, and exclusion of special-category data help ensure compliant data collection.

Are residential proxies GDPR-compliant?

Residential proxies aren’t inherently compliant or non-compliant. Proxy usage for web scraping is lawful in itself; what decides residential proxy compliance is how the IPs are sourced. Because they rely on real user devices, their legitimacy hinges on user consent. A compliant provider can demonstrate lawful collection, a DPA, and transparent data processing.

Can legitimate interest be used for B2B lead generation?

Yes, after a documented LIA that clears the three-part test on every field. Work emails draw less scrutiny than personal phones or private addresses. Article 14 notification still applies, and objections must be added to the suppression list before the next collection run.

How does Article 14 notice work for a large scraped dataset?

Article 14 gives you one month to inform people whose data you collected from another source. The disproportionate-effort exemption exists, but regulators read it narrowly and expect a documented, per-project assessment rather than a blanket claim. Where individual notice genuinely isn’t feasible, publish a privacy notice describing the scraping and record why.

Is robots.txt legally binding during GDPR web scraping?

Not by itself. robots.txt isn’t a GDPR rule or a standalone legal ban on scraping. But it can affect whether people could reasonably expect their publicly available data to be collected. In its June 2025 guidance, the CNIL says controllers should exclude sites that clearly oppose scraping through robots.txt or CAPTCHAs when relying on legitimate interest.

Is scraping data to train AI models allowed under GDPR?

Potentially, yes. The EDPB’s Guidelines 03/2026, issued in July 2026 and open for comment until 30 October, cover exactly this. Consent doesn’t scale, so rely on legitimate interest, publish a clear privacy notice, and let people opt out before collection. The CNIL’s June 2025 sheet adds two rules: stick to freely accessible data and filter out sensitive fields. Special-category data requires an additional Article 9(2) exception, usually where the person manifestly made it public.

Content of the article: