Web Scraping

How to bypass IP ban web scraping in 2026 (guide)

Igor Efimenko
Igor Efimenko
29 September 2026 • 14 min read

A scraper that pulled clean data yesterday returns 403s on every endpoint this morning, and switching IPs before you check the cause can leave the problem unresolved. This guide to bypass IP ban web scraping in 2026 replaces guesswork with a diagnosis: read the signal first, find the layer that failed, then fix it in order. Treat an IP ban as a symptom, not the whole problem. Everything here covers public data only, within the target’s rate policies.

Why sites ban IPs during web scraping

According to the 2026 Thales Bad Bot Report, bots made up more than 53% of all web traffic in 2025, up from 51% a year earlier, so sites screen hard. Before you change anything, name the trigger: an IP ban in web scraping comes from one of four common causes.

Rate-limit violations. Your scraping traffic exceeds what the site tolerates, and it answers with HTTP 429 Too Many Requests. Per RFC 6585, 429 means “the user has sent too many requests in a given amount of time.” The response may include a Retry-After header, but the spec leaves the counting method to the server. So the limit isn’t always per-IP: it can be per-account, per-token, or global.

Bot and anomaly detection. The site decides your traffic doesn’t look human because of headless-browser tells, a library fingerprint, or a scraping pattern that’s too regular. This can surface as a 403 Forbidden, and MDN is careful here: a 403 confirms the server understood the request and refused it. MDN adds that “repeating the request without modification will fail with the same error.” The status alone won’t tell you which signal caused the refusal.

IP reputation. Datacenter and hosting ranges can be scored before any scraping starts, because sites can judge reputation for a whole network range, not just one address. An IP’s standing may depend on its subnet or ASN (Autonomous System Number). So a cloud-range address can start from a lower baseline, while a residential IP, assigned by an internet service provider to a real device, can start higher.

Geo-restrictions. The site serves different content, or blocks you, based on your IP’s country. This is the quiet one: the page can return HTTP 200 with the wrong or empty data, and no error code warns you.

Temporary versus permanent bans

Temporary IP bans in web scraping can last from a few minutes to several days, depending on the site’s policy, and the status alone won’t tell you how long. A 429 may include Retry-After, while a challenge may need extra verification. A restriction tied to an IP, session, credentials, or another server-defined identity may not clear while you wait. Inspect the response body, headers, and access policy before you try to bypass the IP ban.

Some blocks don’t return an error at all. A site can answer with HTTP 200 and a block page, truncated data, or decoy content, a pattern scrapers call a shadow ban. Cloudflare’s AI Labyrinth, for example, lures unauthorized crawlers into AI-generated decoy pages instead of blocking them. Validate content, not just status, and if a ban persists on a permitted target, ask the site about an official API, a data feed, or a review of the block.

How to bypass IP ban web scraping in 2026: a layered approach

Bot detection, in-house or from vendors like Cloudflare, Akamai, DataDome, or HUMAN (merged with PerimeterX), can weigh IP reputation, TLS and HTTP/2 traits, and browser or behavior signals together. When it does, a single change isn’t enough, and if the client’s TLS profile triggered the refusal, a new IP alone may not fix it. So a durable way to bypass IP ban web scraping problems is to read the response first. Find the layer it points to and fix that one, instead of reaching for a new proxy each time.

Before you try to bypass an IP ban in web scraping, read each signal against the table:

Signal

What it confirms

What to check next

First action

429

The client is being rate-limited for now

Retry-After; request rate and concurrency; whether the limit is per-IP, per-account, or global

Honor Retry-After or back off; lower the shared request rate before touching proxies

Immediate 403

The server refused this request

Response body and stated policy; headers and TLS profile; IP origin and geo

Inspect the refusal reason; don’t assume the IP is blocked

403 after many requests

Access refused after sustained traffic

Request rate; IP and subnet reputation; session consistency

Reduce rate and review IP or session before rotating

CAPTCHA or challenge

The site wants human verification for this session

Combined signals: IP, cookies, headers, pace

Review the permitted access route; a challenge is not a cue to auto-solve

402 Payment Required

The site charges crawlers for access, for example through Cloudflare’s pay per crawl (launched July 2025*)

The site’s crawler terms; whether your client is classed as an AI crawler

Stop and check the site’s access terms; don’t retry

HTTP 200, wrong or empty data

The request succeeded at the HTTP layer only

Exit-IP geo; whether a challenge page returned as 200; response content and schema

Validate content, not just status

*On July 1, 2025, Cloudflare began blocking AI crawlers by default on new domains, so check whether your client is identified as one.

Log before you change anything. Record:

  • the URL, timestamp, and status
  • relevant response headers
  • exit IP and geo
  • session identifier and client profile
  • the content-validation result

Here’s one example, not a universal rule: your web scraping job returns HTTP 200, but the prices are wrong for the target country. That isn’t a ban, so check the exit IP’s location, the market the site picked, the currency, and your session settings before you touch the proxy pool. Rotate only once those checks point at the exit IP. The next three sections work the layers in order: proxies, fingerprint, and behavior.

Choose the right proxies, sessions, and concurrency

Match the proxy type and session mode to the workflow: blind rotation alone won’t help you bypass an IP ban in web scraping.

Rotating sessions versus sticky sessions

Use rotating proxies for standalone targets like a product page, a SERP query, or a listing, and test per-request rotation against a fixed budget. Use a sticky session for sequential work, such as paging through one catalog or a multi-step flow that depends on consistent cookies.

The HTTP session holds cookies and state on the client, while the proxy session decides which exit IP you leave from. Changing the exit IP mid-flow can break the sequence the site ties to that address, but it doesn’t delete your cookies.

The example below uses Proxy-Seller’s residential proxy format. A _s_ session ID pins the exit IP and _ttl_ sets how long it holds, while a username without suffixes gets a fresh IP on every request. The httpbin.org/ip endpoint echoes the exit IP, so you can watch rotation happen.

Python

import requests

# Rotating vs sticky: print the exit IP that each request leaves from
USERNAME, PASSWORD, HOST = "USERNAME", "PASSWORD", "res.proxy-seller.com:10000"
ROTATING = f"http://{USERNAME}:{PASSWORD}@{HOST}"                # no suffix = fresh IP on every request
STICKY = f"http://{USERNAME}_s_100_ttl_15m:{PASSWORD}@{HOST}"    # same _s_ ID = same IP up to 15 min

for _ in range(2):  # independent requests: each one can get a new IP
    print(requests.get("https://httpbin.org/ip", proxies={"https": ROTATING}, timeout=20).json())

flow = requests.Session()  # one cookie jar for a multi-page public flow
flow.proxies["https"] = STICKY
for _ in range(2):  # sequential requests: expect the same IP twice
    print(flow.get("https://httpbin.org/ip", timeout=20).json())

Residential, datacenter, ISP, and mobile IPs

There’s no single best type. Choose by how well the target is defended and whether the workflow depends on location.

  • Datacenter proxies: lightly defended targets and high-volume, low-sensitivity collection
  • Residential proxies: defended or geo-specific targets
  • ISP proxies: long sessions that need a static IP with an ISP origin
  • Mobile proxies: carrier-sensitive targets, where blocking one shared carrier IP risks blocking real customers

Proxy-Seller’s residential proxies use ethically sourced IPs and support rotating and sticky sessions. They offer targeting across 220+ locations over HTTP, HTTPS, and SOCKS5.

External benchmark context. In Proxyway’s Proxy Market Research 2026, residential proxies reached a 74.43% median success rate on Amazon, Google, and Instagram, compared with 57.46% for shared datacenter proxies on Amazon and Google.

Match the proxy and session to the job:

Workflow

Session model

Proxy choice

Main risk

Standalone product or listing pages

Rotating

Rotating residential proxies

Per-IP rate limits

Pagination or catalog flow

Sticky per task

Residential with sticky sessions

Breaking session continuity

Geo-specific SERP

Rotating with deep geotargeting

Residential proxies with city- or ISP-level targeting

Wrong localized data

Multi-step public-data flow

Sticky, with a TTL covering the whole flow

Residential with sticky sessions, or static residential (ISP) IPs

Session expiring mid-flow

JS-heavy target

Sticky for the whole browser session

Residential, geo-matched to the browser locale

Browser overhead

Concurrency caps and burned IPs

Cap concurrency per domain and start low: use a handful of parallel connections, and raise the number only while valid responses hold. When an address returns repeated 403s or challenges, quarantine it instead of retrying through it: retries won’t bypass an IP ban.

Set a stop condition for the whole domain, too: if the pool’s failure rate climbs past your threshold, pause the run instead of burning more IPs. Spread the pool across many subnets and ASNs, because sites often block a whole range once a few of its addresses misbehave. Our explainer on IP rotation covers the per-request and interval-based models.

Consumer VPNs rarely help bypass IP bans in web scraping: IPinfo found that 83.6% of commercial VPN traffic runs through hosting networks, the ranges sites score first. Free proxies fare worse, as a 30-month study of 640,600 of them (MADWeb 2024) found only 34.5% active even once, and 16,923 manipulated content.

Residential IPs from $0.30

Stop burning IPs on a pool you haven’t tested.

Proxy-Seller’s residential pool combines 47M+ ethically sourced IPs, per-request rotation, and sticky sessions with a configurable TTL across 220+ locations.

Buy residential proxies

Keep browser fingerprints consistent

A residential IP alone won’t bypass a ban if your scraper fingerprints as a Python client, so aim for one consistent browser and operating system.

User-Agent and header consistency

User-Agent rotation helps only if the rest of the headers agree with it. Picture a User-Agent that claims macOS Safari but arrives with Chrome-style sec-ch-ua client hints and a locale that contradicts the browser’s stated language settings. That’s more suspicious than no rotation at all. Keep these headers aligned:

  • User-Agent
  • Accept
  • Accept-Language
  • Referer
  • sec-ch-ua

Together with the fonts the browser exposes, they should describe one plausible browser on one plausible OS, with a language and timezone that match the exit IP’s country.

TLS and HTTP/2 fingerprints

Before any header is read, your TLS ClientHello exposes cipher suites and extensions, and HTTP/2 settings add more fingerprinting signals. A default requests client uses Python’s TLS stack and stays on HTTP/1.1, while a real Chrome or Firefox presents its own ClientHello and HTTP/2 settings. Sites that fingerprint the handshake (JA3, and its successor JA4) can tell those profiles apart.

You can check this yourself:

  1. Query a TLS-fingerprint echo endpoint with plain requests.
  2. Repeat the call with curl_cffi, the Python binding for curl-impersonate, set to impersonate Chrome.
  3. Compare the normalized JA3N (or the HTTP/2 fingerprint) of the two results.

One caveat comes from the curl_cffi documentation: Chrome 110 permutes the order of TLS extensions, so raw JA3 varies between connections, which is why step 3 compares JA3N. Impersonation aligns one transport layer with a browser profile, but it doesn’t guarantee access, and the curl_cffi FAQ notes that JA3 and Akamai fingerprints are not comprehensive.

Python

import random
from curl_cffi import requests

# Rotate whole browser profiles, not bare UA strings: TLS, HTTP/2, UA, and headers change together
PROFILES = ["chrome", "safari", "firefox"]  # aliases for the newest targets in your curl_cffi build

for page in (1, 2, 3):
    with requests.Session(impersonate=random.choice(PROFILES)) as s:  # one profile per session
        r = s.get(f"https://httpbin.org/headers?page={page}",
                  headers={"Referer": "https://httpbin.org/"}, timeout=20)
        print(r.status_code, r.json()["headers"]["User-Agent"])  # UA matches the chosen profile

Each session picks one full browser profile: the code sets only Referer, and the impersonation layer adds that browser’s TLS handshake, HTTP/2 settings, User-Agent, and default headers. In a local test with curl_cffi 0.16.3, all three profiles reported macOS, but targets change between releases, so pin the version you test. Keep one profile per session, and pair a new profile with a new exit IP.

Headless automation signals

Headless browsers can leak. On sites that check for them, CDP (Chrome DevTools Protocol) automation flags, inconsistent fonts, an unusual viewport, and navigator.webdriver can each mark a session as automated. Aim for a real browser context, with a normal locale, a realistic viewport, and no leftover automation markers. Still, no stealth setup is a guaranteed way to bypass an IP ban in web scraping.

Alter scraper traffic behavior

Even an aligned IP and fingerprint won’t bypass an IP ban in web scraping if the traffic pattern is mechanical.

Randomize timing. Space out your scraping with jittered delays tuned to the target’s published rate policy. As a starting point, not a universally safe rate, a Gaussian delay of random.gauss(3.5, 0.75) keeps about 95% of waits between 2 and 5 seconds. Jitter adds variation, but it doesn’t replace a shared budget, a concurrency cap, or Retry-After.

Set a shared request budget. Per-worker throttling and a per-domain concurrency cap don’t control the total load when several workers hit one host. Apply one shared rate budget to every worker on a domain, or ten workers can each run at a “safe” rate and still overwhelm the site.

Back off on 429, and honor Retry-After. Slow down instead of retrying right away, and respect Retry-After when the server sends it, in seconds or as an HTTP-date.

Python

import random, time, email.utils, requests
# Retry on 429: honor Retry-After (seconds or HTTP-date), else exponential backoff with jitter
def get_with_backoff(url, tries=5, cap=300):
    for attempt in range(tries):
        r = requests.get(url, timeout=20)
        if r.status_code != 429 or attempt == tries - 1:
            return r                       # non-429, or last try still 429 (caller decides)
        ra = r.headers.get("Retry-After", "")
        try:
            wait = float(ra) if ra.isdigit() else \
                   email.utils.parsedate_to_datetime(ra).timestamp() - time.time()
        except (TypeError, ValueError, AttributeError):
            wait = 2 ** attempt            # missing or malformed header
        time.sleep(min(max(0, wait), cap) + random.uniform(0, 1))  # cap limits a huge Retry-After
print(get_with_backoff("https://httpbin.org/status/429", tries=3).status_code)  # 429 after ~3-5 s

Test the parser with both Retry-After forms (120 and Wed, 21 Oct 2026 07:28:00 GMT) using a fixed clock and a mocked time.sleep(). A parser that only calls int() raises ValueError on the date and kills the retry, while this one falls back to exponential delay. The cap keeps a Retry-After: 86400 from parking a worker for a day, and time.sleep() pauses only this worker, so a shared cooldown needs the rate budget above. The demo URL, httpbin.org/status/429, never sends Retry-After, so it runs two backoff pauses and returns the final 429.

Keep session state coherent. Persist cookies within a task. Rotate identity between batches, not mid-flow. Avoid honeypot traps, which are links not meant for real users. Scrape only the pages and elements your task needs, instead of following every anchor on a page.

Schedule around the target’s load. For large runs, schedule collection outside the target’s peak traffic hours to reduce load on the site. Deduplicate URLs before a run and cache pages that rarely change, so retries don’t re-fetch content you already have.

Handle CAPTCHAs, 403s, and 429s without escalating the ban

Each signal needs its own reaction if you want to bypass an IP ban in web scraping without prolonging it.

On a 429, honor Retry-After, lower the rate, and retry only after the wait. Reduce concurrency and the shared rate before you change anything else.

On an immediate 403, inspect the response headers before you act. Per Cloudflare’s documentation, a cf-mitigated: challenge header identifies a Challenge Page, not an ordinary origin response. A critical-ch header, which MDN flags as experimental, lists the Client Hints marked as critical, so compare them with the headers your client sends before you change the proxy.

In a public curl_cffi GitHub discussion (#591, June 2025), a collector of public sports data hit exactly this pattern. The chrome110, chrome120, and safari17_0 profiles all returned 403 challenge responses, while the same URL worked in a regular browser. Our reference on proxy errors breaks down what each status code does and doesn’t tell you.

On a CAPTCHA, treat it as a verification requirement, because it doesn’t identify the underlying cause. Fixing routing or client inconsistencies may reduce challenges on some targets, but repeated CAPTCHAs mean you should review the permitted access route, not auto-solve in a loop. If you’re researching how these systems work, our explainer on how to bypass reCAPTCHA covers the mechanics responsibly.

On an HTTP 200 that looks wrong, validate content, not just status: a challenge page or a geo-wrong page returned with 200 still fails collection.

Note. Endless retries add load and cost without new information. Stop repeated retries when the response remains unchanged. Record the signal and inspect what you sent. Resume only after a permitted change addresses a condition the evidence supports.

Measure what actually changed

Two numbers show whether a fix to bypass an IP ban in web scraping worked, measured on a fixed sample of permitted URLs:

  • valid response rate (VRR = valid responses ÷ total requests)
  • cost per valid response (CPVR = run cost ÷ valid responses)

Count every retry in total requests and run cost, and apply the same content-validation rule before and after the change. Track both numbers on your own permitted URL sample. They’re a comparison method, not a promised result, and the scoped-test checklist in the conclusion turns them into a decision.

TEST BEFORE YOU BUY

Measure the fix, not the promise.

Proxy-Seller scopes a free POC to your proxy type and volume, so you can test the +20–30% VRR and –20–35% CPVR seen in A/B pilots on your own workload.

Talk to our team

Deploy anti-detect frameworks for hardened targets

A full browser stack helps bypass an IP ban in web scraping only when the target renders critical content in JavaScript or runs aggressive client-side checks. HTTP clients can’t help with those checks: the curl_cffi documentation notes that it has no JavaScript runtime and can’t change JavaScript fingerprints. For other public-data web scraping, well-configured HTTP clients and a coherent transport profile come first.

When you do need a browser, options include:

  • Playwright with a stealth plugin such as playwright-stealth
  • Camoufox, an open-source Firefox fork that, per its README, spoofs fingerprint data at the C++ level and exposes a Playwright-compatible Python interface
  • Nodriver, which its README calls the official successor to undetected-chromedriver; it’s fully async and drives Chrome over CDP

None of them guarantees access, so treat each profile as a test variable and re-run a fixed sample after target, browser, or library changes. Route the browser through your proxy, then confirm the exit IP and headers it sends.

Python

from playwright.sync_api import sync_playwright
from playwright_stealth import Stealth  # tested: playwright 1.56, playwright-stealth 2.0.3

# Check what a public test page receives from Chromium with stealth patches, via your proxy
PROXY = {"server": "http://res.proxy-seller.com:10000",
        "username": "USERNAME_s_200_ttl_30m", "password": "PASSWORD"}  # sticky: one IP per browser
with Stealth().use_sync(sync_playwright()) as p:
    browser = p.chromium.launch(headless=True, proxy=PROXY)  # drop proxy= for a first local test
    page = browser.new_context(locale="en-US").new_page()
    page.goto("https://httpbin.org/anything")
    print(page.evaluate("navigator.webdriver"))  # False in a local test; plain headless: True
    print(page.inner_text("body"))               # "origin" = exit IP, plus the headers sent
    browser.close()

In a local check with Playwright 1.56, plain headless Chromium reported navigator.webdriver as true and a HeadlessChrome User-Agent, and playwright-stealth 2.0.3 removed both markers. That removes two obvious tells, not the whole fingerprint, and results can differ on other versions. Version 2 also dropped the stealth_sync(page) call that older tutorials use.

The username pins a sticky session because a browser loads many resources per page, and rotating the IP for each one could scatter them across exit IPs. The echo endpoint confirms your proxy and headers, not how a defended target scores the browser. To confirm your exit IP and geo before a run, see our walkthrough on methods to verify an IP address.

FAQ: bypass IP ban web scraping

Four quick answers for anyone who’s already hit an IP ban in web scraping:

How long does an IP ban last when scraping?

Temporary blocks can clear within minutes or last several days, depending on the site’s rule, and a 429 may state the exact wait in Retry-After. A cooldown may also expire once traffic slows. A challenge may need extra verification, while a ban tied to an address or account may not clear on its own. Inspect the response body, headers, and access policy before you resume.

Can a VPN replace proxies for scraping?

A VPN changes your outbound IP, so it can bypass a one-off IP ban in low-volume, permitted web scraping. At scale, though, a typical consumer VPN gives you less control over individual scraper sessions and provider-managed rotation. You can’t easily get a fresh exit IP for each page or hold separate sticky sessions, and a single shared IP may hit rate limits sooner. Residential and ISP proxy pools are built for that per-session control.

Is scraping with a rotated User-Agent enough to bypass IP ban web scraping blocks?

Not necessarily, because the User-Agent is just one header. The site can still read your TLS and HTTP/2 fingerprint, your other headers, and your scraping pattern. Rotating it alone can even create a mismatch, such as a Windows User-Agent sent with macOS client hints, so change the whole browser profile with it.

Is scraping legal if I respect robots.txt?

Respecting robots.txt is good practice, but it isn’t a complete legal answer on its own. RFC 9309 standardizes the Robots Exclusion Protocol as an advisory signal and states plainly that its rules “are not a form of access authorization.” Legality also depends on the site’s Terms of Service, whether the data is public or login-gated, and your jurisdiction. Collect publicly available data, honor robots.txt and rate policies, avoid personal or gated data, and consult counsel for anything ambiguous.

Conclusion

To bypass IP bans in web scraping, start with the ban trigger, not the workaround, and check the response code and IP reputation first. If the signal is a 429, fix the schedule and rate before anything else. Otherwise, work the layers in order: proxy, then fingerprint, then behavior. Stop at the layer the evidence points to.

Before any change you make to bypass the ban, run a scoped test and let the numbers decide:

  • the same sample of permitted URLs, before and after
  • one variable at a time: hold the shared rate budget fixed when you compare proxies, and make the budget itself the variable when you test a slower rate
  • the same test duration
  • the same validity check on each response (expected fields present, no challenge page behind a 200)
  • compared on valid response rate and cost per valid response, with retries included

PERFORMANCE VALIDATION

Stop changing infrastructure on a guess.

Run Proxy-Seller’s free POC on your own URL sample, scoped to your proxy type and volume, before you commit a budget. ISO/IEC 27001:2022 certified.

Request a free POC

PROXIES FROM $0.02/IP

Ready to put this into practice?

Pick a proxy type and location, and start using it within minutes.

See pricing

Instant delivery

24/7 support

Refund policy


Share:

About the author

Igor Efimenko

Igor Efimenko

Head of Development @ Proxy-Seller
Head of Development at Proxy-Seller with 12+ years in DevOps. Writes on Python scraping, browser automation, and engineering benchmarks — backed by production experience at scale.

Related articles

Web Scraping

How to use a Python rotating proxy for web scraping: full guide

This guide covers how to set up a Python rotating proxy for scraping using Requests, advanced pool strategies, an API alternative, Scrapy, and Selenium.
04 September 2024
Web Scraping

Web Scraping in 2026: Top Proxies to Choose

Web scraping with a proxy is simply an automated way of extracting data from websites. It is used for a variety of tasks including price tracking, market research, content collection, etc.
30 May 2025
Insights

GDPR-compliant web scraping: what in-house teams must verify

GDPR web scraping now draws some of the heaviest data protection enforcement across the EU.
19 September 2026