Web Scraping

How to use a Python rotating proxy for web scraping: full guide

Igor Efimenko
Igor Efimenko
04 September 2024 • 18 min read

When your web scraper starts choking on HTTP 429 errors, knowing how to rotate proxies in Python is the only way to get your data pipeline moving again. To rotate proxies in Python, you spread requests across multiple IP addresses, but rotation alone won't help without proper request pacing and session control. This guide covers how to set up a Python rotating proxy for scraping using Requests, advanced pool strategies, an API alternative, Scrapy, and Selenium.

What you'll learn:

  • What a rotating IP does and when rotation is useful for web scraping
  • How to rotate proxies in Python using Requests (random, round-robin, retries)
  • Advanced strategies: subnet diversity, IP health tracking, weighted rotation, rotation frequency
  • How to use a rotating proxy API instead of managing your own pool
  • How to rotate proxies in Scrapy and in Selenium

What is proxy rotation and why use it for automated web scraping?

Proxy rotation distributes requests across multiple IP addresses instead of concentrating a workload on one route. For the full background on rotating proxies for scraping, what it means, and how IP rotation for scraping in Python works, see our dedicated guide. Websites detect scraping through request rates, repetitive timing, and requests from blocklisted IP addresses. When a site blocks that IP, all requests using it can fail, even if the scraper's code and target URL haven't changed.

This is mitigated through rotation, as this allows the software to switch to a new verified route on each request. This lets the software generate natural traffic patterns while preventing too many connections from one IP address at once. If you need multiple connections for processes such as ranking tracking, price comparison, and ad verification, use 10+ simultaneous connections with rotating residential proxies. They will suffice for debugging and testing purposes, but for any login process, form submission, or any other process involving cookies, the same IP address needs to be used.

How to rotate proxies in Python using the Requests HTTP library

When you rotate proxies in Python with Requests, you control authentication, verification, IP selection order, timeouts, and retry limits directly. No framework middleware needed. If anything goes wrong, you can check the chosen IP and the response, then decide what to fix.

Install required Python libraries and configure your environment

Create a virtual environment before installing Requests, especially when the scraper also depends on Scrapy, Selenium, or another HTTP client. These Python rotating proxy examples use requests; json, random, itertools, and time come from the standard library, so the dependency surface stays small. Pin the Requests version so production uses the behavior you tested locally, and commit the lock file with the project. Rebuild the environment from that lock file in CI to catch dependency drift before it looks like an IP-based failure.

Python

pip install requests

Set up proxy authentication and handle connection credentials

Proxies that need authentication require the Proxy-Authorization header. Faulty, absent, or improperly encoded credentials result in HTTP 407 (Proxy Authentication Required). For HTTPS targets, embed credentials in the proxy URL. The auth parameter in Requests controls target-server authentication, not proxy-tunnel authentication. A production Python rotating proxy should load secrets from environment variables or a secret manager and keep complete proxy URLs out of logs.

Python

import requests
from urllib.parse import quote

proxy_username = quote("your_username", safe="")
proxy_password = quote("your_password", safe="")
proxy_host = "192.168.1.10:8080"

proxy_url = f"http://{proxy_username}:{proxy_password}@{proxy_host}"
proxies = {"http": proxy_url, "https": proxy_url}

response = requests.get(
    "https://httpbin.org/ip", proxies=proxies, timeout=10
)

Load your proxy list from an external configuration JSON file

Hardcoded pool entries become painful as soon as an address, port, or credential changes. Keep the details of each rotating proxy Python uses in an external proxies.json file so you can update addresses and credentials without changing your code. Store each record as a dictionary, then check the top-level list and required keys before returning anything to the caller. Use a small fixture pool locally, keep deployment credentials in protected storage, and fail before a malformed entry can reach the request loop.

Create a proxies.json configuration file:

Python

[
  {
    "proxy_address": "192.168.1.10:8080",
    "proxy_username": "user1",
    "proxy_password": "pass1"
  },
  {
    "proxy_address": "192.168.2.20:8080",
    "proxy_username": "user2",
    "proxy_password": "pass2"
  },
  {
    "proxy_address": "10.0.0.30:3128",
    "proxy_username": "user3",
    "proxy_password": "pass3"
  }
]

Ingest the configuration file into your script using the following function:

Python

import json

REQUIRED_KEYS = {"proxy_address", "proxy_username", "proxy_password"}


def load_proxy_list(filepath="proxies.json"):
    with open(filepath, "r", encoding="utf-8") as file:
        proxies = json.load(file)
    if not isinstance(proxies, list):
        raise ValueError("Proxy configuration must be a list.")
    for index, proxy in enumerate(proxies):
        if not isinstance(proxy, dict):
            raise ValueError(f"Proxy entry {index} must be an object.")
        missing = REQUIRED_KEYS - proxy.keys()
        if missing:
            names = ", ".join(sorted(missing))
            raise ValueError(f"Proxy entry {index} is missing: {names}")
    return proxies

Validate proxy addresses before executing data extraction requests

A proxy might accept the TCP connection but still not work for the specific workload. We need to do more than just check for the open port. Check each proxy against a neutral IP endpoint and then run a smaller targeted test where collection is allowed. Log the result along with the IP address and the time taken. If there are repeated failure attempts, consider a cooldown. If there was only one failure, just log it.

Python

import requests
from urllib.parse import quote


def verify_proxies(proxy_list, test_url="https://httpbin.org/ip"):
    valid_proxies = []
    for proxy in proxy_list:
        address = proxy["proxy_address"]
        username = quote(proxy.get("proxy_username", ""), safe="")
        password = quote(proxy.get("proxy_password", ""), safe="")
        credentials = f"{username}:{password}@" if username and password else ""
        proxy_url = f"http://{credentials}{address}"
        proxies = {"http": proxy_url, "https": proxy_url}
        try:
            response = requests.get(
                test_url, proxies=proxies, timeout=5
            )
        except requests.exceptions.RequestException:
            print(f"Proxy connection failed: {address}")
            continue
        if 200 <= response.status_code < 300:
            valid_proxies.append(proxy)
            print(f"Proxy validated: {address}")
    return valid_proxies

Execute basic random proxy rotation with the random choice module

The simplest proxy selector in Python is random.choice(valid_proxies), assuming that we verified that the list is non-empty. This selector is stateless and easy to follow. Randomness may lead to duplicate addresses and uneven load on short iterations, though. Use round-robin or weighted rotation only when logs show uneven distribution or poor pool quality.

Python

import random
import requests
from urllib.parse import quote


def fetch_with_random_proxy(url, valid_proxies):
    if not valid_proxies:
        raise ValueError("valid_proxies must not be empty.")
    proxy = random.choice(valid_proxies)
    address = proxy["proxy_address"]
    username = quote(proxy.get("proxy_username", ""), safe="")
    password = quote(proxy.get("proxy_password", ""), safe="")
    credentials = f"{username}:{password}@" if username and password else ""
    proxy_url = f"http://{credentials}{address}"
    proxies = {"http": proxy_url, "https": proxy_url}
    return requests.get(url, proxies=proxies, timeout=8)

Configure round-robin proxy rotation using the itertools module

Round-robin selection employs itertools.cycle() to traverse all entries in the Python rotating proxy pool. In a single process, this yields a predictable and even selection pattern without maintaining a separate index. Different workers still make their own cycles unless they share state, which matters only when scaling up. If per-IP rate limits matter, store the selection counter in a shared queue or coordinator.

Python

import itertools
import requests
from urllib.parse import quote


def create_round_robin_cycle(valid_proxies):
    if not valid_proxies:
        raise ValueError("valid_proxies must not be empty.")
    return itertools.cycle(valid_proxies)


def fetch_with_round_robin(url, proxy_cycle):
    proxy = next(proxy_cycle)
    address = proxy["proxy_address"]
    username = quote(proxy.get("proxy_username", ""), safe="")
    password = quote(proxy.get("proxy_password", ""), safe="")
    credentials = f"{username}:{password}@" if username and password else ""
    proxy_url = f"http://{credentials}{address}"
    proxies = {"http": proxy_url, "https": proxy_url}
    return requests.get(url, proxies=proxies, timeout=8)

Handle failed proxy requests and configure automated retry loops

Network failures are common, so establish clear parameters and exit criteria for retries. In the Python code for proxy rotation, each try is counted, and the current proxy is removed from the available list for that request. The retries process stops once the limit is reached. Don't confuse network errors with valid HTTP responses. A common mistake is marking the IP as failed after receiving anything other than 200. For status codes 429, 502, 503, and 504, select a new IP with backoff and randomized jitter.

Python

import random
import time
import requests
from urllib.parse import quote

RETRYABLE_STATUS = {429, 502, 503, 504}


def execute_request_with_retry(url, valid_proxies, max_retries=3):
    available = list(valid_proxies)
    for attempt in range(max_retries):
        if not available:
            break
        proxy = random.choice(available)
        address = proxy["proxy_address"]
        username = quote(proxy.get("proxy_username", ""), safe="")
        password = quote(proxy.get("proxy_password", ""), safe="")
        credentials = f"{username}:{password}@" if username and password else ""
        proxy_url = f"http://{credentials}{address}"
        proxies = {"http": proxy_url, "https": proxy_url}
        try:
            response = requests.get(
                url, proxies=proxies, timeout=6
            )
        except requests.exceptions.RequestException:
            response = None
        else:
            if response.status_code not in RETRYABLE_STATUS:
                response.raise_for_status()
                return response
        available.remove(proxy)
        if attempt + 1 < max_retries:
            delay = min(2 ** (attempt + 1), 8)
            time.sleep(delay + random.uniform(0, 0.5))
    return None

Residential IPs from $0.30

Stop losing budget to invalid responses.

Proxy-Seller provides access to 47M+ residential IPs across 220+ locations, with country, region, city, and ISP targeting. Choose per-request or timed rotation, or use sticky sessions when your workflow needs a consistent IP.

Get residential proxies

Advanced IP rotation strategies for resilient scraping tasks

A basic rotating proxy for scraping will do until logs indicate failure related to the same subnet, ASN, or performance characteristics. When applying IP rotation for scraping in Python, you can detect connection health, apply cooldowns, alternate networks, and select IPs based on past performance. Add controls one by one, because conflicting rules will obfuscate the initial problem. Here are methods of controlling these variables in isolation.

Implement subnet and ASN diversity to prevent network blocks

IPs from distinct networks can end up belonging to the same network prefix or ASN. This creates an undetected failure domain within the Python Rotating Proxy Pool. Assume that the first three octets of an IPv4 are merely a useful /24 classification. It is not a guarantee that both IPs use separate providers or ASNs. The helper function will organize all entries by this prefix and round-robin each group without filtering out any entries. This increases variety in the presence of multiple groups. For ASN diversity, enhance your pool with provider data and a rule for selecting entries based on it.

PRO TIP: If your scraper experiences sudden HTTP 403 blocks despite rotating across dozens of IPs, inspect the Autonomous System Numbers (ASNs). Rotating across 50 IPs within the same /24 subnet provides zero protection against network-level rate limits. Always verify that your IP pool spans diverse subnets and independent ASNs.

Python

from collections import defaultdict, deque


def get_subnet(ip_address):
    # Practical IPv4 /24 prefix; ASN requires verified metadata.
    return ".".join(ip_address.split(".")[:3])


def rotate_proxies_by_subnet(proxy_list):
    groups = defaultdict(deque)
    for proxy in proxy_list:
        ip_address = proxy["proxy_address"].split(":")[0]
        groups[get_subnet(ip_address)].append(proxy)
   
    while groups:
        for subnet in list(groups):
            yield groups[subnet].popleft()
            if not groups[subnet]:
                del groups[subnet]

When there is a requirement for stable sessions, ISP proxies combine provider-based infrastructure with IP addresses announced through consumer ISPs. Before using any diversity provided by networks, first ensure that there is current coverage of subnets and ASNs. You should also analyze a random sampling of received addresses. The selector can use information on the provider's metadata to avoid subsequent routes from the same network. This is possible only when this data is correct and accessible to all workers.

Rotate proxies asynchronously with aiohttp and concurrent workers

Since network requests waste time waiting for responses, aiohttp allows multiple workers to run concurrently. This example distributes IPs in a round-robin fashion and limits the batch to four workers. Each worker processes the next link when it finishes its current request. Use this model for pages that don't depend on cookies.

Python

pip install aiohttp




import asyncio
import json
from itertools import cycle
from urllib.parse import quote
import aiohttp


def build_proxy_url(proxy):
    username = quote(proxy["proxy_username"], safe="")
    password = quote(proxy["proxy_password"], safe="")
    credentials = f"{username}:{password}@" if username or password else ""
    return f"http://{credentials}{proxy['proxy_address']}"


async def fetch_pages(urls, proxy_list, workers=4):
    if not proxy_list or workers < 1:
        raise ValueError(
            "Provide a nonempty proxy pool and at least one worker."
        )
    routes = cycle(build_proxy_url(proxy) for proxy in proxy_list)
    jobs = iter(enumerate(urls))
    results = [None] * len(urls)
    timeout = aiohttp.ClientTimeout(total=30, connect=10)
   
    async with aiohttp.ClientSession(
        timeout=timeout,
        connector=aiohttp.TCPConnector(limit=workers),
        cookie_jar=aiohttp.DummyCookieJar(),
    ) as session:
        async def worker():
            for index, url in jobs:
                async with session.get(
                    url, proxy=next(routes)
                ) as response:
                    response.raise_for_status()
                    results[index] = await response.text()
                   
        async with asyncio.TaskGroup() as tasks:
            for _ in range(min(workers, len(urls))):
                tasks.create_task(worker())
    return results


if __name__ == "__main__":
    with open("proxies.json", encoding="utf-8") as file:
        proxy_list = json.load(file)
    urls = [
        "https://example.com/?page=1",
        "https://example.com/?page=2",
    ]
    pages = asyncio.run(fetch_pages(urls, proxy_list, workers=4))

The worker quota limits simultaneous requests rather than requests per second. Apply a common rate limit to each destination prior to scheduling larger batches. When one worker raises an unhandled exception inside a TaskGroup, Python cancels every other running task in that group and wraps all failures in an ExceptionGroup. In this example, a single HTTP error stops the full batch and the function returns none of the already-collected results. Catch errors inside the worker or collect partial results in a shared list before re-raising. For HTTP 429 and 503 responses, respect the Retry-After header.

Track IP health status and configure automated cooldown periods

However, a timeout is not a death certificate. A single failed attempt should not exclude an IP address. Count consecutive network failures and the most recent failure time. Only call mark_proxy_dead() if the threshold is exceeded. Do not include the IP address during cooldown. Try the IP again afterward, and reset the failure count on success. This ensures temporary issues don't reduce a healthy pool, while allowing bad proxies to recover.

Python

import time

dead_proxies = {}

def is_proxy_dead(proxy_address, cooldown_seconds=300):
    if proxy_address in dead_proxies:
        failure_time = dead_proxies[proxy_address]
        if time.time() - failure_time < cooldown_seconds:
            return True  # Proxy remains in cooldown
        else:
            del dead_proxies[proxy_address]  # Cooldown period completed
    return False

def mark_proxy_dead(proxy_address):
    dead_proxies[proxy_address] = time.time()

Apply weighted random rotation to favor high-performing proxies

The rotation scheme must incentivize good performance based on actual behavior. Rate each proxy for Python rotation based on its cumulative success ratio and average latency, then decrease that rating when it fails. Keep a floor weight so both new and recovering IPs stay eligible. The running average naturally dilutes old results as new samples arrive, but add an explicit decay or rolling window if your pool runs for days. The example code comes with imports and a cooldown helper included.

Python

import random
import time

dead_proxies = {}

def is_proxy_dead(proxy_address, cooldown_seconds=300):
    failed_at = dead_proxies.get(proxy_address)
    if failed_at is None:
        return False
    if time.time() - failed_at < cooldown_seconds:
        return True
    del dead_proxies[proxy_address]
    return False

def calculate_weight(proxy):
    successes = proxy.get("successes", 0)
    failures = proxy.get("failures", 0)
    avg_latency = max(proxy.get("avg_latency", 1.0), 0.05)
    success_rate = (successes + 1) / (successes + failures + 2)
    return max(success_rate / avg_latency, 0.05)

def select_weighted_proxy(proxies):
    healthy = [
        proxy for proxy in proxies
        if not is_proxy_dead(proxy["proxy_address"])
    ]
    if not healthy:
        raise RuntimeError("No healthy proxies are available.")
    weights = [calculate_weight(proxy) for proxy in healthy]
    return random.choices(healthy, weights=weights, k=1)[0]

How often should you rotate proxies during web scraping jobs?

The rotation period doesn't have to be standardized. Request independence and session state have more value than the number itself. When requests are independent and involve public pages, rotate on each request. However, keep the same exit IP for the full session when using cookies. In price tracking, use one request per IP every five seconds if possible, according to website limitations.

How to use a proxy rotation API with a Python Requests session

Validation, cooldown, session maps, and metrics require engineering effort to maintain. A rotating proxy api python implementation can communicate with the proxy provider's managed endpoint. Your code is then responsible only for request rates and validation. This makes implementation easier, but you lose some transparency of routing decisions. The comparison between these two methods should be done on the same parameters.

Self-managed IP pool versus an automated rotating proxy API

While a self-managed pool reveals all addresses and decisions to your program, the rotating proxy API reveals only one endpoint but implements the routing policy of the provider behind the endpoint. The service will either allocate a new IP address, keep a sticky session, or select from a location-specific pool depending on the settings. Prior to substituting the local list with the managed service, it is necessary to document the semantics of sessions, retries, observability, credentials, and locations. When deploying a rotating proxy in Python, a comparison of cost and valid responses should be done on the same targets.

Operational concern

Self-managed proxy pool

Managed rotation endpoint

Pool maintenance

Your code validates and scores each entry

The provider manages the available pool

Routing control

Full control over selection and cooldown logic

Controlled through plan, port, or parameters

Failure visibility

Your logs expose each proxy decision

Visibility depends on provider reporting

Session behavior

Your application stores sticky mappings

Provider-specific sticky-session controls

Engineering effort

Higher build and maintenance cost

Lower application complexity, recurring service cost

Connect to a rotating API endpoint using the Python Requests module

Configure the managed endpoint via requests.Session(). This ensures all connection settings remain in one place and can be safely reused. It is the service's responsibility to choose routes. However, the Python rotating proxy client should still include timeouts, response checking, retries, and exception handling. Credentials should never be embedded in source code, and the full URL of the endpoint should not be logged. If HTTP_PROXY or HTTPS_PROXY environment variables are set, Requests can override session.proxies. Clear them or set session.trust_env = False when using a managed endpoint.

Python

import os
import random
import time
import requests

RETRYABLE_STATUS = {429, 502, 503, 504}

def fetch_via_rotating_api(target_url, max_retries=3):
    proxy_url = os.environ["ROTATING_PROXY_URL"]
    session = requests.Session()
    session.trust_env = False
    session.proxies.update({"http": proxy_url, "https": proxy_url})
    last_error = None
   
    for attempt in range(max_retries):
        response = None
        try:
            response = session.get(target_url, timeout=10)
        except requests.exceptions.RequestException as error:
            last_error = error
        else:
            if response.status_code not in RETRYABLE_STATUS:
                response.raise_for_status()
                return response.text
               
        if attempt + 1 < max_retries:
            retry_after = response.headers.get("Retry-After", "") if response is not None else ""
            delay = int(retry_after) if retry_after.isdigit() else min(
                2 ** (attempt + 1), 8
            )
            time.sleep(delay + random.uniform(0, 0.5))
           
    raise RuntimeError(
        f"Request failed after {max_retries} attempts."
    ) from last_error

How to rotate proxies in Scrapy for automated data pipelines

Python proxy rotation in Scrapy does not limit simultaneous requests to just one per thread. For a deeper look at how Scrapy manages crawls, see our overview of Scrapy web scraping architecture. When using a Scrapy rotate proxy setup, you can let middleware assign proxies to requests, or pick a specific IP address in your spider. Automatic IP rotation for scraping in Python is more appropriate for general concurrent web scraping. Explicit assignment suits a specific URL or geographical location.

Configure automated rotation via the scrapy-rotating-proxies package

A Scrapy IP rotation setup can use the scrapy-rotating-proxies library. It picks out IP addresses for simultaneous requests while ignoring any flagged by its ban detection policy. Install and pin the package, then activate both middleware components in settings.py. The setup options below include an inline pool and a secure file method. Select one, but never put credentials under version control. Always test your ban detection and backoff parameters against your target's responses first.

Python

pip install scrapy-rotating-proxies




# settings.py

HTTPPROXY_ENABLED = True
DOWNLOADER_MIDDLEWARES = {
    "rotating_proxies.middlewares.RotatingProxyMiddleware": 610,
    "rotating_proxies.middlewares.BanDetectionMiddleware": 620,
}

# Option 1: define the pool inline.
ROTATING_PROXY_LIST = [
    "http://user:pass@192.168.1.10:8080",
    "http://user:pass@192.168.2.20:8080",
]

# Option 2: use a protected file instead of ROTATING_PROXY_LIST.
# ROTATING_PROXY_LIST_PATH = "/secure/path/proxies.txt"

ROTATING_PROXY_BACKOFF_BASE = 300

Assign proxies manually per request inside custom Scrapy spiders

A manually assigned rotating IP in Python allows routing based on the particular request rather than general crawler settings. Turn on HTTP proxy capabilities, pick the proxy address, and pass it via meta=. This keeps location, destination, and session routes clear so you can repeat failed requests in tests—Scrapy 2.13 deprecated start_requests() in favor of async def start(). The example below uses start_requests() for backward compatibility with Scrapy 2.12 and earlier. On Scrapy 2.13+, rename the method to async def start(self) and use yield inside it.

Python

import scrapy


class CatalogSpider(scrapy.Spider):
    name = "catalog_spider"
   
    proxy_pool = [
        "http://user:pass@192.168.1.10:8080",
        "http://user:pass@192.168.2.20:8080"
    ]
   
    def start_requests(self):
        urls = ["https://example.com/items/page1", "https://example.com/items/page2"]
        for index, url in enumerate(urls):
            selected_proxy = self.proxy_pool[index % len(self.proxy_pool)]
            yield scrapy.Request(
                url=url,
                callback=self.parse,
                meta={"proxy": selected_proxy}
            )

    def parse(self, response):
        yield {"url": response.url, "html": response.text}

Select auto middleware if you need failover, cooling, and consistent routing of requests during a concurrent crawl. If you have one specific request that must be routed via a predefined route, select manual assignment, such as routing via specific geographic IP addresses in proxy chains without changing the rest of the requests. Both approaches make use of the concurrency, delay, and retry features of Scrapy. Test both under the same load and session requirements as your actual crawler.

ISP proxies from $0.98

Need datacenter speed with residential trust?

Use Proxy-Seller’s ethically sourced ISP proxies when your authorized workflow needs a stable IP across a longer session.

Buy ISP proxies

How to rotate proxies in Selenium with Python for dynamic pages

With Python Selenium rotating proxies for scraping, the unit of rotation transforms from one HTTP request to an entire browser session. Rotating proxies for web scraping using ChromeOptions requires a new WebDriver per route rotation because Chrome sets proxy configuration at browser initialization. A Selenium rotating proxy has a lifecycle where each cookie-dependent journey needs its own driver. This is more time-consuming and uses extra memory. Automated browsing should only be used when JavaScript is really needed; otherwise, use Requests.

Configure rotating proxies for web scraping using ChromeOptions browser sessions

Pass --proxy-server into ChromeOptions before initializing the WebDriver. To get a rotating proxy Selenium Python flow working, initialize the WebDriver for each proxy server. Wait for the desired content, obtain the HTML source, and shut down the driver within a finally block. A new browser means a new cookie jar unless you intentionally load an old profile. Do not rotate within a stateful flow. This is slower than replacing a proxy in a Requests dictionary, but reflects the reality of Chrome.

Python

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

proxy_pool = ["192.168.1.10:8080", "192.168.2.20:8080"]


def scrape_with_selenium(url, proxy_address):
    options = Options()
    options.add_argument(f"--proxy-server=http://{proxy_address}")
    options.add_argument("--headless=new")
   
    driver = webdriver.Chrome(options=options)
    try:
        driver.get(url)
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.TAG_NAME, "body"))
        )
        return driver.page_source
    finally:
        driver.quit()


for proxy in proxy_pool:
    html = scrape_with_selenium("https://httpbin.org/ip", proxy)

ChromeOptions alone doesn't handle proxy authentication. Selenium Wire supports it via seleniumwire_options, but was archived in January 2024. SeleniumBase is a more recent, actively maintained option. The example below passes credentials from environment variables, sets a page load timeout, waits for the target element, and closes the browser even if navigation fails.

Python

import os
from base64 import b64encode
from seleniumwire import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait


def fetch_with_authenticated_proxy(url, ready_selector):
    # PROXY_URL contains the scheme, host, and port, without credentials.
    proxy_url = os.environ["PROXY_URL"]
    username = os.environ["PROXY_USERNAME"]
    password = os.environ["PROXY_PASSWORD"]
   
    credentials = f"{username}:{password}".encode("utf-8")
    token = b64encode(credentials).decode("ascii")
   
    seleniumwire_options = {
        "proxy": {
            "http": proxy_url,
            "https": proxy_url,
            "custom_authorization": f"Basic {token}",
            "no_proxy": "localhost,127.0.0.1",
        },
        "verify_ssl": True,
        "disable_capture": True,
    }
   
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")
   
    driver = webdriver.Chrome(
        options=options,
        seleniumwire_options=seleniumwire_options,
    )
    try:
        driver.set_page_load_timeout(30)
        driver.get(url)
        WebDriverWait(driver, 10).until(
            EC.visibility_of_element_located(
                (By.CSS_SELECTOR, ready_selector)
            )
        )
        return driver.page_source
    finally:
        driver.quit()


html = fetch_with_authenticated_proxy(
    "https://example.com/",
    "h1",
)

Set PROXY_URL to your provider's endpoint (e.g., http://host:port), and supply PROXY_USERNAME and PROXY_PASSWORD through your deployment's secret settings.

Proxy rotation best practices for consistently valid responses

An effective rotation policy helps increase the valid response rate while maintaining the ability to explain failures to an operator investigating an incident. Differentiate between network errors and target responses, maintain session coherence, and then rotate routes. Modify one control at a time and compare VRR, latency, retries, and cost of a valid response to the same endpoint under real-world conditions.

Rotate healthy proxies based on real-time pool metrics

Monitor connection attempts, HTTP status codes, and latencies for each IP address. Log failure counts and timestamps rather than blocking permanently on a first failure. An IP returning a single timeout may recover within seconds; only mark it dead after it crosses a consecutive-failure threshold.

Configure backoff thresholds using Retry-After headers

Honor the Retry-After header when the target returns 429 or 503. If no header is present, use exponential backoff with jitter. Cap retry delays at 300 seconds to prevent indefinite stalls. Never retry on 403; that's an access policy decision, not a transient error.

Avoid excessive rotation during cookie-backed sessions

Don't rotate mid-session when the target tracks state via cookies. Changing IP addresses between sequential steps of a checkout flow or login process breaks the session. Rotate only when starting a new independent request or after the current session ends cleanly.

Use sticky sessions for stateful checkout and login flows

Configure a session timeout that matches the duration of the workflow. For multi-step processes, a sticky session lasting 10–30 minutes is usually sufficient. Proxy-Seller's residential IPs support sticky sessions with a technical maximum of 24 hours when the residential IP stays active for the full duration.

Maintain IP and User-Agent diversity across data pipelines

Rotating IP addresses without rotating User-Agent strings creates an immediate fingerprinting footprint for modern anti-bot systems. Pair each IP route with a consistent browser profile and matching platform headers. When switching to a new IP for an independent request, select a corresponding User-Agent from a validated pool to prevent header-to-network mismatches that signal automation.

Common proxy rotation errors and effective mitigation methods

Error code or symptom

Likely cause

Recommended response

HTTP 407 Proxy Authentication Required

Missing, malformed, or expired proxy credentials

Verify the secret source and URL-encode embedded credentials

HTTP 429 Too Many Requests

The target is limiting the current request pattern

Respect Retry-After, lower concurrency, and review per-IP pacing

HTTP 403 Forbidden

Target access policy, authorization, or request validation

Confirm the request is permitted; don't mark the proxy dead automatically

ConnectTimeout or ProxyError

The proxy connection is slow or unavailable

Try another validated proxy and apply cooldown after repeated failures

Complete Python reference for resilient proxy pool rotation logic

The below Python rotating proxy reference manager brings together the above five elements in a running example. It serves more as an infrastructure than a ready-to-deploy solution, as there is still a need for logs, secrets management, tests, state sharing, and target-specific validations. The code below distinguishes between target failures and connection failures, allows any 2xx response as success, and triggers the cooldown mechanism only after connection failures. The above program additionally processes Retry-After with numeric and HTTP date while ignoring the HTTP 403 or 429 responses as an indicator that the address is down.

Python

import json
import random
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import requests
from urllib.parse import quote

NETWORK_ERRORS = (
    requests.exceptions.ProxyError,
    requests.exceptions.Timeout,
    requests.exceptions.ConnectionError,
)
RETRYABLE_STATUS = {429, 502, 503, 504}
REQUIRED_KEYS = {"proxy_address", "proxy_username", "proxy_password"}

class ResilientProxyManager:
    def __init__(
        self, config_path="proxies.json", cooldown=300,
        failure_threshold=2,
    ):
        self.cooldown = cooldown
        self.failure_threshold = failure_threshold
        self.dead_proxies = {}
        self.network_failures = {}
        self.proxies = self._load_proxies(config_path)

    @staticmethod
    def _load_proxies(path):
        with open(path, "r", encoding="utf-8") as file:
            proxies = json.load(file)
        if not isinstance(proxies, list):
            raise ValueError("Proxy configuration must be a list.")
        for index, proxy in enumerate(proxies):
            if not isinstance(proxy, dict):
                raise ValueError(f"Proxy entry {index} must be an object.")
            missing = REQUIRED_KEYS - proxy.keys()
            if missing:
                names = ", ".join(sorted(missing))
                raise ValueError(f"Proxy entry {index} is missing: {names}")
            proxy.update(successes=0, failures=0, avg_latency=1.0)
        return proxies

    @staticmethod
    def _request_config(proxy):
        address = proxy["proxy_address"]
        username = quote(proxy.get("proxy_username", ""), safe="")
        password = quote(proxy.get("proxy_password", ""), safe="")
        credentials = f"{username}:{password}@" if username and password else ""
        proxy_url = f"http://{credentials}{address}"
        return {"http": proxy_url, "https": proxy_url}

    def validate_pool(self, test_url="https://httpbin.org/ip", attempts=2):
        valid = []
        for proxy in self.proxies:
            proxies = self._request_config(proxy)
            for attempt in range(attempts):
                try:
                    response = requests.get(
                        test_url, proxies=proxies, timeout=5
                    )
                except NETWORK_ERRORS:
                    if attempt + 1 < attempts:
                        time.sleep(1)
                        continue
                    break
                if 200 <= response.status_code < 300:
                    valid.append(proxy)
                    break
        self.proxies = valid
        if not self.proxies:
            raise RuntimeError("No validated proxies are available.")

    def _in_cooldown(self, address):
        failed_at = self.dead_proxies.get(address)
        if failed_at is None:
            return False
        if time.time() - failed_at < self.cooldown:
            return True
        del self.dead_proxies[address]
        return False

    @staticmethod
    def _weight(proxy):
        successes = proxy["successes"]
        failures = proxy["failures"]
        success_rate = (successes + 1) / (successes + failures + 2)
        return max(success_rate / max(proxy["avg_latency"], 0.05), 0.05)

    def _select_proxy(self, exclude_address=None):
        healthy = [
            proxy for proxy in self.proxies
            if not self._in_cooldown(proxy["proxy_address"])
        ]
        if not healthy:
            raise RuntimeError("All validated proxies are in cooldown.")
        alternatives = [
            proxy for proxy in healthy
            if proxy["proxy_address"] != exclude_address
        ]
        candidates = alternatives or healthy
        return random.choices(
            candidates,
            weights=[self._weight(proxy) for proxy in candidates],
            k=1,
        )[0]

    @staticmethod
    def _record(proxy, latency, succeeded):
        key = "successes" if succeeded else "failures"
        proxy[key] += 1
        samples = proxy["successes"] + proxy["failures"]
        proxy["avg_latency"] += (
            latency - proxy["avg_latency"]
        ) / samples

    @staticmethod
    def _retry_delay(response, attempt):
        value = response.headers.get("Retry-After")
        if value:
            try:
                return min(max(float(value), 0), 300)
            except ValueError:
                try:
                    retry_at = parsedate_to_datetime(value)
                    if retry_at.tzinfo is None:
                        retry_at = retry_at.replace(tzinfo=timezone.utc)
                    seconds = (
                        retry_at - datetime.now(timezone.utc)
                    ).total_seconds()
                    return min(max(seconds, 0), 300)
                except (TypeError, ValueError, OverflowError):
                    pass
        base = min(2 ** (attempt + 1), 8)
        return base + random.uniform(0, 0.5)

    def execute_request(self, url, max_retries=3):
        last_address = None
        for attempt in range(max_retries):
            proxy = self._select_proxy(last_address)
            address = proxy["proxy_address"]
            last_address = address
            proxies = self._request_config(proxy)
            started = time.monotonic()
           
            try:
                response = requests.get(
                    url, proxies=proxies, timeout=8
                )
            except NETWORK_ERRORS:
                self._record(proxy, time.monotonic() - started, False)
                failures = self.network_failures.get(address, 0) + 1
                self.network_failures[address] = failures
                if failures >= self.failure_threshold:
                    self.dead_proxies[address] = time.time()
                    self.network_failures[address] = 0
                if attempt + 1 < max_retries:
                    time.sleep(min(2 ** attempt, 4) + random.uniform(0, 0.3))
                continue
               
            self.network_failures[address] = 0
            latency = time.monotonic() - started
           
            if 200 <= response.status_code < 300:
                self._record(proxy, latency, True)
                return response
               
            self._record(proxy, latency, False)
            if response.status_code not in RETRYABLE_STATUS:
                response.raise_for_status()
                return response
               
            if attempt + 1 < max_retries:
                time.sleep(self._retry_delay(response, attempt))
               
        raise RuntimeError(
            f"Failed to fetch {url} after {max_retries} attempts."
        )

if __name__ == "__main__":
    manager = ResilientProxyManager("proxies.json")
    manager.validate_pool()
    response = manager.execute_request("https://httpbin.org/ip")
    print("Extraction response:", response.json())

Ethycally sourced IPs

Scale your scraping pipeline with clean dedicated pools.

Proxy-Seller offers residential and ISP proxies for the Python proxy rotation process requiring location targeting or session stability. During A/B testing, clients experienced an increase of 20% to 30% in the valid response rate over their previous vendor’s proxies.

See pricing

Frequently asked questions about Python proxy rotation methods

These answers cover the questions I see most often from teams setting up Python proxy rotation for the first time.

What is proxy rotation in automated public web scraping workflows?

IP rotation ensures that outbound connections are spread across several IPs from an organized or manually maintained pool. This makes sure there is no dependency on a single IP and also helps isolate rate limiting, timeouts, and reputation-related issues. However, rotation does not ensure availability and does not serve as an alternative to proper pacing.

How does a Python rotating proxy work in Python Requests sessions?

Python rotating proxies for scraping pick a validated item from the pool and apply it to the Requests proxy dictionary before making an eligible request. Random selection is done via random.choice() or an ordered one through itertools.cycle(). Put the request inside a retry loop with limits, measure latency and HTTP code, and do not kill the address just because of a 403 or 429 response from the target.

How do I rotate proxies in Selenium browser automation scripts?

Rotate proxy servers using the --proxy-server command in ChromeOptions before starting the WebDriver. This is because Chrome sets the proxy server on startup; close the old driver and start a new one whenever you switch proxies. Keep one IP for any cookie-related process.

What is the best proxy type for rotating IP network configurations?

The best method depends on your session and geographic requirements. Proxy-Seller's residential proxies suit all requirements, with over 47 million IPs across 220+ locations. You can rotate per request, at fixed intervals, or use sticky sessions.

How many proxies do I need for scalable web scraping workflows?

There is no universal formula for IP count. For instance, a pipeline sending 120 requests per minute needs 10 IP addresses, with one request per IP address in five-second intervals. Adding 20% spare capacity to 10 working IPs gives you 12 total, assuming the website accepts this request frequency.

How can a Python rotating proxy reduce blocks during web scraping?

Using a rotating proxy for scraping in Python can help reduce concentrated load, provided you use it with appropriate pacing, limited concurrency, backoff for retries, valid IPs, and stable sessions. Don't rotate unless necessary, and don't change the IP address every time a 403 error occurs. Collect the publicly available information within your rights.

Conclusion: build a dependable data collection infrastructure

Begin with the minimum Python rotating proxy setup required for the workload: validation, round robin, retries with bounds, and address-level performance metrics. Include subnet variety, weighting, or a proxy management API only if logs show a particular type of failure. Requests is great for simple, straight HTTP pipelines; Scrapy is great for concurrent scraping; and Selenium is great when you need web browser rendering. Test each of these against authorized endpoints. Measure valid responses, latencies, retries, and cost per valid response, all at the same concurrency and session management.

PROXIES FROM $0.02/IP

Ready to put this into practice?

Pick a proxy type and location, and start using it within minutes.

See pricing

Instant delivery

24/7 support

Refund policy


Share:

About the author

Igor Efimenko

Igor Efimenko

Head of Development @ Proxy-Seller
Head of Development at Proxy-Seller with 12+ years in DevOps. Writes on Python scraping, browser automation, and engineering benchmarks — backed by production experience at scale.

Related articles

Web Scraping

Scraping Airbnb Listing Data with Python

The article provides a guide on creating a scraper for data collection on Airbnb using Python. It includes a functional code example for parsing key attributes of listings on the Airbnb website.
23 August 2024
Web Scraping

Guide to Financial Data Web Scraping

The article reviews financial data web scraping, its methods, tools, use cases, and legal aspects for faster market analysis.
10 September 2025
Web Scraping

How to Rotate Proxies in Python Using Requests

This material will explain how to implement such rotation using the popular requests library.
08 September 2025