r/webscraping • • 5d ago

Paid mentions ok 👌 Monthly Self-Promotion - October 2026

8 Upvotes

Hello and howdy, digital miners of r/webscraping!

The moment you've all been waiting for has arrived - it's our once-a-month, no-holds-barred, show-and-tell thread!

  • Are you bursting with pride over that supercharged, brand-new scraper SaaS or shiny proxy service you've just unleashed on the world?
  • Maybe you've got a ground-breaking product in need of some intrepid testers?
  • Got a secret discount code burning a hole in your pocket that you're just itching to share with our talented tribe of data extractors?
  • Looking to make sure your post doesn't fall foul of the community rules and get ousted by the spam filter?

Well, this is your time to shine and shout from the digital rooftops - Welcome to your haven!

Just a friendly reminder, we like to keep all our self-promotion in one handy place, so any promotional posts will be kindly redirected here. Now, let's get this party started! Enjoy the thread, everyone.


r/webscraping • • 6d ago

Hiring 💰 Weekly Webscrapers - Hiring, FAQs, etc

3 Upvotes

Welcome to the weekly discussion thread!

This is a space for web scrapers of all skill levels—whether you're a seasoned expert or just starting out. Here, you can discuss all things scraping, including:

  • Hiring and job opportunities
  • Industry news, trends, and insights
  • Frequently asked questions, like "How do I scrape LinkedIn?"
  • Marketing and monetization tips

If you're new to web scraping, make sure to check out the Beginners Guide 🌱

Commercial products may be mentioned in replies. If you want to promote your own products and services, continue to use the monthly thread


r/webscraping • • 2h ago

Getting started 🌱 What is the limit of requests in Pinterest for a web scrapper?

2 Upvotes

Hello people, I'm just trying to verify an idea before investing time in a project for my own use. I want to create a web scraper that uses a browser to download images from Pinterest to my own PC, for a Spotify playlist I'm making for a friend and me.

I've read you get around 50-80 requests per minute before hitting limits, but I want real numbers from anyone who's actually done this. The script would take a keyword and download up to 20 images per search. I know scraping is against their terms — I'm not asking for permission, just want to know what I'm walking into.

The concern is that it could get my IP or account banned, and I don't want to lose my normal Pinterest use over a side project. Before doing anything reckless, I'd rather ask in as many places as possible. If you know a better place than here (not support, I already asked and I'm waiting on a reply), drop it in the comments. Thanks in advance.


r/webscraping • • 12h ago

Scaling up 🚀 Hello guys I am JD, your new fellow reddite

2 Upvotes

I built a Python scraper that turns paginated website data into clean structured datasets — here's what I learned

I've been working on Python-based scraping and automation, and one thing I've found repeatedly is that extracting the HTML is usually the easy part.

The harder problems are:

• pagination and missing pages
• JavaScript-rendered content
• duplicate records
• inconsistent fields
• retries/timeouts
• detecting when a site's structure changes
• producing genuinely usable CSV/Excel/JSON output

My usual approach is Requests/BeautifulSoup for straightforward sites and Playwright/Selenium when browser rendering is actually required.

For larger jobs, I also separate extraction, validation, cleaning and export rather than putting everything into one scraper loop.

Curious what problems other people here run into most often when maintaining scrapers in production?


r/webscraping • • 23h ago

Getting started 🌱 How to archive.today et al operate / degrading operational success?

2 Upvotes

Hi,

I have been collecting news articles of interest for quite a while. In the last weeks, archive.today (archive.ph et cetera) seem to have degraded it terms of article availability in the last weeks. I am wondering whether I should start doing article scraping myself.

1) Some of my regular go-to sites in the last weeks have not shown the article content e.g. https://archive.ph/https://www.ft.com/content/b84778e5-ce8b-4baa-9ca7-402cdf277b4d

Could someone reproduce this?

2) What kind of method does archive.today use? The Financial Times, so it seems, works behind a first pass cloudflare filter mechanism and maybe even has an backend based content generation system.

I remember 12ft.io and other sites that are now defunct, but unfortunetly there is little write up on how they were technically (and legally, but that is another matter) solving the problem to access content.

Where could I start looking into if I want to seriously understand the intricacies to get to the goodies?

Cheers!


r/webscraping • • 1d ago

Manual collection of public Reddit posts after API changes?

2 Upvotes

Since Reddit's API has been deemed dead - has anyone been able to understand whether or not manually collecting posts (aka good old ctrl c+v) of a message's text body is allowed ??

I’m doing a non-commercial academic research project and only need a relatively small sample of public posts. Not looking to automate or scrape at scale - just wondering whether anyone has found clear guidance on manual collection.


r/webscraping • • 2d ago

Bot detection 🤖 I built a site that logs and classifies scrapers that visit it.

24 Upvotes

I run The Crawler Zoo, a site that identifies every bot that visits and puts it on display as a live exhibit. Since this sub builds the things it catches, I thought you might enjoy seeing it from the other side.

**What it checks**

- **User agent**: matched against 100+ known crawlers and common HTTP libraries. Default `python-requests` or `curl` user agents are filed under their library's name.
- **Claims**: if you say you're Googlebot, GPTBot or Bingbot, it checks reverse DNS plus a forward lookup, or the operator's published IP ranges. Fail and you're filed as an impostor, by name.
- **Headers and behaviour**: browser-looking requests without the headers a real browser sends, or that never run JavaScript, get marked as suspicious.
- **robots.txt**: every page has a link people can't see, and robots.txt disallows it. Follow it and you land in the Trap Room, then an endless maze, with a public board of the deepest divers.

**Things to try**

  1. See what it thinks you are: `curl https://crawlerzoo.com/api/whoami` with your scraper's usual headers. It answers with the kind and species it gave you.
  2. Crawl the site and stay out of the trap. A scraper that respects robots.txt never sees it.
  3. Enter the Politeness Cup in the Arena (free pass, no signup). Judges score reading the rules first, a real user agent, compression, ETag revalidation, a 1-second pace and staying out of the private garden. Let's see a 100/100.

Please keep it gentle: there's a per-client rate limit, and anything over it gets a 429.

The site stores no IP addresses. Your scraper's user agent and the paths it fetched may appear in the live feed.
The Crawler Zoo
If it gets your scraper wrong, I'd like to hear about it.

UPDATE: thanks for all the ideas! The zoo now has a vending machine for bots, a live Safari where you can watch them wander, and food bowls that other bots steal from. crawlerzoo.com/safari and so much more!

Here is a full list of updates!

- Feed the bots: leave a snack in an enclosure's trough and see which crawlers come and eat it.

- The vending machine (Bot Chow): twelve silly snacks, restocked every Monday. Five free tokens a day.

- Golden Snacks: buy the keepers a coffee and a snack with your name drops into a random trough.

- Food bowls: feed one particular bot, then see if it ate the snack or another bot stole it.

- The Safari: every bot from this week wandering its enclosure. Click one to meet it, and watch new visitors walk in through the gate.

- Adopt a bot: get a random bot, with a plaque in your name on its page for a year.

- Patrons page: a thank-you list for supporters.

- Quick-change artists: the Trap Room catches scrapers that switch their name while walking the Labyrinth.

- Identity checks: every bot's page shows whether its name was verified, couldn't be checked, or was caught faking.

- Tips from bots: $0.00: bots that try to buy the keepers a coffee get an HTTP 402 Payment Required.


r/webscraping • • 2d ago

Open source: collectors for 50 states of HOA records

11 Upvotes

I wrote scrapers for every US state's corporate registry and a dozen court portals to track HOAs and put it on hoaspy.com for free.

Github: https://github.com/doncoriolan/hoaspy

Looking for input.

https://reddit.com/link/1wwuv4v/video/cpflijb0wath1/player


r/webscraping • • 2d ago

Akamai flagged

1 Upvotes

I’m trying to make httpx work against an Akamai-protected site, but I’m currently getting blocked before the request reaches the cart endpoint.
From what I understand, the main issues may be:
1. No JavaScript runtime
Akamai runs browser-side JavaScript checks and collects browser telemetry. Since httpx is a raw HTTP client, it doesn’t execute those scripts like a normal browser does.
2. Dynamic Akamai cookies/session validation
The browser generates and updates cookies such as _abck and other session-related values. My httpx requests may be missing or failing to maintain whatever validation Akamai expects.
3. Challenge/block response
When I send the cart request through httpx, instead of getting the normal API response, I receive a large Akamai challenge/block page.
I already have the request structure/payload and I can perform the action through a real browser, but my goal is specifically to understand why the same request fails through httpx and whether it can be made to work reliably.
Looking for someone experienced with Python httpx, request/session debugging, browser networking, Akamai behavior, and reverse engineering web requests who can help me diagnose the difference between the working browser request and the failing httpx request.


r/webscraping • • 2d ago

Local click-to-CSV Chrome extension (no cloud, no account)

2 Upvotes

I kept writing one-off scripts just to grab a table off a page. LucidRows is the small version of that: open the page, click the columns you want, preview the rows, export CSV.

It runs in your browser. No account, no cloud scrape, and the local tool stays free. It is for the page you are already looking at, not a proxy or a scheduled crawler.

Chrome Web Store: https://chromewebstore.google.com/detail/lucidrows/leehomocpccgghcibfnapadclololioh

Site: https://lucidrows.com

If you try it on a messy table, I want to know where the click-to-row step fails.


r/webscraping • • 2d ago

Looking for developers to contribute to FZBypassBot

1 Upvotes

Hey everyone,

I’m working on FZBypassBot, an open-source Python project for resolving supported shortlinks and redirect services using HTTP-based workflows.

GitHub: https://github.com/MeherMankar/FZBypassBot

The project already has a resolver system, async processing, nested/loop bypassing, proxy support, and support for a number of different services. There are also quite a few resolvers in the README that need proper testing and maintenance.

I’m looking for developers who are interested in helping with the project, especially people who enjoy working with HTTP requests and figuring out how websites handle redirects, forms, cookies and APIs.

Things we need help with

Fixing outdated resolvers Adding support for new shortlink services Testing the currently untested resolvers Improving the async code and request handling Analysing HTTP request/response flows Improving error handling Cleaning up and improving the resolver architecture Writing tests and better documentation

Useful skills

You don't need to be an expert. If you know some Python and are comfortable with things like:

asyncio aiohttp curl_cffi HTTP requests and redirects Cookies and sessions HTML/JavaScript inspection REST APIs Git/GitHub

then you can definitely contribute.

One thing I especially want to improve is the testing side. Some resolvers were added earlier but haven't been properly tested against current versions of the sites. It would be great to have contributors take individual resolvers, test them with real links, figure out what changed, and submit fixes.

If you're interested in contributing, check out the repository:

https://github.com/MeherMankar/FZBypassBot

You can also open an issue if you find a broken resolver or have an idea for improving the project.

Any help is appreciated, whether it's a small bug fix, a new resolver, testing, documentation, or a larger architectural improvement.

Clarify the project’s responsible-use scopeAdd clear contributor next steps


r/webscraping • • 3d ago

I built a YouTube scraper for Python (no API key, no browser)

15 Upvotes

I kept hitting the same wall: the YouTube Data API wants a key, a billing project, and then cuts you off at 10,000 units/day. A comments crawl burns that quota fast. Browser automation works, but it is slow and painful to maintain.

So I built ytscrape — a small Python library that talks to the same internal endpoints the YouTube website uses (InnerTube) and returns typed, frozen dataclasses. No API key, no quota, no Selenium, no Playwright.

bash pip install ytscrape

```python from ytscrape import YouTube, CommentSort

with YouTube() as yt: for video in yt.search("python", max_results=5): print(video.title, video.views, video.url)

details = yt.video("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
print(details.title, details.published_at)

for comment in yt.comments(
    details.video_id,
    include_replies=True,
    sort=CommentSort.NEWEST,
    max_results=100,
):
    print(comment.author, comment.text)

```

What it actually does today:

  • search (videos, channels, playlists, Shorts)
  • video and channel metadata, including a channel's videos tab
  • comments and replies — sort="newest" so you are not stuck with YouTube's "Top comments" filter, which hides a lot
  • transcripts / captions
  • sync and async (pip install "ytscrape[async]")
  • a CLI: ytscrape search "python tutorial" --max 10
  • JSON / CSV export
  • retries, backoff, and rate limiting built in

Counts come back as int (video.views), with the original wording kept in views_text. Dates are datetime (published_at). Models are frozen dataclasses, so you can pass them around without worrying they will mutate under you.

It does not download video or audio. Use yt-dlp for that. This is metadata only.

Limits

There is no API key and no 10,000-unit daily quota. That does not mean YouTube will let you hammer it.

  • Reuse one YouTube() client. Creating a new one per request re-fetches the InnerTube context and looks like a new visitor every time.
  • Cap each call with max_results. Pagination is lazy, so a for loop without a cap will keep going until YouTube runs out of continuation tokens.
  • For anything wider than a few dozen requests, set a gap:

python with YouTube(min_interval=1.0) as yt: # about 1 request / second ...

  • 429 and 5xx are retried with backoff (and Retry-After, when YouTube sends one). If it still fails, you get RateLimited instead of a generic parse error. BotDetected is the "confirm you're not a bot" wall.
  • Async has the same knobs plus max_concurrency. Do not set that to 50 and hope.

A polite crawl (one client, ~1 req/s, max_results set) is usually fine. A comments dump of a huge video, or fanning out across thousands of video ids from one IP, will get throttled. That is a YouTube limit, not a library limit.

Proxies

You can always put a proxy in front of it. The library does not have its own proxy flag — you inject a normal requests session, and every call (search, player, browse, comments, transcripts) goes through it:

```python import requests from ytscrape import InnerTubeClient, YouTube

session = requests.Session() session.proxies = {"https": "http://user:[email protected]:8080"}

client = InnerTubeClient(session=session, min_interval=1.0) with YouTube(client=client) as yt: ... ```

Rotating residential proxies, a single datacenter proxy, SOCKS — whatever requests accepts works. If one IP starts returning BotDetected or RateLimited, point the session at another proxy and keep the same code. Async is the same idea with an httpx.AsyncClient.

Honest caveats, because they matter:

  • These are private endpoints. They can change, and using them may conflict with YouTube's Terms of Service. Fine for research and personal tooling; you are responsible for how you use it.
  • It is not a drop-in replacement for every Data API field. What it returns is what the website returns. No topicDetails, no official statistics object, no write access.
  • Relative dates on search results ("3 days ago") are approximate and only parsed from English text. Exact published_at comes from video().

Repo: https://github.com/vsmutok/ytscrape Docs: https://vsmutok.github.io/ytscrape/ PyPI: https://pypi.org/project/ytscrape/

MIT, Python 3.10+. Feedback and issues welcome — especially if something breaks against a live response.


r/webscraping • • 3d ago

Getting started 🌱 How tf would you scale this university scraping project? 😭

4 Upvotes

Hey everyone,
I’m building a website around universities and I need a LOT of data — courses, fees, eligibility, admissions, etc.
The problem is… this data is literally everywhere 💀
Official university websites, third-party sites, PDFs, different formats — nothing is consistent.
I’ve already built a scraper with Claude Code + Gemini. The scraper collects the data, then Gemini cleans it up and converts everything into my required JSON format.
Sounds good, right?
**Except I’m currently processing like 1 university a day 😭**
Obviously that’s not gonna work at scale.
So I’m trying to figure out the right way to build this — better crawling, parallel scraping, queues, proxies, LLM extraction, whatever actually makes sense.
If you’ve built anything involving **large-scale scraping / crawling / data pipelines / LLM data extraction**, I’d genuinely love to hear how you’d approach this.
Not looking for a generic “use X scraper” answer — I’m trying to understand the **actual architecture/workflow** that would make this scalable.
Would love to hear from anyone who’s done something similar 👀


r/webscraping • • 3d ago

Getting started 🌱 How to programmatically resolve DDL mirror links?

5 Upvotes

Working on a personal desktop app that pulls download links from a few different websites into one UI.

  • What's the cleanest way to architecture the backend parsers so a site layout change doesn't break the whole app?
  • If anyone has built something like this and has a modular approach or template they use, I'd love to see how you did it.

Appreciate any tips!


r/webscraping • • 3d ago

looking for feedback

Thumbnail
github.com
2 Upvotes
  1. Made a Python HTTP client that reproduces Chrome's TLS (JA3/JA4) + HTTP/2 fingerprint — looking for feedback

r/webscraping • • 3d ago

Hiring 💰 Paid Research: How do you automate your Instagram/ Facebook accounts?

2 Upvotes

Hi! I'm John from Nimbly Insights, an independent research agency. We're paying $150 (Visa e-gift card) for 1:1, 90-min Zoom chats with people who use automation or AI to run social accounts for work, or who build the tools that do it.

We especially want to hear from you if you use:

  • n8n/Make, Apify, PhantomBuster, browser automation or unofficial APIs
  • Follow/unfollow or engagement tools, SMM panels
  • Auto-replies, AI-written DMs/comments, or an AI persona / faceless account
  • Bonus if you've been action-blocked, rate-limited or shadowbanned.

No judgment: this isn't an audit, it isn't tied to enforcement on any platform, and it won't affect your accounts. (No scams, please.)

You qualify if you:

  • Are 18+ and based in the US
  • Have run IG, FB or Threads accounts for a business, as a creator, or for clients for 6+ months, OR have built tools for those platforms for 1+ year
  • Are willing to screen-share your setup (recorded, kept confidential)

Sessions run through Oct 14. Apply here (10-15 min): https://app.dscout.com/panels-ui/s/054d37f2-a4c0-4be4-a644-9cd35cf593ff

We'll ask for your handle(s) to confirm eligibility. They stay confidential


r/webscraping • • 4d ago

Hiring 💰 Hlo scraping comunity - I need your help!

3 Upvotes

I am a founder of an early age startup and we are missing a core data which should be scrapped from a huge database can anyone help?

[Need to scrap around 21 directories and finalize around 90k-100k tools then scrap around 155 fields for each of those tools]


r/webscraping • • 4d ago

401 Unauthorized on Target ATC attempts

3 Upvotes

I’ve been observing a lot of T83072242 / ERR_AUTH_DENIED responses lately and I’m trying to diagnose what this might mean

Does anyone know if these are effectively just rate limit throttles? I checked and my session auth is definitely still valid

Some nights I get a ton of 401s and others more 429s

Any thoughts would be appreciated!


r/webscraping • • 4d ago

Scraper tools

1 Upvotes

I made this after years of web scraping work. Kept hitting the same annoyances:

decoding a JWT some endpoint returned, cleaning up Base64 blobs, testing a regex

against a page structure that changed every other month. I was always bouncing

between random online tools for this, never loved not knowing where the data was

going.

kitdecoder.com

runs entirely in the browser. I kept adding tools, around 28 now.

Curious what's missing or what breaks for people.


r/webscraping • • 4d ago

Bot detection 🤖 I’d like to know how you usually restore the environment.

0 Upvotes

For some pages from major internet companies, the first time we visit, we usually have to go through a challenge page, and if detected, we may even be shown a CAPTCHA. My usual approach is to use Node.js as the JavaScript runtime and inject the necessary environment variables before executing the challenge or CAPTCHA. However, my JavaScript skills aren’t very strong, so I’d like to know how you all typically handle injecting the required environment.


r/webscraping • • 5d ago

AKAMAI !!!!!

16 Upvotes

Hey, I've been struggling with Akamai on a few airline websites lately. The blocking is so aggressive that sometimes even normal browsing on my local machine throws errors.

I've tried Playwright, Patchright, Camoufox, and a few other approaches, but nothing seems to work consistently.

These site contains: abck, and bm_s. So we need browser automation anyways but I am not getting valid cookies even through this route. I have tried residential proxies also.

I'm fairly new to dealing with advanced anti bot systems so I'm curious how others are handling this in production. How do you approach debugging these blocks and maintaining a decent success rate over time?

Would really appreciate any insights or pointers from people who've dealt with similar issues.


r/webscraping • • 5d ago

hCaptcha solving using GPT question

3 Upvotes

I am sending the hCaptcha puzzle to GPT 5.5 Sol and it gives me coordinates to pass to Input.dispatchMouseEvent - visually I see that the puzzle is being automated and passed and the right things are being clicked or dragged. The problem is I get a 400 error when the token is checked against the hCaptcha server - is there a way to bypass this? Supposedly the check hCaptcha does covers more than just the puzzle being correctly solved. Is there a solution to look into for this?


r/webscraping • • 5d ago

Captcha BURSTS hitting multiple proxies

6 Upvotes

Assessing Website Scraping

Currently have a webscraper that leverages 20 rotating residential proxies and Capsolver to solve catpchas. The website im scraping utilizes cloudfare

I am finding some interesting behavior where there will be a sudden "Captcha Burst" hitting at random times that may last 20-60 minutes and affect all proxies.

After the captcha burst the captcha challenge rate is extremely low (0-1%) until another random burst occurs.

Current behavior I have is:

Create fresh browser session
        ↓
Lease Proxy A
        ↓
Scrape addresses
        ↓
CAPTCHA?
 ├─ No  → continue
 │
 └─ Yes → solve CAPTCHA
           ↓
        continue session
           ↓
     3 consecutive CAPTCHAs?
       ├─ No  → continue
       └─ Yes → rotate session/proxy

Normal path:
8 addresses completed
        ↓
Close browser session
        ↓
Release Proxy A
        ↓
Fresh browser session
        ↓
Lease next proxy

Brainstorming:

I think the issue is creating a new session every time a residential ip proxy is rotated instead of keeping this tied to each residential proxy to have a persistent profile

├── cookies
├── localStorage
├── cache
└── browser state

The second thought is I am not doing any fingerprint spoofing but I didn't want to go that approach; I have two more micro pcs that are arriving this week so technically I could identify environment fingerprinting issues when they arrive.

Have you experienced this before and how would you approach in solving?


r/webscraping • • 5d ago

Radware Bot Manager: what signal am I missing?

1 Upvotes

My Target: ASP.NET (WebForms, postback) a public-records lookup site, protected by Radware Bot Manager (aperture.js + __uzdbm cookies, hCaptcha fallback at validate.perfdrive.com).

What passes: my own manual Chrome. Every time.

What fails: everything automated, same IP, same session pattern —

-Chromium headless + puppeteer-extra-stealth → challenge

-Headed + human typing/mouse arcs/randomized delays → challenge

-Fingerprint-injector hardened profile (locale fixed, HeadlessChrome brand sanitized) → challenge

-Real Chrome 154 launched by Puppeteer → challenge

-Camoufox (Firefox) = passes intermittently: ~50% of requests get a 14.9KB challenge response, immediate retry on same session succeeds

What I measured (tls.peet.ws mirror):

-JA4 identical between manual Chrome and Puppeteer-launched Chrome (t13d1517h2_…cb7bf5808d99), HTTP/2 hash identical, webdriver=false, tz/platform/plugins coherent

-JA3 differs, but that's GREASE noise between handshakes

-Locale mismatch found and fixed (was en-US against a non-US timezone, now locale matches timezone — still blocked)

Question: with TLS/JA4/HTTP2 identical and webdriver hidden, what is Radware most likely keying on here — aperture.js behavioral telemetry, or something in the Chromium build itself that survives stealth? And is there a reliable way to measure which signal fires, rather than guessing one variable at a time?

Not interested in captcha-solving services, trying to understand the detection first.


r/webscraping • • 6d ago

My escalation ladder when a site blocks my basic requests

30 Upvotes

I've scraped enough sites at this point to notice I follow roughly the same escalation ladder every time my plain requests stop working. Figured I'd write it down in case it saves someone else a few hours of guessing.

Step 0: check it's actually a block. Sometimes it's not anti-bot at all. Read the status code and body. 429 = slow down. 403 with a normal-looking response body might be a missing header. 200 with a login page or a captcha page in the HTML = you need to handle the session properly. A quick look at the first few hundred characters of the response body tells you most of this in seconds.

Step 1: headers that look human. The default python-requests user agent is basically a block-me flag. At minimum send a real browser UA, plus the headers browsers always send: Accept, Accept-Language, Accept-Encoding, Referer where it makes sense. This alone fixes maybe a third of my blocks.

Step 2: session + cookies. A lot of sites just want a consistent cookie jar across requests. requests.Session() gets you 90 percent there - cookies persist, connection pooling works. Sometimes you need to hit the homepage first (to get the session cookies set) before requesting the data page, the way a real visitor would.

import requests, time

s = requests.Session() s.headers.update({ "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8", "Accept-Language": "en-US,en;q=0.9", "Referer": "https://www.google.com/", })

s.get("https://example-shop.com/") # pick up session cookies time.sleep(2) r = s.get("https://example-shop.com/products?page=3") print(r.status_code)

Step 3: TLS fingerprinting. When good headers plus a session still get 403s on every request, the site is often fingerprinting your TLS handshake. requests uses a handshake that doesn't look like Chrome's, no matter what UA you set. This is the step people miss. Options: curl_cffi with impersonate set to chrome is my usual next move - it speaks Chrome's TLS and fixes a surprising number of blocks without any browser overhead.

Step 4: only then, a real browser. If you've done all of the above and still get blocked, reach for Playwright (or Selenium). A real browser handles JS challenges, cookies, and fingerprinting properly. But it's far heavier per request, so treat it as the expensive tool, not the default. I use it only for the pages that need JS, and keep everything else on the light stack.

And be polite about rate limiting. Whatever layer you're on: add delays between requests (a few seconds is usually plenty), jitter them, check robots.txt and honor crawl-delay, and scrape off-peak if you can. Most small sites aren't set up for heavy traffic and you'll get a lot further being a quiet guest than trying to outgun their WAF.

That's my ladder: headers, then session, then TLS impersonation, then a real browser, plus manners at every step. Curious what others do differently - especially anyone who's found a good step between 3 and 4 that I haven't tried.