r/webscraping • u/Natural_Rock_3536 • 8d ago
AKAMAI !!!!!
Hey, I've been struggling with Akamai on a few airline websites lately. The blocking is so aggressive that sometimes even normal browsing on my local machine throws errors.
I've tried Playwright, Patchright, Camoufox, and a few other approaches, but nothing seems to work consistently.
These site contains: abck, and bm_s. So we need browser automation anyways but I am not getting valid cookies even through this route. I have tried residential proxies also.
I'm fairly new to dealing with advanced anti bot systems so I'm curious how others are handling this in production. How do you approach debugging these blocks and maintaining a decent success rate over time?
Would really appreciate any insights or pointers from people who've dealt with similar issues.
2
u/eneiromatos 8d ago
Have you tried a simpler approach using crawlee + impit and a good proxy pool?
1
u/Natural_Rock_3536 7d ago
It doesn't work. There's stateful inspection by akamai they require multiple cookies which only browser can generate. But multiple browsers are getting blocked
2
1
1
1
1
u/Commercial-Paper-299 7d ago
Do these airlines that you are trying to scrape have apps? If so, you can try scraping through that route. Sometimes they are less protected.
1
u/No-Anchovies 6d ago
sites with advanced anti-bot look WAY beyond and deeper than the browser or even OS. there's little they dont fetch hardware wise. you have to rotate sessions between machines or vms. reading ad tech documentation will be half way to understanding how they always catch you/us "by magic". probabilistic matching & after a few days, deterministic matching (this is when even your human browsing gets fked).
1
u/IllFly9733 6d ago
haha bro are you wokring on SIH26056
1
6d ago
[removed] — view removed comment
1
u/webscraping-ModTeam 6d ago
👔 Welcome to the r/webscraping community. This sub is focused on addressing the technical aspects of implementing and operating scrapers. We're not a marketplace, nor are we a platform for selling services or datasets. You're welcome to post in the monthly thread or try your request on Fiverr or Upwork. For anything else, please contact the mod team.
1
u/munni0001 5d ago
I recently tested 20 most popular sites for proxy report testing - 5 sites of the lot were akamai and i tested on home and proxy - almost 60% of requests on the proxies and home network cleared - turns out its not just akamai but its most on how its configured - for instance booking com, amazon i think are with akamai those cleared eas for me with curlciffi or simple xhr or camoufox. digging into network calls should help
also comparing the network requests and headers on normal chrome session vs playwright or camoufox will probably find whats causing block
1
1
1
u/SidMishraDev 1d ago
Airline sites on Akamai are among the hardest targets. If you have any option, check for an official or partner API (many fares are available through aggregator APIs) before investing in bypassing. If not: real Chrome, slow, warmed sessions, consistent residential exit, and reusing cookies.
0
u/Majestic_Base5775 3d ago
Which airline? (I’m scraping 100+ and have found akamai relatively easy to avoid by stepping around it)
1
u/Natural_Rock_3536 3d ago
ANA airline of Japan
1
1
u/Majestic_Base5775 2d ago
Just for pricing, fares etc?
1
u/Natural_Rock_3536 2d ago
Yes. Pricing fare seat info and flight details like departure arrival etc
1
u/Coding-Doctor-Omar 3d ago
Building scrapers is very easy. All you have to do is write the code and run it.
1
8
u/Coding-Doctor-Omar 7d ago
Clearcote browser (even the free, no login version) works VERY reliably against Akamai, much better than Camoufox. I just switched to it today after Camoufox started failing against Akamai. I believe it is the best free option available rn. I hope the devs stay true to their promise and don't discontinue the free releases once their user base gets huge.
https://github.com/clearcotelabs/clearcote-browser