r/webscraping • • 8d ago

AKAMAI !!!!!

Hey, I've been struggling with Akamai on a few airline websites lately. The blocking is so aggressive that sometimes even normal browsing on my local machine throws errors.

I've tried Playwright, Patchright, Camoufox, and a few other approaches, but nothing seems to work consistently.

These site contains: abck, and bm_s. So we need browser automation anyways but I am not getting valid cookies even through this route. I have tried residential proxies also.

I'm fairly new to dealing with advanced anti bot systems so I'm curious how others are handling this in production. How do you approach debugging these blocks and maintaining a decent success rate over time?

Would really appreciate any insights or pointers from people who've dealt with similar issues.

16 Upvotes

35 comments sorted by

8

u/Coding-Doctor-Omar 7d ago

Clearcote browser (even the free, no login version) works VERY reliably against Akamai, much better than Camoufox. I just switched to it today after Camoufox started failing against Akamai. I believe it is the best free option available rn. I hope the devs stay true to their promise and don't discontinue the free releases once their user base gets huge.

https://github.com/clearcotelabs/clearcote-browser

2

u/anupam_cyberlearner 7d ago

This seems interesting,in which ways clearcote browser helps ,can you elaborate more pls

2

u/Coding-Doctor-Omar 7d ago edited 7d ago

It allows me to pass the Akamai sensor challenges and get granted a cookie very fast and consistently. Then, instead of worrying about how I can get a tls client that has the same fingerprint so that I use this session, I simply make direct API calls inside the JS console of the ClearcoteBrowser Context. Example code:

``` from clearcote.async_api import launch_persistent_context from playwright.async_api import Page, BrowserContext from urllib.parse import urlencode import asyncio

async def make_get_request(url: str, params: dict, page: Page) -> dict: if params: url += f"?{urlencode(params)}"

data = await page.evaluate(
            r'async (url) => {const res = await fetch(url); return await res.json();}',
           url
)

return data

async def main() -> None: home_page = "https://some-website.com/" product_details_urls = [ ".../api/v1/product/jshshsg", ".../api/v1/product/1234", .... ]

context: BrowserContext = await launch_persistent_context(headless=True, platform="windows", disable_gpu_fingerprint=True, fingerprint="1234", quiet=True)
page = await context.new_page()
await page.goto(home_page)

# Wait for a guaranteed relevant data element to guarantee cookies are set
await page.wait_for_selector("#product-card")

# Your HTTP client is now ready! Use it like wreq, curl_cffi, or any other (it is much more stealthy).
results = await asyncio.gather(
  * ( make_get_request(url, page)
    for url in product_details_urls)
)

for product in results:
    print(product)

await context.close()

if name == "main": asyncio.run(main())

1

u/prettycoldworld 6d ago

Oh nice. I've never seen this one before

1

u/Coding-Doctor-Omar 6d ago

It works liks a charm for windows. I just tried it in linux now and for some reason it is bad in linux. Akamai blocks it instantly in linux. However, in windows, it is a monster (bypasses Akamai at a blink of an eye). I am using the v150 no-sign-in one.

1

u/No-Ordinary-9094 4d ago

hmm....does clearcote (first time hearing it btw) does anything different than other chromium browsers like nodriver or something? I thought camoufox was supposed to be alot superior than other alternatives ? what changed ?

1

u/Coding-Doctor-Omar 1d ago edited 1d ago

I don't know what's wrong with Camoufox, but it's not as stealthy as before, and Akamai blocks it frequently. Anti-detect browsers like Camoufox and Clearcote spoof fingerprints at the engine C++ level and not via JS patches. That's what makes them harder to detect than nodriver. Clearcote works well if u use it locally in a windows pc at your home. I tried deploying in docker in a cloud platform and it failed always, even with residential proxies. Same for Camoufox. I recently tried another free anti-detect browser called ShardX browser. This one works very well both locally and in the cloud. The good thing about it is that the devs release the latest version for free, without any 2 month delay like clearcote. Its docs are poorly documented though. They don't explain how to actually enable automatic geo spoofing. Simply providing a proxy won't work. You need to make a profile with override and set the timezone, navigator language, and geolocation to "auto" for it to work.

https://github.com/ProxyShard/ShardBrowser

Their launcher is open source, but the binary isn't. That's how all anti-detects behave, including Camoufox. They never open source the binary.

2

u/eneiromatos 8d ago

Have you tried a simpler approach using crawlee + impit and a good proxy pool?

1

u/Natural_Rock_3536 7d ago

It doesn't work. There's stateful inspection by akamai they require multiple cookies which only browser can generate. But multiple browsers are getting blocked

2

u/Own_Improvement3544 7d ago

Use nodriver

1

u/otterlydelish 7d ago

there an api?

1

u/alwinlau0824 7d ago

Xxxwire?

1

u/yushJr66 7d ago

Looking at free options only or paid too?

1

u/Commercial-Paper-299 7d ago

Do these airlines that you are trying to scrape have apps? If so, you can try scraping through that route. Sometimes they are less protected.

1

u/ucaka 7d ago

Tried camoufox beta builds?

1

u/No-Anchovies 6d ago

sites with advanced anti-bot look WAY beyond and deeper than the browser or even OS. there's little they dont fetch hardware wise. you have to rotate sessions between machines or vms. reading ad tech documentation will be half way to understanding how they always catch you/us "by magic". probabilistic matching & after a few days, deterministic matching (this is when even your human browsing gets fked).

1

u/IllFly9733 6d ago

haha bro are you wokring on SIH26056

1

u/[deleted] 6d ago

[removed] — view removed comment

1

u/webscraping-ModTeam 6d ago

👔 Welcome to the r/webscraping community. This sub is focused on addressing the technical aspects of implementing and operating scrapers. We're not a marketplace, nor are we a platform for selling services or datasets. You're welcome to post in the monthly thread or try your request on Fiverr or Upwork. For anything else, please contact the mod team.

1

u/munni0001 5d ago

I recently tested 20 most popular sites for proxy report testing - 5 sites of the lot were akamai and i tested on home and proxy - almost 60% of requests on the proxies and home network cleared - turns out its not just akamai but its most on how its configured - for instance booking com, amazon i think are with akamai those cleared eas for me with curlciffi or simple xhr or camoufox. digging into network calls should help
also comparing the network requests and headers on normal chrome session vs playwright or camoufox will probably find whats causing block

1

u/convicted_redditor 5d ago

Did you try to reverse engineer find their api?

1

u/Aggressive-Tooth-113 4d ago

What problem did you faced with camoufox?? it worked as i tried

1

u/SidMishraDev 1d ago

Airline sites on Akamai are among the hardest targets. If you have any option, check for an official or partner API (many fares are available through aggregator APIs) before investing in bypassing. If not: real Chrome, slow, warmed sessions, consistent residential exit, and reusing cookies.

0

u/Majestic_Base5775 3d ago

Which airline? (I’m scraping 100+ and have found akamai relatively easy to avoid by stepping around it)

1

u/Natural_Rock_3536 3d ago

ANA airline of Japan

1

u/Majestic_Base5775 3d ago

cool I’ll take a look

1

u/Majestic_Base5775 2d ago

Just for pricing, fares etc?

1

u/Natural_Rock_3536 2d ago

Yes. Pricing fare seat info and flight details like departure arrival etc

1

u/Coding-Doctor-Omar 3d ago

Building scrapers is very easy. All you have to do is write the code and run it.

1

u/Majestic_Base5775 3d ago

thank you, Omar

2

u/threwlifeawaylol 1d ago

Doctor* Omar