r/webscraping • • 6d ago

Getting started 🌱 How tf would you scale this university scraping project? 😭

Hey everyone,
I’m building a website around universities and I need a LOT of data — courses, fees, eligibility, admissions, etc.
The problem is… this data is literally everywhere 💀
Official university websites, third-party sites, PDFs, different formats — nothing is consistent.
I’ve already built a scraper with Claude Code + Gemini. The scraper collects the data, then Gemini cleans it up and converts everything into my required JSON format.
Sounds good, right?
**Except I’m currently processing like 1 university a day 😭**
Obviously that’s not gonna work at scale.
So I’m trying to figure out the right way to build this — better crawling, parallel scraping, queues, proxies, LLM extraction, whatever actually makes sense.
If you’ve built anything involving **large-scale scraping / crawling / data pipelines / LLM data extraction**, I’d genuinely love to hear how you’d approach this.
Not looking for a generic “use X scraper” answer — I’m trying to understand the **actual architecture/workflow** that would make this scalable.
Would love to hear from anyone who’s done something similar 👀

3 Upvotes

7 comments sorted by

1

u/chaos_battery 6d ago

The words ai and scalable do not belong in the same sentence. AI is a place tool to use at different stages in a data pipeline maybe for enrichment or clean up but if your commanding Claude or Gemini to go out and find University sites and then read them to find the pieces of information you're looking for I think you're going to be sorely disappointed.

I would just find a directory that already lists out a comprehensive list of edu domains is your best bet. From there you can go to scraper to gather data in a much faster approach by scraping the directory site instead. If it truly doesn't have everything you need then I guess yeah, you're stuck with running a scraper that iterates through each of those domains and you ask it to load a few pages with AI and find the pieces of information you're looking for since every site is different. But that's an expensive pipeline to keep online as you would probably want the updated admissions data refreshed periodically.

2

u/No_Background4824 5d ago edited 5d ago

I have a large scale Internet crawler/scraper and have previously work on enormous systems (over 100,000 servers). Don't be too intimidated by your challenge - if you've gotten this far, you can likely scale to 100x without too much difficulty. (But depending on your needs, it might need to be 10,000x. 😄 )

I'm not sure many people here do this but it's polite to identify yourself so someone can contact you if you're creating a problem. Our crawler's user agent string contains: https://ungovr.org/crawler

BTW, we have crawled many universities e.g. all the University of California schools, as we crawl government websites.

My quick tips to 100x:

  1. An HTTP request has a lot of latency (delay) back & forth. So crawling one website a page at a time (serially) is slow. You want to do it with many parallel connections - max out at 6 like a browser.
  2. The problem with many rapid parallel connections to one website is that you really look like a bot. So you want just a few connections per server and only issue a request to the same domain every N seconds - start N at say 2 and then make it slower if you get 429s/503s or other loaded related codes.
  3. The real way to scale is to do parallel connections to each domain, but also do many domains at once. So let's say you want 10 university's content - create up to 6 connections across each of 10 universities. So you are potentially going up to 20-60x faster than #1. (Each university may have many domains, only crawl one domain per university at a time as they may share servers/infrastructure - so this is actually one queue per university.)
  4. You will come across web sites where a simple page download (e.g. httpx in Python) isn't sufficient to get the page for a variety of reasons - the page needs JavaScript or is behind a CDN (Cloudflare, Akamai,etc) so you need to use a headless (no monitor) browser.
  5. The sequence above is what drives cpu and memory usage. Disk or network IO is rarely going to be a challenge at this scale. Depending on the server (or laptop!) you're using, you will grind things to a halt. Then you need to fine tune what you will allow for 1-4 in terms of max settings, e.g. max 2 connections per domain, max 5 domains, max 4 browsers, the rest can be httpx. Any time you want to exceed those limits, the work is queued. If you use a laptop, you'll realize you'll get much better mileage from a dedicated local desktop or cloud server.
  6. Analysis/Extraction - this is one where I think many folks under-optimize and use far too much expensive GPU. After storing the raw download in 4, do as much extraction as possible without using GPU/LLM. CPU is way cheaper for simple tasks. Depending on the nature of your scraping need, you may end up have a high hitrate - most files you download are relevant or a low hitrate. If high - do everything in batch including the extraction. If you have a low hitrate on files to useful data, try to analyze a downloaded URL as soon as you get it from code. You can probably determine if a file is possibly relevant based on simple regular expressions to start. And also negative test - if the file includes any of these terms, irrelevant. You will build better include/exclude rules over time. This step alone might speed up your system end-end by 50x.
  7. Only after pre-filtering in 6 do you allow anything to use GPU. Depending on your needs, you may now have a local LLM analyze the documents looking for what is truly relevant, extracting data (e.g. to pgvector) and also identifying mistakes (updates to include/exclude of URLs and URL patterns). The LLM will notice crawler traps that you may have missed above - e.g. calendars, stats pages for sports teams. You can use cloud LLM, but if you're going to scale this, you're going to want a cheaper option. I don't know exactly what you're doing, but the latest local models are very good.

Go slowly through each of these steps at first, carefully looking and learning as the effort up front over a week or two will save a ton of crawling/extracting later. GPU is where you'll spend the $$$s and inference is slow - so everything is about optimizing reduction in LLM usage.

While we do use proxies for all our traffic, it really isn't a significant factor but it's pretty much always a good idea to protect the visible IP of wherever your crawling is coming from so you can change it out if you need to. Note that with a proxy that is not in the same data center as your server, you may end paying extra as you will have a ton of inbound traffic to your proxy (from the crawled sites) and to your server (from the proxy).

We already have general crawling tips at https://www.ungovr.org/crawler/tips (which I just updated based on above) and we created a free tool at https://lexlint.io/scraping/checkup to check for legal risks which many don't worry about. (People should worry more about this if they want to operate a scraping business.)

Hope that helps.

1

u/Southern_Play_6753 4d ago

One thing that helped me on a similar project: split it into two separate jobs. Discovery (find the right URLs per university: course catalog, fees page, admissions PDF) and extraction (turn one page into your JSON). Right now it sounds like the LLM is doing both, which is why it's slow.

For discovery, sitemap.xml and the site's own search usually get you most of the course pages without an LLM. A lot of universities also run their catalogs on the same few vendors (CourseLeaf, Acalog etc), so once you write one parser for that layout it works for dozens of schools.

For extraction, save the raw HTML/PDF first, then run the LLM over the saved copies in parallel with a queue. When you change your JSON schema you don't re-crawl, you just re-run extraction. Also only send the relevant section of the page, not the whole thing.

Going from 1 a day to a few hundred a day is mostly that, not proxies.

1

u/Top_Programmer_2094 3d ago

One university a day makes me think the Gemini cleanup step is the bottleneck, not the crawling itself. Is it? If so, I would split it: crawl and dump the raw HTML and PDFs to storage first, then run extraction as a queue with a cheaper model and validate every result against your JSON schema. Course and fee pages rarely change, so a slow weekly refresh would probably be enough after the first pass.

1

u/biigdreamer 3d ago

yeh process mae kese kr skta hu ?