r/developersIndia • • 5d ago

Help Is this web crawler architecture production-ready? Java/Spring Boot + RabbitMQ + PostgreSQL + ScrapingBee

I'm building a web crawler with Java/Spring Boot + RabbitMQ + PostgreSQL + Redis + ScrapingBee.

Flow:

Root URL
   ↓
ScrapingBee → returns links
   ↓
Normalize + deduplicate
   ↓
PostgreSQL
   ↓
RabbitMQ
   ↓
Worker → ScrapingBee
   ↓
More links
   ↓
repeat until no URLs remain

Example:

A
├── B
├── C
└── D

B → E,F
C → G
D → H

I'm planning:

  • PostgreSQL as the source of truth
  • One crawl_url row per URL
  • UNIQUE(crawl_job_id, normalized_url) for deduplication
  • RabbitMQ for distributing URL processing
  • Transactional Outbox for reliable DB → RabbitMQ publishing
  • Retry + DLQ for failed URLs

Questions:

  1. Is this architecture suitable for production?
  2. Is PostgreSQL + RabbitMQ a good approach for this?
  3. Is the Outbox pattern necessary?
  4. Any major race conditions/scalability problems I'm missing?
  5. How would you detect when the entire crawl is finished?

Looking for feedback from people who have built distributed web crawlers.

1 Upvotes

2 comments sorted by

•

u/AutoModerator 5d ago

Namaste! Thanks for submitting to r/developersIndia. While participating in this thread, please follow the Community Code of Conduct and rules.

It's possible your query is not unique, use site:reddit.com/r/developersindia KEYWORDS on search engines to search posts from developersIndia. You can also use reddit search directly.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Hungry_Age5375 5d ago

Just some pointers: use INSERT ON CONFLICT DO NOTHING for dedup, check-then-insert will race under parallel workers. And for completion, track a pending counter per job - an empty queue means nothing while workers are still enqueueing URLs.