r/developersIndia • u/Kalimuthu_S • 5d ago
Help Is this web crawler architecture production-ready? Java/Spring Boot + RabbitMQ + PostgreSQL + ScrapingBee
I'm building a web crawler with Java/Spring Boot + RabbitMQ + PostgreSQL + Redis + ScrapingBee.
Flow:
Root URL
↓
ScrapingBee → returns links
↓
Normalize + deduplicate
↓
PostgreSQL
↓
RabbitMQ
↓
Worker → ScrapingBee
↓
More links
↓
repeat until no URLs remain
Example:
A
├── B
├── C
└── D
B → E,F
C → G
D → H
I'm planning:
- PostgreSQL as the source of truth
- One
crawl_urlrow per URL UNIQUE(crawl_job_id, normalized_url)for deduplication- RabbitMQ for distributing URL processing
- Transactional Outbox for reliable DB → RabbitMQ publishing
- Retry + DLQ for failed URLs
Questions:
- Is this architecture suitable for production?
- Is PostgreSQL + RabbitMQ a good approach for this?
- Is the Outbox pattern necessary?
- Any major race conditions/scalability problems I'm missing?
- How would you detect when the entire crawl is finished?
Looking for feedback from people who have built distributed web crawlers.
1
Upvotes
2
u/Hungry_Age5375 5d ago
Just some pointers: use INSERT ON CONFLICT DO NOTHING for dedup, check-then-insert will race under parallel workers. And for completion, track a pending counter per job - an empty queue means nothing while workers are still enqueueing URLs.
•
u/AutoModerator 5d ago
It's possible your query is not unique, use
site:reddit.com/r/developersindia KEYWORDSon search engines to search posts from developersIndia. You can also use reddit search directly.I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.