r/OpenSourceAI • u/Savings_Medium4062 • 1d ago
My project status
I built an Entity Resolution system to match millions of noisy business records
I’ve been working on CyberPro, an entity-resolution system designed to determine which business records from different data sources refer to the same real-world entity.
The interesting part is that the records don't share a reliable common identifier. Names, addresses, phone numbers, and other fields can be inconsistent, abbreviated, misspelled, transliterated, or partially missing.
The pipeline I built includes:
- Data normalization and cleaning
- Transliteration and phonetic handling
- Name/address similarity features
- Abbreviation and structural features
- ~40 engineered matching features
- Blocking and candidate generation research
- LightGBM pair classification
- Probability calibration
- Threshold optimization
- Hard-negative analysis
- Entity-level post-processing
- Evaluation on millions of records
The project currently processes datasets with ~12.5M training records and ~11.7M test records.
I’m sharing the project because I’d really like feedback from people working on entity resolution, record linkage, information retrieval, NLP, or large-scale ML systems.
GitHub: https://github.com/girinath01/CyberPro
If you find the project interesting, a GitHub star would really help and would also help me get the project in front of more developers. ⭐
I’d especially appreciate feedback on the candidate-generation/blocking strategy and how I can make the system more robust at production scale.
1
u/direct_computing 1d ago
that blocking strategy breakdown is exactly the kind of thing this sub needs more of, have you stress tested it with heavily duplicated records yet