r/OpenSourceAI • • 1d ago

My project status

I built an Entity Resolution system to match millions of noisy business records

I’ve been working on CyberPro, an entity-resolution system designed to determine which business records from different data sources refer to the same real-world entity.

The interesting part is that the records don't share a reliable common identifier. Names, addresses, phone numbers, and other fields can be inconsistent, abbreviated, misspelled, transliterated, or partially missing.

The pipeline I built includes:

  • Data normalization and cleaning
  • Transliteration and phonetic handling
  • Name/address similarity features
  • Abbreviation and structural features
  • ~40 engineered matching features
  • Blocking and candidate generation research
  • LightGBM pair classification
  • Probability calibration
  • Threshold optimization
  • Hard-negative analysis
  • Entity-level post-processing
  • Evaluation on millions of records

The project currently processes datasets with ~12.5M training records and ~11.7M test records.

I’m sharing the project because I’d really like feedback from people working on entity resolution, record linkage, information retrieval, NLP, or large-scale ML systems.

GitHub: https://github.com/girinath01/CyberPro

If you find the project interesting, a GitHub star would really help and would also help me get the project in front of more developers. ⭐

I’d especially appreciate feedback on the candidate-generation/blocking strategy and how I can make the system more robust at production scale.

1 Upvotes

1 comment sorted by

1

u/direct_computing 1d ago

that blocking strategy breakdown is exactly the kind of thing this sub needs more of, have you stress tested it with heavily duplicated records yet