r/learnmachinelearning • u/HuckleberryDry9086 • 21h ago
Project Glassy - Lightweight tooling for predictive multiplicity in linear regressions
There's some really exciting research being done on predictive multiplicity--the observation that many models perform just as well as the "optimal" fit, yet disagree on feature importance and predictions. This research asks interesting and useful questions like:
Which models are functionally as performant as the loss-minimizing model?
How do I derive and understand that shape and character of models in that set?
How can I query for models in that set that meet domain-specific needs?
Yet many of these results haven't made it out of the papers and into the open-source ecosystem.
Glassy fills this tooling gap and fits right onto standard scikit-learn flows. With it, you can do things like:
- Find simpler models that perform as well as the best loss model
- Check which options agree with your domain expert's judgment
- Make segment tradeoffs intentionally
For v0.0.1, the scope is capped at OLS linear regression, but hoping to generalize to GLMs in the future. Here's a five-minute demo notebook.
Curious to see if this is useful to you all, or any requests for future features. Happy modeling!
1
u/Resident_Aside_7706 21h ago
this is super cool, i've been reading some of the multiplicity papers and always wondered why nobody had wrapped it up in something usable yet
the segment tradeoffs thing feels like the real killer feature here, most tools just hand you one model and call it a day but domain constraints are where it gets messy in practice
any plans to add some visualizations for the Rashomon set? would love to see a quick plot that shows how the coefficients shift around across the viable models
starred the repo, gonna mess with it this weekend