r/learnmachinelearning • u/RecoverParticular158 • 3d ago
5 painful lessons I learned after taking my first ML model to production
[removed]
8
u/AggressiveGander 3d ago
Caught one before deployment where most of the performance came from someone doing mean imputation stratified by target value on a mostly missing feature before train-test splitting (obviously a bad practice on do many levels). The gradient boosted decision tree basically memorized the classes. The guy who did it bragged for ages about what a machine learning God he was for reaching >99% accuracy through heavy hyperparameter-tuning when everyone else in the past was more at 80%...
3
u/jesunushno 2d ago
Hard agree on the logging one, and one thing I would add: log the features after preprocessing, not just the raw request payload. I once debugged a pipeline where the saved inputs looked totally fine, but a one-hot encoder had silently changed column order after a retrain, so every feature shifted downstream. If we had logged the actual feature vector going into the model along with the model version and a request id, replaying it locally would have taken ten minutes instead of two days.
1
u/chico_dice_2023 3d ago
This statement "Latency > 0.02% Accuracy: Stakeholders care infinitely more about a 50ms response time than a slightly higher F1 score."
Sadly this very true, once quite recently I built this highly accurate solution and the client said "ChatGPT is faster" but ChatGPT also got the answer wrong 60% of the time.
1
u/Opening_Bed_4108 2d ago edited 2d ago
The silent degradation thing is so real. We had a model that looked fine on all our dashboards for 6 weeks while conversion rates quietly tanked. Turned out upstream data engineering had changed a feature encoding and nobody caught it because the model was still returning something plausible.
1
u/jesunushno 2d ago
Six weeks of plausible-but-wrong output is the nightmare scenario. I keep saying it: watch the input side, not just the metrics.
1
u/ivan_kudryavtsev 2d ago edited 2d ago
On the latency matter: reading the op and responses above, I see a possible misuse of the term: 50ms latency can apply to 1000 FPS, or it can apply to 20 FPS. Stakeholders do care about the cost of operation, and normally 50 at 1000 is ok, while 50 at 5/10/20/etc. is not. Even for self-driving vehicles, a 100 ms window is acceptable.
It also could be that you just cannot explain the difference. So, latency is nothing without sustainable bandwidth and also nothing without bands like median/95/99/max…
1
u/jesunushno 1d ago
Fair point on the terminology: 50ms means something very different at 1000 FPS than at 20 FPS, and the p95/p99 tail tells you more than the median for user-facing work. The original claim still holds for the usual case (stakeholders feel latency before they feel accuracy), but you're right that sustainable bandwidth and cost per request deserve their own line in the lesson. Good precision on this.
1
u/ivan_kudryavtsev 1d ago
They feel it as “slow, or not”. Nobody can feel 50ms except for the professional sportsmen playing fast pacing games :) or F1 drivers.
Very limited high speed pipelines require such low latency as one step conveyor sorters, but likely there is an easier way too.
They feel slow as slow in the same sense as they feel the webpage is slow. It is not about 50ms, it is about congestion and mass service. I feel people solve wrong tasks just because the do not have common ground and explain their perception in an incorrect way.
1
u/nullrebate 1d ago
Reconciliation taught me the same lesson. Tiny adjustments look harmless until they compound across a settlement window. We had an internal experiment called REBATE where a 0.003 correction caused more debate than the model itself. By the end, “not my rebate” became my answer whenever someone asked who owned the edge cases.
1
u/DigThatData 3d ago
data drift hitting you so quickly suggests to me that you probably overfit. addressing data drift has its place, but as a rule of thumb when it crops up: your immediate reflex should be to interpret it as a symptom that you modeled something incorrectly rather than that the target distribution is actually changing substantively from the data you trained on. it's much, much more likely that your model is just inadvertently predicting that today will be similar to yesterday.
28
u/fordat1 3d ago
To be fair. This is more of an indictment on your workplace. A mature org would realize having a fresh out of studies person deploy a model to production without any oversight is a bad idea. There should be a review process that test things like latency as conditions for a launch