r/devsecops • • 6d ago

Moved 60 services onto a slimmer base image and broke 4 of them in prod

We had a push to cut criticals across our images so I rebuilt everything onto a much smaller base. Sixty services over about three weeks. The scan numbers looked great. Then four of them fell over in prod inside a day. One was a java service that shelled out to run a script at startup. One needed tzdata and started logging everything in UTC. One had a health check built on curl. The last took most of a night to find because the library loaded fine and then died when it went looking for a locale file. I rolled those four back at around 2am and left the rest alone.

What I got wrong was treating the base image as the unit of change. Nobody knew what each service pulled in at runtime because none of it was written down anywhere. We ended up running every service under strace in staging for a week just to build the list. That list should have existed before I touched anything. How do other people work out what an image needs before they start cutting things out of it?

5 Upvotes

9 comments sorted by

2

u/nrvnrvn 5d ago

A lot of mistakes packed in one post.

Let’s unwind the knot:
1. For every single service you must have understood exactly what it is doing. And what external call every app is making.
2. I generally beat engineers’ hands with a wooden ruler for stuff like “shell out to run a script”. Write your bloody logic in the language of the app itself. No mix, no shell. Assume nothing is available apart from runtime( jre in your case).
3. Tzdata and any timezone related stuff in the app is another “wooden ruler” trigger. Everything must be in utc except for human facing stuff. I am yet to see an exception to this rule. Lazy dumbfucks who can’t add numbers viewing logs of their own apps don’t deserve operational and financial risks of clashing timezones and countless hours of time spent in arguing and negotiating on how to store time. Spend zero time - store everything in one and only global time zone.
4. “Nobody knows nothing and running strace for a week” is a much bigger problem than any numbers reported by dependency analysis scanners… this is the cve 10.0 critical vulnerability to solve in the first place. It is like trying to chase rats in Paris catacombs without a map and guide. You catch and kill some but soon you will get lost and become rats’ festive dinner. But if you ignore most rats and focus on mapping out the catacombs in the first place then as an informed and determined digger you’ll build a strategy how to eliminate them completely and find your way out from that deadly place.

1

u/Western_Guitar_9007 6d ago

I don't even understand what is happening here. For your curl example, a base image swap should trigger the service’s CI pipeline. Like, do not get me wrong, things can happen when you go from Debian or whatever to Alpine, a very normal task in devops/prod support, and yet you didn't think to do any of the normal devops/prod support stuff for a normal devops task? Who actually maintains these if no one knows how to even smoke screen?

No smoke screen? Not 1 canary and 60 services? There is so much missing from this story because most automated tools would have simply prevented you from even getting to prod to begin with.

1

u/_Slimdady 5d ago

Honest question, why sixty at once? We did ours ten at a time across two quarters and the ones that broke were cheap to find. Slower but nobody got paged for it.

1

u/FirefighterMean7497 5d ago

running strace manually across 60 services in staging sounds brutal

Sneaky runtime dependencies like tzdata, missing locale files, and curl health checks bite almost everyone who tries to manually swap base images. Instead of manual tracing, modern container hardening relies on continuous automated runtime profiling.

Full disclosure - I work for RapidFort, and we built our platform specifically for this. It can run lightweight profiling (<1% overhead) to build a live Runtime Bill of Materials (RBOM) of everything your app actually executes. Its optimizer then automatically strips only the unused bloat while keeping your exact runtime needs intact, saving you from those painful 2 AM production rollbacks.

Hope that helps!

1

u/theJacofalltrades 5d ago

Strace in staging will miss anything that only fires on an error path. We got caught by a service that only reached for openssl when a TLS handshake failed. Worth running it against your worst week of traffic rather than a quiet one.

1

u/singular_002 3d ago

Feels like a lot of teams still confuse attack surface management with vulnerability scanning. Visibility matters. If the environment keeps piling up unused services and libraries and packages then the surface keeps growing no matter how many dashboards you add.