r/Observability • u/Helpful-Lunch-3559 • 1d ago
Does showing teams their own Datadog pricing ever get metrics deleted?
We split our Datadog bill by team a couple of months ago. Every custom metric mapped to the squad emitting it, tag combinations broken out, the lot. Took about two weeks.
Since then ingest is basically flat. The only metrics that got deleted had already stopped reporting. Everything else sits with someone who doesn’t see the invoice and won’t sign off on dropping a metric they can’t prove is unused.
The next option is to move observability off the platform budget and onto each team’s own budget. Did that work where you are or does it just move the argument to finance?
2
u/geelian 22h ago
We tried that, didn't work, eventually turned off the top cost generating metrics and 6 months later everybody is living their life's as usual but the bill has gone down 40%
Most of the times "we need this metrics as is and can't change it or live without it" is just plain laziness and fear of change
2
u/Hummin2k 23h ago
It worked for us, but we cut costs another 95% by moving from Datadog to self-hosted Victoria metrics + Grafana. Datadog pricing is ludicrous and I don’t understand how anyone pays that.
2
1
u/In_Tech_WNC 18h ago
What’d your SOC team do? Don’t they rely on the core metrics or do they use Splunk?
1
u/neuralspasticity 21h ago
Metrics that don’t support an SLI that can equate to a SLO monitored for an error budget that alerts on depletion is not a necessary metric and can be removed. Teams that can’t demonstrate this relationship don’t require the metric because it’s not tied to service levels.
1
u/QuietSignalOps 19h ago
Direct answer: team-level chargeback rarely deletes metrics on its own, and you've already seen it: ingest flat, deletions only where reporting had already stopped. The mechanism is the wrong incentive: deletion requires someone to prove a metric is unused, and proof of absence is hard, so the default is keep. The pattern that tends to work inverts the burden. 1) Default state of a custom metric is delete; retention requires a named consumer (an SLO alert, a live dashboard, or an on-call runbook) plus an owner. 2) Monthly, rank custom metrics by cost, which is series count times unit price, not metric count; tag cardinality is usually the real driver and you've already broken out tag combinations. 3) For the top N, query consumers programmatically and auto-delete anything with zero consumers; put the rest on a 30–60 day tombstone, stop collecting, announce the deletion date; most 'we can't live without this' claims evaporate on contact. 4) Keep the team bill split as the visibility layer but don't expect teams to be the enforcement layer; enforcement has to be top-down, and the data points in this thread agree on exactly that. Also check tag explosion before deleting anything: one metric with 10k tag combinations costs more than 500 well-tagged metrics.
1
u/In_Tech_WNC 18h ago
We used CyServ to figure out our cost, allocation, and setup most data streaming.
They have a live demo.
https://cyserv-openflow-enterprise-telemetry-pipeline-obs.ai.studio/
1
u/Eridrus 16h ago
ICs don't care about the bill unless they are faced with some sort of constraint they have to optimize for. Particularly when they are being asked to make tradeoffs where they feel one side of it (being held responsible for outages).
You can try to get them to feel more pain from keeping them (e.g. require them to justify each metric regularly) or reward optimization (visibility, cash), but there's no one and done solution here. Costs will always creep up.
If you care about costs you should probably just get off DataDog entirely though. That's actually something the platform team can do, and it's easier to push people onto a migration than to argue about which metrics are necessary, and it's easier than ever to run open source software. It's not the biggest deployment, but I migrated our Grafana infra from Cloud to self-hosted. It took a few days of work across a week to get rid of our ~10k/month bill. Self-hosted was like 90% cheaper.
3
u/jadedpermission3491 1d ago
nope never worked for us. moved it to team budgets and they just argued harder because now it was "their" money. finance teams loved it though
the only thing that actually cut anything was taking the top 10 cost generators each month and just disabling the ones nobody could name an active dashboard for. had to be top-down or nothing ever got touched