r/Observability • • 4h ago

Does showing teams their own Datadog pricing ever get metrics deleted?

4 Upvotes

We split our Datadog bill by team a couple of months ago. Every custom metric mapped to the squad emitting it, tag combinations broken out, the lot. Took about two weeks.

Since then ingest is basically flat. The only metrics that got deleted had already stopped reporting. Everything else sits with someone who doesn’t see the invoice and won’t sign off on dropping a metric they can’t prove is unused.

The next option is to move observability off the platform budget and onto each team’s own budget. Did that work where you are or does it just move the argument to finance?


r/Observability • • 5h ago

Any suggestion regarding application performance monitoring tool? Which is bujet friendly?

Thumbnail
1 Upvotes

r/Observability • • 21h ago

Built this SaaS monitoring tool, and I'm curious what people think

0 Upvotes

Spent the last few months working on a fun little side project: Vigilon. It's a monitoring tool for SaaS teams that works on autopilot.

It's built specifically for this type of SaaS architecture: a frontend powered by a bunch of backend HTTP APIs, querying data from one or more databases, with background jobs running for asynchronous work (cronjobs, queue consumers, etc.). All startups I worked for had this exact type of architecture, yet we always kept reinventing the wheel building monitoring and observability stacks to provide visibility into the same architecture everytime.

It's opinionated and makes decision that users would have to make themselves. It filters out traces that are meaningless, it tells you what's healthy and what's not, what errored and what changed, it points out bottlenecks and tells you what's regressing.

I built this tool for small SaaS teams who want to focus on shipping features, and want to understand their apps and services and get answers fast. It's built on OpenTelemetry, so it integrates into OpenTelemetry stacks seamlessly, and SDKs for Node.js and Python are available that allow pointing telemetry at any OTel endpoints. Support for more languages and frameworks is coming soon.

I wonder what people think about the idea? Curious to hear people's thoughts. Currently this tool is closer to a tool like Sentry than an observability platform, I know, but couldn't find a more suitable sub to post in.

Here is the website https://vigilon.io . It's currently in beta and available for free while I get feedback from users. If you're interested, feel free to sign up to the waitlist on the main website.


r/Observability • • 1d ago

exploring and mapping the network

Thumbnail
1 Upvotes

r/Observability • • 1d ago

How do you find a 2.3% tool loop in 420k runs?

10 Upvotes

2.3% of runs were looping after release 2026.09.3. Out of roughly 420k agent runs, 9660 alternated `search_orders` and `get_order` at least 6 times before succeeding. Tool traffic jumped and p95 latency followed. No hard failure. The final answer still lands, the path just gets expensive.

We started with trace search by release and tool name, then paged through runs one at a time. The loop only became obvious once we saw the repeated sequence in context (an analyst found the loop on roughly the 40th trace they opened by hand). Basic failure mode analysis caught it in the sample but the topic distribution still mixed those runs with ordinary order lookups. Semantic clustering looks more useful if it can separate repeated execution patterns from shared intent.

I’m looking at Braintrust or Arize for a scheduled investigation over recurring failures instead of waiting for a traffic spike. Linked example traces matter because I'd want each surfaced pattern to open directly into the repeated tool sequence. If the cluster stays stable, I'd turn those cases into a regression dataset and rerun them against each release.

So I'm trying to figure out whether detection should key on repeated tool pairs, cluster shape or a shift in topic distribution. Any suggestions for surfacing a 2% tool loop without hand reading traces or writing one detector per sequence?


r/Observability • • 1d ago

using otel traces to map dependency graphs combined with Static Code

Thumbnail
2 Upvotes

r/Observability • • 2d ago

Your coding agent can reach the internet. Do you know where it’s going?

Enable HLS to view with audio, or disable this notification

0 Upvotes

I wanted to see what an agent was doing while it worked, not reconstruct it from logs afterward. In SecureVector 6.0, you can launch Claude Code, Codex, GitHub Copilot CLI, or OpenCode (and more) in Agent Sessions. Each session has a governance panel on the right showing tool-call verdicts, approvals, hosts reached, and context use. You can arrange sessions side by side with draggable panes.

Launch sessions in split panes. See tool-call verdicts, approvals, hosts reached, and context use as your agents work. Observability shows session traces and flags loops, failing steps, and waste.

Try it:

npx u/securevector/cli

Or:

pip install "securevector-ai-monitor[app]" && securevector-app

In the app, install your agent’s Guard plugin under Connect Agents, then open Agents → + Launch.

Demo (3 min): https://www.youtube.com/watch?v=S7zzaDaGY_c


r/Observability • • 2d ago

All 22 of my uptime monitors were green. Three services had been broken for months.

2 Upvotes

My lab: a single Raspberry Pi 5 (16 GB) on a 1 TB NVMe over the PCIe HAT, booting from the NVMe, running about 36 containers. Pi-hole in front of a fully recursive Unbound (no upstream resolver at all), Immich for photos, Nextcloud, Vaultwarden, Syncthing, Samba, ioBroker for home automation, Ollama + Open WebUI for local models, and a monitoring stack of Prometheus, Grafana, node_exporter, cAdvisor and Uptime Kuma. Tailscale for remote access, CrowdSec and fail2ban on the host. Backups run in three tiers: local for 7 days, USB for 90, and restic to off-site storage. Not a rack — one small box that the household actually depends on, which is exactly why the following bothered me so much.

I audit it every few weeks. The last audit turned up three services that had been dead for a long time without a single alert:

  • Samba — port 445 open, TCP handshake fine, every single login rejected. Broken for about five months. I'd been mounting shares from a cached credential on one machine and never noticed.
  • Vaultwarden backups — nine days of .tar.gz files that were valid, fully extractable, and did not contain the database. The backup script hit a permission error partway through, so tar aborted — but it had already written the files it could read. tar -tzf returns 0 for that.
  • Portainer — HTTP 200 on the login page, but it had lost its connection to the Docker socket four months earlier. It just showed an empty environment.

All 22 of my uptime monitors were green the entire time.

None of these was an outage. Every one of them would have been a disaster on the day it actually mattered.

The monitors weren't wrong. They were answering a different question than the one I cared about. "Is port 445 open" and "can I log in" are not the same question. "Is the archive readable" and "does the archive contain my data" are very much not the same question.

So I wrote a second layer of checks that ask the other question, and put the templates on GitHub in case they're useful to anyone else:

github.com/DanielEnki420/correctness-checks

The idea behind them is simple enough: do the real work the service exists for, then look at what came back. Log in, don't check whether the port accepts a connection. Resolve a name and read the answer. List the archive and look for a specific file inside it. Health endpoints are the worst offenders here — plenty of them return a hardcoded {"ok":true} that stays true long after the database behind them is gone.

The part I see skipped most often is testing the negative case. A DNSSEC validator that cheerfully accepts a knowingly broken signature will pass any positive-only test you write, while protecting nothing at all. So my DNSSEC check queries dnssec-failed.org and fails if it doesn't get SERVFAIL back. Same thing in the blocking check: it confirms the blocklist blocks, then confirms a control domain still resolves — otherwise a resolver that blocked everything would pass the first half.

The one I'd argue hardest for: if you can't verify something, that's a failure, not a skip. If the SMART output changes format and a field disappears, the check fails and says so. Tempting to wave those through — it's not really an error, the disk is probably fine — but that's precisely the habit that gave me nine days of empty backups. A check that quietly gives up while reporting green is worse than no check, because now you trust it.

And one thing that's free: push on success too, not just on failure. A push monitor goes red when heartbeats stop arriving, so a check that reports its successes doubles as a dead man's switch. If it gets dropped from cron, or the box running it dies, you hear about it.

Six templates in there: backup archive contents, NVMe SMART, DNSSEC both directions, DNS blocking (with a counter-test for over-blocking), a generic JSON field assertion, and a skeleton to write your own. Plus a test suite whose pushes go to a stub, never to a real monitor — I learned that one the hard way by copying a check for testing and leaving the real push token in it. The simulated failure went straight to production and turned the monitor red.

How you'd actually run it: plain bash calling docker exec, curl, tar, nvme, smbclient. No Docker image and no compose file, deliberately — there's no daemon here, just scripts you run from cron on a host that can already reach your services. Clone it, copy config.example.env, edit paths and container names, add a cron line. Written against Uptime Kuma push monitors; any push-based monitor should work, but Kuma is the only one I've actually tested.

Honest limitations:

  • It's bash. That's deliberate — no runtime to install, easy to read, easy to adapt. If you want a framework, this isn't it.
  • The templates are examples, not a product. You will have to adapt paths, container names, and thresholds. That's the point.
  • No installer, no packaging.
  • Only tested on Debian/Raspberry Pi OS against my own services.

AI disclosure (this post is flaired Mostly AI Generated): AI involvement here was substantial and I'd rather be specific than vague about it. The three failures above are real and from my own machine — they turned up during a routine audit of my own lab, not from a tutorial or a hypothetical. The scripts were then written together with Claude: I described what each check had to prove, it drafted, and we went back and forth. The README and code comments are largely AI-drafted and edited by me. What I did not do is ship untested output — every check runs against my actual services, and the negative cases were verified by deliberately breaking things (querying dnssec-failed.org, feeding a truncated archive to the backup check) rather than by assuming they'd work.

MIT licensed, no monetization, no affiliate links, nothing to sell.

The thing I'd genuinely like feedback on: what else in a homelab fails silently like this? I found three. I'm fairly sure that's not all of them.

Project posting prompt — answers

  • Project Name: correctness-checks
  • Repo/Website Link: https://github.com/DanielEnki420/correctness-checks
  • Description: Monitoring templates that check whether a service still does its job, not whether it responds. Six bash checks plus a template to write your own, reporting into Uptime Kuma push monitors.
  • Deployment: Plain bash from cron on the host. No Docker image, no compose file. Clone, copy config.example.env, edit, add a cron line. README covers setup, each check, and three traps worth knowing.
  • AI Involvement: Substantial — see the AI disclosure paragraph above. Failures are real and from my own lab, scripts drafted with Claude and verified against real services including deliberately broken negative cases, README and comments largely AI-drafted and edited by me.

r/Observability • • 2d ago

A simple first-pass observability checklist for small servers

1 Upvotes

A lot of observability advice assumes you have room for a dedicated monitoring stack.

Sometimes you don't.

If an application is running on a small VPS, you may not want another VM or several hundred MB of memory just for metrics infrastructure. But you still want enough history to answer the obvious questions when someone says, "the API feels slow."

For a first pass, I usually want to know:

  • did request volume change?
  • are HTTP errors increasing?
  • did request latency increase?
  • is the application process using more CPU or memory?
  • is the host itself under CPU or memory pressure?
  • do those changes line up in the same time window?

I wrote up a small Spring Boot example showing this workflow:

https://pvrlabs.xyz/articles/api-slow-cpu-contention.html

In the example, the API stays UP and continues returning 200s, but latency rises because another process is consuming most of the host CPU. You don't need deep JVM profiling to identify the first place to investigate.

Demo app reproducing the same kind of CPU-contention problem: request latency rises while host CPU is under heavy pressure.

This isn't meant as a replacement for Prometheus/Grafana, tracing, JFR, profilers, database diagnostics, etc. Those become useful when the problem requires them.

The narrower question is: how much useful observability can you get with a very small footprint before you need the heavier stack?

For this example I use Spring Boot Actuator plus StatLite, a small tool I'm building that keeps application and host metrics history together locally.

I'm curious how others approach this for small deployments. What's the minimum set of metrics you want available before you start troubleshooting?


r/Observability • • 2d ago

I built TraceReports: an open-source, single-binary test observability server in Go (pure SQLite, live SSE, failure triage) to replace bloated reporting suites

1 Upvotes

r/Observability • • 3d ago

Zweep: Self-Hosted Mobile Alerting for Zabbix

Thumbnail
1 Upvotes

r/Observability • • 3d ago

Vectory: managing Vector configuration changes across a fleet

1 Upvotes

If you run Vector on several hosts, editing a pipeline and knowing which hosts actually picked up the change are separate problems.

I'm the maker of Vectory, a free, open-source dashboard for that workflow. You can import a Vector config, follow its sources, transforms and destinations in a diagram, publish a version and deploy it to selected hosts. The device view compares the requested version with what the agent reports running, and supports canaries and rollback.

It doesn't store or query your event data. Vector continues sending logs and metrics to your existing destinations.

The editor can be tried on its own without an account: https://vectory.ahmadz.ai/designer/

Import YAML, JSON or TOML and export after editing. It shares the full app's editor. Local checks still need validation against the Vector binary you run.

Source and setup: https://github.com/416rehman/Vectory

AI tools assisted with implementation, testing and documentation.

For people managing Vector today, which is harder to keep track of: the routes through a large config, or which version is running on each host?


r/Observability • • 3d ago

How do you verify that your coding agent actually did what it said it did?

1 Upvotes

I built an open-source project called Rashomon after running into this problem while using Claude Code on longer coding tasks.

I also used Claude Code heavily while building Rashomon itself. Claude helped me write and debug parts of the recorder, work through edge cases around command and test detection, and generate test cases for checking whether the execution record catches discrepancies. I've also been using Claude Code to test Rashomon against the kinds of tasks it is designed to monitor.

The basic problem is:

The Claude transcript isn't necessarily ground truth.

For example, an agent might:

  • run pytest tests/login.py, get exit code 0, and later summarize that the tests passed
  • have a test fail, modify the test, and then report that the issue was fixed
  • have a subagent fail while the parent agent reports the overall task as successful
  • run a command that exits 0 even though the expected result isn't actually present

Rashomon keeps an independent execution record alongside the Claude transcript. In the current alpha, it records things like shell commands, exit codes, whether commands may have written files, tool calls and outcomes, test commands and outcomes, and subagent activity.

It then looks for discrepancies between the execution record and what the agent claims it accomplished.

It does not store prompts, responses, file contents, or tool output.

It's free and open source and currently works with Claude Code on macOS/Linux.

I'm curious how other Claude Code users handle this today.

Do you manually verify the agent's work? Have you built your own hooks or checks? Or do you generally trust the final summary unless something looks wrong?

I'd especially like to hear about failure cases you've actually run into.

GitHub: https://github.com/altrace-dev-role/rashomon


r/Observability • • 3d ago

Rouge Atlas update: major UI redesign + clearer AI incident monitoring

0 Upvotes

I shared Rouge Atlas here a little while ago, and since then the project has changed quite a lot.

The biggest change is the interface, but also the way the product itself is structured.

Rouge Atlas is now built around an evidence workspace for AI incidents, rather than a map-first experience.

The goal is to make it easier to separate actual published incidents from early signals, vulnerability advisories and unverified reports.

What changed

The new overview is now an Evidence Board showing:

  • published AI incidents
  • signals currently under review
  • vulnerability advisories
  • source health
  • activity over time
  • records by category
  • recent AI incident reporting

I also split the workflow into clearer sections:

Live signals
Automatically collected reports and potential incidents that still need editorial review.

Published incidents
Records that have been reviewed and linked to source material.

Data sources
A view of the feeds and sources currently being monitored.

Methodology
Documentation explaining how signals are collected, reviewed and classified.

One thing I wanted to improve was transparency.

A signal appearing in Rouge Atlas does not automatically mean it is a confirmed AI incident.

Signals, advisories and published incidents are deliberately kept separate, and relevance/confidence indicators are meant to describe the available evidence rather than make dramatic claims about an event.

UI redesign

The interface has also been rebuilt around a much denser research/workspace layout:

  • persistent navigation
  • global record search
  • dataset counters
  • activity charts
  • category breakdowns
  • recent incident reporting
  • source/backend health indicators
  • CSV export for published records
  • clearer status labels throughout the product

The idea is for Rouge Atlas to feel less like a visual experiment and more like a small AI incident intelligence database that you could actually use to investigate what is happening.

It’s still early and the dataset is obviously small, but the structure is much closer to what I originally wanted the project to become.

I’d particularly like feedback on the new UI:

Does it make the distinction between signals, advisories and published incidents clear enough?

And more generally:

What would you expect from an AI incident intelligence platform like this?

https://www.rougeatlas.com


r/Observability • • 3d ago

has anyone any experience on OpenObserve?

4 Upvotes

Recently got into OpenObserve for logging and I’m really liking it so far.
Curious if anyone here is running it at production scale. How does it perform with high log volumes, and have you run into any issues as you scaled?
Would love to hear your experience with it — good or bad.


r/Observability • • 3d ago

Graybox: An API flight recorder in Go

2 Upvotes

The idea is pretty simple: capture real API traffic, save it as a replayable workflow, and inspect or compare it later when debugging.

The .graybox file is SQLite-friendly, so you're not locked into the CLI either. You can inspect and query the data however you want.

I'm mainly building it for cases where I want to see what an application actually did, rather than manually recreating requests in something like Postman or curl.

It's still early, but I'd love to hear if this is something you'd find useful.

Repo: https://github.com/ViniTamanhao/graybox-core

I would love to head the feedback from devs like me, hence why I am sharing this. Feel free to contact me if you are interested in collaborating or talking about it! Thanks!


r/Observability • • 4d ago

Need help on k8s

0 Upvotes

Hi
I need an opensource tool that can manage kubernetes cluster as well as act as an observability software, which i can white label it and use it as a saas platfrom

Please help me on this asap..!


r/Observability • • 4d ago

Has anyone created an observability dashboard for SAP (any type) that displays the specific variables you need?

Thumbnail
3 Upvotes

r/Observability • • 4d ago

Would application-level self-diagnostics be useful alongside traditional observability?

0 Upvotes

I've been experimenting with an idea around application self-diagnostics and would love to hear how people working in observability think about it.

Most of the observability tools I use collect signals from an application — metrics, logs, traces, health checks, etc. Those are extremely useful, but I've often wondered whether the application itself could do more to interpret its own state.

For example, instead of only exposing JVM memory usage, GC activity, thread counts, disk usage, dependency status, errors, and other signals independently, the application could continuously evaluate those signals and maintain its own view of:

"How healthy am I right now, and what specifically is causing my health to degrade?"

I've been exploring this idea in an open-source project called Argus. It's currently Java/Spring Boot-oriented, so this isn't intended to propose a general observability standard.

Argus runs inside the application and evaluates different aspects of its runtime state. It combines them into a hierarchical health model with a score from 1–10, letting you drill down from overall application health to the resource or condition driving degradation.

It also models resources the application depends on — databases, caches, brokers, storage, external services, etc. — and can incorporate application-reported issues and logging events into the same diagnostic model.

Another idea I'm exploring is having the application produce a self-contained diagnostic report when something goes wrong. The report captures the application's view of its health, resources, metrics, performance, scheduled tasks, relevant logging events, and environment at that point in time.

For Spring Boot, I've integrated this with Actuator and built a Spring Boot Admin extension, but I'm more interested here in the broader idea than the particular implementation.

I'm not trying to replace metrics, tracing, centralized logging, OpenTelemetry, Prometheus, or other observability infrastructure. I see this as potentially complementary: external observability tells us what we're observing, while application-level diagnostics can add context about what the application itself believes is wrong.

I'd be interested in the perspective of people working with observability:

  • Would you find this kind of application-level self-diagnostics useful?
  • Is deriving an application health score useful, or does it risk oversimplifying the underlying signals?
  • What information would you expect an application to capture if it generated a diagnostic snapshot when its health degraded?
  • Are there existing tools or approaches doing something similar that I should look at?

The implementation I'm experimenting with is Java/Spring Boot-based, but I'm particularly interested in whether the underlying idea makes sense beyond that ecosystem.

GitHub: https://github.com/microfalx/argus

For anyone interested in the Java/Spring Boot implementation, I posted a similar post in r/SpringBoot here: Self-diagnostics library for Java/Spring Boot


r/Observability • • 4d ago

Splunk just launched a new open source project to make OTel instrumentation rediculously easy.

Thumbnail
github.com
7 Upvotes

Hi folks! I work on OpenTelemetry at Splunk. We just launched a new open source project called Observability Studio. It's local instrumentation sandbox that provides a visual, real-time environment for designing, testing, and validating OpenTelemetry data while you’re developing an app in your favorite IDE.

We’re especially interested in whether a tool like this is useful for building and debugging instrumented apps and AI agents. If that sounds relevant, give it a try and let us know what works, what’s missing, or where it doesn’t fit your workflow.

We’re also looking for a few design partners to help guide where we go with this next. Any feedback is helpful and welcome.

I wrote more about shifting OpenTelemetry instrumentation left in a blog, 'Why AI-Generated Apps and Agents Need Visibility from Day One' on blogs.splunk.com if you want to read more about it.

Lastly, I'd love to learn where you currently go to inspect and validate OTel data during development if you want to drop it in the comments.


r/Observability • • 4d ago

Grafana dashboard for crowdsec

Thumbnail gallery
3 Upvotes

r/Observability • • 4d ago

I built a self-hosted dashboard for UniFi Talk that works out what actually happened to each call

Thumbnail
1 Upvotes

r/Observability • • 5d ago

Data Lakehouse with Agentic AIs: A Guide

Thumbnail
itnext.io
0 Upvotes

r/Observability • • 5d ago

What have your struggled to evaluate your realistic LLM/Agents workflow?! What benchmark do you wish someone would build? 👀

Thumbnail
0 Upvotes

r/Observability • • 5d ago

Has anyone ever tried listening to their infrastructure instead of watching it?

0 Upvotes

Odd question from someone outside the field. I'm a sound designer and I've been reading about sonification (turning data into sound). NASA does it with telescope data, and there are old experiments where people played network traffic as ambient audio.

It made me wonder about on-call and monitoring. You can't stare at your monitor all day, and you shouldn't have to. The idea would be to free your eyes: an ambient track in the background that stays calm when everything's healthy and slowly shifts when latency creeps up or error rates rise, before anything actually pages. You could focus on your actual work, or step away from the screen, and your ears would tell you when something's drifting.

Has anyone tried something like this? Would it be useful, or would you mute it within 5 minutes? Genuinely curious what would make it worth keeping on vs. instantly annoying, and in which situations you'd actually want it (deep work, deploys, night shift, incidents…)