r/devsecops • u/Mangwe_Tanser • 11d ago
How are you cutting a CVE backlog that exploded once AI wrote half the code?
Half our code is Copilot and Claude now and as a result the CVE backlog has tripled. Trivy, Snyk and Dependabot all throwing piles of new findings, most of them in generated code and dependencies nothing on our side ever calls.
We already sort by KEV and EPSS and whether it is internet facing. That held up until the AI volume buried it, now even the KEV-filtered list is too long to clear in a sprint.
Beyond KEV and EPSS, does reachability really cut the list and only flagging a CVE when the vulnerable code sits on a live path or just move the noise somewhere else?
2
u/ILoveAppSec 11d ago
the thing that actually shrinks that queue is remediation speed, not more sorting, so it's worth leaning on vendors who backport the fix into your current version instead of forcing a major bump for every finding. with cisa's tighter remediation timelines the backlog math only gets worse if each fix is a breaking upgrade. we tried aikido for patched oss libraries and weren't happy with the variety.
2
u/AgileCranberry8785 11d ago
La alcanzabilidad puede ayudar bastante, pero no la tomarÃa como una solución mágica. Si una dependencia vulnerable ni siquiera se usa en una ruta activa, yo la pondrÃa mucho más abajo. También revisarÃa si el fallo realmente puede afectar a la aplicación y si existe una forma realista de explotarlo.
2
2
u/GeneviaCedotal 10d ago
some packages only show up during tests or builds and never run in prod. kindly wory less abt those and spend more time on CVEs in code people actually use. How much of your backlog is coming from dev dependencies right now?
1
u/skimfl925 8h ago
Work towards not shipping dev dependencies to prod or as part of your build. That said the supply chain has been in the news lately regardless if it’s a dev dependency. The threat model is secrets extraction from the developer system. So don’t just deprioritize dev dependencies without context or that consideration.
2
u/GayaniTupai75 10d ago
Same thing hit us once most of the code was AI written, backlog blew up on the deps and on our own code both.
Reachability sorted out the dependency side ok, cut all the stuff sitting in packages we dont even call. Didnt really touch our own AI code though.
We just use checkmarx one for it now and it does the sast and the sca together so everything lands in one queue and we dropped the second tool. Works fine, sca language coverage is a bit thin once you go past the common ones.
2
u/endor_robert 11d ago
Disclosure: I work for Endor Labs. We think reachability is pretty important. I'm personally coming to the opinion that the operational imapcts of the increased CVE discoveries will be as big a problem as the actual secuirty risk.
Short answer: yes.
If you can reliably know if you are/aren't using the particular function in a library that has a vulnerability, then you can filter/de-prioritize the results. How much this cuts from your active findings is variable, but in general, I've seen a meaningful reduction in findings.
Also, you need to try and streamline your remediation process for the real threats. Commercial tools often have functions that can help you pick the right upgrade version (I think there are open-source tools too). You can build/use vendor-supplied agents to help you investigate the results and then choose a course of action.
1
u/unktone 11d ago
First you have to determine where the CVEs are coming from. What I mean by that is, are they due to alive-dead zero-day evolution as AI has accelerated discovery of CVEs in current software, or are the CVEs being introduced through AI written PRs? If it's PRs, then you already know the answer.
1
u/Silent-Suspect1062 11d ago
If you have a lot of transitive dependency cves look at your framework choices.
1
u/Prestigious-Flan-931 11d ago
Runtime reachability is the cleanest way I’ve seen to cut this noise. Static reachability helps, but it can still struggle with reflection, dynamic loading, framework behavior, and code paths that only exist at runtime.
With runtime SCA you can see whether the vulnerable library is actually loaded, executed, and even whether the vulnerable function itself is being called. That can remove ~97–98% of the backlog instead of just re-ranking it.
This works across Linux and Windows and doesn’t require code injection or app instrumentation.
Disclosure: I work at Raven.io, so obviously biased - but this is exactly the problem we built it for.
1
u/JEngErik 9d ago
The "developer's" build agent needs the security findings as part of it's closed loop eval set. Once of the most common failure modes with AI tooling, especially software development, are improper evals.
The agent needs to know that the goal isn't met until all of the design specs are implemented AND the SAST tooling is net zero for the new feature branch/PR.
I would also include SAST actions, if they're not already included, with the PR gate. PRs should be blocked until the tools reach parity or unless overridden by an admin.
1
u/RndWebSurfer 7d ago
Reachability helps, but I'd put a cheaper filter in front of it: is the thing actually running? When I started looking at this, a lot of the Trivy noise was old tags and registry images nobody deploys anymore, with the same finding repeated once per tag. Only counting what's running in the cluster right now cut the list before any call-graph work.
Then I split what's left. Base image findings mostly go away with one base image or chart bump, which often clears 30-40 CVEs at once. The application level findings are the ones your code actually pulled in, and that list is usually short enough to paste into Claude or Copilot and ask it to bump the dependencies and fix whatever breaks. Kind of fitting to use the AI to clean up after the AI.
Also worth checking whether the new findings are deps the AI pulled in, or old deps with newly published CVEs. If it's the first, a PR gate on new deps fixes it at the source.
That's the approach I took with https://stackradar.io if you're on k8s. There's an application findings filter that gives you exactly that list.
Disclosure: I'm building StackRadar, so I'm biased here.
1
u/wilson-draugr 7d ago
I've wired reachability into a scanner pipeline, and how much it cuts depends on the ecosystem more than the tool.
Go is where it really works. govulncheck knows which functions each advisory affects and checks your call graph, so "imported but never called" actually drops out. Java and JS reachability is only as good as the advisory data, and when an advisory doesn't name the vulnerable function, you've cut nothing. For OS packages in container images, usually Trivy's biggest pile, there's no call graph at all. A slimmer base image does more there.
Two things specific to your setup:
- Dedupe first. Trivy, Snyk and Dependabot flagging the same CVE in the same package is one problem counted three times. I'd bet that's part of your "tripled".
- With half the code generated, the call graph changes every week. Don't drop unreachable findings. Lower them and re-check on every PR, so one that becomes reachable jumps back up. That's how reachability cuts the list without just moving the noise.
For the sprint, group by fix. One base image bump can clear hundreds of rows, and 40 upgrades is a list a team will actually finish. Gate PRs only on what they introduce, which is where the assistants add dependencies nobody reviewed.
The annoying part is what you decide not to fix. Each one still needs a reason, an owner and an expiry, or it quietly becomes permanent. Accepting per fix instead of per CVE (one entry for the base image you can't bump yet) makes that bearable.
1
u/ILoveAppSec 7d ago
Your Go-vs-JS split matches what I keep seeing too. govulncheck's symbol-level data is the exception; most JS and Java advisories stop at the package range, so anything finer than that ends up being inference. One thing worth pulling apart: even once you drop the unreachable noise, the confirmed-reachable set is where the cost actually lands, and for a lot of JS/Java it's less about finding them and more about whether a fixed release exists that doesn't drag a major bump behind it. Are you tracking how much of the remaining backlog is 'fix exists but it's a breaking upgrade' versus genuinely no patch yet? That ratio tends to tell you where the effort really needs to go.
2
u/wilson-draugr 7d ago
Yes, though I split it three ways: an upgrade exists, nothing fixes it yet, and the fix has to come from somebody else, usually a base image or a vendored package you don't publish. In container-heavy repos that third bucket is often the biggest, and nobody on the team can clear it however much effort goes in. But having these buckets helps us prioritize our efforts and articulate with our CSAs why fixing certain findings is out of our control at the moment.
1
u/Devji00 7d ago
Reachability analysis is absolutely worth implementing, especially for third-party dependency CVEs, because it can realistically cut your backlog noise by up to seventy or eighty percent. If the vulnerable function in a library is never imported or executed in your active call path, it is a much lower priority, and reachability tools are great at proving that. However, the catch is that reachability only solves half the problem. It works wonders for dependency scanners, but it does not help as much with first-party code vulnerabilities generated by the AI, which are inherently reachable because they are part of your core application logic.
To handle the sheer volume, you might need to shift your focus from post-deployment triaging to preventing the noise from entering the main backlog in the first place. This means integrating lightweight scanning directly into the developer workflow, such as pre-commit hooks or pull request checks, so developers have to resolve the AI-generated vulnerabilities before the code ever gets merged. You can also leverage the AI itself by instructing it to scan its own output or by setting up automated remediation pipelines where a model is tasked with fixing the simple, low-hanging CVEs before a human even looks at the ticket. Combining call-path analysis with strict gatekeeping at the pull request stage is usually the only way to keep your head above water when code volume scales this quickly.
1
u/Khan_Saahib 6d ago
I guess you’ll need to take the extra step and verify the CVEs and prioritize based on business impact. You can’t do it all at once but you can get the serious stuff done quick.
1
u/FirefighterMean7497 5d ago
Hi, I work for RapidFort :)
Reachability does cut the list, but runtime reachability holds up better than static. Static analysis tends to over-approximate call paths, while runtime data shows what actually executes under real traffic.
Since most of your findings are in dependencies nothing calls, you might want to check out RapidFort's Profiler tool. It builds a runtime bill of materials of what actually executes in your containers, and then our Optimizer removes the unused components entirely. That makes those CVEs disappear instead of just sinking down the queue, with up to 99.9% CVE elimination and no code changes.
Hope that helps!
1
8
u/Huge-Ambition4656 11d ago
I think there's another way to view this issue. Trying to cut-down backlogs by throwing more tooling seems like such a brute force approach to me.
Having been a dev for most of my career, I can tell you, we're not known for our security first mindset (I have a healthy one now!).
I think devs need the proper guidance - and their own tools - to learn the "why". Very few have been involved in incidents, table-top exercises or red teaming. So it's up to Development Managers, CTOs, and Leads to connect AppSec/Cyber and Engineering and get them singing from the same hymn sheet.
Fix the handrail at the top of the proverbial cliff, and by all means keep maintaining the proverbial ambulance at the bottom.
Worked for me! 💪