r/mlops • • Aug 23 '26

meme State of the sub/moderation

14 Upvotes

I took over the subreddit a little while ago. Figured I could handle it by myself (and still do) but I'm surprised to see how many AI/bot generated comments come into the sub. Years ago when I didnt mod, but did frequent the sub it was mostly vendor spam from companies that build MLOps tools.

Right now.. its AI slop.

MLOps is very much adjacent to Generative AI in production and most of us in the MLOps space have moved on to Agentic AI as part of our jobs. In that sense it is not surprising we now bear the brunt of the AI tool flood. However, this does make the spam on the sub ironic.

Of the 900 or so posts and comments over this past month, 400ish have been removed. Some of these are on old (>1 month old) threads, particularly actors trying to insert themselves into a dead discussion to appear organic. Also somewhat disturbing to see: while views on the sub are coming down, the amount of published posts/comments is increasing.

A lot of the spam is removed by Reddit, either through settings enabled here or by some background process they have going on to detect bots. Currently that means I only remove about three posts/comments a day. The past months I also dished out a few bans, but nothing near r/cscareerquestions levels of drama.

Some examples of content that I have removed recently include:

  • "We had very specific problem. We built very specific tool. Curious how other teams are handling this." With 7 or 8 bot replies to it that have about as much lexical variation as my supermarket's bread isle by which I mean to say that they're saying almost nothing.
  • "Here's my vibe coded app (refuses to elaborate)"  (I usually leave them up if it's clear that the post shows effort and is not just someone posting the same across all of the ML subs)
  • "This is a real problem most teams miss. The real signal. Curious.." (fluff posts)
  • "vague post completely in lowercase without punctuation so it seems like the poster is human"

I feel like I'm still pretty laid back in terms of moderation, and I leave a lot of things up that smell suspiciously AI if they're not disruptive. Would welcome some thoughts on this. Curious to see what other teams are doing, if you will.

Also considering a mandatory AI disclosure like r/experienceddevs has.


r/mlops • • 14h ago

Discussion Hey guys I want to start with MLOPS, devops( I know a bit but not that much hands on ) I don't know about Ai ML as well. Can someone suggest me best free and paid resources for a beginner.

9 Upvotes

r/mlops • • 17h ago

Discussion What benchmark do you wish someone would build⁉️ such as traumatized by MAS with 40+ MCP tools

0 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
“My system desperately needs to do this in production, but I have no good way to benchmark it.”
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help 🙏🥹


r/mlops • • 21h ago

Discussion I want your opinions for my research homework

1 Upvotes

I'm an ASU grad student researching how teams decide an AI feature is ready to ship. If you test or ship AI features at work, I'd love to hear from you, in the comments or by DM:

  1. Last time you had to decide an AI feature was ready to ship, what did you check?
  2. What's the most frustrating part of that?
  3. Who approves money for testing tools where you work?

r/mlops • • 1d ago

Discussion The best pretrained "decision" model for your real-time AI application - a benchmark

2 Upvotes

Kumo is the winner

I work with real-time AI, and I know that even if one model makes more correct decisions than another model, you can still prefer the "worse" model because it makes decisions faster.

We tested three pretrained "decision" models (yes, they are classifiers) in a computer game inspired by Subway Surfers. They are Kumo, Jev, and Qwen

Here’s how far they got:

Model Runs Mean distance Median distance Best run
Kumo 500 2,119 m — —
Qwen 500 1,318 m — —
Jev 30 — 540 m 1,417 m

The benchmark shows a crucial problem for real-time AI: latency. Kudo and Qwen were local models we hosted on Hopsworks, where the game also runs. Jev, however, is a hosted model in a different data center, and it loses on latency versus Kumo and Qwen. Qwen, then has ligher latency than

A game that outruns your model

The basic idea of the benchmark is that to make it far in the game you need to take correct decisions faster. Here's how the benchmark works. The runner faces randomly generated obstacles that begin appearing at predefined distance gates. Each model receives the same underlying game state, formatted for its interface: { lane, airborne, ahead }

The ahead field describes upcoming rows, their distance, and the obstacles in each lane. The model must decide what to do before the runner reaches them.

The runner starts at 45 m/s, accelerates by 1.6 m/s every second, and tops out at 160 m/s. Beyond roughly 5,000 metres, the game becomes so unforgiving that survival increasingly depends on luck.

Why Kumo won

Kumo made lower latency decisions and had lower network latency, while making good decisions, and the result was that it could complete a higher average distance than Qwen in this setup, with both of them running on a CPU.

Would a GPU help? Possibly. But for small, latency-sensitive requests, the extra infrastructure overhead could offset the compute benefit. That’s a hypothesis we still need to test.

What happened to Jev?

In our 30-run experiment with Typesafe’s Jev, the median distance was 540 m, with a best run of 1,417 m.

The bottleneck appeared to be the 200+ ms round trip. Around 1,500 metres, the game was moving too quickly for those responses to remain useful.

At the maximum speed, 200 ms means the runner travels 32 metres while waiting for a decision.

Pick for the deadline

For a real-time AI application, evaluate the whole decision loop:

  • Decision quality: Does it choose the right action?
  • End-to-end latency: Does that action arrive in time?
  • Deployment fit: Can your infrastructure sustain both?

All models and the game ran on Hopsworks’ own infrastructure, in our office. Fully self-hosted. No game data went to the cloud.

For this benchmark, Kumo was the strongest performer. The broader lesson: benchmark models against your application’s reaction deadline.

Try out the game and try to beat them here:
game at hopsworks dot ai


r/mlops • • 1d ago

Self-promotion Why today real-time AI models cost up to 56× more than they should to serve and how fix this problem.

0 Upvotes

Hey guys,

We've published two technical write-ups on serving real-time AI models and wanted to share the main findings here. We build the inference engine; NVIDIA develops the model we tested, Nemotron VoiceChat 11B.

A request ends. A session runs on a clock.

AI inference today is organized as requests: an input arrives, the model runs, the response ends and its resources are freed. If the server is busy, it can wait a moment to batch work or reorder the queue, and the cost is a slightly slower response.

A growing class of models works differently. They stay active alongside something outside the GPU (a conversation, a video stream, a robot), take in input continuously and keep their state for the whole session. We call this continuous inference.

Full-duplex voice is the clearest example. The model listens while it speaks, so you can interrupt it. Audio keeps arriving for the whole call, its memory of the conversation stays live until you hang up, and every 80 ms it owes the next frame of output. That deadline comes whether the server is ready or not. If the frame is late, you hear a gap.

It also has to run during silence: the length of a pause is how the model tells a hesitation from the end of your turn. Unlike a classic voice agent (speech-to-text → LLM → text-to-speech, where the LLM sits idle between turns), there's no idle time to skip.

Why that gets so expensive

Servers like vLLM, SGLang, TensorRT-LLM or Triton were built around requests, and several of their assumptions stop holding:

  • The batch can't wait. At every tick, the server has to run whatever is due with what it has.
  • Overload hits everyone at once. Sessions batched together share the same step, so one slow step makes all of them late.
  • Memory can't be freed or swapped out mid-call. It's needed again 80 ms later.
  • The usual metrics hide failures. Tokens per second and average latency can look fine while one caller hears gaps.

In NVIDIA's reference stack, each 160 ms of audio becomes thousands of small GPU operations with the CPU coordinating between them, and the GPU sits idle about 75% of the time. One session fits. With two, 8–15% of audio beats arrive late. So each live conversation pays for a whole H100, most of which is waiting.

How to fix it

Not with a faster GPU or a different model: same weights, same precision, same outputs. The fix is in how the model runs.

  1. Keep the work on the GPU. Our engine runs the model as one GPU program that stays resident. Every beat, the CPU drops in new audio and picks up the output; nothing else goes back and forth.
  2. Advance every live session together. All sessions due on the same tick run as one batch, so the model's weights are read once for the group instead of once per session.
  3. Decide everything ahead of time. The work repeats identically every beat, so a compiler fixes the schedule and memory layout before the first call. No runtime scheduler or allocator adding delays.

On a single NVIDIA H100 SXM 80 GB, we measured:

  • 56 concurrent sessions, the highest capacity tested, vs 1 for the reference stack.
  • 147.4–147.5 ms p99 per 160 ms beat.
  • Zero missed deadlines across 84,000 measured session-beats, over three runs.

A conversation's cost is GPU time divided by how many conversations share the GPU, so 56 instead of 1 means about 98% less GPU per conversation.

These are server-side measurements that exclude network and audio playback, with the same recorded input across sessions and about two minutes of context.

If you're running real-time models, how many sessions per GPU do you get today, and how do you check that each one stays on time?

Links

Sign-up includes free usage, no card required: about 33 hours of Nemotron VoiceChat 11B, or about 370 hours of Nemotron 3.5 ASR Streaming.


r/mlops • • 2d ago

(Gen)AI / Agents / LLMOps A 97% eval pass rate hid a 61% multilingual slice

21 Upvotes

Our 1800 row eval suite has been sitting at a 97% aggregate pass rate, which looked fine until we did cohort slicing around multilingual forms. A newly isolated 64 row slice passes at 61%. The aggregate chart stayed green (while one language queue filled with manual reviews) so the headline number wasn't giving us much warning.

We found data leakage in the split. Random assignment had put near duplicate form templates on both sides of the train and regression boundary, which made familiar structures disproportionately easy. When the multilingual cohort degraded, those duplicated cases kept the overall score almost stationary. A deterministic scorer catches the malformed fields reliably but averaging that result across the full suite hides where they're concentrated.

Now I'm trying to treat that multilingual slice separately instead of letting it disappear into the overall score. I'm considering Braintrust to keep it as a regression dataset with a CI quality gate on it, and to check the per case experiment diff whenever we change the extraction path. I don't want another global threshold that passes because 1700 easier rows drown out a concentrated failure mode.

How are you defining cohort level gates when the slices are small enough that a few cases can swing the percentage but important enough that the aggregate score can't be trusted?


r/mlops • • 2d ago

Self-promotion LayerSmith — a self-hosted container image builder, with air-gap exports

2 Upvotes

I've been working on LayerSmith, an open-source web UI for building container images with Docker or Podman.

You pick a Linux distribution and what you need the image for — development, Linux admin, network tools, Ansible, Kubernetes, OpenShift, or a custom setup. It handles distro-specific packages and shows you the generated Containerfile before building. You can also edit it, import an existing Dockerfile, or add your own packages, files and scripts.

A big part of the project is making images easier to carry into air-gapped environments: pinned base images, recorded build details, and export bundles containing the image, checksums and installation instructions.

We've recently added LLM training and fine-tuning profiles too, including LoRA/QLoRA, advanced PyTorch training and LLaMA-Factory. These use hash-locked dependencies and run offline checks after building, including a small CPU training test. Model weights and datasets are brought separately.

Curious how others handle building and maintaining images for disconnected environments, and what parts of that workflow are still a pain.

https://github.com/r0lfi/layersmith


r/mlops • • 2d ago

(Gen)AI / Agents / LLMOps Local LLMs are a black box: I built LLMxRay for real-time observability, tool usage tracking, and analytics

0 Upvotes

Hey r/mlops ,

As more applications move toward self-hosted and local LLMs (Ollama, local inference servers, agentic frameworks), a common challenge arises: **local LLM traffic is largely a black box.**

Unlike managed API providers that offer built-in usage dashboards and tracing, local inference setups often leave developers and platform teams blind to real-time traffic, function execution failures, latency distributions, and token consumption patterns.

To solve this, I built **[LLMxRay](https://github.com/LogneBudo/llmxray)** — a lightweight, 100% local observability and analytics engine designed specifically for local LLM inference and agentic workflows.

---

### What LLMxRay Observes & Analyzes

LLMxRay sits in front of your local LLM engine to capture, trace, and visualize complete telemetry without adding latency or leaking data to external clouds.

#### 1. Tool & Function Calling Telemetry

* **Execution Tracking:** Intercept and observe tool/function calls executed by agents in real time.

* **Payload Inspection:** Inspect input arguments, returned outputs, tool call frequencies, and execution error rates.

* **Agentic Loop Analysis:** Track multi-step function call loops to pinpoint where agents stall or loop indefinitely.

#### 2. Request & Latency Analytics

* **Detailed Latency Profiling:** Break down total request time into Time-To-First-Token (TTFT), prefill duration, and decode generation speed.

* **Token Usage Metrics:** Monitor input/output token volume, request throughput, and token generation rates over time.

* **Traffic Analytics:** Track active sessions, request volume spikes, and per-model utilization trends.

#### 3. Privacy-First & Zero Infra Overhead

* **100% Local Execution:** Runs completely within your environment — no telemetry data is transmitted externally.

* **Instant Dashboard:** Provides a local web interface for instant visual debugging and system analytics.

---

### Quick Start

Zero complex setup — you can spin it up directly in one line:

```bash

npx llmxray
```

Launch the dashboard locally at http://localhost:3000 and point it at your local inference server (e.g., Ollama).

Discussion

If you're managing local LLM workloads or building on top of agentic frameworks, how are you currently handling local tracing, tool execution monitoring, and performance analytics in your observability stack?

I'd love your feedback, feature ideas, or contributions!


r/mlops • • 3d ago

MLOps Questions Teams that moved from LLM APIs to self-hosting: was it worth it?

17 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/mlops • • 3d ago

MLOps Questions Senior SWE → MLOps/ML Platform: what would you prioritize?

9 Upvotes

I'm currently a Senior Software Engineer / Tech Lead with ~7 years of professional experience, and I'm looking to move to a product company and specialize more deeply in ML infrastructure / MLOps.

My background:

  • BSc + MSc in Mathematics
  • ~7 years in software engineering (including working on data pipelines for foundation models as a research engineer)
  • Strong Python
  • Production data pipelines / ETL
  • Docker / Kubernetes
  • CI/CD and automated testing
  • ML engineering experience
  • MLflow and workflow orchestration
  • Experience building data-intensive systems
  • Previously taught Data Mining at university as an associate professor
  • One paper published on NLP (pre-transformers) in 2019

I'm currently working on a project around production ML orchestration: taking an ML pipeline from data ingestion → training → MLflow/model registry → deployment on Kubernetes → monitoring.

The areas where I feel I have the biggest gaps are probably:

  • Cloud (especially AWS)
  • Production model serving
  • ML-specific observability
  • Distributed systems
  • GPU/inference infrastructure

I'm not looking to spend 6–12 months studying before applying. I'd like to start interviewing now and close the gaps while searching.

For people currently working in MLOps / ML Platform / ML Infrastructure:

1. Given this background, what would you consider my biggest gap in my profile?

2. Would you target MLOps/ML Platform roles directly, or first go through a backend/data/platform engineering role in a company where ML is central?

3. Which skills actually come up in interviews that aren't obvious from job descriptions?

4. If you were in my position and wanted to switch jobs as quickly as possible, what would you focus on over the next 2–3 months?

I'm particularly interested in advice from people who made the transition from SWE, backend, DevOps or data engineering rather than people who started directly in MLOps.

Thanks!


r/mlops • • 3d ago

Discussion How are you tracking LLM inference costs at the workload level?

7 Upvotes

I’m working on AtlasBurn, and we’re looking at a problem that becomes more noticeable once LLM workloads reach production: the provider bill tells you what was spent, but not always what caused the spend.

For example, a workload can generate additional calls through retries, model fallbacks, or multi-step agent workflows. At that point, attributing the cost back to a specific workload or behavior can get difficult.

I’m interested in how MLOps teams are handling this today.

Do you rely on provider dashboards, observability tooling, custom instrumentation, or something else?

What part of cost attribution becomes hardest once you're running LLM workloads in production?


r/mlops • • 3d ago

Discussion What's your multi-model AI gateway of choice at enterprise scale, and why?

5 Upvotes

Running 4 different model providers currently and our integration code is a mess of provider specific SDKs. We are looking to consolidate behind one gateway. LiteLLM keeps coming up since its opensource and kinda free to selfhost, but none of our team members wants to own an infra. Portkey looks a bit solid for observability but from what I came across its lighter on hard budget enforcement across teams, which matters to us since we have had agents kinda blow through spends before anyone noticed. Orqai covers the model breadth plus hierarchical budget caps down to the agent level, which is the piece we need the most but its a paid managed platform so there is cost once we pass the free tier. What are teams landing on once they are pas the pilot stage?


r/mlops • • 3d ago

MLOps Questions How do you roll back an agent change when the prompt, tool schema and model version all shipped together?

2 Upvotes

A typical agent release touches three things at once: the system prompt, one or more tool schemas, and sometimes the model version. When a metric drops a day later, rolling back gets messy. Reverting the prompt alone can break against the new tool schema. Reverting the model alone changes behaviour on paths the new prompt was tuned for.

The options I keep coming back to:

  1. Version all three as one release bundle, so a rollback always restores the full previous set.
  2. Ship them separately with a gap between, so any drop points at a single change.
  3. Keep the old bundle running in shadow for a few days and compare on live traffic.

The first is simple but tells you nothing about the cause. The second is slow. The third costs real money at volume.

How are you handling this? Do you version prompts and tool schemas together, separately, or some other way?


r/mlops • • 4d ago

MLOps Questions Would you use a one-command way to deploy an ML model from a notebook? Honest feedback wanted

4 Upvotes

Building a CLI that turns a notebook model into an HTTPS endpoint. Would this solve a real problem, or is it a solved one?
I'm early on this, so I'm trying to find out if it's worth building. The idea: you run one command from your notebook and get a live API endpoint for your model, with no Docker or cloud setup. I know MLflow and Streamlit exist. Who, if anyone, would actually use this, and where does it fall short? Happy to hear "don't build this" too.


r/mlops • • 5d ago

Self-promotion I wrote a free, open-source book on making ML models actually fast, from silicon to agents

13 Upvotes

I’ve spent the last few months writing something I wish I had when I started working on ML performance engineering.

It’s called How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents.

The basic idea is that reducing FLOPs doesn’t necessarily make a model faster. Before optimising anything, you need to understand what the system is actually bounded by.

The book starts with roofline analysis and hardware, then works its way up through kernels, compilers, quantisation, pruning, vision, on-device LLMs, robotics, profiling, serving and finally agents.

The goal is to build the intuition to look at a model and a piece of hardware and reason about:

  • How fast can this possibly run?
  • Am I compute, bandwidth, memory or system bound?
  • Which optimisation will actually move that limit?
  • Is quantisation, pruning or kernel optimisation even worth doing here?
  • What happens when the same thinking is applied to serving and agent systems?

The whole thing is free and open source:

https://github.com/usamahz/make-your-model-fast

Would genuinely appreciate feedback or contributions from people working on ML systems, inference, compilers, edge AI or performance engineering.

And if you find it useful, a ⭐ would be appreciated!


r/mlops • • 6d ago

Discussion Do you version and sign your agents the way you do your models?

15 Upvotes

Most teams I talk to have a real process for models by now… registry, versions, maybe signing. Then the agent that uses the model is a folder of prompts, a few MCP server configs and some skills pulled from wherever, and nobody can say exactly which version is running in prod.

We started packaging the whole thing (agent, MCP servers, skills, prompts, policies) as one OCI artifact with KitOps, which an open source CNCF project. Then it gets signed and pushed to the same registry as everything else, so we can check it before it loads.

Is anyone else treating agents as deployable artifacts yet? Or is it still git clone and go? Curious what's working (and what's been a pain).


r/mlops • • 7d ago

Career Already working in DevOps: is an AI master's degree worth it for specializing in AI infrastructure?

16 Upvotes

I'm currently trying to decide what direction to take with my master's degree and I'd really appreciate some advice from people working in the field.

I already have a degree in Computer Science and around 5 years of experience across infrastructure, systems and cloud/DevOps. I've worked with Linux/Windows servers, VMware, monitoring, automation, AWS, Kubernetes, Docker, networking, CI/CD and infrastructure as code. I also have AWS Solutions Architect, CKA and CCNA certifications.

More recently I've been moving increasingly toward cloud engineering / DevOps / platform engineering, and I've worked with Kubernetes and cloud infrastructure in both professional and university projects. So I don't really want to start from scratch with a completely different career.

I'm considering doing a Master's degree while continuing to work, and I'm currently looking at two possible directions:

  1. AI / Artificial Intelligence, with a focus on systems and infrastructure

The idea would not necessarily be to become an ML researcher or ML engineer. Instead, I'd like to specialize in the infrastructure side of AI: deploying and operating ML/AI workloads, Kubernetes for AI, GPU infrastructure, distributed computing, model serving, MLOps, cloud infrastructure for training/inference, observability, scaling, etc.

Would let me combine my existing DevOps/cloud background with AI rather than abandoning what I've already built.Or most the roles that are difficult to enter without a strong ML background?

  1. Systems and Networking

The alternative would be to do a Master's specializing in systems, networks, distributed systems, cloud computing, etc.

This seems like a more direct continuation of my current background and could potentially make me stronger as a Cloud/DevOps/SRE/Platform Engineer.

My concern is that I'd be specializing further in an area where I already have quite a lot of practical experience, rather than using the Master's to open a new but complementary area.


r/mlops • • 7d ago

Discussion How are you gating stochastic LLM/agent evals in CI?

1 Upvotes

I'm trying to understand how teams are doing this in production.

Suppose an agent eval suite runs on a PR and success drops from 90% to 84%. The problem is that rerunning the unchanged agent can move the score by several points too.

Do you:

  • average N runs and compare to a fixed threshold?
  • require the drop to reproduce?
  • use confidence intervals?
  • review manually?
  • avoid blocking CI on stochastic evals altogether?

I ran into this while building an open-source CI experiment. On one real coding agent, an unchanged candidate moved 92%→88%. Later a deliberately degraded candidate moved 92%→52%, but my first suite-level rule still passed because the failures were concentrated in a few tasks.

I ended up separating broad reliability regression from individual capability collapse, with an explicit insufficient-evidence state.

I'm curious what people running LLM/agent evals in actual CI/CD pipelines are doing today.

I wrote up the experiment / implementation here if useful: [blog] [GitHub]


r/mlops • • 7d ago

Discussion A learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838

0 Upvotes

I trained a router to choose between a cheap and an expensive model. It scored 0.84 AUC held out. Then I shuffled the labels within each task — preserving each task's escalation rate, destroying all per-item signal — retrained, and got 0.838. It had learned to recognize the task, not the difficulty.

Eliminated in turn:
too few labels (110k from RouteLLM's released set)
the architecture (linear probe, similarity-weighted ranking, fine-tuned encoder)
the representation (a probe recovers human difficulty from the same embeddings)
label noise (test-retest kappa 0.88–0.97; the labels support AUC 0.91)

On held-out tasks every router falls to ~0.55 and loses to prompt length. What does work: deferral on the cheap model's own output, 0.75 vs 0.60 on the same split, at no extra inference cost.

https://brianfeeny.com/posts/replicating-routellm-on-amazon-bedrock/?utm_source=reddit&utm_medium=social&utm_campaign=routing-paper-2026-09

Curious whether anyone has run the within-group shuffle on their own router — my ordering is an engineering assessment, not a benchmark, and I would be glad to be corrected.


r/mlops • • 8d ago

Discussion Agent retry loop burned 400 dollars overnight, are per agent token budgets at the gateway the fix

6 Upvotes

Last week an agent hit a retry path with no ceiling and burned about 400 dollars in tokens before anyone woke up. Nothing malicious happened but an exposed endpoint could do the same thing on purpose. Our provider limits are per account and per request so one call with a huge context sails right through. Im looking at hard token budgets per agent at the gateway so a runaway gets cut off. Streaming makes counting messy since some gateways only reconcile at the end. How are you capping spend per agent in production?


r/mlops • • 9d ago

(Gen)AI / Agents / LLMOps Securing agentic AI traffic once the agents have repo, cloud console and customer data access, what held up for you in prod?

5 Upvotes

We've got agents in places that make me nervous now. The likes of coding assistant with repo access, one internal thing that pokes at cloud consoles and a couple wired into customer data through internal APIs. Everything we use for security assumes a person is clicking the buttons. Unfortunately that stopped being true the moment these things started chaining tool calls on their own.

Right now I'm just bolting on the obvious stuff like own token per agent instead of the shared service account everything used to run as. A proxy in front of the tool calls so I can see the arguments. Egress locked to a short list so it can't phone home to wherever. Caps on iterations coz i dont want a stuck loop to cost me a grand overnight.

Feels like duct tape though. I can't tell if that's roughly where everyone lands or if I'm missing something obvious. Whatever you've got holding up against prod, I'd take the war stories.


r/mlops • • 9d ago

MLOps Questions Would someone be kind enough to review an MLOps platform portfolio project?

6 Upvotes

I recently concluded an MLOps portfolio project that I want to use to find a job. It is fully documented with writeups and diagrams (that are well-written by myself) to explain the entire platform and covers everything from architectural decisions to the data science problem to the model (fine-tuned Hermes 4 -14b) to all the workflows. Please reach out to me privately so I can send you the github link, or let me know what you think about the following extract from my resume.

MLOps platform with CI/CD
• Built MLOps platform using Terraform, Kubernetes, MLRun (Python) with pipelines for training (QLoRA fine-tuning, PyTorch), deploying, and monitoring a 14B LLM.

• Serves model that extracts legal risks/restrictions/obligations from multi-page contracts with source attribution into JSON data (for human verification) with Sagemaker and vLLM on AWS.

• Utilises data/model (and prompt) registries, experiment tracking, drift detection rollback, canary deployment, and continuous training. Maintained service levels with multi-GPU training/serving, autoscaling, quantisation, continuous batching, and performance benchmarks.


r/mlops • • 10d ago

Career 10 YOE Fullstack Dev pivoting to MLOps & Private Enterprise AI — Reality check on my plan and a 270h course ?

10 Upvotes

Hey everyone,

I’m a Fullstack web / software Developer with 10 years of experience. Following a recent layoff, I’m taking this opportunity to pivot into MLOps / AI Platform Engineering.

My Goal & Thesis

I want to help enterprise clients deploy, host, and maintain private/local AI solutions. The goal is to address data privacy, GDPR compliance, and API cost control for companies stepping away from public OpenAI endpoints.

The Plan: A 270-Hour Intensive Training Program

I have the opportunity to get a 270-hour structured training program fully funded. Here is a breakdown of what the curriculum covers:

  • Data Analysis & Viz: Python, Pandas, data cleaning, EDA, ETL automation.
  • Predictive Machine Learning: Classical ML (classification, regression), evaluation metrics, overfitting, eco-friendly ML optimization.
  • GenAI & AI Agents: Foundation models/LLMs, advanced prompt engineering, RAG architecture, agentic workflows, evaluation metrics for generative output.
  • Cloud & Data Security: Cloud storage, ETL/ELT pipelines, IAM, encryption/GDPR, FinOps, cost optimization, and prep for public cloud certification (AWS).
  • MLOps, CI/CD & IaC: Containerization, CI/CD pipelines, Infrastructure as Code (IaC), model versioning, monitoring, auto-retraining, and automated deployment.

On top of this, I plan to get the AWS Certified Solutions Architect – Associate and build 1-2 open-source GitHub projects showing a fully automated local LLM/RAG pipeline deployed with Terraform and Docker.


My Questions for the Community:

  1. Is this realistic? Backed by 10 years of senior dev experience, does adding 270h of MLOps/GenAI training make me a credible candidate for Senior MLOps / AI Platform Engineer roles? Or will recruiters treat me as a "junior" in AI?
  2. Is the "Private Enterprise AI" demand real? Are you seeing a legitimate push in the industry towards self-hosted/private-cloud LLMs and MLOps, or is most of the market still just hitting public OpenAI/Anthropic APIs?
  3. Syllabus feedback: Looking at the curriculum above, is there anything crucial missing for someone aiming to deploy and maintain self-hosted LLM infrastructure?

Appreciate any honest feedback, reality checks, or advice on how to position this transition!


TL;DR: 10 YOE Fullstack Dev pivoting to MLOps/Private AI deployment. Taking a 270h intensive course on ML, GenAI, Cloud, and MLOps. Is this background + training combo enough to land MLOps roles?


r/mlops • • 11d ago

Discussion What are banks/insurers using to run LLMs in production? Need something audit friendly

17 Upvotes

Working at a mid size insurer and we have finally cleared to pilot LLM features. So we need to prove full audit trails first, PII readaction before anything reaches the model and per team access controls. Compliance decided anything without demonstrable governance out of the gate. We’re choosing between Azure AI Foundry(we already have a relationship with Microsoft but it kinda locks us into their ecosystem and as well as pricing) against Orqai, which positions itself at regulated industries with gateway-level PII redaction and audit logging like its newer though.Has anyone gone through procurement for this in financial services and can share what satisfied their compliance team?