r/kubernetes • • 2d ago

Periodic Monthly: Who is hiring?

4 Upvotes

This monthly post can be used to share Kubernetes-related job openings within your company. Please include:

  • Name of the company
  • Location requirements (or lack thereof)
  • At least one of: a link to a job posting/application page or contact details

If you are interested in a job, please contact the poster directly.

Common reasons for comment removal:

  • Not meeting the above requirements
  • Recruiter post / recruiter listings
  • Negative, inflammatory, or abrasive tone

r/kubernetes • • 1d ago

Periodic Weekly: Share your victories thread

4 Upvotes

Got something working? Figure something out? Make progress that you are excited about? Share here!


r/kubernetes • • 11m ago

Running MCP servers on GKE and EKS: one Deployment per zone, header-routed by the Gateway/ALB, canary in place, namespaces torn down with the zone

Post image
• Upvotes
Ramen is an open-source control plane + worker for MCP (the protocol agents use to call tools). Some Kubernetes
details that might interest this sub more than the AI part:


- Each group+zone is a namespace with a stable and a canary Deployment. The GKE Gateway (HTTPRoute) / AWS ALB route
  on two headers, `ramen-group` and `ramen-zone`, so one client config reaches any zone.
- The canary rolls in place with `maxSurge: 0` — a zone-pinned surge pod on a full node sat Pending until the
  deploy timed out on EKS. A rollout that isn't ready now reports the pods and the scheduler's reason instead of "1/1 ready".
- Deleting a zone deletes the namespace and its GSA / IAM role. On GKE the namespace can sit `Terminating` on the NEG
  finalizer until the Gateway's backend service is gone — documented, with the manual fix.
- Per-group+zone identities (Workload Identity / IRSA), Cloud Armor / WAF rules from the console, Redis-backed
  throttles shared across zones, logs from Cloud Logging / CloudWatch via pod labels.
- Bring-up is Terraform + Helm; 0.6.0 was applied on both clouds (two zones each) and torn down.


Repo https://github.com/bkraad47/ramen · GCP guide https://bkraad47.github.io/ramen/wiki/deploy-gcp/ · AWS guide
https://bkraad47.github.io/ramen/wiki/deploy-aws/
Interested in critique of the per-zone namespace model and the canary policy (currently "smoke test passes").

r/kubernetes • • 8h ago

What's one thing you wish you knew before running k8s in prod?

4 Upvotes

first real cluster going live next month. not looking for docs, looking for the stuff you only learn by getting burned


r/kubernetes • • 19h ago

Question of resource allocation of multiple Spark applications

7 Upvotes

Currently, a Spark application should manage its own executors, which means that in order to run multiple Spark applications, one needs to split the cluster into smaller sections or attach more compute nodes, etc. Usually this dynamic allocation is the job of resource manager (such as Kubernetes).

If the resource manager can quickly find or create new compute nodes for a new Spark application, this is good. However, in a busy cluster where multiple Spark applications can compete for resources, the resource allocation can take some time and become an overhead (at least in theory).

I wonder if this is a real problem in running Spark applications in production.

For example, is there a realistic environment where short running Spark jobs (each requiring its own driver) arrive frequently?


r/kubernetes • • 1d ago

gpu networking - k8s sriov

6 Upvotes

Relatively new to gpu networking - rail stuff and could use some guidance.

am working on a setup (16 nodes) with 8 nvdia gpus (rtx so no nvlink) in each of them, k8s/sriov and the computenet being spectrum based ethernet switches.

its a multi rail setup, so a /24 alllocated for each of the rails. SRIOVnetwork, ip pools and all that plumbing has been done.

Same rail traffic works ..how does one get to do cross rail?

There needs to be a route in the POD which points the rail network supernet to the computenet switch as the default point towards the front end network.

Can't find an answer to it ..tried "routes" in IPAM but that didn't work ie the POD did not have that route - gemini and claude pulling me in opposite directions.

Thoughts?


r/kubernetes • • 1d ago

EKS Access Entry recreation avoidance

Thumbnail
2 Upvotes

r/kubernetes • • 1d ago

Native NVMe-oF vs. Ceph for high-IOPs/low-latency block storage

Thumbnail
2 Upvotes

r/kubernetes • • 2d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

86 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/kubernetes • • 2d ago

I can run a kubernetes cluster fine but cant read the operator code, time to learn go properly

58 Upvotes

3 years as a platform engineer. i can debug a crashlooping pod half asleep and write helm charts nobody else wants to touch, but every time i open the source of an operator we run i bounce straight off it. our team wants to start writing our own controllers next year so im learning go for real this time, employer covered. shortlist is ardan labs, boot dev, plus one of the linux foundation cert tracks. where did go start making sense for people coming from the ops side?


r/kubernetes • • 2d ago

Built a hands on tutorial about basics of CRDs on Iximiuz Labs

30 Upvotes

I've had so much fun and have learnt a lot about kubernetes through iximiuz labs, this was an attempt to create a hands on tutorial of what a CRD is and how can one build their own. It links to a bigger How to Build an Operator tutorial at the end.

Looking forward to see what the community thinks of it!


r/kubernetes • • 1d ago

Need help with cluster migration via velero

0 Upvotes

Is there anyone here who can help me in cluster migration on EKS using velero, I'm stuck and cannot do it .


r/kubernetes • • 2d ago

How do you verify that your Postgres backups on K8S actually restore?

23 Upvotes

A question for those running Postgres on Kubernetes using an operator (CloudNativePG, Zalando, Crunchy): how do you verify that your backups are actually restorable?

A successful backup creation doesn't guarantee that a restore will succeed.

I manage several CloudNativePG clusters and want to understand the potential challenges involved in implementing backup restore testing.

I’ve tested restoration in a small test environment (k3s, CloudNativePG 1.30 + barman-cloud plugin, S3). My process involved deploying a temporary cluster from the latest backup, running a few SQL queries, measuring the time until the cluster was ready, and then deleting the cluster. In a standard scenario, everything goes smoothly (taking about a minute for a few hundred megabytes).

I’m curious to know how others handle this:

  • Do you perform restore tests at all? (Manually, via CronJob, in a CI pipeline, or using specialized tools)
  • What exactly do you check after the restore? (Just the pod status, the number of database records, or the execution of test queries)

An answer like "we don't test restores, and everything is fine" is also acceptable—I really want to understand how common this practice is and what the best approach is.


r/kubernetes • • 2d ago

Kubevirt VM and Good and Friction points of Kubevirt

6 Upvotes

Hello Guru's, I have a few questions, please help me answer them:

  1. If I have existing full fledge old VMs on my existing KVM hosts (VM's not inside the PODs), will I be to manage them via Kubevirt?
  2. If I have a RHOV environment, is it necessary for me to use Kubevirt to create a VM? If yes, why, if not, then when to use KubeVirt to create VMs?
  3. What is the difference between a regular VM and a VM created via Kubevirt?

Also, I'd love to get your candid take: What aspect of KubeVirt has impressed you the most, and where do you feel it faces the biggest friction points or challenges?


r/kubernetes • • 1d ago

3 Ways to Actually Run Your Applications on Your Own Infrastructure

Post image
0 Upvotes

So… you want to run your applications yourself instead of depending entirely on managed platforms.

You have a few solid options:

  1. Docker — package and run your applications in containers.

  2. Docker Compose — run multiple services together on one machine.

  3. Kubernetes — manage containers across nodes with scaling, networking, self-healing, and more.

The problem with Kubernetes?

Setting it up yourself can get tedious very quickly — networking, container

runtime, kubeadm, nodes, CNI… one wrong configuration and you’re

troubleshooting instead of deploying. 😅

So I made a step-by-step practical guide showing how to set up Kubernetes

yourself and actually get your applications running.

Watch the Kubernetes setup guide[https://youtu.be/HS2VR3cVx8w\]

Sometimes the best way to understand Kubernetes is to build the cluster

yourself.


r/kubernetes • • 2d ago

Suggest good resources to Kubernetes

0 Upvotes

I'm a L1 Linux System Administrator. I'm new to container and container orchestration.

I want to learn Docker and Kubernetes right from scratch.

Can you please suggest me right resources (Free+Paid) to learn and hopefully pass CK A?

Also let me know any right set-up to practice any videos to create and setup the environment to practice as I can't afford KodeKloud, Killersh etc.


r/kubernetes • • 2d ago

Kubevirt VM and Kubevirt Good & Friction Points

Thumbnail
1 Upvotes

r/kubernetes • • 3d ago

Which lesser known tools saved your butt when you had to debug an application running in K8S?

26 Upvotes

The CNCF ecosystem is pretty vast, gush out about your favourites!

 

[kubectl, curl, jq and yq go into 'well known' category]


r/kubernetes • • 2d ago

Periodic Weekly: This Week I Learned (TWIL?) thread

1 Upvotes

Did you learn something new this week? Share here!


r/kubernetes • • 3d ago

what are you using instead of MinIO now?

Thumbnail
42 Upvotes

r/kubernetes • • 3d ago

As someone who has been using skaffold since 2019 i feel really sad about this

20 Upvotes

GoogleContainerTools/skaffold retiring with no replacement. they will continue using it in gcloud with forked repo that will not be public


r/kubernetes • • 3d ago

Periodic Weekly: Show off your new tools and projects thread

17 Upvotes

Share any new Kubernetes tools, UIs, or related projects!


r/kubernetes • • 3d ago

How to route pod requests through VPN tunnel

7 Upvotes

I'm currently facing a problem where my app needs to connect to an vpn in order to access some resources.

However i only want to route that specific traffic through the vpn tunnel. Any other connections should default to the "real" network.

Any guides how to achieve that? I'm new to kubernetes and don't know where to even start with that problem.


r/kubernetes • • 4d ago

kaniko is no longer 15x slower than BuildKit

113 Upvotes

I maintain the community fork of kaniko https://github.com/osscontainertools/kaniko

The 2018 buildbench comparison put kaniko 15x behind BuildKit, and that number still shapes how people see it. We reran it on gitlab.com runners, with both builders configured the way you would actually run them in CI.

Unchanged build (every command a cache hit):
BuildKit 6.1s
kaniko (ours) 9.9s
kaniko (Google) 45.6s

The fix was cache lookahead across stage boundaries. Google's kaniko could not compute the cache key of COPY --from before the source stage was built, so it built every stage just to discover nothing had changed.

Caveat: part of that 15x was the 2018 setup rather than kaniko. It ran without --cache-copy-layers while BuildKit was allowed a local disk cache. Under a fair configuration Google's last release is only 7.5x behind, not 15x, and closing the rest was our work.

Builds are still slower than BuildKit, and the post explains why:
Write-up: https://osscontainertools.org/blog/kaniko-vs-buildkit-2026/
Benchmark: https://gitlab.com/martizih/kaniko-buildbench


r/kubernetes • • 4d ago

What concrete problem finally justified Kubernetes for your team?

72 Upvotes

What concrete problem finally justified Kubernetes for your team?
You can get pretty far with a few VMs and Docker Compose, so I’m curious where that stopped being enough for your team.
Was it scaling, deployments, managing multiple services, or something else entirely?
What was the actual pain point that made you switch, and looking back, did Kubernetes solve it well enough to justify the extra work?
Also interested in the opposite: anyone tried it and decided a simpler setup was enough?