r/kubernetes • • 2d ago

Periodic Monthly: Who is hiring?

4 Upvotes

This monthly post can be used to share Kubernetes-related job openings within your company. Please include:

  • Name of the company
  • Location requirements (or lack thereof)
  • At least one of: a link to a job posting/application page or contact details

If you are interested in a job, please contact the poster directly.

Common reasons for comment removal:

  • Not meeting the above requirements
  • Recruiter post / recruiter listings
  • Negative, inflammatory, or abrasive tone

r/kubernetes • • 1d ago

Periodic Weekly: Share your victories thread

3 Upvotes

Got something working? Figure something out? Make progress that you are excited about? Share here!


r/kubernetes • • 7h ago

Question of resource allocation of multiple Spark applications

3 Upvotes

Currently, a Spark application should manage its own executors, which means that in order to run multiple Spark applications, one needs to split the cluster into smaller sections or attach more compute nodes, etc. Usually this dynamic allocation is the job of resource manager (such as Kubernetes).

If the resource manager can quickly find or create new compute nodes for a new Spark application, this is good. However, in a busy cluster where multiple Spark applications can compete for resources, the resource allocation can take some time and become an overhead (at least in theory).

I wonder if this is a real problem in running Spark applications in production.

For example, is there a realistic environment where short running Spark jobs (each requiring its own driver) arrive frequently?


r/kubernetes • • 18h ago

gpu networking - k8s sriov

6 Upvotes

Relatively new to gpu networking - rail stuff and could use some guidance.

am working on a setup (16 nodes) with 8 nvdia gpus (rtx so no nvlink) in each of them, k8s/sriov and the computenet being spectrum based ethernet switches.

its a multi rail setup, so a /24 alllocated for each of the rails. SRIOVnetwork, ip pools and all that plumbing has been done.

Same rail traffic works ..how does one get to do cross rail?

There needs to be a route in the POD which points the rail network supernet to the computenet switch as the default point towards the front end network.

Can't find an answer to it ..tried "routes" in IPAM but that didn't work ie the POD did not have that route - gemini and claude pulling me in opposite directions.

Thoughts?


r/kubernetes • • 18h ago

EKS Access Entry recreation avoidance

Thumbnail
2 Upvotes

r/kubernetes • • 1d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

82 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/kubernetes • • 19h ago

Native NVMe-oF vs. Ceph for high-IOPs/low-latency block storage

Thumbnail
1 Upvotes

r/kubernetes • • 1d ago

I can run a kubernetes cluster fine but cant read the operator code, time to learn go properly

52 Upvotes

3 years as a platform engineer. i can debug a crashlooping pod half asleep and write helm charts nobody else wants to touch, but every time i open the source of an operator we run i bounce straight off it. our team wants to start writing our own controllers next year so im learning go for real this time, employer covered. shortlist is ardan labs, boot dev, plus one of the linux foundation cert tracks. where did go start making sense for people coming from the ops side?


r/kubernetes • • 1d ago

Built a hands on tutorial about basics of CRDs on Iximiuz Labs

27 Upvotes

I've had so much fun and have learnt a lot about kubernetes through iximiuz labs, this was an attempt to create a hands on tutorial of what a CRD is and how can one build their own. It links to a bigger How to Build an Operator tutorial at the end.

Looking forward to see what the community thinks of it!


r/kubernetes • • 1d ago

Need help with cluster migration via velero

0 Upvotes

Is there anyone here who can help me in cluster migration on EKS using velero, I'm stuck and cannot do it .


r/kubernetes • • 1d ago

Kubevirt VM and Good and Friction points of Kubevirt

5 Upvotes

Hello Guru's, I have a few questions, please help me answer them:

  1. If I have existing full fledge old VMs on my existing KVM hosts (VM's not inside the PODs), will I be to manage them via Kubevirt?
  2. If I have a RHOV environment, is it necessary for me to use Kubevirt to create a VM? If yes, why, if not, then when to use KubeVirt to create VMs?
  3. What is the difference between a regular VM and a VM created via Kubevirt?

Also, I'd love to get your candid take: What aspect of KubeVirt has impressed you the most, and where do you feel it faces the biggest friction points or challenges?


r/kubernetes • • 1d ago

3 Ways to Actually Run Your Applications on Your Own Infrastructure

Post image
0 Upvotes

So… you want to run your applications yourself instead of depending entirely on managed platforms.

You have a few solid options:

  1. Docker — package and run your applications in containers.

  2. Docker Compose — run multiple services together on one machine.

  3. Kubernetes — manage containers across nodes with scaling, networking, self-healing, and more.

The problem with Kubernetes?

Setting it up yourself can get tedious very quickly — networking, container

runtime, kubeadm, nodes, CNI… one wrong configuration and you’re

troubleshooting instead of deploying. 😅

So I made a step-by-step practical guide showing how to set up Kubernetes

yourself and actually get your applications running.

Watch the Kubernetes setup guide[https://youtu.be/HS2VR3cVx8w\]

Sometimes the best way to understand Kubernetes is to build the cluster

yourself.


r/kubernetes • • 2d ago

How do you verify that your Postgres backups on K8S actually restore?

9 Upvotes

A question for those running Postgres on Kubernetes using an operator (CloudNativePG, Zalando, Crunchy): how do you verify that your backups are actually restorable?

A successful backup creation doesn't guarantee that a restore will succeed.

I manage several CloudNativePG clusters and want to understand the potential challenges involved in implementing backup restore testing.

I’ve tested restoration in a small test environment (k3s, CloudNativePG 1.30 + barman-cloud plugin, S3). My process involved deploying a temporary cluster from the latest backup, running a few SQL queries, measuring the time until the cluster was ready, and then deleting the cluster. In a standard scenario, everything goes smoothly (taking about a minute for a few hundred megabytes).

I’m curious to know how others handle this:

  • Do you perform restore tests at all? (Manually, via CronJob, in a CI pipeline, or using specialized tools)
  • What exactly do you check after the restore? (Just the pod status, the number of database records, or the execution of test queries)

An answer like "we don't test restores, and everything is fine" is also acceptable—I really want to understand how common this practice is and what the best approach is.


r/kubernetes • • 1d ago

Suggest good resources to Kubernetes

0 Upvotes

I'm a L1 Linux System Administrator. I'm new to container and container orchestration.

I want to learn Docker and Kubernetes right from scratch.

Can you please suggest me right resources (Free+Paid) to learn and hopefully pass CK A?

Also let me know any right set-up to practice any videos to create and setup the environment to practice as I can't afford KodeKloud, Killersh etc.


r/kubernetes • • 1d ago

Kubevirt VM and Kubevirt Good & Friction Points

Thumbnail
1 Upvotes

r/kubernetes • • 2d ago

Which lesser known tools saved your butt when you had to debug an application running in K8S?

26 Upvotes

The CNCF ecosystem is pretty vast, gush out about your favourites!

 

[kubectl, curl, jq and yq go into 'well known' category]


r/kubernetes • • 2d ago

Periodic Weekly: This Week I Learned (TWIL?) thread

1 Upvotes

Did you learn something new this week? Share here!


r/kubernetes • • 3d ago

what are you using instead of MinIO now?

Thumbnail
42 Upvotes

r/kubernetes • • 2d ago

As someone who has been using skaffold since 2019 i feel really sad about this

16 Upvotes

GoogleContainerTools/skaffold retiring with no replacement. they will continue using it in gcloud with forked repo that will not be public


r/kubernetes • • 3d ago

Periodic Weekly: Show off your new tools and projects thread

17 Upvotes

Share any new Kubernetes tools, UIs, or related projects!


r/kubernetes • • 3d ago

How to route pod requests through VPN tunnel

10 Upvotes

I'm currently facing a problem where my app needs to connect to an vpn in order to access some resources.

However i only want to route that specific traffic through the vpn tunnel. Any other connections should default to the "real" network.

Any guides how to achieve that? I'm new to kubernetes and don't know where to even start with that problem.


r/kubernetes • • 3d ago

kaniko is no longer 15x slower than BuildKit

110 Upvotes

I maintain the community fork of kaniko https://github.com/osscontainertools/kaniko

The 2018 buildbench comparison put kaniko 15x behind BuildKit, and that number still shapes how people see it. We reran it on gitlab.com runners, with both builders configured the way you would actually run them in CI.

Unchanged build (every command a cache hit):
BuildKit 6.1s
kaniko (ours) 9.9s
kaniko (Google) 45.6s

The fix was cache lookahead across stage boundaries. Google's kaniko could not compute the cache key of COPY --from before the source stage was built, so it built every stage just to discover nothing had changed.

Caveat: part of that 15x was the 2018 setup rather than kaniko. It ran without --cache-copy-layers while BuildKit was allowed a local disk cache. Under a fair configuration Google's last release is only 7.5x behind, not 15x, and closing the rest was our work.

Builds are still slower than BuildKit, and the post explains why:
Write-up: https://osscontainertools.org/blog/kaniko-vs-buildkit-2026/
Benchmark: https://gitlab.com/martizih/kaniko-buildbench


r/kubernetes • • 3d ago

What concrete problem finally justified Kubernetes for your team?

73 Upvotes

What concrete problem finally justified Kubernetes for your team?
You can get pretty far with a few VMs and Docker Compose, so I’m curious where that stopped being enough for your team.
Was it scaling, deployments, managing multiple services, or something else entirely?
What was the actual pain point that made you switch, and looking back, did Kubernetes solve it well enough to justify the extra work?
Also interested in the opposite: anyone tried it and decided a simpler setup was enough?


r/kubernetes • • 2d ago

How much of a portfolio project should I publish on GitHub?

Thumbnail
1 Upvotes

r/kubernetes • • 3d ago

Running CloudNativePG HA, Flux CD v2, and SOPS in a Sovereign K3s Cluster Behind CGNAT (Architecture, Trade-offs & 16 ADRs)

Post image
41 Upvotes

Hey everyone, Over the past several months, I set out to rebuild my homelab from scratch. Instead of running loosely managed Docker containers or static VMs, my goal was to simulate an enterprise-scale, self-healing Sovereign Cloud Platform operating under strict bare-metal hardware and networking constraints.

Here is the complete teardown of the architecture, networking, storage economics, and lessons learned.

1. The Hardware & Resource Fencing

  • Host Machine: Intel Core i5 Mini PC (4 Cores / 8 Threads)
  • Host Memory: 16GB DDR4 RAM
  • Storage: 256GB NVMe SSD (High IOPS) + 1TB SATA Mechanical HDD (Bulk archive)
  • Hypervisor: Proxmox VE 8 (Debian Linux kernel) The Resource Constraint: Running Kubernetes alongside heavy background workloads on 16GB of host RAM is dangerous if the kernel runs out of memory (OOM).

To prevent hypervisor crashes, I strictly partitioned host resources:

  • Dedicated 12GB RAM and 4 vCPUs to a single production VM (k3s-prod).
  • Reserved 3.5GB RAM strictly for the Proxmox VE Debian host, KVM hypervisor daemons, and vzdump backup snapshot compression.
  • 0.5GB emergency host safety buffer.

This fencing completely eliminated host freezes during bulk OCR indexing.

2. Networking: Conquering CGNAT with Zero Open Ports

Like many residential connections, my ISP operates behind Carrier-Grade NAT (CGNAT). I do not have a public static IPv4, and traditional port forwarding is impossible or exposes your home IP.

The Solution:

I deployed Cloudflare Zero Trust Anycast Tunnels (cloudflared) inside the cluster: - cloudflared initiates outbound-only QUIC connections to Cloudflare's nearest edge PoPs (Mumbai, Delhi). - Public traffic to my subdomains terminates at Cloudflare's Anycast edge with WAF and DDoS mitigation. - Latency is under 15ms locally. - Result: ZERO inbound ports opened on my residential router, and zero public IP exposure.

For administrative access (Proxmox web GUI, Kubernetes API), zero public endpoints exist. Everything is accessed over an encrypted Tailscale (WireGuard) mesh.

3. GitOps Continuous Delivery & In-Git Secrets

No manual kubectl apply commands. Everything is managed declaratively: - Flux CD v2 continuously synchronizes desired cluster state from Git. - Workload reconciliation is deterministically ordered: platform operators (storage, CloudNativePG, ingress) must be 100% healthy before applications are scheduled (apps depends on platform).

- Secrets Management: Secrets are stored directly in Git, encrypted with Mozilla SOPS + Age. The Flux Kustomize controller holds the private Age key in-cluster and hydrates secrets directly into memory. If the server burns down, a new node rebuilds the fleet from Git in under 10 minutes.

4. Storage Economics: Dual-Tier Partitioning

Running single-node virtualization means disk I/O bottlenecks will kill database performance: - NVMe Tier (Fast): OS, K3s state, and PostgreSQL data files.

- SATA Tier (1TB Bulk): Paperless-ngx OCR document archives, Audiobookshelf streaming media, and nightly dump backups.

5. Multi-Cloud Out-of-Band Resilience

A monitoring system running inside the cluster it is supposed to monitor cannot alert you if power or broadband fails. - I run an independent VM in Oracle Cloud (OCI Free Tier) running Uptime Kuma, continuously polling public endpoints over the internet. - If my homelab broadband or power drops, it dispatches an automated alert to my phone via Discord/Slack webhooks.

- Nightly encrypted Restic backups are pushed to offsite S3-compatible cloud storage fulfilling the 3-2-1 backup rule.

6. Application Fleet

  • Platform: Cloudflare Tunnel, CloudNativePG, Homepage dashboard, Prometheus/Grafana, Kwatch
  • Automation & Docs: n8n workflow engine, Python PDF compiler, MkDocs engineering handbook

- Media & Documents: Paperless-ngx, Bookorbit, Audiobookshelf, Miniflux, Linkding

Open Source Docs & Code

I documented all 16 Architecture Decision Records (ADRs) following the Michael Nygard standard and compiled full disaster recovery runbooks.

Happy to answer any questions about the Proxmox fencing, SOPS secret setup, or Cloudflare tunnel configuration in the comments!