Kube Builders
@kube
News and links on building infrastructure and Kubernetes clusters curated by the @Learnk8s.io team More K8s news, events, jobs
This tutorial shows how to build a private EKS cluster with zero public API exposure using Terraform It also covers self-hosted OpenVPN as a VPN gateway, NAT masquerade iptables setup, kube-prometheus-stack via internal load balancer, and Route 53 ➤ https://ku.bz/XZ-gvhGdK
This tutorial shows how to expose a self-hosted Kubernetes cluster on Proxmox using a dual HAProxy setup one on the host as an edge gateway and one as an in-cluster ingress controller ➤ https://ku.bz/w3M0Q7fKv
This article explains how to control runaway GPU spend in Kubernetes by adding taints, quotas, labels, Prometheus rules, and admission controls so teams can see who is using expensive GPU workloads and why ➜ https://ku.bz/6cFxnhHGb
VictoriaMetrics MCP Server connects AI tools to VictoriaMetrics APIs and documentation for metric queries, alert analysis, observability debugging, automation, and read-only monitoring workflows ➜ https://ku.bz/Q_chvmrhf
This tutorial explains how to build a Kubernetes operator that automatically manages Synology DSM reverse proxy rules from annotated Services, Ingresses, and Argo CD Applications ➜ https://ku.bz/P015Vw0wn
Migratowl is an AI-powered dependency migration analyzer that upgrades dependencies in isolated sandboxes, runs tests, reads changelogs, explains breakages, and suggests fixes ➜ https://ku.bz/5Ddt6Pjj1
This tutorial shows how to build a high-availability k3s homelab cluster on Proxmox using embedded etcd, kube-vip, Rancher, Traefik, and Ansible automation ➤ https://ku.bz/0_YwB-fCV
SloK lets teams define SLOs as Kubernetes custom resources, generate Prometheus rules, track error budgets, backtest SLO YAML, and view reliability trends in a dashboard ➜ https://ku.bz/XwCd1mhXc
This tutorial shows how to run Hermes Agent on MicroK8s as a stateful gateway, move scheduled briefings into Kubernetes CronJobs, and use PVCs, ConfigMaps, retry jobs, and Helm for a small agent system ➜ https://ku.bz/7ymZBxBMr
This article explains how building a k3s media server with Claude Code exposed both the speed and the limits of AI-first engineering across GitOps, observability, storage tuning, and Kubernetes debugging ➤ https://ku.bz/94Y_G5wtb
KubePlumber tests Kubernetes networking from inside the cluster by checking: - internal DNS, - pod-to-pod traffic, - external DNS, - and bandwidth between nodes ➜ https://ku.bz/nTkFwCt_Y
This case study shows how CoreDNS became an EKS bottleneck under heavy DNS traffic and uses kube-burner test results to explain why NodeLocal DNSCache helped ➜ https://ku.bz/cs-fvstLZ
This article explains why low GPU utilization during LLM inference can be normal by breaking down prefill, decode, memory bandwidth limits, batching, and GKE B200 benchmark numbers ➜ https://ku.bz/zyv99_RgD
This blog post tells how the Render team: - tracked down Kubernetes memory waste caused by many daemonset namespace watches, - fixed config issues, - and freed over 7 TiB of memory across clusters by reducing unnecessary listwatch overhead ➤ https://ku.bz/2vS0QsvjY
This tutorial teaches how to build a local observability stack for AWS EKS logs using Stern for multi-pod tailing, Fluent Bit for log processing with multiline parsing, and Elasticsearch with Kibana for searchable visualization via Docker Compose ➜ https://ku.bz/qK6fTw_SB
This tutorial shows how to serve open source LLMs on Red Hat OpenShift AI with KServe, vLLM, Argo CD, and IBM Fusion using a fully declarative GitOps workflow ➜ https://ku.bz/chQc0HG6n
Goldpinger is a monitoring tool that runs as a DaemonSet and makes calls between pod instances to test connectivity It produces Prometheus metrics for visualization and alerting, while providing a web UI that displays network health and latency ➜ https://ku.bz/WD6zf4B2h
KubeSolo is a single-node Kubernetes distribution optimized for edge, IoT and embedded devices It eliminates clustering and etcd, uses SQLite via Kine, and runs in under 200MB RAM while remaining OCI-compliant and Helm-ready ➤ https://ku.bz/SPpVGdZ5Y
This tutorial shows how to expose a self-hosted Kubernetes cluster on Proxmox using a dual HAProxy setup one on the host as an edge gateway and one as an in-cluster ingress controller ➜ https://ku.bz/w3M0Q7fKv
KubeVPN connects your local machine to a Kubernetes cluster network so you can reach pods and services by name and proxy inbound traffic with service mesh header routing ➤ https://ku.bz/LyJBd4yTF
This article explains how a single-node OpenShift cluster can be turned into a multi-tenant on-premises GPU platform with reservation-based scheduling, MIG partitioning, time slicing, isolated namespaces, and controller-driven self-healing ➤ https://ku.bz/vX9_vc11q
This tutorial shows how to build a high-availability k3s homelab cluster on Proxmox using embedded etcd, kube-vip, Rancher, Traefik, and Ansible automation ➜ https://ku.bz/0_YwB-fCV
This article explains how GKE 1.33 and 1.34 improve node auto provisioning with ComputeClasses, workload specific scaling, parallel node pool creation, and smarter consolidation for targeted autoscaling ➤ https://ku.bz/8Nr9gVXLS
This article explains how building a k3s media server with Claude Code exposed both the speed and the limits of AI-first engineering across GitOps, observability, storage tuning, and Kubernetes debugging ➜ https://ku.bz/94Y_G5wtb
This case study explains how cURL 65 errors and DNS resolution failures on AWS EKS were caused by Linux kernel network limits being exceeded, resolved by increasing `netdev_budget`, `netdev_budget_usecs`, and `netdev_max_backlog` parameters ➤ https://ku.bz/VMSf7zX6P
This article explains how Netflix traced severe container launch slowdowns to Linux mount lock contention, image layer mount storms, and CPU architecture differences while scaling containers on modern Kubernetes infrastructure ➤ https://ku.bz/v1kX9xWXz
Pyrra is a Kubernetes operator that helps you make SLOs with Prometheus manageable, accessible, and easy to use for everyone ➤ https://ku.bz/p6KblzyDW
This blog post tells how the Render team: - tracked down Kubernetes memory waste caused by many daemonset namespace watches, - fixed config issues, - and freed over 7 TiB of memory across clusters by reducing unnecessary listwatch overhead ➜ https://ku.bz/2vS0QsvjY
This article explains how Kubeshark provides packet-level visibility in Kubernetes by capturing live pod traffic, decoding protocols such as HTTP and gRPC, and mapping requests back to workloads for debugging ➤ https://ku.bz/Sg1y678cP
This tool benchmarks Kubernetes log collectors by measuring throughput, CPU, memory, and log loss with a built-in verifier across agents like Vector, Fluent Bit, OpenTelemetry Collector, and Grafana Alloy ➤ https://ku.bz/RzQz-MqdK