Platform & Cluster Engineering
From bare metal to managed cloud — Talos Linux, Red Hat OpenShift, AWS EKS, Azure AKS, and
local k3d/k3s. CNI with Cilium, storage classes with Rook Ceph, certificate automation with
cert-manager, and time sync with PTP. Bootstrapped, declarative, reproducible.
GitOps Delivery with Flux
Git is the single source of truth. Flux source, kustomize, helm, and image-automation controllers
reconcile the cluster to the repo — pulling new images from ECR/ACR, writing tags back to Git,
and self-healing drift. No kubectl apply from a laptop, ever.
CI/CD & Environment Promotion
Commit-to-pod pipelines that build, test, and package OCI images, then promote through dev →
staging → production via branch-mapped Kustomize overlays. “Works in dev” becomes
“behaves identically in prod” because both come from the same declarative source.
Observability & Day-2 Ops
Grafana, Prometheus, and Loki/Alloy wired across every pod so warnings and errors surface to a
single operator view. Runbooks for NATS and Postgres, failover drills, and the unglamorous work
of keeping a system up — including chasing down OOM and log-shipping bugs before they bite.
Stateful Workloads & Data
The hard part most teams avoid on Kubernetes: high-availability PostgreSQL, NATS JetStream with
persistent storage, Valkey, and Rook Ceph for a real storage class. Helm-managed, Flux-deployed,
failover-tested, with automated Flyway schema migration on push.
Autoscaling & Distributed Compute
KEDA event-driven autoscaling from queue depth, Ray for parallelizing ML across the cluster, and
replica strategies for high availability. Scale-to-zero when idle, scale to hundreds of pods when
a billion-calculation job lands.