Sam Ortega

Making infrastructure boring,
one rollout at a time.

edu
Northfield Tech '24
prev
Halcyon · platform
oss
maintainer · kube-shim
cert
CKA · 2025

platform·sre·kubernetes·ai infra

Mar 2025 — Now

Halcyon — Platform Engineer

Own the Kubernetes platform for forty product teams: cluster upgrades as Argo Workflows, GPU node pools with time-slicing, and the on-call rotation that keeps it honest. Cut p99 deploy time from 40 to 9 minutes.

Jan 2024 — Now

kube-shim — Maintainer · open source

A small scheduler extender that lets batch jobs borrow idle GPU capacity. Reviews, releases, and the occasional 3 a.m. bug from a timezone I have never visited.

Jun — Dec 2023

Northwind Analytics — Site Reliability Engineer, intern

Rebuilt the alerting stack on Prometheus and Alertmanager, wrote the runbooks nobody had, and got the pager down from 30 pages a week to 4.

Rollout Sentinelgo · kubernetes · prometheus · operator sdkA Kubernetes operator that watches Prometheus during rollouts and pauses the ones that regress — metrics-driven deployments without a service mesh.Tidewatchkeda · python · fastapi · kubernetes · prophetA predictive autoscaler that forecasts demand from history and scales ahead of the wave with KEDA, so the pods are warm before the traffic arrives.Glasshousekubernetes · kserve · vllm · argo cd · helmA production-shaped LLM inference platform — KServe, vLLM and llm-d behind Envoy Gateway, deployed with Argo CD and Helm, built to learn where serving actually breaks.Forgeargo workflows · kubernetes · aws · go · gitopsAutomated provisioning of isolated e-commerce stores with Argo Workflows — a namespace, database, cache, catalogue and ingress per tenant, from one YAML file to a first order in under seven minutes.
scheduler: respect topology spread for GPU podsexample/kube-shimstore: hedge slow object-storage readsexample/metrics-storefix: cascade delete pipeline versionsexample/pipelinesci: run e2e on /ok-to-test labelexample/trainerdocs: respect system colour scheme on first loadexample/docs
Writing with MDX: callouts, tabs, and snippets
Sharing a GPU without regret
How the scheduler sees a GPU
Why GPUs change the bottleneck
A Kubernetes operator in a weekend
Talk: tail latency and hedged requests