Enterprise production
Python automation, Kubernetes integrations, observability migrations, and distributed-system diagnosis across complex customer environments.
Platform engineering / AI infrastructure / operations
Enterprise production experience, internal developer platform delivery, and a privately operated cloud-native estate converge in one working style: make the platform legible, reusable, observable, and recoverable.
10+
years in production systems
9
k3s nodes operated
25+
GitOps applications
10
CNCF certifications
01 — Why the fit is real
Python automation, Kubernetes integrations, observability migrations, and distributed-system diagnosis across complex customer environments.
Hands-on adoption work across Kubernetes, CI/CD, identity, APIs, and cloud infrastructure for enterprise internal developer platforms.
GitOps, networking, identity, secrets, state, metrics, logs, traces, recovery, and upgrades operated as one connected system.
Private vLLM inference, GPU capacity, retrieval, memory, agent tools, and the telemetry required to operate them safely.
02 — Direct proof
System
Nine-node k3s, Argo CD, Cilium, Longhorn, 102 TB ZFS, and a full observability plane.
Inspect →
Architecture
Approximately fourteen services exposing durable operational context through MCP over SSE.
Inspect →
Operations
Real incidents and near misses converted into causal chains, recovery gates, and stronger platform controls.
Inspect →
Evidence
A public-safe snapshot of real platform state without creating an access path.
Inspect →
03 — Interview walkthrough
Ask me to trace a failure across Kubernetes, identity, networking, storage, observability, and an AI workload. I can show the architecture, the evidence path, the recovery decision, and what changed afterward.
01
Establish impact before selecting a familiar layer to blame.
02
Trace network flow, workload state, storage, identity, and telemetry as one request path.
03
Separate containment from root correction and make the recovery path executable.
04
Leave behind a guardrail, alert, runbook, or platform capability that reduces repeat work.
The short version
I do not stop at understanding the system. I ship the capability and stay close enough to learn whether it worked.
Download matching résumé ↓