Compute · storage · virtualization · orchestration

The Hive.Operate itto learn it.

A self-hosted platform spanning GPU inference, virtualization, 100+ TB of ZFS storage, and a nine-node Kubernetes cluster—built to run AI and cloud-native systems on infrastructure I own end to end.

// platform status

Git reconciled

Local models serving

Workloads observable

ZFS pool online

32

GB GPU memory

102

TB ZFS pool

09

k3s nodes

10

CNCF certs

01 — Why it exists

Fluency through operations.

Cloud-native systems only become intuitive when you live with their failure modes. The goal was to build a durable environment where networking, storage, delivery, identity, observability, and application operations could be learned as one connected system. It is a laboratory, but it is operated with production habits: declarative delivery, observable behavior, repeatable recovery, and documented decisions.

02 — The whole platform

[01]

AI compute

Alpha-5 runs private inference and the surrounding chat, retrieval, memory, search, and tool-serving stack on dedicated GPU hardware.

RTX 5090 · vLLM · LibreChat · Open WebUI · MCP

[02]

Virtualize

Proxmox provides the durable substrate for core services, virtual lab machines, automation targets, and selected cluster capacity.

Proxmox VE · QEMU · LXC · ZFS

[03]

Store

TrueNAS centralizes durable data in a 102 TB ZFS pool and serves application data, shared files, photos, and backups.

TrueNAS · ZFS · NFS · SMB

[04]

Orchestrate

A mixed nine-node Linux fleet runs k3s and containerd, turning the hardware estate into a programmable application platform.

k3s · containerd · Ubuntu · Debian

[05]

Deliver

Git is the control plane. Argo CD continuously reconciles platform services and product workloads into the cluster.

Argo CD · Gitea · AWX · Jenkins

[06]

Operate

Network flows, metrics, logs, traces, secrets, and certificates are visible and managed as parts of one system.

Cilium · Prometheus · Loki · Tempo · Vault

03 — System map

One estate. Distinct jobs. Deliberate boundaries.

Access + edge

Tailscale

Caddy + Traefik

Authentik

AdGuard

Foundation

ProxmoxVMs · LXC · lab environments
TrueNAS102 TB ZFS · NFS · durable data

Alpha-5 / AI

RTX 5090 inference

vLLM · LibreChat · Open WebUI · RAG · vector search · MCP

The Hive / k3s

Nine-node platform

GitOps · products · databases · observability · developer tooling

identity + routingcompute + stateworkload platforms

04 — Alpha-5 / Local AI

The machine at the center.

Alpha-5 is where the homelab stops being an infrastructure exercise and becomes an AI platform. A 24-core Intel Core Ultra 9 285K, 64 GB of system memory, fast local NVMe, and an RTX 5090 with 32 GB of VRAM run private model inference and the services around it.

vLLM serves the models. LibreChat and Open WebUI provide human interfaces. RAG, pgvector, MongoDB, Meilisearch, and MCP services supply retrieval, memory, search, and tools. It is a complete local AI workbench—not a GPU waiting for a prompt.

alpha-5 / hardware profile
CPU
Core Ultra 9 285K / 24 cores
GPU
GeForce RTX 5090 / 32 GB
RAM
64 GB
LOCAL
4 TB NVMe
RUNTIME
Ubuntu · Docker · vLLM

05 — What runs here

A platform for platforms.

The estate hosts the tooling needed to build and operate software, plus real product workloads: local AI, a multi-service practitioner platform, internal MCP services, developer environments, databases, home automation, identity, storage, and observability.

  • Alpha-5 local AI stack: RTX 5090 inference with vLLM, Open WebUI, LibreChat, RAG, vector search, and MCP tooling
  • Proxmox virtualization for Kubernetes, Home Assistant, identity, DNS, reverse proxying, data services, and isolated lab environments
  • TrueNAS with a 102 TB ZFS pool serving shared storage, photos, backups, and application data
  • GitOps delivery for 25+ applications with Argo CD
  • Distributed networking and service exposure with Cilium, Hubble, MetalLB, Traefik, and Kong
  • Persistent workloads on Longhorn and NFS-backed storage
  • Metrics, logs, and traces through Prometheus, Grafana, Loki, Tempo, and OpenTelemetry
  • Platform services including Vault, External Secrets, cert-manager, Gitea, AWX, Jenkins, and Coder
  • Ten CNCF certifications: KCNA, KCSA, PCA, OTCA, CGOA, CAPA, CCA, KCA, CBA, and CNPA

Constraints / Evidence

The conditions shaped the system.

  • Mixed commodity hardware with finite power, memory, storage, and GPU capacity
  • No managed cloud control plane: upgrades, recovery, networking, and state remain operator-owned
  • Stateful workloads span local disks, replicated volumes, NFS, and ZFS failure domains
  • Public evidence must prove operation without exposing topology or creating an access path

9

k3s nodes

A heterogeneous cluster operated as one application platform.

25+

GitOps applications

Desired state continuously reconciled through Argo CD.

102 TB

ZFS storage

Durable application data, files, photos, and backups.

32 GB

GPU memory

Private inference served locally from Alpha-5.

The result

The certifications validate the knowledge. The homelab proves I can design, integrate, and operate the whole system.

Talk infrastructure →
read the operating notes →