Fleet Platform is an end-to-end, fully reproducible private cloud platform built on a 3-node Proxmox cluster (40 cores, 157 GB RAM).
Every layer—from virtual machine provisioning to cluster networking, GitOps controllers, ingress, and wildcard TLS—is declared in code and version-controlled.
- Cattle, Not Pets: The real test of reproducibility is a full teardown. The entire platform—VMs, cluster, GitOps engine, ingress, and wildcard TLS—is continuously verified by rebuilding from empty hypervisors via a single automated script to guarantee zero configuration drift.
- The Endgame: A parallel headless Gazebo robotics simulation farm driven by Argo Workflows (50 deterministic, seeded runs per commit yielding pass/fail regression reports) backed by dedicated GPU acceleration.
Architecture & Build Roadmap
Below is the live tracker for the platform build. Entries are marked with their publication status and link directly to technical write-ups:
| Phase | Milestone / Topic | Focus Areas | Status |
|---|---|---|---|
| Phase 0 | Cluster Foundation | 2-node k3s bootstrap, Proxmox hypervisors, cloud-init templates | 📝 Backfilling write-up |
| Phase 1 | GitOps & Ingress | Ansible AAP bootstrap, HAProxy ingress, Argo CD app-of-apps | 📝 Backfilling write-up |
| Phase 2.1 | TLS & Ingress Security | cert-manager, Cloudflare DNS-01, rate limits, restore race | Read Post → |
| Phase 2.2 | Isolated GitOps Engine | Self-hosted Forgejo server provisioned as code, GitHub mirror | 🔨 In Progress |
| Phase 2.3 | AWS Cloud Footprint | Terraform S3 remote state w/ locking, IAM least-privilege, VPC | 🔨 In Progress |
| Phase 3 | Platform Observability | kube-prometheus-stack via Argo, Grafana, HAProxy ServiceMonitors | ⏳ Planned |
| Phase 4 | Synthetic Workloads & k6 | CPU-burner app, HPA, external AWS EC2 k6 load generation | ⏳ Planned |
| Phase 5 | Zero-Trust Secrets | SOPS + age encryption, gitleaks CI audit, OpenBao / ESO | ⏳ Planned |
| Phase 6 | Cloud-Native Storage | Longhorn distributed storage, MinIO S3-compatible backup tier | ⏳ Planned |
| Phase 7 | GPU Node Integration | Passthrough, nvidia-smi batch jobs, headless Vulkan/EGL render | ⏳ Planned |
| Phase 8 | Robotics Networking | ROS 2 DDS discovery characterization, Zenoh bridge benchmarks | ⏳ Planned |
| Phase 9 | Sim Farm: Gazebo + Argo | 50 seeded parallel runs, deterministic physics, pass/fail report | 🎯 Flagship Capstone |
Published Engineering Logs
Phase 2, Part 1: cert-manager, Rate Limits, and the Restore Race
Note on Build Chronology: This devlog was started mid-flight during the cert-manager & ingress milestone. Foundational build logs for Phase 0 (two-node k3s bootstrap) and Phase 1 (Proxmox/Terraform/Ansible) are currently being backfilled.
cert-manager
Setting up cert-manager was pretty easy with a Helm chart. I did spend a lot of time on this step but it was mostly googling and figuring out what to do about bootstrap secrets which I detail below. The most interesting thing I learned was part of the values file, specifically installCRDs (or crds.enabled, depending on chart version). CRDs, Custom Resource Definitions, are objects that teach the Kubernetes API server about new kinds of resources, like Certificate and ClusterIssuer. They’re not backed by Go directly, they’re schema definitions, and a controller (cert-manager’s own pod, running Go code) is what actually watches for objects of that kind and acts on them. So a CRD without its controller running is just an inert schema. The controller is the operator, the CRD is the vocabulary it teaches the cluster.