Ecosystem:Frank Chinembiri/Techarvest Projects/Mateere Lab (Showcase)
● LIVE TELEMETRY
Governance & Decision Log

Architecture Decision Records (ADRs)

Formal engineering records documenting key architectural decisions made during the design, deployment, and operational hardening of Cloud Lab.

ADR-001

Existing Prometheus Server as Authoritative Telemetry Engine

Accepted
Context & Problem Statement

The cluster requires continuous health, CPU, memory, load, and container resource metrics across 4 physical nodes. We considered installing multiple custom telemetry collectors or querying K3s Metrics Server directly.

Architectural Decision

Use the existing Prometheus server (monitoring/prometheus-server) as the single authoritative telemetry source. A lightweight FastAPI adapter issues targeted PromQL queries with in-memory TTL caching.

Consequences & Trade-offs

Eliminates redundant monitoring agents, protects node CPU/disk I/O, and provides high-precision PromQL rates while decoupling UI clients from raw PromQL complexity.

ADR-002

CloudPanel Non-Invasive Monitoring Strategy

Accepted
Context & Problem Statement

CloudPanel hosts ~30 commercial virtual hosts on Node 1 (node-1-control). We needed full availability, SSL expiry countdown, and error rate tracking without interfering with production PHP/Node processes.

Architectural Decision

Combine Uptime Kuma external probes for availability and SSL certificate countdowns with a high-performance Python log analyzer parsing /home/*/logs/nginx/access.log directly.

Consequences & Trade-offs

Zero agent footprint on host virtual hosts. Strips synthetic healthcheck pings automatically, providing organic traffic figures and SSL expiry alerts.

ADR-003

Unified Server-Side Log Analytics vs. Client-Side Analytics

Accepted
Context & Problem Statement

We needed accurate traffic analytics across both CloudPanel websites and Kubernetes microservices without privacy violations, cookies, or ad-blocker drop-offs.

Architectural Decision

Implement server-side log analysis as the primary traffic tracking engine, supplemented by Umami where client-side journey tracking is explicitly required.

Consequences & Trade-offs

Captures 100% of API consumers, mobile clients, and web visits accurately with zero browser performance overhead and strict IP hashing.

ADR-004

Homepage Declarative Widget Boundary

Accepted
Context & Problem Statement

Homepage supports native integrations for tools like Prometheus, ArgoCD, and Gitea, but lacks built-in cards for multi-node hardware grids, cross-platform traffic, and live sports.

Architectural Decision

Strictly use Homepage native widgets where mature integrations exist (Prometheus, Uptime Kuma, ArgoCD, Gitea). Custom Python endpoints via FastAPI are exclusively used for multi-node vitals, unified traffic tables, sports tickers, and curated RSS news.

Consequences & Trade-offs

Minimizes custom code to under 500 lines of Python while maximizing dashboard density and reliability.

ADR-005

Public Showcase Air-Gapped SSG Deployment Model

Accepted
Context & Problem Statement

The public showcase (lab.techarvest.co.zw) must showcase live platform engineering achievements to hiring managers and recruiters without creating security risks for internal cluster networks.

Architectural Decision

Deploy the showcase as a Next.js Static Site Generation (SSG) application served by standard NGINX. A scheduled snapshot pipeline exports sanitized data from FastAPI into static assets, with zero runtime cluster network connectivity.

Consequences & Trade-offs

Complete air-gap security guarantee. Even under massive public DDoS or complete web container compromise, internal cluster networks and databases are completely unreachable.