How GPU Platform Operations Changed My View of Kubernetes
Why GPU platform operations make Kubernetes a junction of resource allocation, policy, observability, user experience, and operator judgment.
Read →~/brilly/en
Open source, systems, talks, and things learned the hard way.

Why GPU platform operations make Kubernetes a junction of resource allocation, policy, observability, user experience, and operator judgment.
Read →A real re-analysis of a resolved CrashLoopBackOff incident, including missing evidence, collector failures, and operator review.
Read →How a Go backend, FastAPI agent, React frontend, and seven read-only collectors form a reviewable GPU-platform investigation system.
Read →A prototype that connects scattered vendor support cases and technical documents to an operator-verifiable RCA workflow.
Read →What became clearer—and what still needs work—when explaining KubeRCA's ReAct investigation and read-only guardrails to a global audience.
Read →A short note on what Brilly will cover: open source, talks, engineering lessons, and useful traces from the work.
Read →