Serving LLMs Isn't Serving Web Apps: A Production Kubernetes Playbook
How AI infrastructure evolved from raw GPUs to vLLM, LLM-D, and AI gateways — and how a self-hosted model actually wires into K Gateway on production Kubernetes.
3 posts
How AI infrastructure evolved from raw GPUs to vLLM, LLM-D, and AI gateways — and how a self-hosted model actually wires into K Gateway on production Kubernetes.
Moving past IQ benchmarks to understand the three axes that actually matter when deploying AI agents: Capability, Autonomy, and Authority.
How we solved the multi-tenant code execution problem for autonomous AI agents using Kubernetes, gVisor, and dynamic volume isolation.