Software Architect with a networking and data center background (CCNP/CCDP), building hands-on expertise in GPU infrastructure.
Most people in AI infrastructure come from machine learning or software. I came from networks and physical data centers — including a 50-rack, 200 kW facility I designed and delivered. At cluster scale the bottleneck is rarely the kernel; it's the fabric, the scheduler, or the topology.
Currently working through: inference serving (vLLM, batching, KV-cache), multi-GPU scaling and NCCL collectives, GPU orchestration (Slurm vs Kubernetes, MIG, multi-tenancy), and infrastructure economics (cost per token, MFU, TCO).
Repos below are working notebooks: real benchmarks, real numbers, and what didn't work.
