← Back to all projectsAgentic AI & LLM Systems
Low-Latency LLM Serving Benchmark
Reference architecture and reproducible benchmark harness for low-latency LLM serving: a CPU baseline streaming server, real quantization and concurrency measurements, Kubernetes autoscaling manifests, and a documented vLLM/continuous-batching production path.
LLM ServingvLLMQuantizationKubernetes
View source on GitHub