← Back to all projectsAgentic AI & LLM Systems

Low-Latency LLM Serving Benchmark

Reference architecture and reproducible benchmark harness for low-latency LLM serving: a CPU baseline streaming server, real quantization and concurrency measurements, Kubernetes autoscaling manifests, and a documented vLLM/continuous-batching production path.

LLM ServingvLLMQuantizationKubernetes
View source on GitHub