Inference Pipeline Optimization
Notes on getting more throughput out of the same GPUs by keeping requests batched and reusing cached attention state instead of recomputing it per request. Full write-up and benchmark in progress — see the Inference Engineering research track for the current state of this work.