open-speech-serve

Serving speech models well is harder than running a single transcription. open-speech-serve is a Docker-first toolkit for benchmarking Whisper across HF Transformers, faster-whisper, vLLM, SGLang, and TensorRT-LLM/Triton — load tests, concurrency, and streaming latency (TTFS).