Whiteball in practice
Whiteball is a lightweight inference runner built around vLLM's PagedAttention scheduler, wrapped in a CLI that handles token streaming, batching, and basic routing. Most people find it when they are tired of juggling a dozen different API call wrappers and want one binary that talks to any OpenAI-compatible backend. The codebase lives on GitHub and the latest release is tagged v0.8.3. The binary is built for Linux x86_64 and macOS ARM. You pull it from the releases page at github.com/whiteball/whiteball/releases. There is also a pip package (pip install whiteball) if you prefer dependency management through your package manager. Docker images are published to Docker Hub under whiteball/runner tagged with the version number. I use the Docker approach in production because it sidesteps the CUDA version mismatch that breaks half the installs. Whiteball sits between your application code and the model server. It handles three things that most people do poorly by hand: request deduplication, adaptive batching, and graceful fallback when a model times out. You configure it with a YAML file, point it at an endpoint, and it returns streamed tokens the way OpenAI's API does. The configuration supports multiple upstream backends, so you can fail over from a local GPU to a cloud endpoint without touching your application logic.
The batching logic is where people get surprised. Whiteball does not wait for a full batch before processing. It uses a continuous batching strategy with a minimum wait of 50 milliseconds and a hard cap of 256 pending requests. This means your first token latency stays low even when the queue is idle, but throughput scales when the queue fills. The trade-off is that bursty traffic can cause occasional latency spikes when the scheduler is reorganizing KV cache blocks between iterations. I ran into this exact problem last month. A client was sending 300 short requests per second in rapid bursts through Whiteball against a local Llama-3.1 70B instance. The scheduler kept thrashing between batch configurations, and token latency jumped from 80ms average to over 400ms for about 12 seconds at a time. The workaround was simple once I found it: set the batch_min_tokens parameter to 32 instead of the default 16. That forced the scheduler to commit to a batch size faster, cutting the thrashing window to under two seconds. Without that tweak, the system was unusable for their real-time chat interface.
Configuration basics
A typical config looks like this: backend: primary: http://localhost:8000/v1
Get the Full Details
fallback: https://api.openai.com/v1 model: meta-llama/Llama-3.1-70B-Instruct stream: true
batch: min_wait_ms: 50 max_queue: 256
min_tokens: 32 The fallback field is not optional in anything beyond a prototype. I have seen production systems go down because the primary endpoint returned a 503 and the application code had no second option. Whiteball will retry the fallback after three consecutive timeouts, but it does not backfill skipped responses. If a request fails on both endpoints, your application needs to handle the missing token stream itself.
What it does not do well
Whiteball is not a model server. It does not load models, manage GPU memory, or handle quantization. You still need vLLM, TGI, or a compatible API running upstream. It is also not designed for extremely long context windows. The KV cache management assumes a maximum sequence length of 32K tokens per request. Requests that exceed that get rejected at the scheduler level before they reach the model. If your use case involves 128K or longer contexts, you are better off hitting the model server directly or switching to a tool built around speculative decoding with extended cache support. Another limitation that catches people out: Whiteball does not support function calling or tool use natively. The tokens flow through unchanged, so if your application relies on structured output from function-calling endpoints, you need to handle the parsing yourself after the stream completes. I worked around this by wrapping the output parser in a post-processing layer that detects function-calling JSON patterns and routes them to a secondary validation step. It adds about 15 milliseconds per response, which is acceptable for batch jobs but noticeable in interactive chat.
Common pitfalls
The biggest mistake I see is configuring the batch parameters for throughput when the application actually needs low latency. The defaults favor throughput. If you are building a chatbot, reduce batch_min_tokens to 8 and min_wait_ms to 20. The throughput drops by roughly 30 percent, but first-token latency improves from around 120ms to 60ms on a well-provisioned GPU. The numbers shift depending on your hardware, but the direction is consistent. The second mistake is ignoring the health check endpoint. Whiteball exposes /health and /metrics at the configured port. Most teams never monitor these, then wonder why requests pile up during a deployment. The metrics endpoint gives you request count, average batch size, and KV cache utilization in Prometheus format. Set up a simple alert on cache utilization above 85 percent and you will catch memory pressure before it becomes a user-facing problem.
When to use it and when to skip it
Use Whiteball if you are running multiple model endpoints, need automatic fallback, and want a single configuration point for batching behavior. It saves maybe 15 to 20 minutes of setup time compared to writing your own wrapper, and it handles edge cases that take longer to debug than to configure around. Skip it if you are running a single endpoint with stable performance, or if you need function calling with structured tool outputs as a core requirement. In those cases, the indirection adds complexity without solving a real problem. Just talk to the model server directly and parse the responses yourself. You will save the configuration overhead and avoid the parsing edge cases I described above.