PhyseaWiki How AI actually works physea.ai →

Serving & runtimes

Which runtime serves a model to many users at once?

vLLM is built for serving, not just personal use. Techniques like PagedAttention and continuous batching let a single model handle many requests at once, and it exposes an OpenAI-compatible API.

Last updated 2026-07-25 · Physea Labs

The first three tools are aimed at one person, or a few, using a model on their own machine. vLLM is aimed at the next step: serving a model to many users at the same time. Its own description is “a fast and easy-to-use library for LLM inference and serving,” and the README highlights state-of-the-art serving throughput.[1]

Two ideas do most of the work. PagedAttention borrows directly from a much older idea: virtual memory and paging in operating systems. Instead of reserving one contiguous chunk of memory per request sized for the worst case, it splits each request’s KV cache into fixed-size blocks that can sit anywhere in memory and get allocated on demand as the reply grows, the way an OS pages memory in chunks rather than reserving a program’s full working set up front.[2] The vLLM team measured that older, contiguous-allocation systems waste 60 to 80% of memory to fragmentation and over-reservation; block-based allocation cuts that down to near-optimal, with waste limited to the last, partly-filled block of each sequence.[2] Continuous batching is the second idea: it lets the server weave many in-flight requests together instead of handling them strictly one after another. The README lists both as key features, alongside chunked prefill and prefix caching.[1] The practical result is that one machine can answer far more requests per second than it could by running prompts one at a time.

Like the others, vLLM offers an OpenAI-compatible API server, so the same client code works against it.[1] That keeps the swap honest: you can build against Ollama on a laptop and move to vLLM when you need to serve real traffic, without changing how your app talks to the model.

Ollama / LM Studio / llama.cpp – Built for one person– Simple setup– Mostly one request at a time vLLM – Built for many users– PagedAttention + continuous batching– Requests answered concurrently
PagedAttention and continuous batching solve one problem from two directions, memory efficiency and request scheduling, which is why vLLM serves so much more traffic per machine.

A rough rule of thumb. Reach for Ollama or LM Studio when it is just you. Reach for vLLM when a model needs to sit behind a service and answer for a crowd.

References

  1. vLLM README — vLLM project
  2. vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention — vLLM Blog