Every token served.
No page wasted.
PageServe is an LLM inference engine built from scratch: continuous batching over a paged KV cache, prefix caching, preemption and speculative decoding, in readable PyTorch. Every step is one packed forward pass. No vLLM. No generate().
- Continuous batching
- Paged attention
- Prefix caching
- Speculative decoding
- Swap / recompute preemption
iteration
batch
queue
tok/s
kv blocks
28.3×
continuous batching vs HF sequential
27.5 → 779.9 tok/s
2.8×
vs HF static batching
276.6 → 779.9 tok/s
1.9×
prefix caching · shared system prompt
464.6 → 884.9 tok/s
1.5×
n-gram speculation · copy-heavy output
924.6 → 1382.9 tok/s
NVIDIA T4 (Colab) · Qwen2.5-1.5B-Instruct fp16 · offline ablation · output tokens/sFull T4 results →
All the work, one forward pass.
Every scheduler step packs the prefill chunks of new prompts and one decode token per running sequence into a single forward pass. Requests join and leave the batch between steps.
Pick this step’s work
- decodes: 1 token each (+k spec)
- prefill chunks ≤ token budget
- admit if blocks allow (FCFS)
- prefix-cache hits reused
One [1, T] batch
- concat every scheduled token
- explicit position_ids
- slot mapping per token
- no padding through MLPs
One forward pass
- HF model, paged attention
- writes K/V into the pool
- reads context via block tables
- logits only where sampled
Tokens out, blocks back
- append / verify tokens
- stream them (SSE)
- publish full blocks to cache
- finish + free
System architecture
One number drives scheduling. Each sequence tracks num_computed_tokens, the tokens that already have K/V in the pool. A prefill chunk, a decode step and a recompute are the same operation over the uncomputed tokens, which is why one forward pass can mix them. Attention finds each token's slot as block_table[pos // 16] * 16 + pos % 16.
Four requests. Two very different waits.
4 concurrent requests, 50 tokens each, replayed from the v2 (legacy) engine on an Apple M2. Both views share one wall-clock scale: watch where the playhead is when the last request finishes. v3 numbers from the same laptop are below the chart.
Req 4 waited 7.98 s for its first token.
v3 · offline, 128 requests
HF sequential 27.5 tok/s
779.9 tok/s
v3 · vs HF static batching
static batching 276.6 tok/s
779.9 tok/s
v3 · serving capacity (peak)
Phase 1 server 26.7 tok/s
733.9 tok/s
v3 cards: NVIDIA T4 (Colab) · Qwen2.5-1.5B-Instruct fp16 · output tokens/s
Try it: inject a 5th request
In sequential mode it has to wait for the whole running generation.
Virtual memory for attention: pages, not slabs.
Fixed 16-token blocks in one pre-allocated pool, and attention reads and writes it directly through per-sequence block tables. On CUDA the pool is sized from free GPU memory; the grid below is a small demo pool.
allocated
4
free
252
pool used
2%
Memory efficiency
Naïve · max_seq_len pre-allocated
PageServe · paged, demand-allocated
Engineering decision
Reserving max_length of KV per sequence up front leaves most of it unused, because output lengths are unpredictable, so few sequences fit.
PageServe's BlockAllocator gives each sequence a block table of 16-token blocks in one device pool. A custom attention backend plugged into the HuggingFace model scatters new K/V into those slots and attends through the table, so there is no per-sequence past_key_values. Full blocks are content-hashed and ref-counted, so requests with a shared prefix share them.
class BlockAllocator: # block tables, ref counts, prefix cache def allocate(seq_id, num_blocks=1) -> list[int] def ensure_capacity(seq_id, total_tokens) -> list[int] def register_computed_blocks(seq_id, token_ids, num_computed_tokens) def free(seq_id) -> int # attention_wrapper.py — where each token's K/V lives in the poolslot = block_table[pos // 16] * 16 + pos % 16Out of blocks? Preempt, don't drop.
FCFS admission, LIFO preemption. A new request waits for free blocks instead of evicting anyone. When a running sequence can’t grow, the most recently arrived one is preempted: swapped to pinned CPU memory if it’s decoding, otherwise dropped and recomputed later.
Small demo pool: real pools are sized from free GPU memory.
$ waiting for memory pressure…
Without preemption
- ✕ Out-of-memory error
- ✕ Request dropped
- ✕ Work lost
- ✕ Client retries
With PageServe
- ✓ Newest request preempted
- ✓ Swap or recompute
- ✓ Older requests keep going
- ✓ Swapped work resumes first
Victim selection: newest first (LIFO)
FCFS admission plus LIFO preemption means old requests are never starved. A victim that is still prefilling isn't swapped: its blocks are dropped and it re-prefills later, usually cheaply, because its blocks stay in the prefix cache. Nothing new is admitted while swapped sequences wait.
Reuse the past, guess the future.
Prefix caching skips work the engine has already done. Speculative decoding does several tokens of work per step. Both return exactly what greedy decoding would.
Automatic prefix caching
radix tree over blocksFull blocks are content-hashed and ref-counted. A new request walks its prompt block by block and shares every hit instead of recomputing it. Freed blocks keep their contents in an LRU pool until memory is needed.
hₙ = hash(hₙ₋₁, 16 tokens) — a hash names the whole prefix up to that block
T4 · output tok/s
464.6 → 884.9· 93% prefix hit rate
Speculative decoding
n-gram · draft modelDecode is bandwidth-bound, so verifying 5 tokens costs about as much as generating 1. A proposer guesses k tokens and the target checks them in one pass. Output is identical to greedy decoding.
n-gram proposes 4 drafts → target scores [last, d1..d4] in one pass → 5 tokens this step
T4 · output tok/s
924.6 → 1382.9· 93% accepted · copy-heavy output
Caveat · Free-form chat gives n-gram drafts little to copy (27% accepted), so there is no gain. Speculation pays off for copy-heavy output (summaries, RAG, code edits), not as a default.
Streaming + OpenAI API
SSE on the native /generate, plus an OpenAI-compatible /v1/completions. A client disconnect frees its KV memory.
SSE · /v1/completions
KV pool sized for you
On CUDA the engine profiles free GPU memory and sizes the block pool to GPU_MEMORY_UTILIZATION (0.85 by default).
KV_NUM_BLOCKS=auto
Honest metrics
TTFT includes queueing, ITL is the real gap between streamed chunks, p50/p90/p95/p99. Load is Poisson or closed-loop, measured from the scheduled send time.
no coordinated omission
Figures: NVIDIA T4 (Colab) · Qwen2.5-1.5B-Instruct fp16 · offline ablation. The draft-token animation is illustrative.
Twelve phases, one engine.
PageServe was built phase by phase, retracing how production inference engines evolved. Open any phase for the problem it hit and how it was solved.
Problem
One request at a time. A lock serializes all generation, so new arrivals block until the previous request fully completes.
Solution
The ground-truth baseline: a naive prefill → decode loop around HuggingFace, with streaming. Its TTFT now includes the time spent waiting for the lock.
sourceengine/sequential.pyProblem
Head-of-line blocking: a 500-token request starves a 5-token request for its entire duration.
Solution
A background loop does iteration-level scheduling: plan → execute → apply. Requests join and leave the running batch between steps.
sourceengine/scheduler.pyProblem
Concurrent clients overwhelm the scheduler: connections drop or race.
Solution
FIFO RequestQueue with timeouts, cancellation and backpressure. Returns 503 when the queue is full and 504 when a request waits too long.
sourceengine/request_queue.pyProblem
Long prompts stall every running decode, so time-to-first-token spikes for everyone.
Solution
A per-step prefill token budget plus chunked prefill: long prompts are split across steps so they cannot stall running decodes.
sourceengine/scheduler.pyProblem
No visibility into how much memory KV caches really need.
Solution
Bytes-per-token sizing from the model config (with the correct head_dim per family) and logical KV accounting per sequence.
sourceengine/kv_cache_config.pyProblem
Reserving max_length of KV per sequence leaves most of it unused, so few sequences fit.
Solution
Fixed 16-token blocks and per-sequence block tables, with ref counts and content hashes so full blocks can be shared.
sourceengine/block_allocator.pyProblem
Logical blocks need physical storage that every layer can address.
Solution
One pre-allocated device pool shaped [layers, blocks, 16, H, D]. On CUDA it is sized automatically by profiling free GPU memory.
sourceengine/paged_kv_cache.pyProblem
In v2 each sequence kept its own HF cache and ran its own forward pass. The pool was only a shadow copy, so there was no real GPU batching.
Solution
A custom attention backend registered with HuggingFace writes new K/V into the pool and attends through block tables. Every step is one packed [1, T] forward pass mixing prefill chunks and decode tokens.
sourceengine/attention_wrapper.pyProblem
When the pool is full, a running sequence can’t grow. Crashing or dropping requests is not acceptable.
Solution
FCFS admission, LIFO preemption: the most recently arrived sequence is preempted. If it is decoding it is swapped to pinned CPU memory, otherwise its blocks are dropped and it re-prefills later.
sourceengine/cpu_swap_manager.pyProblem
v2 metrics flattered the engine: TTFT excluded queueing, and "per-token latency" was forward time, not inter-token latency.
Solution
One /metrics surface: TTFT (including queueing), TPOT, real ITL, E2E p50–p99, prefix hit rate, speculation acceptance, step stats, KV utilisation and SLO compliance.
sourceengine/metrics_aggregator.pyProblem
Measuring after a client-side semaphore hides queueing (coordinated omission).
Solution
Open-loop Poisson or closed-loop load with a streaming client. Latency is measured from the scheduled send time, so a server that falls behind is charged for it.
sourcerun_load_test.pyProblem
Shared prompts were recomputed, decode stayed bandwidth-bound, and there was no streaming or standard API, or GPU numbers.
Solution
Hash-chained prefix caching; n-gram and draft-model speculative decoding (identical to greedy); SSE streaming and an OpenAI-compatible /v1/completions; an offline + serving benchmark suite, a Colab notebook and a GPU smoke test.
sourceengine/spec_decode.py
Next steps
Fused paged-attention kernel
Triton or FlashInfer instead of gather + matmul, reading the pool in place.
CUDA graphs for decode
Remove Python and launch overhead at small batch sizes.
Rejection sampling
So speculative decoding also applies to sampled (non-greedy) requests.
Tensor parallelism
For models that don’t fit on one GPU.
Tested where it runs. Honest where it hasn’t.
Verified means it ran and matched: tests, exact-match checks and real-model runs. Pending GPUs have a planned model and a notebook ready, but no numbers yet. Nothing on this site is extrapolated.
Apple M2
CPU · Local
- Model
- Tiny test models
- dtype
- fp64
- Date
- 2026-09-27
141 tests pass. Output is token-for-token identical to HF generate() under every feature combination.
- ✓141 tests passing
- ✓Token-for-token match vs HF generate()
- ✓Chunked prefill · batching · prefix caching
- ✓Swap + recompute preemption
- ✓n-gram + draft speculation
Apple M2
MPS · Apple GPU · Local
- Model
- Qwen/Qwen2-0.5B
- dtype
- fp16
- Date
- 2026-09-27
Runs on the Apple GPU. Small local smoke runs below — not headline GPU results.
- ✓Real-model runs on MPS
- ✓Offline ablation smoke run
NVIDIA T4
CUDA · Colab
- Model
- Qwen/Qwen2.5-1.5B-Instruct
- dtype
- fp16
- Date
- 2026-09-27
Full Colab notebook run on commit f168063. KV pool auto-sized to 9 GB (331k tokens). T4 has no bf16, so it runs fp16.
- ✓Smoke test: all 6 engine features exact on CUDA
- ✓fp16 accuracy equal to HF (99.7% top-1 vs fp32)
- ✓0 failed requests · 0 preemptions · no server errors
- ✓28× HF sequential throughput (offline)
NVIDIA L4
CUDA · Colab
- Planned model
- Qwen/Qwen2.5-3B-Instruct
- dtype
- bf16
- Date
- —
Benchmarks coming.
benchmarks coming · no numbers until they’re measured
NVIDIA A100
CUDA · Colab
- Planned model
- Qwen/Qwen2.5-7B-Instruct
- dtype
- bf16
- Date
- —
Benchmarks coming.
benchmarks coming · no numbers until they’re measured
Software versions tested
The engine is tested against both stacks.
Numbers you can reproduce.
Seeded workloads; every result records GPU, versions and git commit. Tables mirror what bench_offline.py and bench_serving.py write. Pending GPUs show empty tables until real Colab runs land. The NVIDIA T4 results come from a full notebook run; the raw files are in benchmarks/published/.
NVIDIA T4 · Qwen/Qwen2.5-1.5B-Instruct · fp16
Full Colab notebook run on commit f168063. KV pool auto-sized to 9 GB (331k tokens). T4 has no bf16, so it runs fp16.
Offline ablation · bench_offline.py
batching / random
128 requests · prompts 256–512 tokens · outputs 128–256 (fixed) · all sent at t=0 · the two one-at-a-time systems run the first 12
| system | output tok/s | speedup | TTFT p50 ms | TTFT p99 ms | TPOT p50 ms | ITL p99 ms | peak mem MB | GPU util % |
|---|---|---|---|---|---|---|---|---|
| HF sequential | 27.5 | 1.00× | 36.4 | 594.3 | 33.3 | 52.6 | 4,047 | 54% |
| HF static batching | 276.6 | 10.04× | – | – | – | – | 5,384 | 98% |
| Engine, no batching | 25.2 | 0.91× | 39,005 | 82,176 | 39.8 | 61.5 | 12,066 | 45% |
| Continuous batching | 779.9 | 28.31× | 7,544 | 19,658 | 63.8 | 366.6 | 12,144 | 81% |
Note · TTFT is high by design here: all 128 requests arrive at once, so later ones queue — see the serving sweep for latency under realistic load. With batching off, the engine is ~9% slower per token than HF (the price of gathering KV through block tables); batching is where it wins.
prefix / shared_prefix
128 requests · 1,024-token shared system prompt + a short unique question · outputs 128–256 (fixed)
| system | output tok/s | speedup | TTFT p50 ms | TTFT p99 ms | TPOT p50 ms | ITL p99 ms | peak mem MB | GPU util % | notes |
|---|---|---|---|---|---|---|---|---|---|
| Continuous batching | 464.6 | 1.00× | 16,641 | 39,056 | 107.6 | 412.1 | 12,145 | 91% | – |
| + prefix caching | 884.9 | 1.90× | 4,733 | 15,244 | 58.7 | 87.4 | 12,141 | 83% | prefix hit 93% |
spec / repetitive
128 copy-a-passage requests · outputs end at EOS
| system | output tok/s | speedup | TTFT p50 ms | TTFT p99 ms | TPOT p50 ms | ITL p99 ms | peak mem MB | GPU util % | notes |
|---|---|---|---|---|---|---|---|---|---|
| Continuous batching | 924.6 | 1.00× | 2,500 | 15,124 | 53.7 | 144.4 | 12,144 | 76% | – |
| + n-gram speculation | 1,383 | 1.50× | 2,583 | 10,328 | 26.5 | 383.5 | 12,125 | 77% | accept 93% |
| + draft-model speculation | 779.3 | 0.84× | 4,033 | 19,134 | 59.5 | 596.1 | 12,167 | 64% | accept 90% |
Note · Draft-model speculation (Qwen2.5-0.5B drafting for the 1.5B target) is slower on a T4: the draft costs a third of the target and runs k steps per verification. It is meant for large targets, e.g. 7B on an A100.
spec / chat
128 chat questions · outputs end at EOS
| system | output tok/s | speedup | TTFT p50 ms | TTFT p99 ms | TPOT p50 ms | ITL p99 ms | peak mem MB | GPU util % | notes |
|---|---|---|---|---|---|---|---|---|---|
| Continuous batching | 1,040 | 1.00× | 3,039 | 11,608 | 46.5 | 67.9 | 12,069 | 59% | – |
| + n-gram speculation | 966.6 | 0.93× | 3,053 | 12,487 | 48.2 | 83.9 | 12,065 | 56% | accept 27% |
| + draft-model speculation | 696.8 | 0.67× | 3,649 | 15,925 | 61.3 | 278.3 | 12,166 | 49% | accept 60% |
Note · Free-form chat gives n-gram drafts little to copy (27% accepted), so there is no gain. Speculation pays off for copy-heavy output (summaries, RAG, code edits), not as a default.
Speedup is relative to the first system in each table, as in bench_offline.py.
Serving sweeps · bench_serving.py · streaming · Poisson arrivals · latencies in ms
workload / random · fixed output lengths
200 requests per load level (24 for the sequential server) · prompts 256–512 tokens · outputs 128–256 (fixed) · Poisson arrivals
| config · load | req/s | out tok/s | TTFT p50 | TTFT p99 | TPOT p50 | TPOT p99 | ITL p99 | E2E p50 | goodput % | GPU util % | GPU mem MB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Sequential (Phase 1) · 2/s | 0.14 | 26.7 | 77,163 | 152,159 | 37.5 | 39.8 | 56.5 | 85,008 | 4% | 52% | 3,708 |
| Sequential (Phase 1) · 4/s | 0.14 | 26.7 | 81,002 | 160,047 | 37.6 | 39.3 | 57.6 | 87,261 | 4% | 53% | 3,730 |
| Sequential (Phase 1) · 8/s | 0.14 | 26.7 | 80,808 | 163,491 | 37.6 | 39.7 | 57.1 | 88,134 | 4% | 52% | 3,730 |
| Sequential (Phase 1) · 16/s | 0.13 | 26.7 | 86,686 | 167,278 | 37.3 | 39.0 | 57.1 | 93,044 | 4% | 52% | 3,732 |
| Sequential (Phase 1) · inf | 0.14 | 26.7 | 80,133 | 162,404 | 37.7 | 39.7 | 56.8 | 87,243 | 4% | 53% | 3,732 |
| Continuous batching · 2/s | 1.8 | 339.6 | 112.3 | 233.4 | 49.8 | 54.3 | 105.1 | 9,542 | 100% | 59% | 12,746 |
| Continuous batching · 4/s | 3.1 | 603.9 | 145.8 | 1,231 | 60.9 | 72.4 | 171.5 | 12,254 | 100% | 70% | 12,746 |
| Continuous batching · 8/s | 3.7 | 703.3 | 5,415 | 15,751 | 74.4 | 79.6 | 181.7 | 19,707 | 33% | 72% | 12,748 |
| Continuous batching · 16/s | 3.7 | 733.9 | 11,535 | 28,170 | 75.8 | 83.6 | 270.6 | 26,784 | 32% | 73% | 12,808 |
| Continuous batching · inf | 3.7 | 708.4 | 18,086 | 41,654 | 75.9 | 91.8 | 215.1 | 33,103 | 11% | 71% | 12,808 |
| + prefix caching · 2/s | 1.8 | 337.8 | 111.0 | 251.4 | 51.2 | 56.2 | 105.0 | 9,709 | 100% | 58% | 12,746 |
| + prefix caching · 4/s | 3.1 | 600.5 | 148.2 | 1,367 | 62.3 | 73.5 | 172.5 | 12,469 | 100% | 69% | 12,746 |
| + prefix caching · 8/s | 3.7 | 697.5 | 5,117 | 16,006 | 74.3 | 81.5 | 195.6 | 19,763 | 33% | 71% | 12,748 |
| + prefix caching · 16/s | 3.6 | 711.1 | 12,164 | 29,475 | 78.4 | 86.4 | 228.5 | 27,539 | 32% | 70% | 12,808 |
| + prefix caching · inf | 3.7 | 706.2 | 18,072 | 42,080 | 77.3 | 92.4 | 301.3 | 33,440 | 13% | 70% | 12,808 |
Note · The sequential server tops out at 0.14 req/s, so every request after the first few waits in line (TTFT ≈ 80 s). Continuous batching serves 3.7 req/s and keeps 100% goodput up to 4 req/s. Prefix caching matches plain batching here — this workload has no shared prefixes, so it is a check that the cache costs nothing when it cannot help.
workload / shared_prefix · fixed output lengths
120 requests per load level · 1,024-token shared system prompt + a short question · outputs 128–256 (fixed)
| config · load | req/s | out tok/s | TTFT p50 | TTFT p99 | TPOT p50 | TPOT p99 | ITL p99 | E2E p50 | goodput % | GPU util % | GPU mem MB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Continuous batching · 4/s | 2.2 | 417.1 | 1,021 | 8,760 | 117.3 | 135.7 | 338.1 | 23,743 | 0% | 84% | 12,804 |
| Continuous batching · 8/s | 2.3 | 446.0 | 4,778 | 21,385 | 108.9 | 144.5 | 406.4 | 31,825 | 0% | 85% | 12,804 |
| Continuous batching · 16/s | 2.3 | 449.8 | 9,206 | 30,980 | 110.2 | 147.7 | 422.1 | 36,496 | 0% | 86% | 12,804 |
| + prefix caching · 4/s | 2.8 | 523.6 | 107.3 | 310.5 | 54.1 | 60.0 | 99.2 | 10,476 | 100% | 65% | 12,722 |
| + prefix caching · 8/s | 3.4 | 668.6 | 144.3 | 4,611 | 65.1 | 71.6 | 109.2 | 14,247 | 63% | 66% | 12,744 |
| + prefix caching · 16/s | 3.7 | 731.2 | 189.2 | 11,688 | 69.8 | 74.7 | 124.0 | 17,912 | 53% | 68% | 12,744 |
Note · With the system prompt cached (94% of prompt tokens hit), TTFT p50 at 4 req/s drops from 1,021 ms to 107 ms and goodput goes from 0% to 100%.
workload / chat · outputs end at EOS
120 chat requests per load level · outputs end at EOS · both configs have prefix caching on, so the only difference is speculation
| config · load | req/s | out tok/s | TTFT p50 | TTFT p99 | TPOT p50 | TPOT p99 | ITL p99 | E2E p50 | goodput % | GPU util % | GPU mem MB |
|---|---|---|---|---|---|---|---|---|---|---|---|
| + prefix caching · 2/s | 1.6 | 269.8 | 74.1 | 124.5 | 43.6 | 47.7 | 67.0 | 7,132 | 100% | 46% | 12,662 |
| + prefix caching · 4/s | 2.7 | 478.6 | 78.5 | 147.4 | 47.4 | 51.4 | 72.6 | 8,012 | 100% | 50% | 12,704 |
| + prefix caching · 8/s | 4.0 | 716.1 | 114.2 | 1,008 | 53.8 | 57.7 | 88.2 | 9,424 | 100% | 50% | 12,704 |
| + n-gram speculation · 2/s | 1.6 | 268.2 | 82.2 | 142.6 | 41.4 | 51.5 | 73.3 | 6,691 | 100% | 46% | 12,542 |
| + n-gram speculation · 4/s | 2.7 | 474.5 | 89.1 | 161.7 | 46.1 | 57.8 | 80.9 | 7,668 | 100% | 47% | 12,604 |
| + n-gram speculation · 8/s | 4.1 | 735.7 | 104.5 | 522.2 | 52.7 | 66.3 | 105.1 | 8,934 | 100% | 49% | 12,624 |
Note · n-gram drafts are accepted only ~27% of the time on free-form chat: a small TPOT gain at 2 req/s (41.4 vs 43.6 ms), no real gain at higher load.
TTFT
Scheduled send time → first streamed token. Includes queueing.
TPOT
(last token − first token) / (tokens − 1), per request.
ITL
Every gap between streamed chunks, pooled across requests.
E2E
Scheduled send time → last token.
Goodput
Share of requests meeting TTFT ≤ 2 s and TPOT ≤ 100 ms.
GPU util
NVML “time a kernel was running”: how busy the GPU is, not FLOP efficiency.
Latency is measured from the scheduled send time, so a server that falls behind is charged for it. Batching and prefix-caching runs use ignore_eos so every system generates the same number of tokens; speculative-decoding runs let outputs end at EOS instead (greedy speculation is exact, so every config emits the same text, while forced output past EOS makes models loop and would overstate the gain).
TTFT · 4 concurrent requests (v2)
| Metric (v2, M2) | Phase 1 · sequential | v2 · batched | Δ |
|---|---|---|---|
| TTFT under 4-way load | 1,418 ms | 20.9 ms | −98.5% |
| Wall clock (all done) | 5,673.6 ms | 3,990.4 ms | −29.6% |
| Aggregate throughput | ~42.0 tok/s | 50.1 tok/s | +19.3% |
| Speedup vs serial | 1.00× | 1.42× | |
| Single-request TTFT / decode | — | 16.52 ms / 97.25 tok/s |
| bench_phases.py (v2) | Phase 1 | Phase 2 |
|---|---|---|
| TTFT (mean) | 168.7 ms | 52.6 ms |
| Total latency | 668.1 ms | 555.2 ms |
| Throughput | 83.2 tok/s | 49.8 tok/s |
| Speedup | 1.00× | 1.41× |
| Chunked prefill (v2) | value |
|---|---|
| 402-token prompt | 4 × 128-token chunks |
| Full prefill (blocking) | 415.1 ms |
| First chunk only | 306.6 ms |
| Decode stall saved | 108.5 ms |
| PR #1 fixes (v2) | Before | After | Δ |
|---|---|---|---|
| TTFT (single req) | 19.8 ms | 16.5 ms | -17% |
| Throughput (single req) | 80.8 tok/s | 97.3 tok/s | +20% |
| 4-way concurrent speedup | 1.32× | 1.42× | +0.10× |
| Tests passing | unknown | 84 / 84 |
Sources: benchmarks/legacy/bench_direct_results.json, bench_phases_results.json, baseline_metrics.json, benchmark_comparison.md · Apple M2 (MPS) · Qwen2-0.5B · fp16.
Stream it, or talk OpenAI.
Native /generate with SSE streaming, an OpenAI-compatible /v1/completions, /health and /metrics, on port 8001. Client disconnects free KV memory. The Phase 1 baseline still runs on :8000.
1# Start the engine on :80012MODEL_NAME=Qwen/Qwen2.5-0.5B-Instruct \3 python -m uvicorn inference_engine.server.app_v2:app --port 80014 5# POST /generate with "stream": true → Server-Sent Events6curl -N localhost:8001/generate -H 'content-type: application/json' \7 -d '{"prompt": "Explain paged attention in one sentence.",8 "max_new_tokens": 64, "stream": true}'9 10# data: {"token_ids": [...], "text": "Paged"} ← one event per step11# data: {"token_ids": [...], "text": " attention"}12# ...13# data: {"done": true, "ttft_ms": ..., "tpot_ms": ..., "finish_reason": ...}14# data: [DONE]Endpoints
- POST
/generateNative · SSE - POST
/v1/completionsOpenAI - GET
/healthLiveness - GET
/metricsp50–p99
Config defaults
CUDA · MPS/CPU
- MAX_BATCH_SIZE
- 64 · 8
- PREFILL_BUDGET_TOKENS
- 2048 · 512
- PREFILL_CHUNK_SIZE
- 512 · 128
- KV_BLOCK_SIZE
- 16
- KV_NUM_BLOCKS
- auto
- GPU_MEMORY_UTILIZATION
- 0.85
- ENABLE_PREFIX_CACHING
- 1
- PREEMPTION_MODE
- swap
- SPECULATIVE_METHOD
- off
- NUM_SPECULATIVE_TOKENS
- 4
Quick start
No GPU? Run it on Colab↗$ git clone …/vLLM_Inference_Engine.git$ cd vLLM_Inference_Engine$ pip install -r requirements.txt$ MODEL_NAME=Qwen/Qwen2.5-0.5B-Instruct python -m uvicorn inference_engine.server.app_v2:app --port 8001