v3141 tests passing→

Every token served.
No page wasted.

PageServe is an LLM inference engine built from scratch: continuous batching over a paged KV cache, prefix caching, preemption and speculative decoding, in readable PyTorch. Every step is one packed forward pass. No vLLM. No generate().

  • Continuous batching
  • Paged attention
  • Prefix caching
  • Speculative decoding
  • Swap / recompute preemption
scheduler.py — continuous batching
a1b2
decoding
c3d4
decoding
·
free slot
·
free slot
·
free slot

iteration

0000

batch

2/4

queue

0

tok/s

180

kv blocks

5/256

28.3×

continuous batching vs HF sequential

27.5 → 779.9 tok/s

2.8×

vs HF static batching

276.6 → 779.9 tok/s

1.9×

prefix caching · shared system prompt

464.6 → 884.9 tok/s

1.5×

n-gram speculation · copy-heavy output

924.6 → 1382.9 tok/s

NVIDIA T4 (Colab) · Qwen2.5-1.5B-Instruct fp16 · offline ablation · output tokens/sFull T4 results →

01Scheduler loop

All the work, one forward pass.

Every scheduler step packs the prefill chunks of new prompts and one decode token per running sequence into a single forward pass. Requests join and leave the batch between steps.

01Plan

Pick this step’s work

  • decodes: 1 token each (+k spec)
  • prefill chunks ≤ token budget
  • admit if blocks allow (FCFS)
  • prefix-cache hits reused
02Pack

One [1, T] batch

  • concat every scheduled token
  • explicit position_ids
  • slot mapping per token
  • no padding through MLPs
03Execute

One forward pass

  • HF model, paged attention
  • writes K/V into the pool
  • reads context via block tables
  • logits only where sampled
04Apply

Tokens out, blocks back

  • append / verify tokens
  • stream them (SSE)
  • publish full blocks to cache
  • finish + free

System architecture

packed forward passKV via block tablestokens backpreemption
plan: allocate blocks · reuse prefix hitsblock tablespreempt: swapswap_inwrite K/V · read context via block tablessampled tokensstream tokens (SSE)ClientSSE · OpenAI APIRequestQueueFIFO · backpressureSchedulerplan → execute → applyModelRunnerONE packed pass · [1, T]Paged attentionHF attention backendBlockAllocatorblock tables · prefix cachePagedKVCache[layers, blocks, 16, H, D]CPUSwapManagerpinned host RAM

One number drives scheduling. Each sequence tracks num_computed_tokens, the tokens that already have K/V in the pool. A prefill chunk, a decode step and a recompute are the same operation over the uncomputed tokens, which is why one forward pass can mix them. Attention finds each token's slot as block_table[pos // 16] * 16 + pos % 16.

02Phase comparison

Four requests. Two very different waits.

4 concurrent requests, 50 tokens each, replayed from the v2 (legacy) engine on an Apple M2. Both views share one wall-clock scale: watch where the playhead is when the last request finishes. v3 numbers from the same laptop are below the chart.

0s2s4s6s8s
Req 1
Req 2
Req 3
Req 4
0.00s
queuedprefilldecodingfirst token

Req 4 waited 7.98 s for its first token.

v3 · offline, 128 requests

HF sequential 27.5 tok/s

779.9 tok/s

28.3×

v3 · vs HF static batching

static batching 276.6 tok/s

779.9 tok/s

2.8×

v3 · serving capacity (peak)

Phase 1 server 26.7 tok/s

733.9 tok/s

27.4×

v3 cards: NVIDIA T4 (Colab) · Qwen2.5-1.5B-Instruct fp16 · output tokens/s

Try it: inject a 5th request

In sequential mode it has to wait for the whole running generation.

Waiting
Admitted
Decoding
Done
03Phase 6–8 · Paged KV cache

Virtual memory for attention: pages, not slabs.

Fixed 16-token blocks in one pre-allocated pool, and attention reads and writes it directly through per-sequence block tables. On CUDA the pool is sized from free GPU memory; the grid below is a small demo pool.

demo pool — 256 blocks
↖ hover any block to inspect it
seq-a1b2seq-c3d4seq-e5f6seq-g7h8free

allocated

4

free

252

pool used

2%

Memory efficiency

Naïve · max_seq_len pre-allocated

45% of reserved slots used55% wasted padding

PageServe · paged, demand-allocated

88% of reserved slots used≤ 15 empty slots / seq

Engineering decision

Reserving max_length of KV per sequence up front leaves most of it unused, because output lengths are unpredictable, so few sequences fit.

PageServe's BlockAllocator gives each sequence a block table of 16-token blocks in one device pool. A custom attention backend plugged into the HuggingFace model scatters new K/V into those slots and attends through the table, so there is no per-sequence past_key_values. Full blocks are content-hashed and ref-counted, so requests with a shared prefix share them.

engine/block_allocator.py
class BlockAllocator:  # block tables, ref counts, prefix cache    def allocate(seq_id, num_blocks=1) -> list[int]    def ensure_capacity(seq_id, total_tokens) -> list[int]    def register_computed_blocks(seq_id, token_ids, num_computed_tokens)    def free(seq_id) -> int # attention_wrapper.py — where each token's K/V lives in the poolslot = block_table[pos // 16] * 16 + pos % 16
04Phase 9 · Preemption

Out of blocks? Preempt, don't drop.

FCFS admission, LIFO preemption. A new request waits for free blocks instead of evicting anyone. When a running sequence can’t grow, the most recently arrived one is preempted: swapped to pinned CPU memory if it’s decoding, otherwise dropped and recomputed later.

memory pressure — replay
Demo data
waiting—
GPU block pool · 32 blocks94% used
swap_out ▼ ▲ swap_in
CPU swap pool (pinned) · 16 blocks0% used

Small demo pool: real pools are sized from free GPU memory.

scheduler.log — tail -f

$ waiting for memory pressure…

Without preemption

  • ✕ Out-of-memory error
  • ✕ Request dropped
  • ✕ Work lost
  • ✕ Client retries

With PageServe

  • ✓ Newest request preempted
  • ✓ Swap or recompute
  • ✓ Older requests keep going
  • ✓ Swapped work resumes first
⚖

Victim selection: newest first (LIFO)

FCFS admission plus LIFO preemption means old requests are never starved. A victim that is still prefilling isn't swapped: its blocks are dropped and it re-prefills later, usually cheaply, because its blocks stay in the prefix cache. Nothing new is admitted while swapped sequences wait.

05New in v3

Reuse the past, guess the future.

Prefix caching skips work the engine has already done. Speculative decoding does several tokens of work per step. Both return exactly what greedy decoding would.

Automatic prefix caching

radix tree over blocks

Full blocks are content-hashed and ref-counted. A new request walks its prompt block by block and shares every hit instead of recomputing it. Freed blocks keep their contents in an LRU pool until memory is needed.

h₁ref 2
h₂ref 2
h₃ref 2
req A · private blocks
req B · private blocks

hₙ = hash(hₙ₋₁, 16 tokens) — a hash names the whole prefix up to that block

T4 · output tok/s

464.6 → 884.9· 93% prefix hit rate

1.9×

Speculative decoding

n-gram · draft model

Decode is bandwidth-bound, so verifying 5 tokens costs about as much as generating 1. A proposer guesses k tokens and the target checks them in one pass. Output is identical to greedy decoding.

last✓✓✓✓+1 bonus

n-gram proposes 4 drafts → target scores [last, d1..d4] in one pass → 5 tokens this step

T4 · output tok/s

924.6 → 1382.9· 93% accepted · copy-heavy output

1.5×

Caveat · Free-form chat gives n-gram drafts little to copy (27% accepted), so there is no gain. Speculation pays off for copy-heavy output (summaries, RAG, code edits), not as a default.

Streaming + OpenAI API

SSE on the native /generate, plus an OpenAI-compatible /v1/completions. A client disconnect frees its KV memory.

SSE · /v1/completions

KV pool sized for you

On CUDA the engine profiles free GPU memory and sizes the block pool to GPU_MEMORY_UTILIZATION (0.85 by default).

KV_NUM_BLOCKS=auto

Honest metrics

TTFT includes queueing, ITL is the real gap between streamed chunks, p50/p90/p95/p99. Load is Poisson or closed-loop, measured from the scheduled send time.

no coordinated omission

Figures: NVIDIA T4 (Colab) · Qwen2.5-1.5B-Instruct fp16 · offline ablation. The draft-token animation is illustrative.

06Development history

Twelve phases, one engine.

PageServe was built phase by phase, retracing how production inference engines evolved. Open any phase for the problem it hit and how it was solved.

  1. Problem

    One request at a time. A lock serializes all generation, so new arrivals block until the previous request fully completes.

    Solution

    The ground-truth baseline: a naive prefill → decode loop around HuggingFace, with streaming. Its TTFT now includes the time spent waiting for the lock.

    sourceengine/sequential.py
  2. Problem

    Head-of-line blocking: a 500-token request starves a 5-token request for its entire duration.

    Solution

    A background loop does iteration-level scheduling: plan → execute → apply. Requests join and leave the running batch between steps.

    sourceengine/scheduler.py
  3. Problem

    Concurrent clients overwhelm the scheduler: connections drop or race.

    Solution

    FIFO RequestQueue with timeouts, cancellation and backpressure. Returns 503 when the queue is full and 504 when a request waits too long.

    sourceengine/request_queue.py
  4. Problem

    Long prompts stall every running decode, so time-to-first-token spikes for everyone.

    Solution

    A per-step prefill token budget plus chunked prefill: long prompts are split across steps so they cannot stall running decodes.

    sourceengine/scheduler.py
  5. Problem

    No visibility into how much memory KV caches really need.

    Solution

    Bytes-per-token sizing from the model config (with the correct head_dim per family) and logical KV accounting per sequence.

    sourceengine/kv_cache_config.py
  6. Problem

    Reserving max_length of KV per sequence leaves most of it unused, so few sequences fit.

    Solution

    Fixed 16-token blocks and per-sequence block tables, with ref counts and content hashes so full blocks can be shared.

    sourceengine/block_allocator.py
  7. Problem

    Logical blocks need physical storage that every layer can address.

    Solution

    One pre-allocated device pool shaped [layers, blocks, 16, H, D]. On CUDA it is sized automatically by profiling free GPU memory.

    sourceengine/paged_kv_cache.py
  8. Problem

    In v2 each sequence kept its own HF cache and ran its own forward pass. The pool was only a shadow copy, so there was no real GPU batching.

    Solution

    A custom attention backend registered with HuggingFace writes new K/V into the pool and attends through block tables. Every step is one packed [1, T] forward pass mixing prefill chunks and decode tokens.

    sourceengine/attention_wrapper.py
  9. Problem

    When the pool is full, a running sequence can’t grow. Crashing or dropping requests is not acceptable.

    Solution

    FCFS admission, LIFO preemption: the most recently arrived sequence is preempted. If it is decoding it is swapped to pinned CPU memory, otherwise its blocks are dropped and it re-prefills later.

    sourceengine/cpu_swap_manager.py
  10. Problem

    v2 metrics flattered the engine: TTFT excluded queueing, and "per-token latency" was forward time, not inter-token latency.

    Solution

    One /metrics surface: TTFT (including queueing), TPOT, real ITL, E2E p50–p99, prefix hit rate, speculation acceptance, step stats, KV utilisation and SLO compliance.

    sourceengine/metrics_aggregator.py
  11. Problem

    Measuring after a client-side semaphore hides queueing (coordinated omission).

    Solution

    Open-loop Poisson or closed-loop load with a streaming client. Latency is measured from the scheduled send time, so a server that falls behind is charged for it.

    sourcerun_load_test.py
  12. Problem

    Shared prompts were recomputed, decode stayed bandwidth-bound, and there was no streaming or standard API, or GPU numbers.

    Solution

    Hash-chained prefix caching; n-gram and draft-model speculative decoding (identical to greedy); SSE streaming and an OpenAI-compatible /v1/completions; an offline + serving benchmark suite, a Colab notebook and a GPU smoke test.

    sourceengine/spec_decode.py

Next steps

↯

Fused paged-attention kernel

Triton or FlashInfer instead of gather + matmul, reading the pool in place.

◎

CUDA graphs for decode

Remove Python and launch overhead at small batch sizes.

⚄

Rejection sampling

So speculative decoding also applies to sampled (non-greedy) requests.

⇶

Tensor parallelism

For models that don’t fit on one GPU.

07Where it's been tested

Tested where it runs. Honest where it hasn’t.

Verified means it ran and matched: tests, exact-match checks and real-model runs. Pending GPUs have a planned model and a notebook ready, but no numbers yet. Nothing on this site is extrapolated.

Apple M2

CPU · Local

Verified
Model
Tiny test models
dtype
fp64
Date
2026-09-27

141 tests pass. Output is token-for-token identical to HF generate() under every feature combination.

  • ✓141 tests passing
  • ✓Token-for-token match vs HF generate()
  • ✓Chunked prefill · batching · prefix caching
  • ✓Swap + recompute preemption
  • ✓n-gram + draft speculation

Apple M2

MPS · Apple GPU · Local

Verified
Model
Qwen/Qwen2-0.5B
dtype
fp16
Date
2026-09-27

Runs on the Apple GPU. Small local smoke runs below — not headline GPU results.

  • ✓Real-model runs on MPS
  • ✓Offline ablation smoke run

NVIDIA T4

CUDA · Colab

Verified
Model
Qwen/Qwen2.5-1.5B-Instruct
dtype
fp16
Date
2026-09-27

Full Colab notebook run on commit f168063. KV pool auto-sized to 9 GB (331k tokens). T4 has no bf16, so it runs fp16.

  • ✓Smoke test: all 6 engine features exact on CUDA
  • ✓fp16 accuracy equal to HF (99.7% top-1 vs fp32)
  • ✓0 failed requests · 0 preemptions · no server errors
  • ✓28× HF sequential throughput (offline)

NVIDIA L4

CUDA · Colab

Pending
Planned model
Qwen/Qwen2.5-3B-Instruct
dtype
bf16
Date
—

Benchmarks coming.

benchmarks coming · no numbers until they’re measured

NVIDIA A100

CUDA · Colab

Pending
Planned model
Qwen/Qwen2.5-7B-Instruct
dtype
bf16
Date
—

Benchmarks coming.

benchmarks coming · no numbers until they’re measured

▶

Run it on Colab

Pick a T4, L4 or A100 runtime and Run all: unit tests, the GPU smoke test (exact-match on the GPU), the offline ablation and the serving sweep, then charts and a zip of the results.

Open PageServe_Colab.ipynb↗

Software versions tested

The engine is tested against both stacks.

transformers 4.51 + torch 2.6transformers 5.12 + torch 2.12transformers 5.16 + torch 2.11 + CUDA 12.8
08Measured results

Numbers you can reproduce.

Seeded workloads; every result records GPU, versions and git commit. Tables mirror what bench_offline.py and bench_serving.py write. Pending GPUs show empty tables until real Colab runs land. The NVIDIA T4 results come from a full notebook run; the raw files are in benchmarks/published/.

NVIDIA T4 · Qwen/Qwen2.5-1.5B-Instruct · fp16

Full Colab notebook run on commit f168063. KV pool auto-sized to 9 GB (331k tokens). T4 has no bf16, so it runs fp16.

Offline ablation · bench_offline.py

batching / random

128 requests · prompts 256–512 tokens · outputs 128–256 (fixed) · all sent at t=0 · the two one-at-a-time systems run the first 12

systemoutput tok/sspeedupTTFT p50 msTTFT p99 msTPOT p50 msITL p99 mspeak mem MBGPU util %
HF sequential27.51.00×36.4594.333.352.64,04754%
HF static batching276.610.04×––––5,38498%
Engine, no batching25.20.91×39,00582,17639.861.512,06645%
Continuous batching779.928.31×7,54419,65863.8366.612,14481%

Note · TTFT is high by design here: all 128 requests arrive at once, so later ones queue — see the serving sweep for latency under realistic load. With batching off, the engine is ~9% slower per token than HF (the price of gathering KV through block tables); batching is where it wins.

prefix / shared_prefix

128 requests · 1,024-token shared system prompt + a short unique question · outputs 128–256 (fixed)

systemoutput tok/sspeedupTTFT p50 msTTFT p99 msTPOT p50 msITL p99 mspeak mem MBGPU util %notes
Continuous batching464.61.00×16,64139,056107.6412.112,14591%–
+ prefix caching884.91.90×4,73315,24458.787.412,14183%prefix hit 93%

spec / repetitive

128 copy-a-passage requests · outputs end at EOS

systemoutput tok/sspeedupTTFT p50 msTTFT p99 msTPOT p50 msITL p99 mspeak mem MBGPU util %notes
Continuous batching924.61.00×2,50015,12453.7144.412,14476%–
+ n-gram speculation1,3831.50×2,58310,32826.5383.512,12577%accept 93%
+ draft-model speculation779.30.84×4,03319,13459.5596.112,16764%accept 90%

Note · Draft-model speculation (Qwen2.5-0.5B drafting for the 1.5B target) is slower on a T4: the draft costs a third of the target and runs k steps per verification. It is meant for large targets, e.g. 7B on an A100.

spec / chat

128 chat questions · outputs end at EOS

systemoutput tok/sspeedupTTFT p50 msTTFT p99 msTPOT p50 msITL p99 mspeak mem MBGPU util %notes
Continuous batching1,0401.00×3,03911,60846.567.912,06959%–
+ n-gram speculation966.60.93×3,05312,48748.283.912,06556%accept 27%
+ draft-model speculation696.80.67×3,64915,92561.3278.312,16649%accept 60%

Note · Free-form chat gives n-gram drafts little to copy (27% accepted), so there is no gain. Speculation pays off for copy-heavy output (summaries, RAG, code edits), not as a default.

Speedup is relative to the first system in each table, as in bench_offline.py.

Serving sweeps · bench_serving.py · streaming · Poisson arrivals · latencies in ms

workload / random · fixed output lengths

200 requests per load level (24 for the sequential server) · prompts 256–512 tokens · outputs 128–256 (fixed) · Poisson arrivals

config · loadreq/sout tok/sTTFT p50TTFT p99TPOT p50TPOT p99ITL p99E2E p50goodput %GPU util %GPU mem MB
Sequential (Phase 1) · 2/s0.1426.777,163152,15937.539.856.585,0084%52%3,708
Sequential (Phase 1) · 4/s0.1426.781,002160,04737.639.357.687,2614%53%3,730
Sequential (Phase 1) · 8/s0.1426.780,808163,49137.639.757.188,1344%52%3,730
Sequential (Phase 1) · 16/s0.1326.786,686167,27837.339.057.193,0444%52%3,732
Sequential (Phase 1) · inf0.1426.780,133162,40437.739.756.887,2434%53%3,732
Continuous batching · 2/s1.8339.6112.3233.449.854.3105.19,542100%59%12,746
Continuous batching · 4/s3.1603.9145.81,23160.972.4171.512,254100%70%12,746
Continuous batching · 8/s3.7703.35,41515,75174.479.6181.719,70733%72%12,748
Continuous batching · 16/s3.7733.911,53528,17075.883.6270.626,78432%73%12,808
Continuous batching · inf3.7708.418,08641,65475.991.8215.133,10311%71%12,808
+ prefix caching · 2/s1.8337.8111.0251.451.256.2105.09,709100%58%12,746
+ prefix caching · 4/s3.1600.5148.21,36762.373.5172.512,469100%69%12,746
+ prefix caching · 8/s3.7697.55,11716,00674.381.5195.619,76333%71%12,748
+ prefix caching · 16/s3.6711.112,16429,47578.486.4228.527,53932%70%12,808
+ prefix caching · inf3.7706.218,07242,08077.392.4301.333,44013%70%12,808

Note · The sequential server tops out at 0.14 req/s, so every request after the first few waits in line (TTFT ≈ 80 s). Continuous batching serves 3.7 req/s and keeps 100% goodput up to 4 req/s. Prefix caching matches plain batching here — this workload has no shared prefixes, so it is a check that the cache costs nothing when it cannot help.

workload / shared_prefix · fixed output lengths

120 requests per load level · 1,024-token shared system prompt + a short question · outputs 128–256 (fixed)

config · loadreq/sout tok/sTTFT p50TTFT p99TPOT p50TPOT p99ITL p99E2E p50goodput %GPU util %GPU mem MB
Continuous batching · 4/s2.2417.11,0218,760117.3135.7338.123,7430%84%12,804
Continuous batching · 8/s2.3446.04,77821,385108.9144.5406.431,8250%85%12,804
Continuous batching · 16/s2.3449.89,20630,980110.2147.7422.136,4960%86%12,804
+ prefix caching · 4/s2.8523.6107.3310.554.160.099.210,476100%65%12,722
+ prefix caching · 8/s3.4668.6144.34,61165.171.6109.214,24763%66%12,744
+ prefix caching · 16/s3.7731.2189.211,68869.874.7124.017,91253%68%12,744

Note · With the system prompt cached (94% of prompt tokens hit), TTFT p50 at 4 req/s drops from 1,021 ms to 107 ms and goodput goes from 0% to 100%.

workload / chat · outputs end at EOS

120 chat requests per load level · outputs end at EOS · both configs have prefix caching on, so the only difference is speculation

config · loadreq/sout tok/sTTFT p50TTFT p99TPOT p50TPOT p99ITL p99E2E p50goodput %GPU util %GPU mem MB
+ prefix caching · 2/s1.6269.874.1124.543.647.767.07,132100%46%12,662
+ prefix caching · 4/s2.7478.678.5147.447.451.472.68,012100%50%12,704
+ prefix caching · 8/s4.0716.1114.21,00853.857.788.29,424100%50%12,704
+ n-gram speculation · 2/s1.6268.282.2142.641.451.573.36,691100%46%12,542
+ n-gram speculation · 4/s2.7474.589.1161.746.157.880.97,668100%47%12,604
+ n-gram speculation · 8/s4.1735.7104.5522.252.766.3105.18,934100%49%12,624

Note · n-gram drafts are accepted only ~27% of the time on free-form chat: a small TPOT gain at 2 req/s (41.4 vs 43.6 ms), no real gain at higher load.

TTFT

Scheduled send time → first streamed token. Includes queueing.

TPOT

(last token − first token) / (tokens − 1), per request.

ITL

Every gap between streamed chunks, pooled across requests.

E2E

Scheduled send time → last token.

Goodput

Share of requests meeting TTFT ≤ 2 s and TPOT ≤ 100 ms.

GPU util

NVML “time a kernel was running”: how busy the GPU is, not FLOP efficiency.

Latency is measured from the scheduled send time, so a server that falls behind is charged for it. Batching and prefix-caching runs use ignore_eos so every system generates the same number of tokens; speculative-decoding runs let outputs end at EOS instead (greedy speculation is exact, so every config emits the same text, while forced output past EOS makes models loop and would overstate the gain).

TTFT · 4 concurrent requests (v2)

Sequential
1,418 ms
Phase 1 · lock-serialized
Batched (v2)
20.9 ms
Metric (v2, M2)Phase 1 · sequentialv2 · batchedΔ
TTFT under 4-way load1,418 ms20.9 ms−98.5%
Wall clock (all done)5,673.6 ms3,990.4 ms−29.6%
Aggregate throughput~42.0 tok/s50.1 tok/s+19.3%
Speedup vs serial1.00×1.42×
Single-request TTFT / decode—16.52 ms / 97.25 tok/s
bench_phases.py (v2)Phase 1Phase 2
TTFT (mean)168.7 ms52.6 ms
Total latency668.1 ms555.2 ms
Throughput83.2 tok/s49.8 tok/s
Speedup1.00×1.41×
Chunked prefill (v2)value
402-token prompt4 × 128-token chunks
Full prefill (blocking)415.1 ms
First chunk only306.6 ms
Decode stall saved108.5 ms
PR #1 fixes (v2)BeforeAfterΔ
TTFT (single req)19.8 ms16.5 ms-17%
Throughput (single req)80.8 tok/s97.3 tok/s+20%
4-way concurrent speedup1.32×1.42×+0.10×
Tests passingunknown84 / 84

Sources: benchmarks/legacy/bench_direct_results.json, bench_phases_results.json, baseline_metrics.json, benchmark_comparison.md · Apple M2 (MPS) · Qwen2-0.5B · fp16.

09HTTP API

Stream it, or talk OpenAI.

Native /generate with SSE streaming, an OpenAI-compatible /v1/completions, /health and /metrics, on port 8001. Client disconnects free KV memory. The Phase 1 baseline still runs on :8000.

stream.shbash
1# Start the engine on :80012MODEL_NAME=Qwen/Qwen2.5-0.5B-Instruct \3  python -m uvicorn inference_engine.server.app_v2:app --port 80014 5# POST /generate with "stream": true → Server-Sent Events6curl -N localhost:8001/generate -H 'content-type: application/json' \7  -d '{"prompt": "Explain paged attention in one sentence.",8       "max_new_tokens": 64, "stream": true}'9 10# data: {"token_ids": [...], "text": "Paged"}        ← one event per step11# data: {"token_ids": [...], "text": " attention"}12# ...13# data: {"done": true, "ttft_ms": ..., "tpot_ms": ..., "finish_reason": ...}14# data: [DONE]

Endpoints

  • POST/generateNative · SSE
  • POST/v1/completionsOpenAI
  • GET/healthLiveness
  • GET/metricsp50–p99

Config defaults

CUDA · MPS/CPU

MAX_BATCH_SIZE
64 · 8
PREFILL_BUDGET_TOKENS
2048 · 512
PREFILL_CHUNK_SIZE
512 · 128
KV_BLOCK_SIZE
16
KV_NUM_BLOCKS
auto
GPU_MEMORY_UTILIZATION
0.85
ENABLE_PREFIX_CACHING
1
PREEMPTION_MODE
swap
SPECULATIVE_METHOD
off
NUM_SPECULATIVE_TOKENS
4

Quick start

$ git clone …/vLLM_Inference_Engine.git
$ cd vLLM_Inference_Engine
$ pip install -r requirements.txt
$ MODEL_NAME=Qwen/Qwen2.5-0.5B-Instruct python -m uvicorn inference_engine.server.app_v2:app --port 8001
No GPU? Run it on Colab↗