{"slug":"sglang","title":"SGLang","description":"SGLang (sgl-project/sglang, sglang.io) — a high-performance serving framework for LLMs and multimodal/diffusion models: RadixAttention prefix caching, a frontend DSL, an OpenAI-compatible server, and a deep advanced-features + diffusion stack. Pro tier adds full per-flag/per-hardware references.","tags":["sglang","inference","serving","llm","radixattention","cuda"],"category":"inference","scope":{"covers":"SGLang (verified against main fetched 2026-08-24; newest release v0.5.18): install, request/APIs, offline engine, sampling, architecture/RadixAttention, frontend DSL, server arguments, model gateway, parallelism/disaggregation, hierarchical caching, speculative decoding, quantization, structured outputs, LoRA, attention backends, observability, diffusion serving/optimization, supported hardware/models, developer/benchmarking. XL (Pro): full 435-flag server-args + 321 env-var references, model-gateway ops, per-hardware deep tuning, and exhaustive parallelism/caching/decoding/quant references.","notCovered":"Model-specific training; the exhaustive per-model cookbook (hundreds of entries — mapped, not ingested); behavior of a specific pinned release vs main.","currentAs":null},"lastUpdated":null,"documentCount":28,"raw_base":"/raw/sglang/","html_base":"/wiki/sglang/","xl":{"documentCount":14,"subscribe":"/pro","auth":"Authorization: Bearer <api key>"},"documents":[{"path":"README.md","title":"LLM Wiki","type":null,"updated":null,"gated":false},{"path":"wiki-xl/index.md","title":"XL Edition Index","type":"index","updated":"2026-08-24","gated":true},{"path":"wiki/index.md","title":"SGLang Knowledge Base","type":null,"updated":null,"gated":false},{"path":"wiki-xl/reference/benchmarking-and-profiling.md","title":"Benchmarking and Profiling — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/environment-variables.md","title":"Environment Variables — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/hardware-amd-rocm.md","title":"AMD ROCm on SGLang — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/hardware-ascend-npu.md","title":"Huawei Ascend NPU on SGLang — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/hardware-intel-xpu-cpu.md","title":"Intel XPU and CPU on SGLang — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/hardware-nvidia.md","title":"NVIDIA GPUs on SGLang — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/hardware-tpu-and-others.md","title":"TPU and Other Platforms on SGLang — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/hicache-full.md","title":"Hierarchical KV Caching (HiCache) — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/model-gateway.md","title":"Model Gateway — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/parallelism-full.md","title":"Parallelism and Disaggregation — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/quantization-matrix.md","title":"Quantization Matrix — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/server-arguments-full.md","title":"Server Arguments — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki-xl/reference/speculative-decoding-methods.md","title":"Speculative Decoding Methods — Full Reference","type":"reference","updated":"2026-08-24","gated":true},{"path":"wiki/concepts/architecture-and-radixattention.md","title":"Architecture and RadixAttention","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/attention-backends-and-cuda-graph.md","title":"Attention Backends and CUDA Graph","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/developer-and-benchmarking.md","title":"Developer Guide and Benchmarking","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/diffusion-optimization.md","title":"SGLang Diffusion Optimization Stack","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/diffusion-serving.md","title":"SGLang Diffusion Serving","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/frontend-dsl.md","title":"Frontend DSL","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/hierarchical-caching.md","title":"Hierarchical Caching","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/installation.md","title":"Installation","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/lora-and-model-loading.md","title":"LoRA and Model Loading","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/observability-and-determinism.md","title":"Observability and Determinism","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/offline-engine.md","title":"Offline Engine","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/parallelism-and-disaggregation.md","title":"Parallelism and Disaggregation","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/quantization.md","title":"Quantization","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/references-and-faq.md","title":"References, FAQ, and Troubleshooting","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/router-and-model-gateway.md","title":"Router and Model Gateway","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/sampling-parameters.md","title":"Sampling Parameters","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/sending-requests.md","title":"Sending Requests","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/server-apis.md","title":"Server APIs","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/server-arguments.md","title":"Server Arguments","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/sglang-overview.md","title":"SGLang Overview","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/speculative-decoding.md","title":"Speculative Decoding","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/structured-outputs-and-tool-calling.md","title":"Structured Outputs and Tool Calling","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/supported-hardware.md","title":"Supported Hardware","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/concepts/supported-models.md","title":"Supported Models","type":"concept","updated":"2026-08-24","gated":false},{"path":"wiki/log.md","title":"Change Log","type":null,"updated":null,"gated":false},{"path":"wiki/summaries/release-digest.md","title":"SGLang Release Digest (v0.5.x line)","type":"summary","updated":"2026-08-24","gated":false}]}