# SGLang > SGLang (sgl-project/sglang, sglang.io) — a high-performance serving framework for LLMs and multimodal/diffusion models: RadixAttention prefix caching, a frontend DSL, an OpenAI-compatible server, and a deep advanced-features + diffusion stack. Pro tier adds full per-flag/per-hardware references. > Covers: SGLang (verified against main fetched 2026-08-24; newest release v0.5.18): install, request/APIs, offline engine, sampling, architecture/RadixAttention, frontend DSL, server arguments, model gateway, parallelism/disaggregation, hierarchical caching, speculative decoding, quantization, structured outputs, LoRA, attention backends, observability, diffusion serving/optimization, supported hardware/models, developer/benchmarking. XL (Pro): full 435-flag server-args + 321 env-var references, model-gateway ops, per-hardware deep tuning, and exhaustive parallelism/caching/decoding/quant references. > Not covered: Model-specific training; the exhaustive per-model cookbook (hundreds of entries — mapped, not ingested); behavior of a specific pinned release vs main. - [LLM Wiki](/raw/sglang/README.md) - [SGLang Knowledge Base](/raw/sglang/wiki/index.md) - [Architecture and RadixAttention](/raw/sglang/wiki/concepts/architecture-and-radixattention.md) - [Attention Backends and CUDA Graph](/raw/sglang/wiki/concepts/attention-backends-and-cuda-graph.md) - [Developer Guide and Benchmarking](/raw/sglang/wiki/concepts/developer-and-benchmarking.md) - [SGLang Diffusion Optimization Stack](/raw/sglang/wiki/concepts/diffusion-optimization.md) - [SGLang Diffusion Serving](/raw/sglang/wiki/concepts/diffusion-serving.md) - [Frontend DSL](/raw/sglang/wiki/concepts/frontend-dsl.md) - [Hierarchical Caching](/raw/sglang/wiki/concepts/hierarchical-caching.md) - [Installation](/raw/sglang/wiki/concepts/installation.md) - [LoRA and Model Loading](/raw/sglang/wiki/concepts/lora-and-model-loading.md) - [Observability and Determinism](/raw/sglang/wiki/concepts/observability-and-determinism.md) - [Offline Engine](/raw/sglang/wiki/concepts/offline-engine.md) - [Parallelism and Disaggregation](/raw/sglang/wiki/concepts/parallelism-and-disaggregation.md) - [Quantization](/raw/sglang/wiki/concepts/quantization.md) - [References, FAQ, and Troubleshooting](/raw/sglang/wiki/concepts/references-and-faq.md) - [Router and Model Gateway](/raw/sglang/wiki/concepts/router-and-model-gateway.md) - [Sampling Parameters](/raw/sglang/wiki/concepts/sampling-parameters.md) - [Sending Requests](/raw/sglang/wiki/concepts/sending-requests.md) - [Server APIs](/raw/sglang/wiki/concepts/server-apis.md) - [Server Arguments](/raw/sglang/wiki/concepts/server-arguments.md) - [SGLang Overview](/raw/sglang/wiki/concepts/sglang-overview.md) - [Speculative Decoding](/raw/sglang/wiki/concepts/speculative-decoding.md) - [Structured Outputs and Tool Calling](/raw/sglang/wiki/concepts/structured-outputs-and-tool-calling.md) - [Supported Hardware](/raw/sglang/wiki/concepts/supported-hardware.md) - [Supported Models](/raw/sglang/wiki/concepts/supported-models.md) - [Change Log](/raw/sglang/wiki/log.md) - [SGLang Release Digest (v0.5.x line)](/raw/sglang/wiki/summaries/release-digest.md) ## XL edition (Pro) > Deeper per-item reference pages. Requires a subscription (/pro) — send > `Authorization: Bearer ` on the /raw/ URLs below. - [XL Edition Index](/raw/sglang/wiki-xl/index.md) — Pro - [Benchmarking and Profiling — Full Reference](/raw/sglang/wiki-xl/reference/benchmarking-and-profiling.md) — Pro - [Environment Variables — Full Reference](/raw/sglang/wiki-xl/reference/environment-variables.md) — Pro - [AMD ROCm on SGLang — Full Reference](/raw/sglang/wiki-xl/reference/hardware-amd-rocm.md) — Pro - [Huawei Ascend NPU on SGLang — Full Reference](/raw/sglang/wiki-xl/reference/hardware-ascend-npu.md) — Pro - [Intel XPU and CPU on SGLang — Full Reference](/raw/sglang/wiki-xl/reference/hardware-intel-xpu-cpu.md) — Pro - [NVIDIA GPUs on SGLang — Full Reference](/raw/sglang/wiki-xl/reference/hardware-nvidia.md) — Pro - [TPU and Other Platforms on SGLang — Full Reference](/raw/sglang/wiki-xl/reference/hardware-tpu-and-others.md) — Pro - [Hierarchical KV Caching (HiCache) — Full Reference](/raw/sglang/wiki-xl/reference/hicache-full.md) — Pro - [Model Gateway — Full Reference](/raw/sglang/wiki-xl/reference/model-gateway.md) — Pro - [Parallelism and Disaggregation — Full Reference](/raw/sglang/wiki-xl/reference/parallelism-full.md) — Pro - [Quantization Matrix — Full Reference](/raw/sglang/wiki-xl/reference/quantization-matrix.md) — Pro - [Server Arguments — Full Reference](/raw/sglang/wiki-xl/reference/server-arguments-full.md) — Pro - [Speculative Decoding Methods — Full Reference](/raw/sglang/wiki-xl/reference/speculative-decoding-methods.md) — Pro