{"slug":"dgx-spark","title":"NVIDIA DGX Spark","description":"DGX Spark (GB10 Grace Blackwell, 128GB unified memory) for local AI: setup, DGX OS, unified memory & ARM64, llama.cpp/Ollama/vLLM/LM Studio/ComfyUI on Spark, agents, clustering, and a cross-project troubleshooting casebook.","tags":["dgx-spark","nvidia","gb10","local-ai","unified-memory","arm64"],"category":"inference","scope":{"covers":"NVIDIA DGX Spark as a local-AI machine, verified against DGX OS 7.5.0 (2026-07-18): hardware & the GB10 unified-memory model, first boot & setup, DGX OS & updates/recovery, containers & NGC, ARM64/aarch64 realities, running llama.cpp / Ollama / vLLM / LM Studio / ComfyUI on Spark, agents on Spark (incl. the Hermes playbook), multi-Spark clustering & benchmarking, fine-tuning & NVFP4 quantization, the full official playbooks catalog (66), a cross-project troubleshooting casebook (llama.cpp/Ollama/vLLM issues with hardware provenance), and an inference stack picker.","notCovered":"DGX Station (GB300) and other DGX systems (Station playbooks are catalogued but excluded from Spark guidance); generic per-tool depth covered by the dedicated llama-cpp / vllm / ollama / lmstudio / comfyui wikis; CUDA programming beyond Spark-specific kernel playbooks; purchasing/pricing advice; and Windows/WSL (Spark runs DGX OS).","currentAs":"2026-07-18 (7.5.0)"},"lastUpdated":"2026-07-18","documentCount":21,"raw_base":"/raw/dgx-spark/","html_base":"/wiki/dgx-spark/","xl":{"documentCount":35,"subscribe":"/pro","auth":"Authorization: Bearer <api key>"},"documents":[{"path":"README.md","title":"LLM Wiki","type":null,"updated":null,"gated":false},{"path":"wiki-xl/index.md","title":"XL Edition Index","type":"index","updated":"2026-07-18","gated":true},{"path":"wiki/index.md","title":"Knowledge Base Index","type":"index","updated":"2026-07-18","gated":false},{"path":"wiki-xl/reference/benchmarking.md","title":"Deep Reference — Performance Benchmarking Methodology on DGX Spark","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/casebook-llama-cpp.md","title":"Casebook: llama.cpp on DGX Spark — Per-Issue Deep Dives","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/casebook-ollama.md","title":"Casebook: Ollama on DGX Spark — Per-Issue Deep Dives","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/casebook-vllm.md","title":"Casebook: vLLM on DGX Spark — Per-Issue Deep Dives","type":"reference","updated":"2026-07-19","gated":true},{"path":"wiki-xl/reference/playbook-cli-coding-agent.md","title":"CLI Coding Agent — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-comfy-ui.md","title":"Comfy UI — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-connect-three-sparks.md","title":"Connect Three Sparks — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-connect-to-your-spark.md","title":"Connect to Your Spark — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-connect-two-sparks.md","title":"Connect Two Sparks — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-cuda-x-data-science.md","title":"Playbook Deep Reference — CUDA-X Data Science (cuDF / cuML)","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-cutile-kernels.md","title":"Playbook Deep Reference — cuTile Kernels (TileGym)","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-dgx-dashboard.md","title":"DGX Dashboard — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-flux-finetuning.md","title":"Playbook Deep Reference — FLUX.1 Dreambooth LoRA Fine-Tuning","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-hermes-agent.md","title":"Hermes Agent — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-isaac.md","title":"Playbook Deep Reference — Isaac Sim and Isaac Lab","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-jax.md","title":"Playbook Deep Reference — Optimized JAX on Spark","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-live-vlm-webui.md","title":"Live VLM WebUI — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-llama-cpp.md","title":"llama.cpp on Spark — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-llama-factory.md","title":"Playbook Deep Reference — LLaMA Factory","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-lm-studio.md","title":"LM Studio on Spark — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-multi-agent-chatbot.md","title":"Multi-Agent Chatbot — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-multi-modal-inference.md","title":"Multi-modal Inference (TensorRT) — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-multi-sparks-through-switch.md","title":"Connect Multiple Sparks Through a Switch — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-nccl.md","title":"Playbook Deep Reference — NCCL for Two Sparks","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-nemo-fine-tune.md","title":"Playbook Deep Reference — NeMo AutoModel Fine-Tuning (nemo-fine-tune)","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-nemoclaw.md","title":"Playbook Deep Reference — NemoClaw (Local Agent Sandbox + Applications)","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-nemotron.md","title":"Playbook Deep Reference — Nemotron Model Family on DGX Spark","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-nim-llm.md","title":"NIM on Spark — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-nvfp4-quantization.md","title":"Playbook Deep Reference — NVFP4 Quantization","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-ollama.md","title":"Ollama on Spark — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-open-webui.md","title":"Open WebUI with Ollama — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-pytorch-fine-tune.md","title":"Playbook Deep Reference — Fine-Tune with PyTorch (single and dual Spark)","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-unsloth.md","title":"Playbook Deep Reference — Unsloth on DGX Spark","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki-xl/reference/playbook-vllm.md","title":"vLLM on Spark — Full Playbook Reference","type":"reference","updated":"2026-07-18","gated":true},{"path":"wiki/concepts/agents-on-spark.md","title":"Agents on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/clustering.md","title":"Clustering Multiple DGX Sparks","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/comfyui-and-image-gen.md","title":"ComfyUI and Image Generation on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/containers-and-ngc.md","title":"Containers and NGC","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/dgx-os-and-updates.md","title":"DGX OS and Updates","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/fine-tuning-and-quantization.md","title":"Fine-Tuning and Quantization on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/first-boot-and-setup.md","title":"First Boot and Setup","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/llama-cpp-on-spark.md","title":"llama.cpp on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/lm-studio-on-spark.md","title":"LM Studio on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/ollama-on-spark.md","title":"Ollama on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/overview-and-hardware.md","title":"Overview and Hardware","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/system-operations.md","title":"System Operations","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/unified-memory-and-arm64.md","title":"Unified Memory and ARM64 on DGX Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/concepts/vllm-on-spark.md","title":"vLLM on Spark","type":"concept","updated":"2026-07-18","gated":false},{"path":"wiki/entities/playbooks-catalog.md","title":"DGX Spark Playbooks Catalog","type":"entity","updated":"2026-07-18","gated":false},{"path":"wiki/log.md","title":"Activity Log","type":"log","updated":null,"gated":false},{"path":"wiki/summaries/release-digest.md","title":"DGX Spark Release Notes Digest","type":"summary","updated":"2026-07-18","gated":false},{"path":"wiki/syntheses/stack-picker.md","title":"Stack Picker: Choosing an Inference Stack on Spark","type":"synthesis","updated":"2026-07-18","gated":false},{"path":"wiki/syntheses/troubleshooting-casebook.md","title":"Troubleshooting Casebook: llama.cpp, Ollama, and vLLM on Spark","type":"synthesis","updated":"2026-07-18","gated":false}]}