# How to Run a Private Remote AI Assistant on DGX Spark

Canonical source: [https://isaiuseful.com/remote-spark](https://isaiuseful.com/remote-spark)

<a id="main-content"></a>

Remote operator · Linux compute

Operate DGX Spark from desktop, browser or messaging clients while keeping its gateway, model and retrieval ports private.

- [Choose a control surface](#operators)

- [Build the Spark side](#spark-setup)

Official NVIDIA starting points

## Pick a DGX Spark recipe.

These four official NVIDIA playbooks adapt our [workflow recipes](https://isaiuseful.com/guides.html.md#chooser) , [local-model stack](https://isaiuseful.com/local-models.html.md#runtimes) and [remote-operation runbook](#operators) to DGX Spark. Pick one here, then use our guides for model choice, private access and operating safeguards.

- [Run OpenClaw with a local LLM 30 min Install the local-first agent and connect it to a private OpenAI-compatible model endpoint.](https://build.nvidia.com/spark/openclaw)

- [Run NemoClaw with a local LLM 30 min Build an OpenClaw assistant in an OpenShell sandbox with local vLLM inference.](https://build.nvidia.com/spark/nemoclaw)

- [Run Hermes Agent with a local LLM 30 min Connect the terminal-first Nous Research agent to a model served locally with vLLM.](https://build.nvidia.com/spark/hermes-agent)

- [Open WebUI with Ollama 15 min Use the browser-based chat interface with Ollama and a model running on Spark.](https://build.nvidia.com/spark/open-webui)

- [Official DGX Spark collection **Explore every NVIDIA Spark playbook** Open build.nvidia.com/spark](https://build.nvidia.com/spark)

**Pocket-ready**

phone steers; state stays on Spark
OPENCLAW OR HERMES
**Insider**

for the newest Windows isolation
MXC SESSION PREVIEW
**Localhost**

services cross an SSH tunnel
NO OPEN ADMIN PORTS

Compatibility first

## How can you control DGX Spark remotely?

The operator device can be a phone, laptop or desktop without moving the gateway, memory, models or jobs off the always-on Linux host.

Native companion

### Windows 11

Use OpenClaw Windows Hub for Command Center, notifications and an optional Windows node. The current bleeding-edge containment route adds Windows Insider and MXC; ordinary remote chat does not move the gateway onto Windows.

Native companion

### macOS

Use the OpenClaw menu bar app in Remote mode. It can own the SSH tunnel, health checks and Web Chat while the gateway remains loopback-bound on Spark.

CLI or desktop

### Linux

Use the Linux companion, CLI, browser Control UI or an SSH tunnel. It is also the closest match to Spark when you need to reproduce commands locally.

**iPhone and Android:** use a paired companion, the browser UI or a configured messaging channel. OpenClaw documents iOS and Android as clients of the Gateway, not replacements for it, so Spark keeps the session and compute while the phone carries chat, status, approvals and only the device capabilities you explicitly enable.

- [iPhone companion →](https://docs.openclaw.ai/platforms/ios)

- [Android companion →](https://docs.openclaw.ai/platforms/android)

- [Compare every control surface →](#operators)

### Hardware + price warning · checked 24 July 2026 Do not buy DGX Spark on the assumption that Windows support will arrive later.

DGX Spark and the pre-release RTX Spark family publish strikingly similar headline numbers, up to one petaflop and up to 128 GB unified memory, but NVIDIA currently lists them as different product categories: DGX Spark is a Linux companion system, while RTX Spark is a Windows primary system. Public documentation does not establish that their drivers, firmware or boards are interchangeable, and NVIDIA has announced no Windows driver, Windows image or upgrade path for DGX Spark.

Buy DGX Spark only if DGX OS works for the machine's useful life. Future Windows support is possible in theory, but today it is speculation. There is also no public evidence for claims about NVIDIA's commercial motive; the support gap itself is the decision-relevant fact.

Indicative MSRP · NVIDIA Marketplace Germany
**€4,800**

NVIDIA's own DGX Spark listing displayed this price and was out of stock when checked. Use it as a reference for the NVIDIA-branded system, not a ceiling for partner products. OEMs choose their own configurations and selling prices.

- [Check the current NVIDIA Marketplace price →](https://marketplace.nvidia.com/de-de/enterprise/personal-ai-supercomputers/dgx-spark/)

- [NVIDIA platform comparison →](https://developer.nvidia.com/local-ai)

- [DGX Spark software requirements →](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/software-requirements.html)

- [Pre-release RTX Spark example →](https://www.microsoft.com/en-us/surface/devices/surface-rtx-spark-dev-box)

### NVIDIA AI Enterprise check · GB10 / DGX Spark Do not treat preinstallation, NIM access or a Spark badge as a production entitlement.

A 90-day NVIDIA AI Enterprise - DGX Spark evaluation exists only when it is purchased, requested or explicitly issued on the NVIDIA Entitlement Certificate (EC); not every OEM offer includes it. Record the EC/order line, GPU metric, start and end date, support route and renewal before deployment. It is an evaluation, not a public free-forever entitlement to the complete supported production suite.

Free components remain useful for development: Omniverse and NVIDIA AI Workbench are free, and the standard NVIDIA Developer Program NIM route is for development, research and test up to 16 GPUs with community support. Production self-hosting generally needs NVIDIA AI Enterprise.

- [DGX Spark 90-day evaluation and EC activation →](https://docs.nvidia.com/dgx/dgx-spark/nvaie-quickstart.html)

- [NVIDIA AI Enterprise licensing guide →](https://docs.nvidia.com/ai-enterprise/planning-resource/licensing-guide/latest/licensing.html)

- [NIM developer-program route and production boundary →](https://forums.developer.nvidia.com/t/nvidia-nim-faq/300317)

- [Omniverse license and support route →](https://docs.omniverse.nvidia.com/dev-guide/latest/common/NVIDIA_Omniverse_License_Agreement.html)

- [AI Workbench introduction →](https://docs.nvidia.com/ai-workbench/user-guide/latest/overview/introduction.html)

See the category boundary

### RTX Spark is being introduced as a Windows PC family, not as a DGX Spark software update.

These NVIDIA and Microsoft videos show the platform framing, an announced product, example workloads and one model-routing layer. Treat them as vendor demonstrations: they clarify positioning, but they do not prove interchangeability, delivery timing or performance for your workflow.

- [Video: NVIDIA RTX Spark Reinvents Windows PCs for the Age of Personal AI](https://www.youtube.com/watch?v=H4nJo-oqAro)

Platform overview · NVIDIA

### A Windows PC category for local AI

Use this overview to understand NVIDIA's intended product category. Keep the announced positioning separate from tested application support on a shipping system.

- [Video: Introducing Surface RTX Spark Dev Box](https://www.youtube.com/watch?v=VlAI1_JkXL4)

Product example · Microsoft

### Surface makes the Windows distinction concrete

The Surface RTX Spark Dev Box is a separate announced Windows product. Its existence is not evidence that a DGX Spark can be converted into one.

- [Video: Architectural Design With Agents on NVIDIA RTX Spark](https://www.youtube.com/watch?v=a6fUvL9gYAQ)

Workload demo · NVIDIA

### Judge the workflow, then test the stack

The architectural-design demo shows how NVIDIA imagines agents using the platform. It is a useful workflow reference, not a benchmark or compatibility matrix.

- [Video: Debugging with a Local Agent While You Get Coffee, Powered by NVIDIA RTX Spark](https://www.youtube.com/watch?v=WCRNR1Ve9s0)

Agent workflow · NVIDIA

### Keep the local agent working while you step away.

The demo shows a Hermes agent monitoring communication channels, prioritizing an urgent software issue and coordinating debugging and quality assurance on RTX Spark. It demonstrates NVIDIA's intended always-on local workflow, not measured autonomy, task success or hardware performance.

- [Video: Announcing NVIDIA RTX Spark | GTC Taipei 2026 Keynote by CEO Jensen Huang](https://www.youtube.com/watch?v=11Y3B33oCLE)

Launch context · NVIDIA

### Hear the promise, then compare DGX Station.

The keynote gives the broad RTX Spark context. If 128 GB is the constraint rather than the solution, the DGX Station guide maps the 748 GB tier, realistic model fits and the current benchmark gap.

- [Compare DGX Spark with Station →](https://isaiuseful.com/dgx-station.html.md)

- [Video: Get Started with Open Model Routing | Nemotron Labs](https://www.youtube.com/watch?v=ZQ3EeU-hYSA)

Nemotron Labs · model routing

### Route by task only after the local baseline works.

NVIDIA's walkthrough shows a router selecting among models in a configured pool. On Spark, add that layer only after one private endpoint passes the task acceptance set. Record which choices stay local: a router can introduce new model, provider, logging, cost and data-boundary decisions.

- [Review NVIDIA's open routing description →](https://developer.nvidia.com/topics/ai/nemotron)

- [Inspect the hosted router boundary →](https://docs.nvidia.com/nemoclaw/user-guide/openclaw/inference/hosted-inference/set-up-model-router)

<a id="spark-reviews"></a>

Independent reviews + benchmarks

## Spark is a capacity machine with a measured bandwidth ceiling.

Hands-on third-party tests agree on the shape even when engines, quants and prompts differ: 128 GB unlocks models that ordinary small systems cannot load, while 273 GB/s LPDDR5X limits dense single-stream decode.

Hands-on review

### LMSYS measured the trade.

Its early-access Spark ran GPT-OSS 20B MXFP4 at 49.7 decode tok/s, but Llama 3.1 70B FP8 at 2.7 decode tok/s. Llama 3.1 8B scaled to 368 aggregate decode tok/s at batch 32. NVIDIA supplied early access, and the authors warn that software results can age.

- [Read the LMSYS methodology and tables →](https://www.lmsys.org/blog/2025-10-13-nvidia-dgx-spark/)

Independent comparison

### Dense 70B was capacity-first, not fast.

A separate published comparison reported 4.67 tok/s for Llama 3.3 70B, 38.03 tok/s for Qwen3 Coder and 60.33 tok/s for GPT-OSS 20B on Spark. Treat these as workload-specific results, not universal product scores.

- [Inspect the comparison and conditions →](https://www.pcgamer.com/hardware/graphics-cards/nvidias-little-gold-box-of-pure-ai-power-the-dgx-spark-is-finally-out-and-the-comparison-with-amds-much-cheaper-strix-halo-chip-is-looking-a-little-fugly/)

Cluster evidence

### Two nodes need the right parallelism.

StorageReview tested Dell, GIGABYTE and HP pairs over the 200 Gb fabric and found OEM performance within a narrow band. For batched inference at practical concurrency, its pipeline-parallel layout mattered more than small chassis differences.

- [Read the two-node review →](https://www.storagereview.com/review/nvidia-dgx-spark-cluster-review-distributed-inference-on-dell-gigabyte-and-hp)

Independent video reviews

### Three longer practitioner views to put beside the numbers.

These videos add independent system context. Keep their software versions, workloads and methodology attached to any performance observation, and use the normalized table below for direct comparisons.

- [Video: Deep Dive into Nvidia's DGX Spark GB10](https://www.youtube.com/watch?v=Lqd2EuJwOuw)

Independent deep dive · Level1Techs

### Put the GB10 platform under a practitioner’s lens.

Level1Techs examines DGX Spark as a system rather than a specification sheet. Treat its observations as independent context and keep software versions and workload conditions attached to any result.

- [Video: Getting the Same Results with Smaller "Cheaper" Dual Sparks AI as the More Expensive Clusters](https://www.youtube.com/watch?v=tmcn1-jFLWY)

Independent cluster test · Level1Techs

### Test when two Sparks can match a larger cluster.

Level1Techs compares a dual-Spark setup with more expensive cluster configurations. Keep model, quantization, parallelism, concurrency and software versions attached to the result; it does not show that two Sparks replace every larger system.

- [Video: NVIDIA DGX Spark: From “Inference Box” to Dev Rig (What It Actually Is) | Ep 2](https://www.youtube.com/watch?v=0CI19dXmOws)

Independent perspective · Domesticating AI

### Frame Spark as a development rig, not only an inference box.

This longer independent discussion focuses on what the machine is and how its development role differs from a simple inference appliance. Pair the framing with the measured limits below before buying.

Measured snapshot · do not mix rows

### Engine, precision, batch and model architecture explain the spread.

“Tokens per second” without those fields is not a benchmark. Prefill and decode are separate, and aggregate batch throughput is not the speed each interactive user sees.

| Source | Model and format | Workload | Measured Spark result | What it supports |
| --- | --- | --- | --- | --- |
| LMSYS · Oct 2025 | GPT-OSS 20B · MXFP4 · Ollama | Single-stream decode | 49.7 tok/s | A small sparse model can be comfortably interactive. |
| LMSYS · Oct 2025 | Llama 3.1 70B · FP8 · SGLang | Batch 1 decode | 2.7 tok/s | Loading a large dense model is not the same as serving it quickly. |
| LMSYS · Oct 2025 | Llama 3.1 8B · FP8 · SGLang | Batch 32 aggregate decode | 368 tok/s | Batching can use compute that one bandwidth-bound stream leaves idle. |
| Third-party comparison · Oct 2025 | Llama 3.3 70B | Single prompt | 4.67 tok/s | A second setup reproduces the slow dense-70B shape, not the exact LMSYS number. |
| Community recipe · Jul 2026 | GLM-4.7 355B · NVFP4 · custom vLLM · TP=2 | Two Sparks · 65,536 max sequence length · single-stream decode | About 17.5 tok/s | Two 128 GB nodes can serve the full checkpoint, but only through a pinned patched stack. |
| Level1 practitioner · Aug 2026 | DeepSeek-V4-Flash-0731 304B · custom vLLM · NVFP4 KV · TP=2 | Warm decode · `stream:false` | 67.5 mean tok/s · 219-227 aggregate at c6 | Speculative-decode acceptance and concurrency make prompt shape part of the result. |
| Level1 practitioner · Aug 2026 | DeepSeek-V4-Flash-0731 304B · same service | 40 min · c4 mixed agent traffic · 532 requests | 87.2 aggregate · 21.8 per stream · 0 request errors | The pinned service survived this bounded soak; it does not establish long-term production reliability. |
| NVIDIA SANA · Aug 2026 | MiniMax H3 · pruned FP8 · Sol Engine | 5 s video · 480p · 124 frames · 50 steps | 181.3 s optimized · 3.92× | A specialized FP8 runtime makes H3 fit and materially faster; generation remains far from real time. |

### Community recipes · checked 23 August 2026 Full GLM and DeepSeek fit two Sparks, but only on pinned custom stacks.

A July practitioner recipe serves the full **GLM-4.7 355B** Salyut1 NVFP4 checkpoint, about 188 GB across 41 shards, on two DGX Sparks over the built-in ConnectX-7 fabric. The author reports about 17.5 tok/s with vLLM TP=2, CUDA graphs enabled and a configured 65,536-token ceiling. The setup passed the author's coherent-completion and tool-call checks plus an exact needle-recall check around 21K tokens; that is not a full-64K quality validation. It is not a stock install: the recipe uses a development vLLM image, an open loader-fix pull request, a no-Ray launch path, explicit GB10 NCCL handling, capped batch settings and 0.90 memory utilization. The author also reports about 9.5 minutes to load the checkpoint.

An August Level1Techs report measures the official **DeepSeek-V4-Flash-0731 304B** checkpoint on two Sparks with TP=2 over dual 200 Gb/s RoCE. Here, NVFP4 describes the KV cache rather than the model weights. On a custom vLLM fork at the model's calibrated 1,048,576-token ceiling, with six sequence slots, five speculative tokens, warm kernels and non-streamed accounting, the author measured 84.1 tok/s peak, 67.5 mean decode, 219.2-226.9 aggregate tok/s at concurrency six and about 2,690 prefill tok/s at 100K. A separate 40-minute mixed-temperature, mixed-budget soak at concurrency four produced 87.2 aggregate and 21.8 per stream across 532 requests without a request error.

**Do not turn the RTX comparison into a quality ranking.** The forum's temperature-zero diagnostic found stack-specific counting and arithmetic differences, but its expanded artifacts were not publicly downloadable when checked. DeepSeek recommends temperature 1.0 and top-p 0.95 for agentic local use. The public recipe also changes quickly and contains results for both the official 0731 checkpoint and an earlier preview. Keep the checkpoint, commit, container, patches, prompt, sampling, warm-up, stream mode and concurrency attached to every number.

- [Read the Level1Techs GLM recipe thread →](https://forum.level1techs.com/t/full-glm-4-7-355b-nvfp4-at-64k-on-two-dgx-sparks-working-recipe/252212)

- [Inspect the full GLM configuration and author checks →](https://forums.developer.nvidia.com/t/full-glm-4-7-355b-nvfp4-at-64k-context-on-2x-dgx-spark-gb10-working-recipe-vllm-tp-2/375690)

- [Inspect the still-open GLM loader-fix pull request →](https://github.com/eugr/spark-vllm-docker/pull/307)

- [Read the dual-Spark DeepSeek measurements →](https://forum.level1techs.com/t/dual-sparks-in-nvfp4-vs-4x-rtx-pro-6000-with-native-deepseek-v4-0731-quants-and-speed/253539)

- [Inspect the moving two-Spark recipe →](https://github.com/tonyd2wild/DeepSeek-v4-Flash-0731-DSpark-1M-NVFP4-KV-2x-DGX-Spark)

- [Check the official model card and sampling guidance →](https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731)

**Buying implication:** Spark's best LLM fit is usually a 20-35B daily model or a larger sparse MoE with few active parameters, not the largest dense checkpoint that can be forced into memory. Re-run your exact engine after every major software update. [See where DGX Station changes the boundary →](https://isaiuseful.com/dgx-station.html.md#comparison)

### Spark owner tests · checked 15 August 2026 Muse Glimmer fits one Spark, and DFlash lifts decode into the 20s and 30s.

Meta’s Apache-2.0 **Muse Glimmer 30B** is a dense local-agent model with tool use, coding and optional image input. Its official GGUF release provides 16.76 GB and 19.65 GB language-model builds, plus a separate 1.40 GB perception encoder and optional 1.63 GB DFlash drafter. That leaves ample room inside a 128 GB Spark-class system for runtime and a useful context budget, although the advertised 131,072+ model limit is not a promise that the maximum context will be fast or memory-efficient.

Early owner runs provide a useful speed range, not a controlled benchmark. On one DGX Spark, a llama.cpp run reported about 10.5 tok/s conventional decode and 36-38 tok/s with the official DFlash drafter at 15 speculative tokens; a separate vLLM run rose from 5-8 tok/s to about 23 tok/s at the same setting. Workload, draft acceptance, quant, context depth and engine build can move those numbers.

The llama.cpp test also reported about 700 tok/s prefill at short context, about 390 tok/s deep into an 832K-token prompt, and 3/3 needle retrieval at 97K, 188K, 415K and 832K after an 8× YaRN and metadata override. Treat that as an extended-context experiment rather than a new guaranteed model limit. Its two-Spark RPC layer split slowed decode to 25-28 tok/s, so a model this size is better run as one instance per Spark unless a measured workload proves otherwise. Pin the exact engine revision and rerun your own tool-call, recovery, context and latency harness.

- [Inspect Meta’s official GGUF files →](https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF/tree/main)

- [Read Meta’s launch and RTX 5090 speed conditions →](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model)

- [Inspect the DGX Spark llama.cpp and extended-context owner test →](https://www.reddit.com/r/LocalLLaMA/comments/1vl9adk/i_ran_muse_glimmer_1m_context_all_tests_passed/)

- [Inspect the DGX Spark vLLM DFlash owner test →](https://www.reddit.com/r/LocalLLM/comments/1vm4j0i/muse_glimmer_30b_on_dgx_spark_using_dflash_is/)

- [Compare memory fits and vendor benchmarks →](https://isaiuseful.com/local-models.html.md#muse-glimmer)

<a id="minimax-h3-spark"></a>

### Vendor video benchmark · checked 15 August 2026 MiniMax H3 proves the Spark fit, but 3.92× still means three minutes for five seconds.

NVIDIA’s SANA team ran the 33B dense **MiniMax H3** audio-video generator on one DGX Spark at 832×480, 24 fps, 124 frames and 50 denoising steps. Sol Engine reduced end-to-end wall time from 710.6 to 181.3 seconds, a reported 3.92× speedup. The full recipe combines kernel work with approximate Sol-Attn sparse attention and cross-step caching, so the accompanying near-lossless quality assessment is part of the vendor experiment, not proof that every prompt is unchanged. NVIDIA’s separate GeForce RTX 5090 result uses 1344×768 and must not be treated as an RTX Spark benchmark or compared by raw time.

The ordinary BF16 FL2VA package is about 134.2 GiB before activations, so NVIDIA’s GB10 runtime uses a pruned FP8 DiT, FP8 Qwen3-VL conditioner and the released VAEs, with one resident model process. MiniMax’s downloadable local system produces 768p H3-Base output; the recommended Context-IR preprocessing and Regenerate-2K stage remain hosted. There is also a procurement-level license gate: the community license excludes the EU, UK, United States and South Korea, with a separate application route for those territories.

- [Inspect NVIDIA’s H3-on-device results →](https://nvlabs.github.io/Sana/Sol-Engine/H3-OnDevice/)

- [Inspect the GB10 runtime and reproducible config →](https://github.com/NVlabs/Sana/tree/sol-engine/models/minimax_h3/GB10)

- [Read the MiniMax H3 model card →](https://huggingface.co/MiniMaxAI/MiniMax-H3)

- [Read the MiniMax H3 license →](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE)

- [Request an excluded-territory license →](https://platform.minimax.io/h3-license)

- [Compare the server and cloud routes →](https://isaiuseful.com/cloud-models.html.md#minimax-h3)

### Native NVFP4 hardware · speculative capacity DGX Spark has Blackwell FP4 hardware. The serving kernel is still a gate.

NVFP4 is promising here because it can shrink weights and reduce memory traffic while preserving more quality than a crude four-bit conversion. For planning, reserving 20-25% of Spark's 128 GB and assuming a mixed checkpoint at roughly 5.0-5.2 bits per parameter gives an estimated **150-165B parameter** single-node fit. That is below NVIDIA's “up to 200B” capacity ceiling because this estimate leaves useful room for the runtime and cache.

Two linked Sparks provide a speculative **295-330B parameter** planning band with the same reserve. NVIDIA has demonstrated Qwen-235B in NVFP4 on two systems and markets an up-to-405B model ceiling, but neither statement guarantees your model, context, kernel or interactive speed.

- [NVIDIA's dual-Spark NVFP4 example →](https://developer.nvidia.com/blog/?p=111120)

- [Understand NVFP4 and Blackwell support →](https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)

- [Compare Spark, Station and datacenter capacity →](https://isaiuseful.com/cloud-models.html.md#nvfp4)

DGX Spark-class OEM systems

## Eight GB10 listings. One workload test.

The active products share the compact GB10, 128 GB coherent-memory class, but not necessarily NVIDIA's DGX Spark product name, factory image, storage, regional price or support terms. HP advised directly that ZGX Nano is discontinued, so it remains here as a clearly labelled historical reference.

![Acer logo](https://cdn.simpleicons.org/acer/white)

### Veriton GN100

GB10 AI mini workstation.

- [Open Acer system page →](https://www.acer.com/gb-en/desktops-and-all-in-ones/veriton-workstations/veriton-gn100-ai-mini-workstation?utm_source=nvidia)

![ASUS logo](https://cdn.simpleicons.org/asus/white)

### ASUS Ascent GX10

GB10 desktop AI supercomputer.

- [Open ASUS system page →](https://www.asus.com/networking-iot-servers/desktop-ai-supercomputer/ultra-small-ai-supercomputers/asus-ascent-gx10/?utm_source=nvidia)

![Dell logo](https://cdn.simpleicons.org/dell/white)

### Dell Pro Max with GB10

GB10 micro workstation.

- [Open Dell system page →](https://www.dell.com/en-ie/shop/desktop-computers/dell-pro-max-with-gb10/spd/dell-pro-max-fcm1253-micro/xcto_fcm1253_emea#features_section)

![GIGABYTE logo](https://upload.wikimedia.org/wikipedia/commons/d/d5/Gigabyte_Technology_Logo.svg)

### GIGABYTE AI TOP ATOM

GB10 AI TOP system.

- [Open GIGABYTE system page →](https://www.gigabyte.com/AI-TOP-PC/GIGABYTE-AI-TOP-ATOM)

![HP logo](https://cdn.simpleicons.org/hp/white)

### HP ZGX Nano AI Station

GB10 nano AI station.

- [Open the HP reference page →](https://www.hp.com/us-en/workstations/zgx-nano-ai-station.html?utm_source=nvidia)

![Lenovo logo](https://upload.wikimedia.org/wikipedia/commons/0/05/Lenovo-Logo.svg)

### Lenovo ThinkStation PGX

GB10 AI development workstation.

- [Open Lenovo system page →](https://www.lenovo.com/ie/en/p/workstations/thinkstationp/lenovo-thinkstation-pgx-sff/len102s0023)

![MSI logo](https://cdn.simpleicons.org/msi/white)

### MSI EdgeXpert MS-C931

GB10 AI supercomputer.

- [Open MSI system page →](https://ipc.msi.com/product_detail/Industrial-Computer-Box-PC/AI-Supercomputer/EdgeXpert-MS-C931?utm_source=nvidia)

![PNY logo](https://upload.wikimedia.org/wikipedia/commons/4/4f/PNY_Technologies_logo.svg)

Offer checked 10 August 2026

### PNY DGX Spark

Authorized DGX Spark channel route. The current NVIDIA/PNY offer advertises an optional free 90-day NVIDIA AI Enterprise - DGX Spark evaluation after registration and entitlement; it is time-limited, offers community-driven support only, and is not a production-support entitlement.

- [Open PNY system page →](https://www.pny.com/dgx-spark)

- [NVIDIA AI Enterprise - DGX Spark 90-day evaluation →](https://docs.nvidia.com/dgx/dgx-spark/nvaie-quickstart.html)

- [NVIDIA Marketplace DGX Spark offer →](https://marketplace.nvidia.com/en-us/enterprise/personal-ai-supercomputers/dgx-spark/)

- [NVIDIA trial support caveat →](https://www.nvidia.com/content/dam/en-zz/Solutions/dgx-spark/workstation-print-gtc26-nvaie-spark-solution-overview-5004550-r7.pdf)

**Shortlist + quote check:** first compare the exact model, engine, quant, context and concurrency. Then verify the configured SSD, DGX OS or factory image, ConnectX accessories and two-node kit, thermals, power supply, regional availability, support lifecycle, warranty and VAT-inclusive price. Start with the [independent Spark benchmark shape](#spark-reviews) , then repeat the same harness on the OEM unit you can actually buy.

Try before you buy

## Rent the remote workload before buying a Spark or larger system.

Before committing four, five or six figures to a Spark or larger deployment, rent the candidate workload and prove the whole operator path. A rental is a test drive, not a guarantee of GPU availability, equivalent hardware, benchmark transfer or price.

**Exercise the real route.** Run the exact model, engine, quantization, context, concurrency and acceptance harness, including the remote-control and tunnel pattern you need; retain the container, settings, logs, latency, throughput, memory and task-quality results.

 **Referral disclosure:** this isaiuseful.com link is a Runpod referral link. As checked 10 August 2026, eligible new first-time users must sign up through it with Google SSO and load their first $10: European users receive $5 credit, while non-European users receive a weighted $5-$500 credit (Runpod says most are $10 or less). Terms can change. If eligible, isaiuseful.com receives its referral bonus and earns Runpod credits on actual usage for the first six months: 3% of Pod spend and 5% of Serverless spend. Using the link supports the site.

- [Rent a Runpod test environment through this referral link →](https://runpod.io?ref=l40ix174)

- [Read Runpod’s current referral terms →](https://docs.runpod.io/accounts-billing/referrals)

DGX Spark product views

## Hear NVIDIA's platform promise. Keep it separate from measured results.

The keynote explains the platform story, the PNY film shows one product implementation, and the assistant and [GraphRAG](https://isaiuseful.com/rag.html.md#graph) sessions demonstrate intended workflows. All four are vendor presentations, use the independent review block above for measured model performance.

- [Video: NVIDIA GTC Spring 2025 Keynote: Introducing NVIDIA DGX Spark](https://www.youtube.com/watch?v=6p4U1kSiegg)

Launch context · NVIDIA

### Introducing DGX Spark at GTC

The keynote segment sets out NVIDIA's original product positioning and target users. Compare those claims with current documentation and the independent throughput results on this page.

- [Video: NVIDIA DGX Spark | A Grace Blackwell AI Supercomputer on your desk](https://www.youtube.com/watch?v=kZRMshaNrSA)

PNY product view · NVIDIA

### PNY DGX Spark in NVIDIA's product film

The short film shows the PNY implementation and intended compact appliance experience. It does not establish model quality, sustained tokens per second, thermals or equivalence with other OEM systems.

- [Video: Build Your Own AI Assistant with Hugging Face on NVIDIA DGX Spark](https://www.youtube.com/watch?v=dMpLCGvE2A0)

Build tutorial · NVIDIA

### Turn the appliance into an assistant workflow.

This concise NVIDIA walkthrough shows a Hugging Face assistant build on DGX Spark. Use it for workflow ideas and product setup context, not as independent performance evidence.

- [Video: DGX Spark Live: Process Text for GraphRAG With Up to 120B LLM](https://www.youtube.com/watch?v=uQtzjAvJMlE)

GraphRAG demo · NVIDIA Developer

### Use the memory pool for a larger retrieval pipeline.

The NVIDIA Developer session demonstrates text processing for [GraphRAG](https://isaiuseful.com/rag.html.md#architecture) with an LLM of up to 120B parameters. It shows an intended capacity-led workflow, not a standardized throughput benchmark.

<a id="operators"></a>

Choose the operator

## Use the device already in your hands.

A phone is enough for chat, status and approvals. Desktop- or mobile-node permissions are optional additions for workflows that genuinely need to touch the operator device.

Windows 11

### Windows Hub + MXC preview

Best native OpenClaw diagnostics and Windows-node experience. Use the Insider path below when you want the current session-isolation work; use SSH, Web Chat or messaging when the Windows machine is only an operator.

- [Follow the Insider setup →](#windows-setup)

- [Windows Hub →](https://github.com/openclaw/openclaw-windows-node)

macOS

### Menu bar app in Remote mode

Point the signed macOS companion at `user@SPARK_IP` . Its default remote mode manages a strict-host-key SSH tunnel, health checks and Web Chat without starting a second gateway on the Mac.

- [Official remote-mode guide →](https://docs.openclaw.ai/platforms/mac/remote)

Linux

### Companion, CLI or browser

Use the Linux desktop companion when you want a tray and Canvas, or forward port 18789 and open the Control UI on localhost. The CLI is the simplest path on a minimal workstation.

- [Official Linux guide →](https://docs.openclaw.ai/platforms/linux)

**Phone or desktop companion**

Best local notifications, status and optional device-node capabilities.

**Web or CLI**

Universal path through an SSH tunnel; no desktop-node permissions required.

**Messaging**

Telegram, Discord, WhatsApp or another configured channel for everyday remote conversations.

**Hermes surfaces**

Use its TUI, dashboard or messaging gateway from any client while Hermes stays on Spark.

Remote access · private by default

### Use the narrowest path that fits the job.

Keep Spark services loopback-bound and authenticate every client. NVIDIA Sync is the managed route for Windows, macOS and Ubuntu; direct SSH is the universal fallback; Sunshine with Moonlight is the concrete remote-desktop path when the work truly needs the full Linux interface.

**01 · Sync** - Managed SSH, application launches and tunnels.
**02 · SSH** - Direct terminal access and explicit port forwards.
**03 · Desktop** - Sunshine host, Moonlight client and a headless virtual display.

- [Repair Sunshine + Moonlight after installation →](#sunshine-moonlight-repair)

- [Open the human/AI repair runbook →](https://isaiuseful.com/downloads/dgx-spark-sunshine-moonlight-repair.md)

- [NVIDIA Sync user guide →](https://docs.nvidia.com/sync/latest/index.html)

- [Compare and set up the three routes →](#spark-remote-access)

<a id="windows-setup"></a>

Windows 11 · bleeding edge

## Insider is mandatory for the newest containment path.

OpenClaw Windows Hub can run on ordinary Windows builds. This guide targets the newer MXC session-isolation and agent-policy work Microsoft is still shipping through Windows Insider.

01

### Give the preview a recoverable Windows installation.

Use a secondary device or separate system image when possible. Enable BitLocker or device encryption, Secure Boot, Windows Hello and Defender; create a tested recovery drive and backup before enrolling. Run the companion as a standard user, not from a shared or daily administrator account.

02

### Join Windows Insider Experimental on a retail-aligned core.

Experimental is where Microsoft says actively developed features appear first. Stay on the 25H2 or 26H1 line offered for your hardware rather than the separate Future Platforms option. After updating, confirm the actual build with `winver` ; MXC currently documents build 26300.8553 as the minimum for its `isolation_session` backend.

- [Current channel definitions →](https://blogs.windows.com/windows-insider/2026/04/10/improving-your-windows-insider-experience/)

- [MXC build matrix →](https://github.com/microsoft/mxc)

03

### Verify the containment feature, not merely the OS label.

Install the current feature and platform updates, then check the OpenClaw release notes for MXC support and confirm that the requested backend is available. An Insider badge alone proves nothing. Microsoft currently describes OpenClaw's Windows node and gateway as an MXC integration, but the MXC repository also warns that its early-preview profiles are still overly permissive in known cases.

**Preview rule**

If MXC or the expected session backend is unavailable, stop. Do not silently fall back to unrestricted execution and call the result hardened.

04

### Use the canonical signed installer and verify it.

Download the x64 or ARM64 asset and checksum file from the project's latest release. Compare the SHA-256 value before opening it, retain SmartScreen and Defender checks, and reject a binary whose publisher or digest does not match.

```
Get-FileHash .\OpenClawTray-Setup-x64.exe -Algorithm SHA256
```

- [Latest release →](https://github.com/openclaw/openclaw-windows-node/releases/latest)

- [Official setup guide →](https://github.com/openclaw/openclaw-windows-node/blob/main/docs/SETUP.md)

05

### Connect to the existing Spark gateway.

Choose the remote or existing-gateway route in Windows Hub. Keep the gateway on Spark bound to loopback and let the Hub manage an SSH tunnel, or use a private Tailnet with authentication. The local WSL gateway is a fallback for people without an always-on host, not the default topology here.

06

### Pair, then allow only harmless Windows commands.

Approve the Windows node from the Spark gateway. Begin with notifications and device information/status. Leave command execution, screen capture, camera, location, speech and browser control disabled until a named workflow needs them and both the gateway policy and MXC policy deny everything else.

```
openclaw devices list
openclaw devices approve <device-id>
```

**Starter allowlist**

`system.notify · device.info · device.status`

OpenClaw requires exact command names. If `system.run` is later enabled, keep the separate Windows node execution policy default-deny as well.

07

### Test the deny path before startup automation.

Run ten harmless actions, inspect the activity and MXC diagnostics, disconnect the gateway, reject an unexpected pairing request, attempt access to a deliberately denied file and domain, and confirm that every disabled node capability fails. Re-run this after Windows, Hub, MXC, OpenShell or gateway updates.

Insider policy

### Required for this bleeding-edge path; not proof of safety.

Stable Windows can run OpenClaw and MXC's lighter process backend, but the current session-isolation backend and several Windows AI connector/workspace policies are Insider-era features. That makes Insider non-optional for the setup described here. It also makes the setup a lab: Microsoft explicitly says current MXC profiles have known over-permissive cases and should not yet be treated as complete security boundaries. Keep Windows permissions, OpenClaw allowlists, network isolation, backups and human approval in place.

- [OpenClaw + MXC announcement →](https://blogs.windows.com/windowsdeveloper/2026/06/02/windows-platform-security-for-ai-agents/)

- [MXC preview warning →](https://github.com/microsoft/mxc)

- [Insider policy surface →](https://learn.microsoft.com/en-us/windows/client-management/mdm/policy-csp-windowsai)

Quality-of-life layer

## Add convenience without hiding authority.

OpenClaw extensions execute inside a trust boundary. Install fewer, inspect them and pin versions where the package route supports it.

Native companions

### Hub, menu bar or Linux tray

Use the platform companion for health, chat and notifications without moving the gateway. Enable a desktop node only when the agent must act on that device; operator access and node authority are separate choices.

- [Compare platforms →](https://docs.openclaw.ai/platforms)

Remote access

### Managed SSH tunnel or Tailscale

Prefer the Hub's SSH-tunnel support or Tailscale Serve over firewalling the gateway port open. Keep the gateway bound to loopback and authenticate every client.

- [OpenClaw remote access →](https://docs.openclaw.ai/gateway/remote)

High-trust browser tool

### OpenClaw Chrome extension

Use only when the workflow must operate an already signed-in browser tab. Share an explicit OpenClaw tab group, keep the relay on loopback and assume the agent can act with that tab's account permissions.

- [Extension security model →](https://docs.openclaw.ai/tools/chrome-extension)

Skills + plugins

### ClawHub, workspace-local first

Search by the exact capability you need, inspect source and scan state, then install into one workspace before promoting it. A popular package is still executable code, not a permission boundary.

- [ClawHub quickstart →](https://docs.openclaw.ai/clawhub/quickstart)

- [Manage plugins →](https://docs.openclaw.ai/plugins/manage-plugins)

Execution isolation

### OpenShell + platform controls

Put tool execution behind NVIDIA OpenShell where its backend fits, then layer the host controls: MXC preview on Windows, Bubblewrap or LXC on Linux, and Seatbelt-backed containment on macOS. Test denied files and destinations on the actual host.

- [OpenShell →](https://github.com/NVIDIA/OpenShell)

- [MXC status →](https://github.com/microsoft/mxc)

Operational habit

### One capability ledger

Record the package, version, owner, data it can read, actions it can take, token scopes and removal test. Re-run the ten-action test after every gateway, node or plugin update.

- [Use the ten-run method →](https://isaiuseful.com/guides.html.md#ten-run-evaluation)

<a id="topology"></a>

Preferred topology

## Your devices operate. Spark remembers and computes.

The agent does not disappear when a laptop sleeps. Gateway state, model serving and scheduled work remain together on the always-on host.

> Visual: Remote operator and DGX Spark architecture

**Visual reading order:**
1. **01 · Any operator** **Windows · macOS · Linux** Companion app, browser, CLI, mobile or messaging; desktop-node powers remain optional
2. SSH tunnel 18789 + 8000
3. **02 · DGX Spark** **OpenClaw or Hermes** Loopback-bound, authenticated, always-on sessions, memory, channels and tool orchestration
4. **03 · DGX Spark** **vLLM + optional Qdrant** OpenAI-compatible inference and private retrieval services; no direct LAN exposure
5. **04 · DGX Spark** **Training workspace** Stopped while serving needs the memory; versioned datasets, adapters, evaluations and manifests

### Default placement: gateway on Spark

OpenClaw's own model is one gateway with many clients: the gateway owns sessions, authentication profiles, channels and state. Hermes is similarly comfortable on a remote Linux host with its dashboard and messaging gateway exposed only through the private access layer. A local laptop gateway is the fallback for someone without an always-on server, not the target architecture here.

Local first · cloud by exception

## Scrub here. Escalate only the derivative.

When the laptop model cannot finish the hard part, keep the original and the identity map local. Send a stronger model only the smallest reviewed artifact that policy permits.

Vendor case study

### Bayer taught Phi the exceptions.

Microsoft says Bayer fine-tuned a small Phi model on proprietary crop-protection labels, regulatory rules and expert-authored Q&A. Labels can exceed 100 pages; Bayer reports early complex questions falling from days or weeks to under 30 seconds.

- [Read the Microsoft customer story →](https://www.microsoft.com/en/customers/story/25255-bayer-azure-phi)

Vendor case study

### Discovery Bank split one job into five.

Discovery fine-tuned five variants across Azure OpenAI 4o-mini and 4.1-mini for company language, SQL shape and workflow templates. It reports average response time dropping from five or six seconds to 1.5-2 seconds.

- [Read the Microsoft customer story →](https://www.microsoft.com/en/customers/story/26157-discovery-bank-azure-openai-in-foundry-models)

Laptop implication

### Teach the boundary before the model.

A downloaded model can flag candidate names, secrets, clauses and contextual identifiers without giving the source file to a model provider. It can also finish simple extraction or comparison locally. Escalation begins only when a harder model adds enough value to justify a reviewed data crossing.

- [Run the one-document test →](#offline-scrub)

> Visual: Local scrub and cloud escalation workflow

**Visual reading order:**
1. **01 · Local** **Classify** Route the source red, amber or green before any model sees it.
2. **02 · Local** **Detect twice** Use deterministic checks plus a local model for direct and contextual identifiers.
3. **03 · Local** **Replace + minimize** Create stable typed tokens; prefer a task brief over a redacted full copy.
4. **04 · Human gate** **Review the exact artifact** Approve recipient, purpose, region, retention and unresolved spans.
5. **05 · Route** **Finish or escalate** Use local output when it passes; otherwise send only an approved derivative.
6. **06 · Local** **Rejoin carefully** Inspect the answer and restore approved values without exposing the token map.

**The boundary that matters**

Redaction is not deletion and pseudonymization is not anonymity. The original, findings manifest and re-identification map stay local. Credentials are removed rather than tokenized, and the cloud service never receives the map. If context still identifies the person, client, transaction or project, the file remains red.

<a id="offline-scrub"></a>

The offline setup

### One document. One local model. No fallback.

During a connected maintenance window, install LM Studio and download one instruction-tuned local model, its runtime and any local embedding model the attachment workflow requires. Test with a synthetic file, close the app, move the authorized copy outside synced folders, then switch off Wi-Fi and unplug Ethernet before reopening it.

1. 01 **Load locally.** Choose only the downloaded model; start with an 8,192-token context and temperature 0.
2. 02 **Remove side doors.** Turn off tools, MCP, web search, plugins, cloud models and automatic fallback.
3. 03 **Keep the server closed.** Leave it off; if an integration needs it, bind only to `127.0.0.1` with authentication.
4. 04 **Return locations, not prose.** Ask first for exact span, page, category, reason, confidence and a proposed token.

**Local candidate-finding prompt**

`Work only with the attached local document. Do not use tools, web search, external sources or a remote model. Find DIRECT_IDENTIFIER, CREDENTIAL_OR_SECRET, REGULATED_DATA, COMMERCIAL_CONFIDENTIAL, QUASI_IDENTIFIER and UNCERTAIN spans. Return exact text, page or section, reason, confidence and a stable token such as [PERSON-001]. Do not rewrite yet. Do not infer missing identities.`

**Do not certify from document chat**

[RAG](https://isaiuseful.com/rag.html.md#model) may retrieve only selected passages. Long documents require page-by-page or overlapping-chunk inspection, deterministic secret and identifier checks, a fresh rescan of the derivative and human approval. Any unreadable or skipped content stops cloud routing.

- [LM Studio offline-operation documentation →](https://lmstudio.ai/docs/app/offline)

- [LM Studio network binding warning →](https://lmstudio.ai/docs/developer/core/server/serve-on-network)

Red · never ordinary cloud

### Keep the job local.

Credentials, identity evidence, health or KYC records, privileged advice, protected investigations, safety-critical material, raw client or board files, and anything with unclear authority or processor terms. Use approved private infrastructure or no model.

Amber · reviewed derivative

### Minimize, rescan, approve.

Internal contracts, narratives or reports that can lose direct identifiers and distinctive combinations without losing the task. Name the exact service, purpose, region, retention, logging, training and deletion terms before sending.

Green · approved content

### Still send less.

Published material, genuinely synthetic tests, checked templates or content explicitly approved for this external purpose. Inspect comments, metadata, tracked changes, hidden sheets, notes and embedded objects too.

Grok Build · July 2026

### The model obeyed. The product still moved the repository.

An independent wire analysis of Grok Build 0.2.93 recovered a never-read tracked file and Git history from a separate uploaded bundle after the prompt said not to open files. The researcher later reported that xAI disabled the path server-side. The test established transmission and storage in that setup, not training, employee access or universal behavior. Its durable lesson is that a prompt controls the task, not necessarily packaging, traces, sync or upload: “no training,” “no retention,” “no human review,” “no upload” and “local-only” are five different claims.

- [Inspect the original wire analysis →](https://gist.github.com/cereblab/dc9a40bc26120f4540e4e09b75ffb547)

Microsoft's lock-in test

### Exclusive is private. It is not automatically portable.

Microsoft says Azure-hosted direct-model inputs, outputs, embeddings and training data are not made available to model providers or used to improve foundation models without permission, and that a customer's fine-tuned model is exclusive to that customer. Those are useful privacy commitments. They do not by themselves make the resulting weights exportable or the workflow provider-independent.

- [Read the current Foundry data terms →](https://learn.microsoft.com/en-us/azure/foundry/responsible-ai/openai/data-privacy)

### Keep the learning outside the endpoint.

Store rules, source provenance, prompts, schemas, corrections, permission policy and an untouched evaluation set in exportable formats. Treat a fine-tune as a replaceable build artifact, not the only copy of what the company learned.

**The 30-day question**

Could we replace the model in 30 days using only the artifacts we can export today, then prove the replacement against the same holdout?

The guide + skill

## Take the boundary with you.

The walkthrough includes the full LM Studio pilot, red/amber/green router, verification gates, laptop-to-enterprise thresholds and portability test. The companion Codex skill prepares the derivative and review package, then stops before transmission.

- [Download the full guide](https://isaiuseful.com/downloads/sensitive-document-hybrid-guide.md)

- [Download the router skill](https://isaiuseful.com/downloads/scrub-sensitive-documents.zip)

- [Build local document search](https://isaiuseful.com/guides.html.md#private-documents)

<a id="spark-setup"></a>

DGX Spark build

## Build a service you can recover.

Use NVIDIA's Spark-specific playbooks. Generic CUDA, x86 container or desktop instructions can fail on ARM64/Blackwell even when the project itself supports NVIDIA GPUs.

01

### Baseline DGX OS.

Record the installed DGX OS release, driver, CUDA stack and firmware. Apply NVIDIA-recommended updates through DGX Dashboard, then verify health before adding services. Do not replace the OS with generic Ubuntu or Windows.

- [Update guide →](https://docs.nvidia.com/dgx/dgx-spark/os-and-component-update.html)

- [Release notes →](https://docs.nvidia.com/dgx/dgx-spark/release-notes.html)

- [CUDA libraries reference →](https://docs.nvidia.com/cuda-libraries/index.html)

<a id="spark-remote-access"></a>

02

### Set up private remote access.

After updating DGX OS, create a named non-root account, install an SSH key and choose the least-powerful route that covers the operator's work. Keep it on a trusted LAN or Tailnet, retain host firewall rules and never publish SSH or remote-desktop ports directly to the internet.

**NVIDIA Sync**

Managed SSH, application launches and forwarding; optional Tailscale off-LAN. No jump or bastion hosts.
**Direct SSH**

Terminal administration and explicit forwards. Confirm the host fingerprint and keep strict verification.
**Sunshine + Moonlight**

Full DGX OS desktop with hardware encoding and a software virtual display, including headless operation without an HDMI dummy plug.
**Community installer warning · checked 22 August 2026:** treat the linked repository as a bootstrap, not a reliable one-command finish. Its current checked revision still forces `DISPLAY=:0` , chooses a fixed Sunshine output and writes host options rejected by Sunshine `2026.516.143833` . If you use `./install.sh` , reboot and run `./after-install.sh` , follow the repair section before pairing Moonlight, or give the standalone Markdown runbook to Codex or another terminal-capable AI agent. For access beyond the LAN, use a private Tailnet and confirm a direct peer-to-peer path.

- [Follow the on-page post-install repair →](#sunshine-moonlight-repair)

- [Open the human/AI Markdown runbook →](https://isaiuseful.com/downloads/dgx-spark-sunshine-moonlight-repair.md)

- [Inspect the community installer →](https://github.com/seanGSISG/dgx-spark-sunshine-setup)

- [Download a Moonlight client →](https://moonlight-stream.org/)

- [Read the official Sunshine documentation →](https://docs.lizardbyte.dev/projects/sunshine/)

- [Set up NVIDIA Sync →](https://docs.nvidia.com/sync/latest/index.html)

- [Review NVIDIA's supported access options →](https://docs.nvidia.com/dgx/dgx-spark/system-overview.html)

03

### Start with the field-tested official Qwen3.8 FP8 route.

Use the [Qwen3.8-27B FP8 trial below](#qwen38-vllm) as the default Spark vLLM path, then compare another model only after the endpoint, tool loop and edit result pass a fixed acceptance set. Pin the container and official model revision, set explicit context, cache and concurrency budgets, and keep the endpoint on loopback.

- [Open the Qwen3.8 FP8 field route →](#qwen38-vllm)

- [Spark vLLM recipe →](https://build.nvidia.com/spark/vllm/instructions)

- [Community Spark vLLM Docker setup →](https://github.com/eugr/spark-vllm-docker)

- [vLLM docs →](https://docs.vllm.ai/en/stable/)

04

### Add retrieval only when the workflow needs it.

Start with Qdrant only if filters, hybrid search, persistence or multiple collections justify a service. Store the original document and page metadata with every chunk; back up source documents and collection configuration, then test a restore.

- [Use the private-document recipe →](https://isaiuseful.com/guides.html.md#private-documents)

05

### Place OpenClaw or Hermes on Spark.

Install one primary agent through its supported Linux path, point it at vLLM's loopback OpenAI-compatible endpoint and keep its sessions, memory and channels on persistent storage. OpenClaw uses its gateway service; Hermes can expose its dashboard and messaging gateway. Run both only when you have intentionally separated their state, ports and permissions.

- [OpenClaw on Linux →](https://docs.openclaw.ai/platforms/linux)

- [Hermes Agent →](https://github.com/NousResearch/hermes-agent)

06

### Lock down tools and data services.

Bind the gateway, vLLM and Qdrant to loopback unless a private container network requires otherwise. Run high-authority agent tools through an OpenShell sandbox where supported, then test its filesystem and outbound-network denies. Expose only the application endpoint the chosen operator route needs.

- [OpenShell →](https://github.com/NVIDIA/OpenShell)

- [Secure Qdrant →](https://qdrant.tech/documentation/operations/security/)

07

### Separate serving from training.

Do not let a training job silently evict or starve the always-on model. Use explicit service and training modes, drain requests, stop vLLM when the recipe needs the memory, checkpoint to persistent storage and restore serving from a known configuration.

08

### Prove recovery.

Back up only what cannot be recreated: gateway configuration and keys, source data, dataset manifests, adapters, evaluation sets and service definitions. Rebuild one clean service from the manifest before calling the system production-ready.

NVIDIA Sync + Docker · checked 24 August 2026

## Know what Sync installed before you update Open WebUI.

NVIDIA Sync is the launcher and SSH tunnel, not the container runtime. The Open WebUI container, bundled Ollama and both persistent data volumes live on the Spark.

Runtime

### The container is on Spark.

NVIDIA's remote recipe pulls `ghcr.io/open-webui/open-webui:ollama` and creates a Docker container named `open-webui` . Spark port `12000` maps to container port `8080` ; Sync forwards the same port to `localhost:12000` on the operator device.

Persistent state

### The data is in two named volumes.

`open-webui` is mounted at `/app/backend/data` for accounts, chats, settings and uploads. `open-webui-ollama` is mounted at `/root/.ollama` for Ollama model data. Docker manages their host paths; inspect and back them up instead of editing the mountpoints directly.

Launcher lifecycle

### Sync starts and stops the existing container.

The custom script creates the container only when it is absent. Later launches restart the same container, and closing the launcher stops it. Custom scripts are saved per remote device, so this launcher follows the same Spark across Sync clients; it is not copied automatically to every Spark.

Inspect first

### Ask Docker what exists, do not guess a folder.

```
docker ps -a --filter 'name=^/open-webui$'
docker inspect open-webui --format '{{.Config.Image}}'
docker inspect open-webui --format \
  '{{range .Mounts}}{{println .Name "->" .Destination}}{{end}}'
docker volume inspect open-webui open-webui-ollama

# Check the state before using docker exec.
running=$(docker inspect open-webui --format '{{.State.Running}}')
printf 'running=%s\n' "$running"
if [ "$running" = true ]; then
  docker exec open-webui ollama --version
else
  printf '%s\n' 'Ollama version check skipped: open-webui is stopped.'
fi
```

The volume inspection output includes Docker's current host mountpoint. Treat that path as implementation state, not a user-managed application folder.

`docker exec` requires a running container. The inspection block now reports the stopped state without failing. If you need the bundled Ollama version while `docker ps -a` shows `Exited (0)` , start Open WebUI from NVIDIA Sync first (preferred), or use the one-shot check below. A stopped container does not mean either named volume is missing.

When the Sync launcher is stopped, an optional one-shot check restores the prior stopped state:

```
was_running=$(docker inspect open-webui --format '{{.State.Running}}')
if [ "$was_running" != true ]; then docker start open-webui >/dev/null; fi
docker exec open-webui ollama --version
if [ "$was_running" != true ]; then docker stop open-webui >/dev/null; fi
```

Safe refresh

### Pull, remove the old container and let Sync recreate it.

Pulling the moving `:ollama` tag changes the local image, but it does not replace an existing container. Before an update, review the Open WebUI release notes and back up at least the application volume; the Ollama model-volume backup is optional but can be very large. Replace `YYYY-MM-DD` with the actual backup date.

```
mkdir -p "$PWD/open-webui-backups"
docker run --rm \
  -v open-webui:/data:ro \
  -v "$PWD/open-webui-backups:/backup" \
  alpine tar czf /backup/open-webui-YYYY-MM-DD.tar.gz /data

# Optional: preserve downloaded Ollama models as well.
docker run --rm \
  -v open-webui-ollama:/data:ro \
  -v "$PWD/open-webui-backups:/backup" \
  alpine tar czf /backup/ollama-models-YYYY-MM-DD.tar.gz /data
```

Stop Open WebUI with the `x` beside its Sync launcher. In a Sync terminal, pull the image and remove only the stopped container. Do not remove either named volume.

```
docker pull ghcr.io/open-webui/open-webui:ollama
docker rm open-webui
```

Launch Open WebUI from Sync again. Its script now creates a new container from the pulled image and reattaches both volumes. Then inspect startup and the bundled runtime:

```
docker ps --filter 'name=^/open-webui$'
docker logs --tail 100 open-webui
docker exec open-webui ollama --version
```

**Two update caveats:** the NVIDIA launch script does not persist a `WEBUI_SECRET_KEY` , so recreating the container can invalidate active browser sessions even though accounts and chats remain. For a shared instance, pin a versioned `-ollama` tag, preserve a stable secret and test migrations before switching. Never use `docker volume rm` as part of an update.

- [Open WebUI update and backup guide →](https://docs.openwebui.com/getting-started/updating/)

- [Docker named-volume behavior →](https://docs.docker.com/engine/storage/volumes/)

Bundled Ollama

### Treat its version as part of the image.

The combined image carries its own Ollama build. It can lag the newest standalone release, and current Ollama model metadata can declare a minimum required runtime version. If a new model refuses to pull or load, compare `docker exec open-webui ollama --version` with the model requirement before diagnosing the GPU. Updating the combined image is the clean route; replacing Ollama inside a running container creates an unrepeatable installation.

**Browser field note · August 2026:** on one Sync-deployed Spark, Open WebUI's model-pull control worked in Safari but not in Chromium-based browsers. This is a field observation, not a browser support matrix. If the pull action is missing or inert, hard-refresh, try Safari, or bypass the UI with `docker exec open-webui ollama pull <model:tag>` .

- [Ollama minimum-version metadata →](https://docs.ollama.com/modelfile#requires)

- [NVIDIA's Sync recipe →](https://build.nvidia.com/spark/open-webui/sync)

- [How Sync custom applications are stored →](https://docs.nvidia.com/sync/latest/applications.html)

Standalone Ollama

### Run a separate Ollama service when a model needs a newer runtime.

When a model's minimum required version outpaces the bundled build, decoupling Ollama from Open WebUI is a practical lower-complexity step before a vLLM migration. That is an operational comparison, not a vendor support or cost guarantee. Decoupling is what upgrades Ollama, not the tag switch: a moving `:main` tag tracks Open WebUI on its own schedule, while the host Ollama service tracks the standalone release independently.

Install standalone Ollama with the official installer. On the Spark, NVIDIA's recipe uses that same installer, and it detects the GB10 Blackwell GPU.

```
curl -fsSL https://ollama.com/install.sh | sh
```

Stop the newly installed service before copying. Create the destination for the `ollama` service account, copy the model store, restore ownership, then verify one real GPU-backed inference before switching Open WebUI. The old named volume remains untouched as the immediate rollback copy.

```
sudo systemctl stop ollama
sudo install -d -o ollama -g ollama /usr/share/ollama/.ollama
docker run --rm \
  -v open-webui-ollama:/data:ro \
  -v /usr/share/ollama/.ollama:/store \
  alpine cp -a /data/. /store/
sudo chown -R ollama:ollama /usr/share/ollama/.ollama
sudo systemctl start ollama

ollama list
model_tag='replace-with-model-tag'
ollama run "$model_tag" 'Reply with exactly: migration verified'
ollama ps
```

Then stop the Sync launcher, rename the retained combined container so the `open-webui` name is free, and start the ordinary image against the host service. Replace `YYYY-MM-DD` and the secret placeholder before running the commands. Docker host networking keeps Open WebUI on host port 12000 and lets it reach loopback-bound Ollama without exposing port 11434 to the LAN.

```
docker stop open-webui
docker rename open-webui open-webui-bundled-rollback-YYYY-MM-DD

docker run -d \
  --network=host \
  -e PORT=12000 \
  -e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
  -e WEBUI_SECRET_KEY='<replace-with-a-long-random-secret>' \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart unless-stopped \
  ghcr.io/open-webui/open-webui:main
```

**Verify before deleting anything:** check `docker logs --tail 100 open-webui` , sign in through Sync, confirm model discovery, send one inference request and restart the launcher. Keep `open-webui-bundled-rollback-YYYY-MM-DD` and `open-webui-ollama` until that acceptance check passes.

- [Ollama Linux installer →](https://docs.ollama.com/linux)

- [Ollama service model store →](https://docs.ollama.com/faq)

- [NVIDIA's Spark Ollama playbook →](https://build.nvidia.com/spark/live-vlm-webui/instructions)

- [Docker container rename →](https://docs.docker.com/reference/cli/docker/container/rename/)

- [Open WebUI host Ollama connection →](https://docs.openwebui.com/troubleshooting/connection-error/)

- [Open WebUI env configuration →](https://docs.openwebui.com/reference/env-configuration/)

<a id="qwen38-vllm"></a>

Default recommended local-agent trial · field-tested 25 August 2026

### Start with Qwen3.8-27B NVFP4, a BF16 language-model head and DFlash 2.

The current DGX Spark run pairs `RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead` with the `incoai/Qwen3.8-27B-DFlash2` drafter in a digest-pinned SGLang container. Two top-level VS Code agent tasks produced 161 successful local model calls. SGLang returned 166 successful chat-completion responses when five separate readiness and acceptance calls are included, with no inference HTTP failure in the recorded run.

Across 1,242 server windows with one active request, generation throughput averaged 20.4 tok/s and had an 18.6 tok/s median. Another 102 windows with two active requests averaged 36.6 aggregate tok/s. The DFlash windows averaged a 3.39-token acceptance length and a 34.2% acceptance rate. VS Code round-trip latency had a 20.1-second median and a 278.8-second 95th percentile across the 161 calls; that client timing includes prompt handling, reasoning, generation, transport and harness overhead.

During the VS Code interval, SGLang's prefill records represented 14,316,770 prompt tokens: 290,146 newly processed tokens and 14,026,624 reused cache tokens, for a 98.0% cached share. Across 84 complete uncached 2,048-token chunks, input throughput had a 1,005 tok/s median and a 1,239.1 tok/s mean. High cache reuse is useful for an agent loop, but it also means this session is not a cold-prompt benchmark.

The recipe below pins the tested container digest and the locally observed target and draft revisions. It binds the endpoint to loopback, limits the container to 96 GB and leaves system memory outside the service. Re-test before changing any pin, memory fraction or speculative setting.

```
cache_root="${XDG_CACHE_HOME:-${HOME}/.cache}"
mkdir -p "$cache_root/huggingface" "$cache_root/sglang-qwen38"

docker run -d \
  --name sglang-qwen38 \
  --restart unless-stopped \
  --gpus all \
  --network host \
  --ipc host \
  --shm-size 16g \
  --memory 96g \
  --memory-swap 96g \
  --ulimit memlock=-1 \
  --ulimit stack=67108864 \
  -e HF_HOME=/root/.cache/huggingface \
  -e SGLANG_CACHE_DIR=/root/.cache/sglang \
  -v "$cache_root/huggingface:/root/.cache/huggingface" \
  -v "$cache_root/sglang-qwen38:/root/.cache/sglang" \
  --entrypoint "" \
  lmsysorg/sglang@sha256:616a3e97f45191af975896cfa644279096cb31bd408a071c2e99ca7209c3cafe \
  python3 -m sglang.launch_server \
    --trust-remote-code \
    --model-path RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead \
    --revision 009632fef96dd349150baa780c984e62e70e91fe \
    --served-model-name qwen38-dev \
    --context-length 262144 \
    --tp-size 1 \
    --host 127.0.0.1 \
    --port 30000 \
    --mem-fraction-static 0.60 \
    --kv-cache-dtype auto \
    --attention-backend flashinfer \
    --chunked-prefill-size 2048 \
    --max-running-requests 4 \
    --mamba-radix-cache-strategy extra_buffer_lazy \
    --mamba-ssm-dtype float32 \
    --max-mamba-cache-size 16 \
    --cuda-graph-max-bs-decode 4 \
    --disable-prefill-cuda-graph \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2 \
    --speculative-draft-model-revision dedf8df68adfb1afeaf7b7480c0a0243108177b4 \
    --speculative-num-draft-tokens 8 \
    --speculative-draft-model-quantization unquant \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --enable-metrics \
    --enable-cache-report \
    --sleep-on-idle
```

The tested startup loaded 22.68 GB for the target and 3.64 GB for the draft. It allocated 19.48 GB of FP8 target KV cache and 12.18 GB of BF16 draft KV cache for 638,179 tokens, then reported 41.52 GB available after CUDA graph capture. End-to-end engine startup took 165.2 seconds. The 638,179-token pool is shared across active requests; four configured running requests do not mean four simultaneous 262,144-token contexts.

Check health and the served model before connecting an interface. Then run a bounded acceptance task that must read once, edit once and verify the diff. Stop if the same tool request repeats, if no edit appears, or if the harness changes the requested model name.

```
curl -fsS http://127.0.0.1:30000/health
curl -fsS http://127.0.0.1:30000/v1/models
```

**Evidence boundary:** 20.4 tok/s is the arithmetic mean of SGLang's reported generation-throughput values for 1,242 windows with one active request, not a standalone vendor benchmark; 18.6 tok/s is the median. The 1,005 tok/s prefill value is the median of 84 full uncached chunks. The run establishes local compatibility and useful agent-loop behavior, not general model quality or production reliability.

- [Inspect the tested NVFP4 target →](https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead)

- [Inspect the DFlash 2 drafter →](https://huggingface.co/incoai/Qwen3.8-27B-DFlash2)

- [Read SGLang speculative-decoding guidance →](https://docs.sglang.ai/advanced_features/speculative_decoding.html)

Open WebUI + SGLang

## Decouple the interface after Qwen passes the acceptance task.

Once the loopback SGLang endpoint is proven, run the ordinary Open WebUI image and connect it to the private OpenAI-compatible URL. Keep the `open-webui` application volume; retire the bundled-Ollama image and its model volume only after exporting anything you still need.

- [Connect an OpenAI-compatible endpoint](https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai-compatible/)

- [Open SGLang's speculative-decoding guide](https://docs.sglang.ai/advanced_features/speculative_decoding.html)

<a id="sunshine-moonlight-repair"></a>

Sunshine + Moonlight · verified repair

## Bind Sunshine to the real X11 desktop, not an example display.

A user service can be active while capture and every encoder probe fail. Fix the graphical-session environment first, select the connected numeric Sunshine display ID, then judge success from encoder discovery.

**Compatibility warning · checked 22 August 2026:** the community installer's current checked revision still hardcodes `DISPLAY=:0` , defaults `output_name = 0` and generates settings rejected by the reproduced Sunshine release. Use it only after inspecting the scripts, then apply this post-install correction. The verified desktop used `:1` , but that is evidence against hardcoding, not a new default.

- [Inspect the checked service override →](https://github.com/seanGSISG/dgx-spark-sunshine-setup/blob/5666b64c698a549869275a32af21d3af317f8ed6/templates/sunshine-override.conf)

- [Inspect the checked config template →](https://github.com/seanGSISG/dgx-spark-sunshine-setup/blob/5666b64c698a549869275a32af21d3af317f8ed6/templates/sunshine.conf.template)

| Concept | What it provides | What it does not provide |
| --- | --- | --- |
| Physical-monitor mode | A real connected output in a logged-in GNOME X11 desktop. | A desktop before the user logs in. |
| Virtual/headless mode | An Xorg virtual display without an HDMI dummy plug. | A desktop session by itself; GDM autologin or another session creator is still required. |
| User-manager lingering | A systemd user manager before login. | An X server, GNOME desktop or capturable display. |
| Graphical autologin | An unattended X11 desktop after boot. | The same local-login security as an interactive sign-in; enable it deliberately. |

01 · Discover

### Verify the logged-in X11 session.

Run these commands in a terminal inside the logged-in GNOME desktop. A plain SSH shell may not carry the graphical values.

```
echo "$XDG_SESSION_TYPE"
echo "$DISPLAY"
echo "$XAUTHORITY"
xrandr --listmonitors
systemctl --user show-environment |
  grep -E '^(DISPLAY|XAUTHORITY|XDG_SESSION_TYPE)='
```

Continue only with the display belonging to the logged-in X11 desktop. GDM can use `:0` for its greeter and `:1` for the user session. Never force either value in generic instructions.

02 · Inherit

### Publish the live X11 values on every graphical login.

Add this exact line to `~/.xprofile` , then run it once from the current graphical terminal:

```
dbus-update-activation-environment --systemd DISPLAY XAUTHORITY
```

Remove every hardcoded `Environment="DISPLAY=..."` shown by `systemctl --user cat sunshine.service` . In the user-service drop-in, keep the session ordering and replace any old pre-start check with this one:

```
[Unit]
After=graphical-session.target
Wants=graphical-session.target
StartLimitBurst=5
StartLimitIntervalSec=60

[Service]
ExecStartPre=
ExecStartPre=/bin/sh -c 'i=0; while [ "$$i" -lt 60 ]; do env="$$(systemctl --user show-environment)"; printf "%%s\n" "$$env" | grep -q "^DISPLAY=" && printf "%%s\n" "$$env" | grep -q "^XAUTHORITY=" && exit 0; i=$$((i + 1)); sleep 1; done; echo "DISPLAY/XAUTHORITY not available in systemd user environment" >&2; exit 1'
```

The format string must be `%%s` . systemd treats a single `%` as a specifier, so `%s` can be expanded before the shell sees it.

03 · Configure

### Keep only supported GB10 host settings.

Back up `~/.config/sunshine/sunshine.conf` . Remove `resolutions` , `nvenc_tuning_info` , `bitrate` , `channels` , `nvenc_multipass` , `nvenc_rc` , `fps` and `codec` when the tested release reports them as unrecognized. Set requested resolution, frame rate, bitrate and codec in Moonlight. Preserve site-specific `csrf_allowed_origins` locally.

```
adapter_name = /dev/dri/renderD128
capture = x11
encoder = nvenc
min_threads = 2
nvenc_preset = 1
nvenc_spatial_aq = 1
nvenc_twopass = quarter_res
nvenc_vbv_increase = 0
output_name = <numeric ID reported as connected>
qp = 20
```

- [Sunshine configuration reference →](https://docs.lizardbyte.dev/projects/sunshine/latest/md_docs_2configuration.html)

04 · Select + verify

### Read the output ID from Sunshine's own startup log.

```
journalctl --user -u sunshine.service -n 200 --no-pager |
  grep "Detected display"
```

Choose the numeric `id: <N>` on the row that says `connected: true` . Do not copy an example number, assume HDMI is ID 0, or use a connector string such as `USB-C-2` ; the reproduced Sunshine build converted that string into an invalid display number.

```
systemctl --user daemon-reload
systemctl --user enable sunshine.service
systemctl --user restart sunshine.service

systemctl --user is-enabled sunshine.service
systemctl --user is-active sunshine.service

journalctl --user -u sunshine.service --since "2 minutes ago" --no-pager |
  grep -E \
  "Detected display|Configuring selected display|Found .*encoder|System tray|Fatal|Unable to"
```

Success means the connected ID is selected and the log finds `h264_nvenc` , `hevc_nvenc` and `av1_nvenc` . If `cuda::cuda_t doesn't support any format other than AV_PIX_FMT_NV12` appears during probing but all three encoders are found afterward, it is not the startup failure.

**Tray troubleshooting:** correct `DISPLAY` and `XAUTHORITY` first, then confirm GNOME has an AppIndicator extension such as `ubuntu-appindicators@ubuntu.com` . Use `system_tray = disabled` only for a genuinely headless system that does not need an icon.

Manual or AI-assisted

## Use the same evidence-gated repair.

The standalone Markdown begins with read-only discovery, backs up the existing service and configuration, defines stop conditions for an AI operator, includes the exact systemd escaping, and ends with rollback plus log-based acceptance checks.

- [Open the repair runbook](https://isaiuseful.com/downloads/dgx-spark-sunshine-moonlight-repair.md)

- [Inspect the community installer](https://github.com/seanGSISG/dgx-spark-sunshine-setup)

- [Get Moonlight](https://moonlight-stream.org/)

<a id="spark-training"></a>

Fine-tuning on Spark

## Use the smallest recipe that answers the question.

NVIDIA publishes PyTorch and NeMo fine-tuning playbooks for Spark. Their example model sizes describe tested recipes, not universal capacity or speed guarantees.

| Question | First method | NVIDIA Spark example | Artifact to keep | Gate |
| --- | --- | --- | --- | --- |
| Does the data pipeline work? | Small LoRA run | 8B LoRA path | Adapter + run manifest | Loss is sane and sample outputs improve without obvious memorization. |
| Can a larger model learn the behavior? | LoRA or QLoRA | 70B LoRA / 70B QLoRA paths | Adapter, tokenizer config and holdout results | Beats the unchanged base model on the untouched test set. |
| Is full-weight training justified? | Full SFT only after adapter evidence | 3B full SFT path | Checkpoint + reproducible environment | Material gain over LoRA, retention tests pass and operating cost is acceptable. |

**Official starting points:** [PyTorch fine-tuning](https://build.nvidia.com/spark/pytorch-fine-tune/instructions) · [NeMo fine-tuning](https://build.nvidia.com/spark/nemo-fine-tune/instructions) · [dataset and evaluation checklist](https://isaiuseful.com/guides.html.md#finetuning) .

No Spark yet

## Keep the topology; shrink the host.

Run the gateway on the most reliable machine you already own: WSL 2 on Windows, a launchd service on macOS, or a systemd user service on Linux. Keep it loopback-bound, use a small local model or hosted endpoint and operate it from the same companion, browser and messaging surfaces. This proves the workflow; it does not reproduce Spark's unified memory or sustained service role.

- [Choose an OS path](https://docs.openclaw.ai/platforms)

- [WSL networking](https://learn.microsoft.com/en-us/windows/wsl/networking)

- [Compare hardware](https://isaiuseful.com/guides.html.md#hardware)

Quick answers

## Frequently asked questions about DGX Spark

Power, unified memory, networking, remote access and troubleshooting answers checked against current NVIDIA documentation.

### What is the expected DGX Spark power draw?

Budget for the included **240W** power supply. NVIDIA specifies a **140W TDP** for the GB10 SoC and reserves **100W** for ConnectX-7, Wi-Fi, storage, USB-C and other system components. Its EU technical disclosure reports **233.2W maximum** and **38W idle** under the stated test method. The power shown by `nvidia-smi` is not whole-system power at the wall.

- [NVIDIA power requirements →](https://docs.nvidia.com/dgx/dgx-spark/hardware.html#power-requirements)

- [NVIDIA measured-power disclosure →](https://docs.nvidia.com/dgx/dgx-spark/compliance.html)

### Why does `nvidia-smi` report “Memory-Usage: Not Supported”?

This is expected. DGX Spark’s integrated GPU shares unified system memory instead of having dedicated framebuffer memory, so `nvidia-smi` does not provide the usual aggregate VRAM-usage field. Use `top` , `htop` , `free -h` or DGX Dashboard for system-memory monitoring. Open the dashboard locally at `http://localhost:11000` ; remote access needs NVIDIA Sync or an SSH tunnel.

- [NVIDIA known issue →](https://docs.nvidia.com/dgx/dgx-spark/known-issues.html#nvidia-smi-reports-memory-usage-not-supported)

- [DGX Dashboard access →](https://docs.nvidia.com/dgx/dgx-spark/dgx-dashboard.html)

### Why does an application run out of memory below the 128 GB capacity?

Unified-memory applications can report less allocatable memory than the system can reclaim, and some software does not yet account correctly for swap or cache. Start with `free -h` , the process allocations and swap use. Linux normally reclaims clean cache automatically. For a controlled diagnostic only, not routine memory management, sync writes and drop reclaimable cache with:

```
sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches'
```

Dropping caches can create significant I/O and CPU work; stop the affected workload first and remeasure before treating cache as the cause.

- [NVIDIA unified-memory guidance →](https://docs.nvidia.com/dgx/dgx-spark/known-issues.html#guidance-for-reporting-memory-resources-with-unified-memory-architecture)

- [Linux drop-caches warning →](https://docs.kernel.org/admin-guide/sysctl/vm.html#drop-caches)

### Why does NVIDIA Sync say “Host already exists” after I deleted a device?

A stale SSH `Host` alias may remain. Back up the file, remove only the exact stale `Host` block, then add the device again. Current Sync documentation uses `~/.ssh/config` on macOS and Ubuntu and `C:\Users\<username>\.ssh\config` on Windows. Older Sync builds may also have a managed file at the paths below.

```
Windows: C:\Users\<username>\AppData\Local\NVIDIA Corporation\Sync\config\ssh_config
macOS:   /Users/<username>/Library/Application Support/NVIDIA/Sync/config/ssh_config
Linux:   /home/<username>/.config/NVIDIA/Sync/config/ssh_config
```

- [Review NVIDIA Sync SSH aliases →](https://docs.nvidia.com/sync/latest/direct-connections.html#importing-an-existing-ssh-configuration)

### Is GPUDirect RDMA supported on DGX Spark?

No. DGX Spark’s unified-memory architecture does not support GPUDirect RDMA, `nvidia-peermem` , `dma-buf` or GDRCopy for CUDA device allocations. A portable application should query `CU_DEVICE_ATTRIBUTE_GPU_DIRECT_RDMA_SUPPORTED` and `CU_DEVICE_ATTRIBUTE_DMA_BUF_SUPPORT` , then use a supported fallback. For Linux `ibverbs` applications, NVIDIA suggests `cudaHostAlloc` memory registered with `ibv_reg_mr` .

- [NVIDIA DGX Spark porting guidance →](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/cuda.html#gpudirect-rdma)

### How do I set up DGX Spark after moving it to another location?

Connect Ethernet first. Try the Spark’s `spark-hostname.local` address; if mDNS is unavailable, find its wired IP in the router. Connect with NVIDIA Sync or SSH, identify the Wi-Fi device, join the new network without putting the password in shell history, and note the Wi-Fi address before unplugging Ethernet:

```
nmcli device status
sudo nmcli --ask device wifi connect "<wifi-name>"
ip -4 address show
```

- [NVIDIA Sync direct connections →](https://docs.nvidia.com/sync/latest/direct-connections.html)

- [NVIDIA network setup guidance →](https://docs.nvidia.com/dgx/dgx-spark/first-boot.html)

### Why are CPU-bound NVCC processes slow?

CUDA sources built through CMake can miss OpenMP compiler support. Pass `-Xcompiler=-fopenmp` for CUDA compilation, rebuild and remeasure the CPU-bound section.

```
# CMakeLists.txt
target_compile_options(mytarget PRIVATE
  $<$<COMPILE_LANGUAGE:CUDA>:-Xcompiler=-fopenmp>
)

# Command line
nvcc -Xcompiler=-fopenmp ...
```

- [NVIDIA DGX Spark compiling guide →](https://docs.nvidia.com/dgx/dgx-spark-porting-guide/porting/compilation.html)

### Can I reset a lost BIOS or UEFI Administrator password?

There is no documented user reset for a lost firmware Administrator password. Contact NVIDIA hardware support for a Founders Edition system or the device manufacturer for an OEM model. Do not assume that an OS recovery or “Restore Defaults” in UEFI clears the security credential.

- [Choose the correct support route →](https://docs.nvidia.com/dgx/dgx-spark/support.html)

- [DGX Spark UEFI security settings →](https://docs.nvidia.com/dgx/dgx-spark-uefi/security-tab.html)

### Why is the ConnectX-7 module missing from `lspci` ?

The January 2026 DGX OS release added ConnectX-7 hot-plug power management. With it enabled, the adapter can stay powered down until a QSFP cable is inserted, saving up to 18W. Connecting the cable activates the adapter and increases system power and temperature. To keep ConnectX-7 active without hot-plug power saving, remove the marker; recreate it to re-enable the feature:

```
sudo rm -f /etc/nvidia/cx7-hotplug-enabled
sudo touch /etc/nvidia/cx7-hotplug-enabled
```

- [NVIDIA January 2026 release note →](https://docs.nvidia.com/dgx/dgx-spark/release-notes.html#january-2026-release)

### How many DGX Spark systems can I cluster?

Current NVIDIA Sync supports **two or three systems** with direct 200 Gb/s QSFP cabling, or **up to four systems through a switch** . Use an approved cable and the NVIDIA Sync Cluster Assistant or the matching two-node, three-node or switched NVIDIA playbook; four devices are not supported as a direct-cabled topology.

- [ConnectX-7 cabling and playbooks →](https://docs.nvidia.com/dgx/dgx-spark/spark-clustering.html)

- [NVIDIA Sync Cluster Assistant →](https://docs.nvidia.com/sync/latest/cluster-assistant.html)

The operating rule
Keep the agent on the server. Grant every client capability separately.

- [Choose your operator](#operators)

- [Build the Spark host](#spark-setup)

- [Choose the rest of the stack](https://isaiuseful.com/tools.html.md)
