04 / AI infrastructure Merged upstream
Local AI SystemsA self-hosted AI modelfor my coding tools.
My coding tools use GLM-5.3-Flash, a large AI model that runs on two NVIDIA DGX Spark computers I own. I set it up with a community recipe, a public guide for this model and these computers. The recipe now includes two of my contributions.
- Role
- Operator and recipe contributor
- Built
- September 2026
- Software
- vLLM, Docker, Linux, Python
- Runs on
- Two NVIDIA DGX Spark computers
The request was 262,144 tokens long. Tokens are the word pieces a model reads. With the fix, 95% of repeats started within 2.98 s.
The problem
Startup could stall, and a repeated long request took over three minutes.
Processor
Graphics chip
Shared memory
The modelTemporary file copiesFree memory
Model start
–
Starting
Can stallwith no free memory left
Each repeat of a long request
First time191.7 s
Same request again191.7 s
0% of the finished work reused
Each computer's processor and graphics chip share one memory.
Loading leaves temporary file copies that fill the free memory.
With no free memory left, startup can stall.
Each repeat of a long request took 191.7 s again.
Five design choices fix both problems.
What goes wrong, in technical terms
A DGX Spark has one pool of memory that the CPU and the GPU share. Loading the model's weights floods that pool with file cache. The GPU driver can then fail to find a free block during vLLM's startup profile, and the vLLM process spins.
The recipe's prefix cache reported zero hits, even for an exact repeat of the same prompt. So every repeat of a 262,144-token prompt was processed again from the start, which takes about three minutes.
The approach
I run the community recipe and build the parts around it.
My partOther people's work
Coding tools send requests to one shared address, an API.
Two computers each hold half the model and answer together.
The answer returns to the tool through the same API.
Other people made the model, the software that runs it, and the recipe. I built the shared API, health checks, automatic startup, recovery, and the fixes on this page.
What runs where, in technical terms
I run the community recipe for GLM-5.3-Flash at tensor parallel 2: vLLM in Docker, 4-bit NVFP4 weights, an fp8 KV cache, and the DFlash2 speculative drafter. NVFP4 is NVIDIA's 4-bit floating-point format. The KV cache holds the attention keys and values of tokens the model has already read. One OpenAI-compatible API serves my coding tools.
Decisions
Five design choices fixed the stalled starts and the slow repeats.
Shared memory
Guard onGuard off
The modelTemporary file copiesFree memory
Model start
Loading
Would failif memory were held back
Started8 GiB kept free
Reloads from disk
–
0in a full test
Repeat answers in
–
191.7 sfirst time
2.98 swas 191.7 s
Power
78 Wbefore the limit
47–50 Wwas 78 W, same speed
A guard clears file copies to keep 8 GiB free.
Holding memory back would fail startup. I kept the default.
Once loaded, the guard stops and a practice request runs.
My cache fix reuses work: repeats start in 2.98 s.
Limiting the graphics chip's clock cut power to 47–50 W.
With all five, the model starts, a full test runs with no reloads, repeats start in 2.98 s, and the graphics chip draws 47–50 W.
How each choice works, in technical terms
Guard free memory during the load. The load fails when the free memory blocks run out. A per-node guard drops caches and compacts memory whenever free memory falls under 8 GiB. It checks every second, from launch until the server is ready.
Leave the kernel reserve alone. Raising the kernel's free-memory reserve looks like a fix. But vLLM's startup check reads available memory, and the reserve lowers it by about 6.2 GiB, so the check fails before the model loads.
Stop the guard before it causes faults. Major page faults came from two places: the first use of the OpenSSL encryption libraries, and the guard's own cache drops, which evicted mapped weight pages. I stop the worker's guard once the head node is ready, then run one warm-up pass. After that, both nodes had zero major faults through a full qualification run at 262,144-token context.
Patch the cache at the smallest point. The drafter keeps its own short sliding-window cache. vLLM treated every cache group as a drafter group, so the draft window's short hit cut the shared hit to zero. My patch flags only the drafter group and never lets it shorten the hit that the model's own groups agree on.
Cap the GPU clock without losing speed. With the cap, decoding stayed at 55.6 tokens per second. Power fell from 78 W to 47–50 W, and temperature from 86 °C to 69–71 °C.
Evidence
The community recipe now includes my cache fix and memory report.
Results of one full test
- 98.6%of the finished work reused on a repeatwas 0%
- 2.98 suntil the answer to a repeat startswas 191.7 s
- 55.6tokens per second of output
Community recipe by tonyd2wild
- MergedMy cache fixPull request 18 · September 16, 2026
- MergedMy memory reportPull request 19 · September 16, 2026
vLLM, the software that runs the model
- In reviewLower memory use while loadingPull request 55103 · extra memory in a test at one eighth of full size
These results come from one full test on both computers.
The recipe's maintainer, tonyd2wild, accepted two of my proposed changes.
A third change, to vLLM, is in review.
The recipe includes two of my changes. A third change, to vLLM, the software that runs the model, is in review.
Full test results, in technical terms
| Measure | Result |
|---|---|
| Prefix-cache hit rate, repeated prompt | 0.986, from 0 before the fix |
| Cached tokens on the repeat | 258,048 |
| Time to first token, cold | 191.7 s |
| Time to first token, cached (p95) | 2.98 s |
| Decode speed, one stream (p50) | 55.6 tokens/s |
| Drafter acceptance | 0.928 |
The test sent 262,144‑token requests one at a time. The recipe's maintainer merged my two pull requests on September 16, 2026. In a test at one eighth of full size, the vLLM change cut the extra memory used during loading from 1,011 MiB to 3 MiB.
On my memory report
"This is the most useful GB10 memory write-up anyone has sent us"