04 / AI infrastructure Merged upstream

Local AI SystemsA self-hosted AI modelfor my coding tools.

My coding tools use GLM-5.3-Flash, a large AI model that runs on two NVIDIA DGX Spark computers I own. I set it up with a community recipe, a public guide for this model and these computers. The recipe now includes two of my contributions.

Role
Operator and recipe contributor
Built
September 2026
Software
vLLM, Docker, Linux, Python
Runs on
Two NVIDIA DGX Spark computers
Time until the answer starts when the same long request is sent again

The request was 262,144 tokens long. Tokens are the word pieces a model reads. With the fix, 95% of repeats started within 2.98 s.

Startup could stall, and a repeated long request took over three minutes.

Shared memory and the cache, before the fixesIllustrative

Five design choices fix both problems.

What goes wrong, in technical terms

A DGX Spark has one pool of memory that the CPU and the GPU share. Loading the model's weights floods that pool with file cache. The GPU driver can then fail to find a free block during vLLM's startup profile, and the vLLM process spins.

The recipe's prefix cache reported zero hits, even for an exact repeat of the same prompt. So every repeat of a 262,144-token prompt was processed again from the start, which takes about three minutes.

I run the community recipe and build the parts around it.

How a request travels through the systemIllustrative
How a request travels through the systemA coding tool sends a request to a shared API. The API passes it to two computers that run one model together, and the answer returns through the API to the coding tool. TWO COMPUTERS, ONE MODEL CODINGTOOLS SHARED API COMPUTER ACOMPUTER B

My partOther people's work

Other people made the model, the software that runs it, and the recipe. I built the shared API, health checks, automatic startup, recovery, and the fixes on this page.

What runs where, in technical terms

I run the community recipe for GLM-5.3-Flash at tensor parallel 2: vLLM in Docker, 4-bit NVFP4 weights, an fp8 KV cache, and the DFlash2 speculative drafter. NVFP4 is NVIDIA's 4-bit floating-point format. The KV cache holds the attention keys and values of tokens the model has already read. One OpenAI-compatible API serves my coding tools.

Five design choices fixed the stalled starts and the slow repeats.

Shared memory during startup, with each fixIllustrative

With all five, the model starts, a full test runs with no reloads, repeats start in 2.98 s, and the graphics chip draws 47–⁠50 W.

How each choice works, in technical terms

Guard free memory during the load. The load fails when the free memory blocks run out. A per-node guard drops caches and compacts memory whenever free memory falls under 8 GiB. It checks every second, from launch until the server is ready.

Leave the kernel reserve alone. Raising the kernel's free-memory reserve looks like a fix. But vLLM's startup check reads available memory, and the reserve lowers it by about 6.2 GiB, so the check fails before the model loads.

Stop the guard before it causes faults. Major page faults came from two places: the first use of the OpenSSL encryption libraries, and the guard's own cache drops, which evicted mapped weight pages. I stop the worker's guard once the head node is ready, then run one warm-up pass. After that, both nodes had zero major faults through a full qualification run at 262,144-token context.

Patch the cache at the smallest point. The drafter keeps its own short sliding-window cache. vLLM treated every cache group as a drafter group, so the draft window's short hit cut the shared hit to zero. My patch flags only the drafter group and never lets it shorten the hit that the model's own groups agree on.

Cap the GPU clock without losing speed. With the cap, decoding stayed at 55.6 tokens per second. Power fell from 78 W to 47–⁠50 W, and temperature from 86 °C to 69–⁠71 °C.

The community recipe now includes my cache fix and memory report.

Test results, and where my changes wentMeasured

Results of one full test

  • 98.6%of the finished work reused on a repeatwas 0%
  • 2.98 suntil the answer to a repeat startswas 191.7 s
  • 55.6tokens per second of output

Community recipe by tonyd2wild

vLLM, the software that runs the model

The recipe includes two of my changes. A third change, to vLLM, the software that runs the model, is in review.

Full test results, in technical terms
Qualification run, two DGX Spark nodes, September 2026
MeasureResult
Prefix-cache hit rate, repeated prompt0.986, from 0 before the fix
Cached tokens on the repeat258,048
Time to first token, cold191.7 s
Time to first token, cached (p95)2.98 s
Decode speed, one stream (p50)55.6 tokens/s
Drafter acceptance0.928

The test sent 262,144‑token requests one at a time. The recipe's maintainer merged my two pull requests on September 16, 2026. In a test at one eighth of full size, the vLLM change cut the extra memory used during loading from 1,011 MiB to 3 MiB.

On my memory report

"This is the most useful GB10 memory write-up anyone has sent us"

The recipe's maintainer tonyd2wild, on GitHub

Contact

Get in touch.

I'm open to full-time AI product engineering roles.

Email hello@shayanbianconi.com or see more of my work on GitHub.

Email me