All writing

Articles

Local inference, AI infrastructure, MSO and telecom architecture, and field notes from the systems I build and the work I do.

Messing With SGLang and Qwen3.8: 40 Percent More Decode on One RTX PRO 6000

I started directing Codex at Qwen3.8 because 74 tokens per second apparently was not enough. Several speculative paths eventually produced 108.75 tokens per second on one RTX PRO 6000. Then a 29 percent long-prefill regression sent us through HiCache and NIXL so the server could stop rebuilding prefixes it already knew.

Read →