Local Inference · September 12, 2026
Nineteen days after publishing Pennyroyal, another user reported 4.8 billion prompt tokens, 55 million generated tokens, and more than a week of real work on a 300 W Max-Q. Then a second report arrived from the 27B configuration.
Read →Local Inference · August 24, 2026
I started directing Codex at Qwen3.8 because 74 tokens per second apparently was not enough. Several speculative paths eventually produced 108.75 tokens per second on one RTX PRO 6000. Then a 29 percent long-prefill regression sent us through HiCache and NIXL so the server could stop rebuilding prefixes it already knew.
Read →Local Inference · August 19, 2026
My RTX PRO 6000 can consume 600 watts, but local AI rarely asks for all of it. A tuned 450-watt profile preserved practical Qwen performance, while a completed 994,987-token DeepSeek V4 Flash request averaged just 202.78 watts.
Read →Local Inference · August 18, 2026
A 96 GB RTX PRO 6000 led to transient-power reboots, a third power supply, dual Xeons, external GPUs, a Dremel, blood, and a 1,500-watt MiniMax M3 experiment that still came up six gigabytes short.
Read →MSO & Telecom · August 8, 2026
Taalas physically embodied Llama 3.1 8B in silicon, reached roughly 17,000 tokens per second for one user, and was acquired by AMD before I finished writing about it.
Read →Local Inference · August 4, 2026
ROCm worked on the first try. Vulkan was substantially faster, and the hardest failures belonged to everything around the card.
Read →Local Inference · August 3, 2026
Five mapped expert layers made room for speculative decoding, a correction tier, and near-million-token context.
Read →Local Inference · August 2026
The origin story, hardware escalation, runtime experiments, failures, and corrections behind a private local AI system.
Read →