
John Pezzulli · A Chief Technologist for MSO customers · Telecom presales
Twenty years connecting engineering depth to customer and business outcomes.
My career runs from carrier-grade systems engineering through MSO solution leadership, modernization, field CTO work, and telecom presales. The common thread is making complicated technology useful to the people who have to fund it, build it, operate it, and live with the consequences.

From Architecting and Selling AI to Building a Fairly Insane Edge Node in My Living Room
How “I just need a little more VRAM” became a private AI working partner, multiple runtime lanes, and one ridiculous living-room edge node.
Read the story →01 · Local Inference
Local Inference and Data Sovereignty Home System
4.8 Billion Prompt Tokens Later: Somebody Else Put My Custom SGLang Runtime to Work
Nineteen days after publishing Pennyroyal, another user reported 4.8 billion prompt tokens, 55 million generated tokens, and more than a week of real work on a 300 W Max-Q. Then a second report arrived from the 27B configuration.
Messing With SGLang and Qwen3.8: 40 Percent More Decode on One RTX PRO 6000
Several speculative paths moved Qwen3.8 from 74.06 to 108.75 tokens per second on one RTX PRO 6000. A 29 percent long-prefill regression then sent the runtime through HiCache and NIXL so it could stop rebuilding prefixes it already knew.
I Lowered My RTX PRO 6000’s Power Ceiling by 150 Watts. Local AI Barely Noticed.
My RTX PRO 6000 can consume 600 watts, but local AI rarely asks for all of it. A tuned 450-watt profile preserved practical Qwen performance, while a completed 994,987-token DeepSeek V4 Flash request averaged just 202.78 watts.
I Bought the Ferrari of GPUs to Stop Compromising. Then the Server Demanded a Blood Sacrifice.
A 96 GB RTX PRO 6000 led to transient-power reboots, a third power supply, dual Xeons, external GPUs, a Dremel, blood, and a 1,500-watt MiniMax M3 experiment that still came up six gigabytes short.
AMD AI PRO R9700: The card that just works. My stack around it didn’t.
ROCm worked on the first try. Vulkan was substantially faster, and the hardest failures belonged to everything around the card.
Running High-Quality DeepSeek V4 Flash on One 96 GB GPU
Five mapped expert layers made room for speculative decoding, a correction tier, and near-million-token context.
Building a fairly insane edge node in my living room
The origin story, hardware escalation, runtime experiments, failures, and corrections behind a private local AI system.
02 · MSO & Telecom
Industry work and architecture at scale.
The Most Interesting Physical AI Isn’t a Robot. It’s a Model Etched Into Silicon.
Taalas physically embodied Llama 3.1 8B in silicon, reached roughly 17,000 tokens per second for one user, and was acquired by AMD before I finished writing about it.
Intel-Dell Verified Reference Configuration for vCMTS on Red Hat OCP
A validated configuration for scalable, high-performance cloud-native cable access infrastructure.