<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>MSO Expert</title>
    <link>https://msoexpert.com/</link>
    <description>Local AI inference, MSO infrastructure, telecom architecture, and enterprise technology from John Pezzulli.</description>
    <language>en-us</language>
    <lastBuildDate>Sat, 12 Sep 2026 15:17:45 GMT</lastBuildDate>
    <atom:link href="https://msoexpert.com/feed.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>4.8 Billion Prompt Tokens Later: Somebody Else Put My Custom SGLang Runtime to Work</title>
      <link>https://msoexpert.com/articles/pennyroyal-real-workload-field-report/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/pennyroyal-real-workload-field-report/</guid>
      <pubDate>Sat, 12 Sep 2026 15:17:45 GMT</pubDate>
      <description>Nineteen days after publishing Pennyroyal, another user reported 4.8 billion prompt tokens, 55 million generated tokens, and more than a week of real work on a 300 W Max-Q.  Then a second report arrived from the 27B configuration.</description>
    </item>
    <item>
      <title>Messing With SGLang and Qwen3.8: 40 Percent More Decode on One RTX PRO 6000</title>
      <link>https://msoexpert.com/articles/qwen38-dflash2-rtx-pro-6000/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/qwen38-dflash2-rtx-pro-6000/</guid>
      <pubDate>Tue, 25 Aug 2026 02:39:43 GMT</pubDate>
      <description>Several speculative paths moved Qwen3.8 from 74.06 to 108.75 tokens per second on one RTX PRO 6000.  A 29 percent long-prefill regression then sent the runtime through HiCache and NIXL so it could stop rebuilding prefixes it already knew.</description>
    </item>
    <item>
      <title>I Lowered My RTX PRO 6000’s Power Ceiling by 150 Watts.  Local AI Barely Noticed.</title>
      <link>https://msoexpert.com/articles/rtx-pro-6000-600w-local-ai-efficiency/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/rtx-pro-6000-600w-local-ai-efficiency/</guid>
      <pubDate>Thu, 20 Aug 2026 02:00:24 GMT</pubDate>
      <description>My RTX PRO 6000 can consume 600 watts, but local AI rarely asks for all of it.  A tuned 450-watt profile preserved practical Qwen performance, while a completed 994,987-token DeepSeek V4 Flash request averaged just 202.78 watts.</description>
    </item>
    <item>
      <title>I Bought the Ferrari of GPUs to Stop Compromising. Then the Server Demanded a Blood Sacrifice.</title>
      <link>https://msoexpert.com/articles/rtx-pro-6000-server-demanded-blood-sacrifice/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/rtx-pro-6000-server-demanded-blood-sacrifice/</guid>
      <pubDate>Tue, 18 Aug 2026 14:48:30 GMT</pubDate>
      <description>A 96 GB RTX PRO 6000 led to transient-power reboots, a third power supply, dual Xeons, external GPUs, a Dremel, blood, and a 1,500-watt MiniMax M3 experiment that still came up six gigabytes short.</description>
    </item>
    <item>
      <title>The Most Interesting Physical AI Isn’t a Robot. It’s a Model Etched Into Silicon.</title>
      <link>https://msoexpert.com/articles/the-model-is-the-silicon/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/the-model-is-the-silicon/</guid>
      <pubDate>Sat, 08 Aug 2026 16:33:03 GMT</pubDate>
      <description>Taalas physically embodied Llama 3.1 8B in silicon, reached roughly 17,000 tokens per second for one user, and was acquired by AMD before I finished writing about it.</description>
    </item>
    <item>
      <title>AMD AI PRO R9700: The card that just works. My stack around it didn’t.</title>
      <link>https://msoexpert.com/articles/amd-ai-pro-r9700-rocm-vulkan/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/amd-ai-pro-r9700-rocm-vulkan/</guid>
      <pubDate>Tue, 04 Aug 2026 22:11:07 GMT</pubDate>
      <description>ROCm worked on the first try. Vulkan was substantially faster, and the hardest failures belonged to everything around the card.</description>
    </item>
    <item>
      <title>Running High-Quality DeepSeek V4 Flash on One 96 GB GPU</title>
      <link>https://msoexpert.com/articles/making-deepseek-v4-flash-fit-one-96gb-gpu/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/making-deepseek-v4-flash-fit-one-96gb-gpu/</guid>
      <pubDate>Tue, 04 Aug 2026 00:47:58 GMT</pubDate>
      <description>Five mapped expert layers made room for speculative decoding, a correction tier, and near-million-token context.</description>
    </item>
    <item>
      <title>From Architecting and Selling AI to Building a Fairly Insane Edge Node in My Living Room</title>
      <link>https://msoexpert.com/articles/living-room-ai-edge-node/</link>
      <guid isPermaLink="true">https://msoexpert.com/articles/living-room-ai-edge-node/</guid>
      <pubDate>Mon, 03 Aug 2026 17:08:34 GMT</pubDate>
      <description>The origin story, hardware escalation, runtime experiments, failures, and corrections behind a private local AI system.</description>
    </item>
  </channel>
</rss>

