<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>swesty.net</title><description>Field notes on LLM serving, fine-tuning, GPU performance, and agentic AI — measured on real hardware.</description><link>https://blog.swesty.net/</link><item><title>vLLM on the DGX Spark, the Hard Way: Seven Months of GB10 Field Notes</title><link>https://blog.swesty.net/blog/vllm-on-dgx-spark-gb10-field-notes/</link><guid isPermaLink="true">https://blog.swesty.net/blog/vllm-on-dgx-spark-gb10-field-notes/</guid><description>Seven months of empirical GB10 (sm_121) vLLM findings: why &apos;Marlin is mandatory&apos; has a load-bearing exception (it&apos;s your checkpoint, not your GPU), the two-condition rule for native FP4, the MTP illegal-memory-access saga and how the W4A4/CUTLASS configuration avoided it, unified-memory OOM math, and measured numbers for every claim.</description><pubDate>Sun, 19 Jul 2026 00:00:00 GMT</pubDate><category>AI Infrastructure</category><category>LLM Serving</category></item></channel></rss>