<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Neurometric Blog]]></title><description><![CDATA[A substack for Neurometric - AI system orchestration and workflow SLMs ]]></description><link>https://blog.neurometric.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png</url><title>Neurometric Blog</title><link>https://blog.neurometric.ai</link></image><generator>Substack</generator><lastBuildDate>Wed, 19 Aug 2026 21:08:58 GMT</lastBuildDate><atom:link href="https://blog.neurometric.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[neurometric]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[neurometric@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[neurometric@substack.com]]></itunes:email><itunes:name><![CDATA[neurometric]]></itunes:name></itunes:owner><itunes:author><![CDATA[neurometric]]></itunes:author><googleplay:owner><![CDATA[neurometric@substack.com]]></googleplay:owner><googleplay:email><![CDATA[neurometric@substack.com]]></googleplay:email><googleplay:author><![CDATA[neurometric]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Gemma 4b vs Gemini Flash: You Don't Need Frontier Models For Tool Calling Workflows]]></title><description><![CDATA[What our Acebench analysis showed about the models]]></description><link>https://blog.neurometric.ai/p/gemma-4b-vs-gemini-flash-you-dont</link><guid isPermaLink="false">https://blog.neurometric.ai/p/gemma-4b-vs-gemini-flash-you-dont</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Tue, 11 Aug 2026 21:19:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!twzy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most of what makes an AI assistant useful isn&#8217;t prose &#8212; it&#8217;s calling tools: booking the meeting, adding the cart item, pulling the shipping estimate. When these products fail, it&#8217;s rarely bad writing. It&#8217;s the wrong function, a mangled date, or three calls when one would do.</p><p>AceBench measures exactly that: ~2,000 hand-annotated tasks drawn from ~4,500 synthetic APIs across eight domains (finance, health, travel, tech, entertainment, and more), published January 2025 and later accepted at EMNLP.</p><p>Three things make it worth a practitioner&#8217;s attention:</p><p><strong>Grading is mechanical and brutal.</strong> Right function, right call count, right arguments &#8212; or it&#8217;s wrong. No partial credit, no AI judge. One bad field voids the form. Harsher than real life, but reproducible.</p><p><strong>It comes in tiers.</strong> Normal (clear requests), Special (vague), Agent (multi-turn). We ran Normal, English only &#8212; this says nothing about ambiguity or agent loops.</p><p><strong>It&#8217;s designed to be taken apart.</strong> Tasks isolate specific failure modes: value types, near-duplicate functions, mid-conversation drops. A leaderboard number says which model wins overall. A benchmark that comes apart says whether a cheap model is good enough for <em>your</em> work &#8212; usually the real question.</p><h2>What we ran</h2><p>Five models, 772 tasks each, 3,860 attempts, English Normal split:</p><ul><li><p><code>gemini-3.6-flash</code>, <code>gemini-3.1-flash-lite</code> &#8212; Google&#8217;s hosted models</p></li><li><p><code>gemma-4-12B-it</code>, <code>gemma-4-E4B-it</code>, <code>gemma-4-E2B-it</code> &#8212; open-weight, self-served (big/small/very small)</p></li></ul><p>Same harness and tasks for everyone. Tool calls return a bare acknowledgement &#8212; no feedback, one shot per call.</p><p><strong>Key point:  the open model on hardware we control tied Google&#8217;s hosted model, using about an eighth of the words.</strong></p><h2>The scoreboard</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!twzy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!twzy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 424w, https://substackcdn.com/image/fetch/$s_!twzy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 848w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1272w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!twzy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png" width="1292" height="460" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:460,&quot;width&quot;:1292,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66008,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!twzy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 424w, https://substackcdn.com/image/fetch/$s_!twzy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 848w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1272w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The top two are a real tie: 45 tasks the hosted model got and the open one missed, 35 the other way. That&#8217;s the signature of two equally capable models, not a better and a worse one. Run it again and the order could flip. Every other gap in the table is real &#8212; but the tie at the top is between a hosted frontier model and a 12B open model on a single GPU.</p><h2>What it costs</h2><p>Scoring the same isn&#8217;t interesting. Scoring the same <em>this cheaply</em> is.</p><p>ModelWords per correct answer*:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8OpP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8OpP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 424w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 848w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1272w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8OpP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp" width="1306" height="476" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:476,&quot;width&quot;:1306,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:16632,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8OpP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 424w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 848w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1272w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>*tokens (~&#190; word each), total output</p><p>The hosted model generates ~8x more text per correct answer than the 12B model it&#8217;s tied with &#8212; almost all invisible &#8220;thinking&#8221; (655,000 tokens across the run). The other four models did none of that: read, call, stop.</p><p>And the thinking isn&#8217;t buying anything. When the hosted model got a task wrong, it generated <em>more than twice</em> as much text as when it got one right. Extra effort here is a sign of being stuck, not a way out. On a benchmark where the winning move is two steps, there&#8217;s not much to think about.</p><p>Generated tokens are the expensive half of any pricing page, and what determines latency. So: same accuracy, a fraction of the tokens, and a deployment you own instead of rent.</p><h2>The catch</h2><p>Cheap on tokens &#8800; cheap in practice:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lNL9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lNL9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 424w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 848w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1272w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lNL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png" width="1334" height="562" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:562,&quot;width&quot;:1334,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:73710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lNL9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 424w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 848w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1272w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The 12B model that ties Google&#8217;s best was 4x slower per answer than the API it ties, despite generating far less text. The smaller E4B, on the same GPU, hit 5.6s at ~91% of the top score. Note: these numbers are default vLLM, full BF16, on a single L40S &#8212; plenty of room to push both further with better hardware.</p><h2>The smallest model is better than its score</h2><p><code>gemma-4-E2B-it</code> came last at 67%. Look at <em>how</em> it failed and half the gap disappears.</p><p>Its top mistake wasn&#8217;t picking the wrong tool &#8212; it was making the right call plus extra, already-completed ones from earlier in the conversation:</p><blockquote><p><strong>Called:</strong> ask about gift etiquette &#8594; send the gift &#8594; schedule the meeting <strong>Wanted:</strong> schedule the meeting</p></blockquote><p>It knew the current turn. It just also re-did two things already done. Graded: zero.</p><p>72 of its 255 failures end with exactly the right call, buried under repeated history. Score only the current turn and it jumps from 67% to 76% &#8212; close to Google&#8217;s smaller model. That&#8217;s a prompt fix, not a bigger model &#8212; and it barely moves the other four (0-3 failures each of this kind).</p><h2>Where bigger models still earn their keep</h2><p>Three places:</p><p><strong>Nested arguments.</strong> A tool wanting <code>{"journey": {"from": "Shanghai", "to": "Hangzhou", "times": {...}}}</code> trips up everyone &#8212; best models barely clear 60%, versus 85-96% for flatter arguments.</p><p><strong>Several calls at once.</strong> Same tool, three times, different details: the hosted model pulls ahead ~5 points. Coordinating calls is harder than making one.</p><p><strong>Long conversations</strong> &#8212; the sharpest split. Four of five models <em>improve</em> turn over turn as context narrows the options. The smallest model collapses:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4T2I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4T2I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 424w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 848w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1272w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4T2I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png" width="1290" height="594" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:594,&quot;width&quot;:1290,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77120,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4T2I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 424w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 848w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1272w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That last column is what matters in production: a conversation only works if every turn lands. 87% per-turn becomes 74% overall.</p><h2>Two models beat one</h2><p>Since every model saw the same tasks, we can check what a pair covers:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QlGi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QlGi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 424w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 848w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1272w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QlGi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png" width="1300" height="376" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:376,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45669,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QlGi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 424w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 848w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1272w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Pairing the hosted model with the open one adds 4.5 points; pairing it with Google&#8217;s other model adds just 2.1. The two hosted models fail on the same tasks &#8212; the open one fails on different ones. If you can check an answer and retry, the self-served model is the better (and cheaper) second opinion.</p><h2>So what do you do with this</h2><p>If the task is &#8220;here&#8217;s what I want, here are the tools, go&#8221; &#8212; use a small self-hosted model. There&#8217;s no ambiguity for extra reasoning to resolve. A 12B model matches the frontier here; even 4B gets within ~90%.</p><p>If the work needs complex structures, coordinated multi-calls, or long conversations, the gap reopens &#8212; fastest for the smallest models.</p><p>Check what a scoreboard actually measures before trusting it: one category here was scoring a missing input, not the model, so we dropped it. Another score was understated nine points over a technicality about which calls &#8220;count&#8221; for a turn. Neither shows up in a leaderboard number &#8212; and both change what you&#8217;d buy.</p>]]></content:encoded></item><item><title><![CDATA[New Podcast Episode - TMax: Closing the Frontier Gap With Open Data]]></title><description><![CDATA[A conversation with Yash Sharma, Director of AI Research at Neurometric AI]]></description><link>https://blog.neurometric.ai/p/new-podcast-episode-tmax-closing</link><guid isPermaLink="false">https://blog.neurometric.ai/p/new-podcast-episode-tmax-closing</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Fri, 07 Aug 2026 18:24:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week, Cooper and Yash (Neurometric AI) unpack TMax, a paper out of AI2 (the Allen Institute for AI) that takes a genuinely different approach to closing the gap between small open-weight models and frontier labs &#8212; not by training a bigger model, but by open-sourcing the entire recipe: data, code, and technique.</p><p>The paper focuses on terminal agents &#8212; models given nothing but command-line access as a tool, which turns out to be enough to tackle an enormous range of tasks, from software engineering to security work to scheduling. Using this recipe, AI2 boosted a Qwen model&#8217;s Terminal-Bench score from roughly 25% to 31%, closing meaningful ground on Claude Haiku&#8217;s roughly 33%.</p><p>What made this one worth an episode wasn&#8217;t just the score bump. It was the design choices behind it.</p><p>Watch on YouTube: <a href="https://youtu.be/4U0gqv9sSn4?si=ejv04gktGbTOMJM8">https://youtu.be/4U0gqv9sSn4?si=ejv04gktGbTOMJM8</a></p><p>Listen:<a href="https://tokenengineering.podbean.com/"> https://tokenengineering.podbean.com</a></p><p>Paper: TMax: A recipe for terminal agents &#8212;<a href="https://arxiv.org/abs/2606.23321"> https://arxiv.org/abs/2606.23321</a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Glass Is Half... Correct? Half Our SLM Benchmark 'Failures' Contained The Right Answer]]></title><description><![CDATA[An analysis of small model failures on CRM arena]]></description><link>https://blog.neurometric.ai/p/the-glass-is-half-correct-half-our</link><guid isPermaLink="false">https://blog.neurometric.ai/p/the-glass-is-half-correct-half-our</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 05 Aug 2026 18:55:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DhE3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What we ran</h2><p>Three models, two tool interfaces, 340 CRM tasks each &#8212; 2,040 agent rollouts in total.</p><p>The benchmark is CRMArena, run through the Harbor/dockworker harness. Every task asks a question about a read-only Salesforce org (&#8221;in May 2021, which state had the quickest case closures?&#8221;, &#8220;which knowledge article does this quote violate?&#8221;) and grades the submitted answer by exact match. Seventeen task categories, split evenly across a B2B and a B2C org.</p><p>The interesting design choice is that the same 340 questions are posed twice, behind two different tool surfaces:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DhE3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DhE3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 424w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 848w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1272w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DhE3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png" width="1324" height="434" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:434,&quot;width&quot;:1324,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:78759,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DhE3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 424w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 848w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1272w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Identical questions, identical ground truth &#8212; we verified this holds on every task where both suites produced a graded answer. So the two suites isolate one variable: does the model do better writing SQL, or picking the right pre-built tool and filling in its arguments?</p><p>The three models: <strong>gemini-3.6-flash</strong> and <strong>gemini-3.1-flash-lite</strong> (hosted), and <strong>gemma-4-E4B-it</strong> (open weights, served locally on vLLM). One is a small local model; two are hosted. Every trial used the same <code>noshell</code> agent scaffold, whose only way to answer is to call a terminal <code>submit_answer</code> tool.</p><h2>The scoreboard</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!37I2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!37I2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 424w, https://substackcdn.com/image/fetch/$s_!37I2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 848w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1272w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!37I2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png" width="1312" height="348" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:348,&quot;width&quot;:1312,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53643,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!37I2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 424w, https://substackcdn.com/image/fetch/$s_!37I2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 848w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1272w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Read it and you&#8217;d conclude the ordering is obvious and the small model isn&#8217;t close: 39.7% against 67.6% is a 28-point gap. Every pairwise difference here is statistically significant (p &#8804; 0.009).</p><p>Then you look at how the failures happen, and the picture inverts.</p><h2>Half of the small model&#8217;s &#8220;failures&#8221; contain the right answer</h2><p>Not every zero is a wrong answer. A trial scores zero if the answer file was never written at all &#8212; and the scaffold only writes it when the model calls <code>submit_answer</code>. Classifying all 2,040 trials by how they failed:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AbqI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AbqI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 424w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 848w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1272w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AbqI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png" width="1348" height="778" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:778,&quot;width&quot;:1348,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:98995,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AbqI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 424w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 848w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1272w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>More than half of gemma&#8217;s SQL trials never submitted anything. And they didn&#8217;t crash or time out &#8212; in 171 of them the model finished its work, wrote the answer out in prose, and simply never called the tool. Something like:</p><blockquote><p>The <code>case_metrics</code> call returned an average closure time of 4.2 days for CA, the lowest of any state. The answer is CA.</p></blockquote><p>Graded: zero.</p><p>So we tested the obvious question &#8212; were those prose answers right? Ground truth isn&#8217;t recorded for ungraded trials, but because the same 340 tasks appear in both suites, we could recover the expected answer for every one of them and grep the final message for it.</p><p>Of gemma&#8217;s 250 prose non-submissions, <strong>44&#8211;52% contained the correct answer</strong>. The range is the strict and loose reading of the same check: the loose count is any trial whose prose contains the expected value; the strict count additionally requires that the model wasn&#8217;t hedging across a list of candidates (&#8804;1 other record ID mentioned). Both bound the same conclusion.</p><p>That reframes the scoreboard as a lower bound. Fixing one scaffold behaviour &#8212; get the model to call the tool &#8212; moves gemma to:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!w90h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!w90h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 424w, https://substackcdn.com/image/fetch/$s_!w90h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 848w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1272w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!w90h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png" width="1326" height="282" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:282,&quot;width&quot;:1326,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!w90h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 424w, https://substackcdn.com/image/fetch/$s_!w90h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 848w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1272w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>A model that looked like it scored 26% on SQL was doing work worth about 50%. Its measured number was roughly half its actual competence, and every point of that gap is a formatting bug.</p><p>The same correction barely moves the hosted models &#8212; gemini-3.6-flash has exactly zero prose non-submissions, and gemini-3.1-flash-lite has one. They always call the tool. What we were measuring, for a third of the benchmark, was instruction-following on the harness contract, not CRM reasoning.</p><p>There&#8217;s a cleaner way to see it. Restrict to trials that submitted anything, and the reasoning quality behind the scoreboard separates from the plumbing:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RiZl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RiZl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 424w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 848w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1272w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RiZl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png" width="1330" height="346" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4436c62-6974-44e8-a378-003463c52213_1330x346.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:346,&quot;width&quot;:1330,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47261,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RiZl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 424w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 848w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1272w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On answers it actually submits, the small local model is within two points of the hosted flash-lite model &#8212; a difference well inside the noise at this sample size. The 11-point headline gap between them is almost entirely tool-calling discipline.</p><h2>What it costs</h2><p>This is where the small model stops being a curiosity. Token totals are the whole run; the per-win column divides by correct answers, which is the number that matters if you&#8217;re paying for throughput.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nOs-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nOs-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 424w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 848w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1272w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nOs-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png" width="1322" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33631410-c178-489a-a32c-2baa2db91b67_1322x850.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1322,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:122754,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nOs-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 424w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 848w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1272w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Gemma on the tool API buys a correct answer for 45.4k tokens. gemini-3.6-flash needs 196k for the same thing &#8212; <strong>4.3&#215; more</strong>. On SQL it needs 598k, or <strong>13&#215; gemma&#8217;s best configuration</strong>.</p><p>The driver is visible in the reasoning-token column. gemini-3.6-flash spent 1.36M reasoning tokens on the API suite and 2.59M on SQL. The other two models spent none. That&#8217;s what the extra accuracy is bought with: 3.6-flash&#8217;s median trial emits 3,978 output tokens on the API suite against gemma&#8217;s 501 &#8212; an 8&#215; difference in generated text per attempt.</p><p>And time to completion &#8212; median wall time of the agent execution phase alone, excluding container build and verification:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LtyB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LtyB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 424w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 848w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1272w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LtyB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png" width="1350" height="616" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:616,&quot;width&quot;:1350,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:83023,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LtyB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 424w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 848w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1272w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here the small model does not win, and it&#8217;s worth being precise about why. gemma is 12.6s per attempt against 3.6-flash&#8217;s 30.0s &#8212; faster per attempt &#8212; but because it converts fewer attempts into correct answers, it lands at 73.8s per win versus 51.8s. And this column mixes model speed with serving infrastructure: a local vLLM instance against Google&#8217;s production endpoints. It is not a property of the models, and it&#8217;s the one metric here we&#8217;d throw out of a purchasing decision. Tokens are the honest efficiency measure; wall-clock is an artifact of where each model happened to be running.</p><p>One more pattern worth noting: every model spends more tokens on the trials it gets wrong &#8212; 1.5&#215; to 2.1&#215; its passing median. Failure is not cheap. Effort is a symptom of being lost, not a route out of it, which means the token cost of a wrong answer exceeds the token cost of a right one across the board.</p><h2>Where each model actually fails</h2><p>The three models have almost nothing in common in their failure profiles.</p><p><strong>gemma-4-E4B-it</strong> &#8212; a formatting problem wearing a capability problem&#8217;s clothes. 250 of its failures are prose-instead-of-tool-call; roughly half contain the right answer. Its second tendency is over-caution: 50 trials answered <code>None</code> (&#8221;nothing matches&#8221;) when a real record existed. It abstains too readily and it won&#8217;t call the terminal tool. Both are addressable without touching the model.</p><p><strong>gemini-3.1-flash-lite</strong> &#8212; quietly broken generations. 34 API trials (10%) ended with an empty completion: a single EOS token, no text and no tool call. Not prose, not a wrong answer &#8212; nothing at all. That signature points at serving or sampling rather than the prompt, and it&#8217;s the cheapest 10% anyone in this comparison could recover. Beyond that its losses are ordinary wrong answers (21&#8211;23%), the highest wrong-answer rate of the three.</p><p><strong>gemini-3.6-flash</strong> &#8212; runs out of budget, and won&#8217;t say &#8220;none.&#8221; It never fails to submit when it finishes, but it frequently doesn&#8217;t finish: 60 SQL trials (18%) and 17 API trials (5%) hit the 25-step cap mid-work. Its failing trials burn a median 9,747 output tokens against 6,452 when passing &#8212; it iterates on queries that never land. Its other weakness is the mirror image of gemma&#8217;s: 63 trials invented a value where <code>None</code> was correct. The strongest model is the one most likely to manufacture an answer rather than concede there isn&#8217;t one.</p><p>That last contrast is the most useful thing here for anyone building on these models. The two abstention errors point in opposite directions &#8212; gemma says &#8220;none&#8221; when an answer exists, 3.6-flash asserts an answer when none does &#8212; so there is no single prompt that fixes both. It&#8217;s a calibration problem per model, not a benchmark-wide one.</p><h2>The interface matters, and not the way you&#8217;d guess</h2><p>Because the same questions appear behind both tool surfaces, we can ask whether SQL or semantic tools suit each model better as a paired comparison &#8212; the same task, two interfaces &#8212; using an exact McNemar test on the tasks where the two disagree.</p><p>For gemini-3.6-flash, the only model whose two runs had matched step budgets, the tool API wins clearly: +8.5 points, p &lt; 0.001, with 47 tasks solved only through the API against 18 solved only through SQL.</p><p>But the aggregate hides something better. Across the models, between 65 and 133 of the 340 tasks flip outcome between the two interfaces. For gemini-3.1-flash-lite the headline rates are nearly identical (49.1% vs 50.9%) while 94 tasks flip &#8212; 44 solved only via SQL, 50 only via the API, with just 123 solved by both. <strong>The interface doesn&#8217;t change how many tasks it gets right; it changes which third of the benchmark it gets right.</strong> A model that looks interface-indifferent on the scoreboard is nothing of the kind, and running both interfaces and taking either success would score far above either alone.</p><p>Category detail shows where this bites. gemini-3.6-flash gets <code>sales-amount-understanding</code> right 60% of the time through semantic tools and 5% through SQL &#8212; but that collapse is not an inability to write the query. 18 of its 20 SQL attempts hit the step cap mid-work. Where a semantic tool answers &#8220;total order amount by owner in this window&#8221; in one call, SQL requires discovering the schema, joining line items to orders, and windowing on the right date column &#8212; and it ran out of budget getting there. <code>lead-routing</code>, by contrast, it solves 100% either way.</p><h2>Not all 17 categories are the same benchmark</h2><p>Treating CRMArena as one number hides the most actionable result in the run. Here is every category, worst-first, across all six runs (<code>36</code> = gemini-3.6-flash, <code>31</code> = gemini-3.1-flash-lite, <code>gm</code> = gemma-4-E4B-it):</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sMca!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sMca!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 424w, https://substackcdn.com/image/fetch/$s_!sMca!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 848w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1272w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sMca!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png" width="1334" height="1474" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1474,&quot;width&quot;:1334,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:207655,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sMca!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 424w, https://substackcdn.com/image/fetch/$s_!sMca!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 848w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1272w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7rD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7rD8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 424w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 848w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7rD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png" width="1350" height="1038" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1038,&quot;width&quot;:1350,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:142926,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7rD8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 424w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 848w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Group the 17 categories by what they actually ask for, and a pattern appears that the headline rates completely obscure:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!llyz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!llyz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 424w, https://substackcdn.com/image/fetch/$s_!llyz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 848w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1272w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!llyz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png" width="1330" height="944" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:944,&quot;width&quot;:1330,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:126035,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!llyz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 424w, https://substackcdn.com/image/fetch/$s_!llyz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 848w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1272w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On policy judgement, model scale buys almost nothing. Reading a chat transcript and deciding whether an agent breached policy, or whether a lead is qualified: gemini-3.6-flash manages 37.5%, gemma manages 32.5%. The small model is at 87% of the best model&#8217;s performance. Compare that to aggregate metrics, where the same small model reaches only 57% of the frontier score and the spread across runs hits 85 points on <code>monthly-trend-analysis</code>.</p><p>The reason is visible in the failure modes. On <code>policy-violation-identification</code>, half of gemini-3.6-flash&#8217;s API failures are missed_abstention &#8212; 8 trials asserting a violation where the correct answer was &#8220;none&#8221; &#8212; against 8 passes. gemma&#8217;s failures on the same category split 4 missed_abstention and 4 false_abstention. This category isn&#8217;t primarily testing reasoning; it&#8217;s testing whether a model will decline to answer. That&#8217;s a calibration property, and calibration doesn&#8217;t scale with capability the way multi-step numeric work does.</p><p>So the practical read is: the categories where a small model is nearly as good are the judgement ones, and the categories where it falls off a cliff are the multi-step aggregations. If your workload is &#8220;read this conversation and classify it,&#8221; a 4B local model is a serious candidate. If it&#8217;s &#8220;compute this metric across three joins and a date window,&#8221; it is not.</p><p>Two more things the table says:</p><p><strong>invalid-config is the one genuine capability cliff.</strong> gemini-3.6-flash gets 50% on both interfaces; every other run scores 5&#8211;10%. A 45-point gap that survives both tool surfaces is the clearest evidence in this run of something the smaller models simply cannot do &#8212; matching a quote&#8217;s configuration against the knowledge article it violates. Notably it&#8217;s not a formatting artifact: gemma&#8217;s failures here are 8 wrong answers alongside 7 prose non-submissions, so even the recovered ceiling stays low.</p><p><strong>quote-approval defeats everything.</strong> 8% pooled, topping out at 20%, and gemini-3.6-flash scores 5% and 0%. No model, no interface, no step budget helps. When the strongest model in a comparison scores 5% on a category, that&#8217;s usually a signal about the task or the grader rather than the models &#8212; and it&#8217;s the first thing we&#8217;d re-examine before drawing conclusions about the remaining headroom.</p><p>And where does the small model actually beat a hosted model head-to-head? On the tool API, gemma matches or beats gemini-3.1-flash-lite in 5 of 17 categories: <code>policy-violation-identification</code> (35% vs 15%, +20 points), <code>wrong-stage-rectification</code> (40% vs 35%), <code>invalid-config</code> (10% vs 5%), and ties on <code>case-routing</code> and <code>lead-routing</code>. Four of those five are judgement or routing tasks, not aggregations &#8212; the same pattern again.</p><h2>What we&#8217;d take away</h2><p>The headline number on an agent benchmark is a joint measurement of the model and the scaffold around it, and for small models the scaffold dominates. gemma-4-E4B-it looked like a 26% model on SQL. It was doing work worth about 50%, and the missing half was one unmade tool call. Anyone comparing a small open model against a hosted frontier model on a leaderboard number, without looking at the failure modes underneath it, is partly measuring which model was better at following the harness&#8217;s calling convention.</p><p>Small models win on the axis that gets left off the leaderboard. At 45.4k tokens per correct answer against 196k, gemma is 4.3&#215; cheaper per unit of useful output than the model that beats it by 28 points &#8212; and it gets there with zero reasoning tokens against 1.36M. If your workload tolerates 50% task accuracy, or you can put a verifier behind it and retry, the small model is the better economic choice by a wide margin.</p><p>But the honest version of &#8220;small models can win&#8221; is narrower than the slogan. gemma did not beat gemini-3.6-flash in a single one of the 17 categories, on either interface. What it did was reach parity with a hosted flash-lite model on reasoning quality &#8212; 56.0% against 57.9% on submitted answers &#8212; at a fraction of the token cost, while running locally. That&#8217;s the claim the data supports: not that small models beat frontier models, but that the gap to the tier below frontier is mostly plumbing, and the plumbing is cheap to fix.</p><p>The category breakdown sharpens that into a selection rule. The gap is not uniform across work: on policy judgement the small model is at 87% of the frontier model, on multi-step aggregation it is at 57%, and on one category (<code>invalid-config</code>) it is nowhere near. &#8220;Can a small model do this?&#8221; has no general answer, but it has a reliable one per task family &#8212; and the families where the answer is yes are the judgement-shaped ones, which is the opposite of where most people assume a small model will struggle.</p><p>The fixes are unglamorous and they&#8217;re all in the harness: make the terminal tool call unmissable for models that write prose, investigate the empty completions, raise the step cap for models that iterate, and calibrate abstention per model rather than globally. None of them require a bigger model.</p>]]></content:encoded></item><item><title><![CDATA[Neurometric Is Now On TrustedRouter]]></title><description><![CDATA[Hosted generic and task specific SLMs are available.]]></description><link>https://blog.neurometric.ai/p/neurometric-is-now-on-trustedrouter</link><guid isPermaLink="false">https://blog.neurometric.ai/p/neurometric-is-now-on-trustedrouter</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 29 Jul 2026 19:28:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Neurometric is now a <a href="https://trustedrouter.com/providers/neurometric">listed provider on TrustedRouter</a>. If you already route inference through TrustedRouter, our models are available to you today &#8212; no new account, no new SDK, no new base URL.</p><h2>What Is TrustedRouter?</h2><p>TrustedRouter is an inference gateway in the same family as OpenRouter: one OpenAI-compatible API, one base URL, and hundreds of models and routes behind it. You point your existing client at the gateway and swap model strings instead of rewriting integrations every time you want to try a different provider.</p><p>What makes it interesting is the trust posture. TrustedRouter keeps no record of what you route &#8212; no prompt or output logs, by design, and the gateway is attested and fails closed rather than silently degrading. It publishes a live status page and a continuously sampled performance leaderboard, so latency and uptime claims are measured rather than asserted. There&#8217;s also a documented migration path from OpenRouter for teams that want to move without a rewrite.</p><p>For anyone who has spent time explaining to a security team why prompts containing customer data are being logged by a third party, that combination is worth something.</p><h2>What We&#8217;re Serving Today</h2><p>We&#8217;re live with three small models on prepaid routes:</p><ul><li><p><code>ibm-granite/granite-4.1-8b</code> &#8212; 32K context, $0.0525/1M prompt and $0.105/1M completion</p></li><li><p><code>qwen/qwen3-vl-8b-instruct</code> &#8212; 262K context, vision-capable</p></li><li><p><code>qwen/qwen3-vl-8b-thinking</code> &#8212; 32K context, reasoning traces</p></li></ul><p>Early measured numbers on our routes: 863 ms p50 time-to-first-token on Granite and 100% uptime across the sampling window. Our listed policy note is accurate about what we do and don&#8217;t guarantee &#8212; upstream trace logging is disabled, prompts and completions are not written to observability or object storage, and we retain only aggregate request counts and token totals. That&#8217;s a no-store posture, not contractual ZDR, and we&#8217;d rather say so plainly than let people assume more than we&#8217;ve committed to.</p><h2>Task-Specific Models Are Coming Next</h2><p>Small general-purpose models are the entry point, not the thesis. The reason we build the way we do is that most production workloads aren&#8217;t &#8220;chat&#8221; &#8212; they&#8217;re a narrow, repetitive task where a tuned 8B model matches or beats a frontier model at a fraction of the cost.</p><p>Three task-specific models from our model marketplace we expect to list on TrustedRouter in the coming weeks:</p><ol><li><p><strong>Text-to-SQL</strong> &#8212; natural language to correct, schema-aware queries against a known database, tuned for join accuracy rather than conversational fluency.</p></li><li><p><strong>Structured extraction</strong> &#8212; pulling typed JSON out of invoices, contracts, and claims documents against a supplied schema, with predictable failure behavior on missing fields.</p></li><li><p><strong>Support intent classification</strong> &#8212; routing inbound tickets to the right queue and priority, the kind of high-volume, low-token task where per-call cost dominates everything else.</p></li></ol><p>Each is graded against our proprietary task benchmark corpus, so you can see how a specialist performs on your task class before you route production traffic to it.</p><p>If you want to try the current models, grab a key at TrustedRouter and use <code>neurometric</code> as the provider. If you have a task class you&#8217;d like us to build for, tell us &#8212; that&#8217;s how the roadmap gets set.</p>]]></content:encoded></item><item><title><![CDATA[Navigating the Enterprise AI Labyrinth: Build, Buy, Route, or Wait]]></title><description><![CDATA[How to operate when the world is changing so fast.]]></description><link>https://blog.neurometric.ai/p/navigating-the-enterprise-ai-labyrinth</link><guid isPermaLink="false">https://blog.neurometric.ai/p/navigating-the-enterprise-ai-labyrinth</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 29 Jul 2026 00:20:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At Neurometric, we spend our days in the trenches with enterprise technology leaders helping them save inference costs and optimize model usage. In most companies the pressure to &#8220;do something with AI&#8221; is intense, but the rush to production often leads to bloated budgets, architectural dead ends, and a staggering accumulation of technical debt. When the dust settles, the organizations that succeed are not the ones that deploy the most models; they are the ones that apply a rigorous, intellectually honest framework to how they adopt them.</p><p>To cut through the hype, organizations must fundamentally rethink their deployment strategies. In our experience, the most critical decision a technical leader makes isn&#8217;t which specific foundational model to choose, but rather the strategic posture they take toward the underlying business task. We have observed that enterprise AI strategy ultimately boils down to a fundamental set of decisions. There are usually four options on the table, and they require a level of candor that is often missing from vendor pitches and internal strategy meetings.</p><h3>The Four Options </h3><p>When evaluating a new AI capability or use case, technology leaders must choose between four distinct paths: Buy, Build, Route, or Wait.</p><p><strong>Buy</strong> is the right answer much more often than technical leaders want it to be. Engineers are naturally wired to create, and the allure of constructing a bespoke AI system is incredibly strong. However, if a vendor has already solved a generic enterprise problem&#8212;like drafting marketing copy, summarizing meeting notes, or triaging customer support tickets&#8212;purchasing that solution is almost always the superior economic choice.</p><p><strong>Build</strong> is the right answer only <strong>when the task at hand </strong><em><strong>is</strong></em><strong> the core business</strong>. If the AI system is going to directly drive your competitive advantage in the marketplace, you cannot outsource it.</p><p><strong>Route</strong> is the correct posture when the task is high-volume and the models are fungible. If you are processing millions of identical queries where the subtle nuances of a massive frontier model are unnecessary, routing queries to the most efficient model available is the only way to scale without destroying your margins.</p><p>Finally, there is <strong>Wait</strong>. Wait is the honest default. It is the right decision more often than anyone wants to admit, and yet it is almost never proposed in a strategy meeting. There is immense career risk in telling a CEO to wait on AI. But for highly volatile use cases, or problems where the underlying foundational models are currently struggling but rapidly improving, waiting six months for the ecosystem to mature is often vastly superior to burning capital on a brittle, premature V1.</p><h3>The &#8220;Build Test&#8221;</h3><p>If you are leaning toward building a custom AI solution, you must subject your proposal to the &#8220;Build Test.&#8221; Building bespoke AI is an expensive, resource-intensive endeavor that requires long-term commitment. At Neurometric, we advise clients to build only if they can definitively answer &#8220;yes&#8221; to all three of the following criteria:</p><p>First, is the task core to your differentiation? The AI must do something that separates you from your competitors. If it is merely an operational efficiency that every other company in your sector will eventually adopt, it fails this test.</p><p>Second, do you possess proprietary data or a unique workflow that a vendor structurally cannot access? You need an unfair advantage. If you are building a model using the exact same public datasets and standard enterprise tools as the major SaaS vendors, they will eventually commoditize your creation. You must have a moat built on data or processes that are uniquely yours.</p><p>Third, can you staff the maintenance of this system for the next three years? Building the model is merely the starting line. Models drift, APIs change, underlying data distributions shift, and security vulnerabilities emerge. You are not just funding a build phase; you are funding a permanent product team.</p><p>If you meet two out of these three criteria, it is not a &#8220;maybe.&#8221; Two out of three means you default to Buy. The economics of maintaining a sub-scale, non-differentiated AI system will slowly drain your engineering resources.</p><h3>Model Selection: The Frontier Premium</h3><p>Once you have decided how to acquire the capability, you must choose the right engine. The market is currently bifurcated between massive &#8220;frontier&#8221; models and smaller, specialized, or distilled models.</p><p>Frontier models are the bleeding-edge giants of the industry. They possess incredible reasoning capabilities and vast world knowledge. You should reserve these models strictly for tasks requiring deep judgment, handling high ambiguity, and managing low-volume, high-stakes scenarios. If an AI is reviewing a complex legal contract for a multi-million dollar merger, you want the frontier model.</p><p>Conversely, small, specialized, or distilled models are the workhorses of the modern enterprise. These should be deployed for high-volume, well-specified, and latency-sensitive tasks. When you are processing tens of thousands of basic data extraction requests per hour, you do not need an AI that can write a sonnet or pass the bar exam.</p><p>The divergence in economics here is huge, often separated by one to two orders of magnitude. A distilled model can easily be 10x to 100x cheaper per token than a frontier model, while returning responses in a fraction of the time. Yet, we routinely see organizations using frontier models for absolutely everything simply because it requires only one API integration. They are paying a 30&#215; premium for the privilege of architectural laziness.</p><h3>Routing as an Operating Discipline</h3><p>To capture the economic benefits of smaller models without sacrificing quality, organizations must adopt routing as a core operating discipline.</p><p>The biggest mistake teams make is routing at the application level&#8212;deciding that an entire application will use just one model. Instead, you must route per task. A single customer service application might use a cheap, fast model to identify the language of an incoming ticket, a specialized model to extract the customer&#8217;s account number, and only invoke a frontier model if the ticket requires a complex, nuanced apology for a service failure.</p><p>Implementing this requires establishing a strict quality floor for every specific task. You determine the minimum acceptable accuracy, and then you dynamically route the workload to the absolute cheapest model that clears that floor. Because the open-source and proprietary model landscapes are evolving at a breakneck pace, this is not a set-it-and-forget-it architecture. You must re-benchmark your routing logic monthly. The model that was the most cost-effective in January might be entirely obsolete by April.</p><h3>Vendor Risk and AI Sovereignty </h3><p>Finally, organizations must wake up to the reality of vendor risk in the AI space. It is real, it is severe, and it is currently vastly underweighted in enterprise risk assessments.</p><p>When you build a product highly dependent on a third-party API, you are at the mercy of their roadmap. Deprecation schedules are often brutally short, forcing expensive emergency migrations when a vendor decides to sunset a specific model version. Price changes can occur overnight, destroying the unit economics of your application. Rate limits can be abruptly tightened, throttling your application&#8217;s ability to scale during a surge in user demand.</p><p>Furthermore, data privacy terms are a constantly shifting target. An update to a vendor&#8217;s terms of service could suddenly allow them to use your enterprise data to train their next generation of models, violating your internal compliance policies.</p><p>Most insidiously, a model can change its behavior right underneath you without any version bump. Vendors continuously tweak, align, and &#8220;improve&#8221; their models behind the scenes. A prompt that reliably output perfect JSON formatting on Tuesday might suddenly start wrapping its output in conversational pleasantries on Thursday, breaking your entire data pipeline.</p><p>To mitigate this, technology leaders must push back during procurement. Do not accept standard click-wrap agreements for critical AI infrastructure. You must explicitly contract for extended notice periods regarding deprecations, strict rate limit guarantees, immutable data terms, and rigid version control to ensure the model you test is the exact model you run in production.</p><p>The AI ecosystem is moving incredibly fast, but the fundamental laws of enterprise software engineering and business economics have not been suspended. By rigorously applying the Build Test, rightsizing your model selection, establishing a disciplined routing architecture, and aggressively managing vendor risk, you can navigate this landscape successfully. The goal isn&#8217;t to deploy AI the fastest; the goal is to deploy it in a way that actually works for your business.</p><p>If you want to shield yourself from the chaos of the model selection and routing ecosystem, Neurometric&#8217;s core platform evaluates, chooses, and routes models automatically.  <a href="https://studio.neurometric.ai/">Try it</a> for free.</p>]]></content:encoded></item><item><title><![CDATA[What Does A Token Engineering Platform Do?]]></title><description><![CDATA[The new tool for the most important job in AI]]></description><link>https://blog.neurometric.ai/p/what-does-a-token-engineering-platform</link><guid isPermaLink="false">https://blog.neurometric.ai/p/what-does-a-token-engineering-platform</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Tue, 14 Jul 2026 09:36:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If your company runs AI in production, you&#8217;re buying intelligence by the token. And if you&#8217;re like most enterprises, you have almost no tooling to manage that spend the way you manage your cloud spend.</p><p>This should feel familiar. When companies moved to the cloud, an entire layer of cost management and optimization tooling emerged around it &#8212; because once compute became a metered utility, optimizing that meter became a discipline with real dollars attached. AI inference is following the same path, except the stakes are arriving faster. A production AI workload at enterprise scale &#8212; say, a million model calls a day &#8212; can swing by millions of dollars a year based on decisions most teams made once, early, and never revisited.</p><p>Token engineering is the discipline that fixes this. A token engineering platform is the system that operationalizes it.</p><h2>What Is Token Engineering?</h2><p>Token engineering is the practice of treating tokens as an engineered resource: measured, benchmarked, routed, and continuously optimized. It&#8217;s a systems discipline, not a procurement exercise.</p><p>The common misconception is that token engineering means &#8220;use a cheaper model.&#8221; It doesn&#8217;t. It means optimizing every AI workload across three dimensions simultaneously: <strong>cost, speed, and reliability</strong>. Sometimes the right answer is a smaller, cheaper model. Sometimes it&#8217;s a faster one. Sometimes it&#8217;s the frontier model, but with a compressed prompt and an aggressive caching layer in front of it. The point is that the answer is different for every task, and it changes constantly.</p><p>Three forces make this urgent right now. First, model proliferation: frontier LLMs, open-weight models, and small language models (SLMs) now number in the hundreds, with meaningful new releases every month. Second, price variance: the cost of completing the same task can vary by 100x or more depending on which model, technique, and hardware you choose. Third, the capability crossover: for a growing share of enterprise tasks, purpose-built SLMs now match or beat frontier models at a fraction of the cost.</p><p>Meanwhile, most teams pick a model at the start of a project, hardcode it, and move on. Every month that decision goes unexamined, the gap between what they pay and what they should pay gets wider.</p><h2>The Core Components of a Token Engineering Platform</h2><p>Managing this problem manually doesn&#8217;t scale. A token engineering platform automates it. Here&#8217;s what a complete platform looks like &#8212; and how the Neurometric platform implements each piece.</p><h3>1. Workload Evaluation and Benchmarking</h3><p>You cannot optimize what you cannot measure, and public leaderboards won&#8217;t save you. Generic benchmarks tell you how models perform on academic tasks, not on <em>your</em> workloads &#8212; your customer service intents, your document extraction formats, your code review standards.</p><p>A token engineering platform continuously evaluates models and token engineering techniques against your actual tasks. When a new model ships, you shouldn&#8217;t have to wonder whether it&#8217;s better for your use case; the platform should tell you, with graded evidence. Neurometric&#8217;s Harbor evaluation infrastructure does exactly this, with more than 15,000 graded tasks and thousands of benchmark runs powering every recommendation the platform makes.</p><h3>2. Task-Level Routing</h3><p>Evaluation tells you which model is best for which task. Routing acts on it &#8212; automatically, per request, in production.</p><p>This is the decision engine at the heart of the platform. Most applications don&#8217;t have one workload; they have dozens of distinct tasks hiding inside a single product, each with different cost, latency, and quality requirements. Routing at the application level means paying frontier prices for tasks a model one-tenth the cost handles perfectly. Routing at the task level means every request goes to the cheapest model that meets its quality bar. Our Task Level Router keeps this workload-to-model mapping current as models, prices, and your traffic all change &#8212; so the routing decision you&#8217;d make today doesn&#8217;t quietly decay into the wrong decision six months from now.</p><h3>3. Automated SLM Creation and Fine-Tuning</h3><p>Sometimes no existing model sits at the right point on the cost/quality frontier for your task. The frontier model is overkill and overpriced; the small models miss your quality bar. Historically, the answer was a fine-tuning project: hire ML engineers, build a data pipeline, spend a quarter.</p><p>A token engineering platform makes this a feature, not a project. When evaluation data shows that a task is a candidate for a purpose-built model, the platform can distill and fine-tune an SLM automatically, validate it against your benchmarks, and slot it into the routing layer. The economics are hard to ignore: a task-specific SLM can run at 1/50th the cost of a frontier model while matching its accuracy on that narrow task. Our Auto-SLM Creator turns what used to be a specialized data science effort into a platform capability.</p><h3>4. Hardware and Deployment Optimization</h3><p>Which model you run is only half the equation. Where you run it matters just as much.</p><p>The same open-weight model has wildly different economics depending on whether you consume it through an API, self-host it on dedicated GPUs, run it quantized on cheaper hardware, or batch it for throughput over latency. At scale, the API-versus-self-hosting decision alone can be worth seven figures annually &#8212; and the right answer flips as your volume grows and hardware prices move. A token engineering platform models these deployment economics continuously and recommends the optimal placement for each model in your stack, so the decision gets revisited by software instead of forgotten by people.</p><h3>5. Prompt Compression, Rewriting, and Caching</h3><p>The cheapest token is the one you never send.</p><p>Before a request ever reaches a model, there are three opportunities to shrink it: compress the prompt to strip redundancy, rewrite it for token efficiency without losing intent, and cache aggressively &#8212; both exact-match and semantic &#8212; so repeated or near-repeated requests never hit the model at all. Individually these look like small percentage gains. At enterprise volume they compound into real money: at a million calls a day, a 20% reduction in tokens per call is not a rounding error, it&#8217;s a budget line. The platform applies these techniques automatically and only where evaluation shows they don&#8217;t degrade quality.</p><h3>6. Cost Attribution</h3><p>None of the above works as a one-time exercise. Optimization needs a feedback loop, and that loop starts with knowing exactly where your tokens go.</p><p>A token engineering platform attributes token spend at the level your business actually thinks in: per task, per team, per product feature, per customer. This is the FinOps layer for AI. It&#8217;s what lets you answer questions like &#8220;what does our IVR summarization actually cost per call?&#8221; or &#8220;which feature&#8217;s AI spend grew 40% last quarter, and was that traffic or inefficiency?&#8221; Without attribution, every optimization is a guess. With it, the platform can show you &#8212; in dollars &#8212; what each routing decision, SLM deployment, and caching policy is saving.</p><h3>7. Governance, Budgets, and Reliability Controls</h3><p>Finally, the layer that makes all of this safe to run in production: spend limits per team or application, fallback chains when a provider degrades, SLA enforcement on latency and quality, and defined degradation policies for when things go wrong.</p><p>This is the difference between a developer tool and an enterprise platform. Your finance team gets budget enforcement. Your platform team gets reliability guarantees. Your compliance team gets an audit trail of which model handled which request and why.</p><h2>How It All Fits Together</h2><p>These aren&#8217;t seven point solutions bolted together. They&#8217;re a flywheel: <strong>benchmark &#8594; route &#8594; optimize &#8594; attribute &#8594; re-benchmark.</strong></p><p>Evaluation data drives routing decisions. Routing data reveals which tasks are candidates for purpose-built SLMs. Deployment optimization changes the economics that feed back into routing. Cost attribution measures the impact of all of it and surfaces the next opportunity. Every component makes the others smarter, and the loop runs continuously &#8212; which matters, because the model landscape changes monthly and your traffic changes daily.</p><p>This is also the honest answer to the build-versus-buy question. Any strong engineering team can build one of these components. Very few can build all seven, and almost none can afford to <em>maintain</em> all seven against a model market that reprices and re-ranks itself every few weeks. Nobody builds their own cloud cost management platform anymore. The same logic is arriving for tokens.</p><h2>In Summary</h2><p>AI inference is becoming one of the largest new line items in enterprise technology budgets, and it&#8217;s currently one of the least managed. Token engineering is the discipline that closes that gap. A token engineering platform &#8212; evaluation, routing, automated SLM creation, deployment optimization, prompt optimization, cost attribution, and governance &#8212; is how you run it at scale.</p><p>If you&#8217;re spending real money on model calls and can&#8217;t say with confidence that each task is running on the right model, at the right price, on the right hardware, that&#8217;s the gap Neurometric was built to close. [Get in touch to benchmark your workloads -sales@neurometric.ai]</p>]]></content:encoded></item><item><title><![CDATA[Tokenmaxxing to Tokenminning: The Case for Token Engineering]]></title><description><![CDATA[In this episode, Rob May and Calvin Cooper unpack why "token engineering" is on track to become a discipline every AI company needs, on the same trajectory as DevOps or UX before it.]]></description><link>https://blog.neurometric.ai/p/tokenmaxxing-to-tokenminning-the</link><guid isPermaLink="false">https://blog.neurometric.ai/p/tokenmaxxing-to-tokenminning-the</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Sat, 11 Jul 2026 18:04:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>What's in this episode:</p><ul><li><p><strong>The tokenminning.com origin story &#8212; (</strong><a href="https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html">As seen in the NYT</a><strong>)</strong> The Tokenminning Manifesto by Neurometric is gaining traction, with coverage in the NYT. As enterprises prove out AI use cases while blowing through their token budgets, token engineering is becoming a first-order concern.</p></li></ul><ul><li><p>Specialized models vs. frontier models &#8212; Small, task-specific models can almost always beat a frontier model on a given task, even at a fraction of the parameter count. The catch: that specialized model can only do that one thing well. This is a feature, not a bug.</p><p></p></li><li><p>The new Token Engineering Platform &#8212; Neurometric's answer to a fragmented tooling landscape. It brings SLM fine-tuning, distillation, and token caching into a single system so teams get one view across their entire model stack.</p></li></ul><ul><li><p>&nbsp;COGS vs. OPEX &#8212; how Neurometric segments its customer base, and why companies whose AI spend hits gross margin (not just headcount-adjacent budgets) are the ones moving fastest toward token efficiency.</p></li></ul><p></p><p>Listen to the full episode:<strong> <a href="https://tokenengineering.podbean.com/">https://tokenengineering.podbean.com/</a></strong></p><p><strong>Watch on YouTube: <a href="https://youtu.be/JHFeraXq3RU?si=fzzRcEBuPo3vbF6A">https://youtu.be/JHFeraXq3RU?si=fzzRcEBuPo3vbF6A</a></strong></p><p></p><p>**Resources mentioned:**</p><ul><li><p>Tokenminning Manifesto:<a href="http://tokenminning.com/"> tokenminning.com</a></p></li><li><p> Neurometric AI: <a href="http://neurometric.ai/">neurometric.ai</a></p></li><li><p>NYT Article: <a href="https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html">https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html</a></p></li></ul><p></p>]]></content:encoded></item><item><title><![CDATA[Introducing The Neurometric Token Engineering Platform]]></title><description><![CDATA[optimize your inference system]]></description><link>https://blog.neurometric.ai/p/introducing-the-neurometric-token</link><guid isPermaLink="false">https://blog.neurometric.ai/p/introducing-the-neurometric-token</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Thu, 25 Jun 2026 11:08:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2oWO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When we started Neurometric we spoke to many large enterprises about what they were doing with AI and how they were building out their systems.  Most were not very far along but the few that were had all independently come to the same architecture.</p><p>They all started with one large frontier model, and as their inference costs grew, they started to peel off their high volume simpler workloads and set up &#8220;task specific endpoints.&#8221;  Workloads like customer sentiment analysis of support tickets, named entity extraction from documents, email summarization, these don&#8217;t need frontier intelligence.  By setting up an endpoint with a small model for just those workflows, they saw lower costs and faster latency.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2oWO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2oWO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 424w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 848w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1272w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2oWO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp" width="1456" height="1167" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/beb30d01-738b-442d-b053-332956b4519e_1916x1536.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1167,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:71524,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/203535954?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2oWO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 424w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 848w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1272w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The problem is, there are thousands of models out there, and hundreds of techniques to optimize them - from where you run the model to prompting and token management techniques.  And that is all on top of the complicated issue of which model works best for your use case in the first place.</p><p>Today we are announcing that Neurometric has pulled all of these tools together in a single platform that makes it easy to make token engineering decisions.  With our platform you can:</p><ul><li><p>Monitor and measure AI workloads</p></li><li><p>Optimize workloads for cost or latency</p></li><li><p>Build custom SLMs for workloads that benefit from that approach</p></li><li><p>Evaluate and test many models, techniques, and hosting platforms</p></li></ul><p>Our customers typically see an 80% drop in inference charges and a 4x improvement in latency using the Neurometric platform.  </p><p>If you want take your AI optimization to the next level, hire a token engineer, and use the Neurometric platform.  Reach out to us at sales@neurometric.ai if you want to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Tokenmaxxing is out, tokenminning is in]]></title><description><![CDATA[The New York Times (Eli Tan) reported on a shift companies are now making after a year of unchecked AI spending.]]></description><link>https://blog.neurometric.ai/p/tokenmaxxing-is-out-tokenminning</link><guid isPermaLink="false">https://blog.neurometric.ai/p/tokenmaxxing-is-out-tokenminning</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Fri, 19 Jun 2026 15:15:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The New York Times (Eli Tan) reported on a shift companies are now making after a year of unchecked AI spending.</p><p></p><p>Rob May, our CEO, told the Times: CEOs who couldn&#8217;t measure AI savviness defaulted to &#8220;who&#8217;s using the most tokens.&#8221; Volume over efficiency was never going to hold once the bills landed.</p><p></p><p>Uber blew through its full-year AI budget in four months. Meta is capping usage after an &#8220;exponential increase&#8221; in costs. AT&amp;T&#8217;s shows companies can save up to 90% by using less powerful models for most tasks.</p><p></p><p>This is the insight we built Neurometric on.&nbsp;</p><p></p><p>Read the full article: <a href="https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html">https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html</a></p>]]></content:encoded></item><item><title><![CDATA[Automating the SLM Development Loop — Pioneer Agent Paper Breakdown]]></title><description><![CDATA[The bottleneck in deploying small language models isn&#8217;t training.]]></description><link>https://blog.neurometric.ai/p/automating-the-slm-development-loop</link><guid isPermaLink="false">https://blog.neurometric.ai/p/automating-the-slm-development-loop</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Tue, 16 Jun 2026 22:05:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The bottleneck in deploying small language models isn&#8217;t training. It&#8217;s data curation, failure diagnosis, regression avoidance, and iteration control.</p><p>That&#8217;s the central argument of the Pioneer Agent paper from Fastino Labs &#8212; and it maps pretty cleanly onto what we&#8217;ve been building at Neurometric.</p><p>In this episode of Inference Time Tactics, Director of AI Research Yash Sharma breaks it down with co-founder Calvin Cooper: what Pioneer Agent actually is, how it uses Claude Sonnet as an ML engineer in a box, and what the results tell us about where SLMs go next.</p><p></p><p>Key topics:</p><p>&#8212; Cold start data curation with zero customer traces</p><p>&#8212; Why naive retraining fails with noisy production data (and how an agentic loop fixes it)</p><p>&#8212; 83-point benchmark improvements &#8212; and why the 1.6-point cases matter just as much</p><p>&#8212; Regression rollback, hyperparameter search, and the heuristics baked into the system</p><p>&#8212; Where we think SLMs can go beyond &#8220;simple tasks&#8221;</p><p></p><p>Watch the full episode on YouTube: <a href="https://youtu.be/PiIrywMGsAA?si=gs_3nDwgS7QbvwQG">https://youtu.be/PiIrywMGsAA?si=gs_3nDwgS7QbvwQG</a></p><p></p><p>Or listen in on any platform: <a href="https://inferencetimetactics.podstream.com/">https://inferencetimetactics.podstream.com</a></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.neurometric.ai/subscribe?utm_source=email&r=&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.neurometric.ai/subscribe?utm_source=email&r="><span>Subscribe</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Discipline of Token Engineering: Why Your AI Infrastructure Is Bleeding Cash]]></title><description><![CDATA[And how to fix it.]]></description><link>https://blog.neurometric.ai/p/the-discipline-of-token-engineering</link><guid isPermaLink="false">https://blog.neurometric.ai/p/the-discipline-of-token-engineering</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Sat, 13 Jun 2026 17:54:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Remember when building an LLM application simply meant wrapping a basic system prompt around a single frontier API key? Those days are officially over. In 2026, we find ourselves in the middle of a massive architectural shift. We are no longer just building chatbots; we are deploying complex, multi-step agent fleets. Yet this evolution has exposed a glaring operational vulnerability: every company is now a token company, but almost none of them have a token engineer.</p><p>Currently, enterprise infrastructure spend is run with zero discipline. We regularly see frontier models costing $5 to $25 per million tokens assigned to commodity work that $0.10 to $0.50 small models handle easily on benchmarks. This creates a 50x to 250x price spread, paid on every request, every single day. To survive this efficiency gap, a new engineering discipline has emerged: <strong>Token Engineering</strong>. It is the systematic optimization of which model runs which task&#8212;at what size, with what prompt structure, and at what cost.</p><h2>The Agentic Scale Problem</h2><p>The root of this cost crisis is that our foundational software design patterns have completely changed. Over 70% of routed inference traffic now comes from autonomous agents and CLI tools&#8212;not human chat interfaces. While a standard chat turn consumes just a few thousand tokens, a single agentic task regularly chews through 100K to 1M tokens as it loops, reasons, and self-corrects.</p><p>This volume shift has caused platform-wide token volume to explode by more than 10x in roughly a year, with 16 to 18 trillion tokens per week routed on OpenRouter alone. When workloads scale to this magnitude, token waste becomes an engineering failure, not a model failure.</p><h2>The Five Pillars of Token Waste</h2><p>When auditing modern agent pipelines, token waste typically boils down to five core engineering oversights:</p><ul><li><p><strong>Oversized Models</strong>: Allocating expensive frontier models to simple extraction, classification, and formatting tasks that small models win on benchmarks.</p></li><li><p><strong>Prompt Bloat</strong>: Deploying unversioned, unmeasured prompts that carry thousands of redundant tokens into every single call.</p></li><li><p><strong>No Caching Strategy</strong>: Re-sending massive chunks of static context on every request instead of caching it at 10% to 20% of the standard price.</p></li><li><p><strong>Sequential Sprawl</strong>: Running agent steps serially with full context when steps could be decomposed, parallelized, and right-sized across lean endpoints.</p></li><li><p><strong>Blind Retries</strong>: Retrying unexpected failures on the same expensive model with no confidence scoring or cheaper fallback path.</p></li></ul><h2>Why Human Optimization Fails</h2><p>Fixing these leaks manually is a noble goal, but token engineering simply does not scale as a human job. First, we face continuous <strong>Model Churn</strong>. Major model releases land weekly. Look at Gemma 4: it went from non-existent to routing 240 billion tokens per week&#8212;roughly one-third of all small-model traffic on OpenRouter&#8212;in a mere 70 days. No human team can re-benchmark thousands of model-by-task combinations on that clock.</p><p>Second, we hit <strong>Price Churn</strong>. Providers reprice continuously; caching, batching, and changing infrastructure economics shift the optimal choice even when models remain static. The half-life of an optimization is measured in weeks; a hand-tuned pipeline is stale before the next sprint ends. The engineer has to be automated.</p><h2>The FinOps Parallel</h2><p>Every era of computing waste eventually forces the transition from a heroic manual task into an automated platform practice. When web applications struggled with uptime, we turned manual on-call duties into Site Reliability Engineering (SRE). When cloud spend spiraled out of control, cloud waste created FinOps platforms that natively paid for themselves. Today, token spend is the fastest-growing cost line in software, and it demands its own dedicated engineering platform.</p><p>As LLM infrastructure commoditizes, value is rapidly migrating away from foundational providers and straight to the orchestration layer&#8212;the core intelligence about which intelligence to use. By decoupling your software workflows from rigid API keys and transitioning to dynamic task endpoints, you can let automation drive delivery costs down via leaner prompts, continuous model re-matching, and purpose-built small models. In this new landscape, remember: every standard inference vendor makes more money when your system wastes tokens. True operational maturity belongs to the engineering teams who build systems that actively eliminate the bleed.</p><p>If you want help, Neurometric offers a platform ideal for token engineering.  Contact us to chat more about it.</p>]]></content:encoded></item><item><title><![CDATA[What We Learned About The Harbor Framework From More than 5,700 Benchmark Runs.]]></title><description><![CDATA[The infrastructure you run on matters a lot.]]></description><link>https://blog.neurometric.ai/p/what-we-learned-about-the-harbor</link><guid isPermaLink="false">https://blog.neurometric.ai/p/what-we-learned-about-the-harbor</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 10 Jun 2026 18:13:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We ran a large-scale benchmark study to evaluate AI agent performance across a broad model catalog. What we expected to produce was a leaderboard. What we got instead was a lesson in how evaluation infrastructure can quietly invalidate your results before you ever read them.</p><div><hr></div><h2>The Numbers</h2><p><strong>5,791 benchmark runs. 15,750 individual graded tasks.</strong></p><p>That&#8217;s a meaningful dataset. But the headline isn&#8217;t a model score &#8212; it&#8217;s that the framework couldn&#8217;t reliably finish.</p><p><strong>53% of runs errored out.</strong> An errored run produces nothing: zero tasks completed, no usable result. More than half our compute returned empty-handed, not because the models failed, but because the harness crashed.</p><p>This matters beyond the obvious waste. When your error rate crosses 50%, you&#8217;re no longer sampling performance &#8212; you&#8217;re sampling infrastructure luck.</p><div><hr></div><h2>The Grading Problem</h2><p>The errors weren&#8217;t just crashes. They were silent distortions inside the results that <em>did</em> come back.</p><p>In <strong>765 cases</strong>, the agent produced the exactly correct answer &#8212; and Harbor logged it as an error anyway. That&#8217;s <strong>one in five of all correct results</strong> misclassified as failures. A grading system that can&#8217;t recognize its own correct answers isn&#8217;t grading. It&#8217;s noise with a spreadsheet attached.</p><p>The implication: any model rankings produced under these conditions would systematically understate performance, with the degree of understatement varying arbitrarily across runs. You can&#8217;t normalize your way out of that.</p><div><hr></div><h2>Dataset Coverage</h2><p>Harbor ships with 80 datasets. We got <strong>42 of them to run at all</strong> &#8212; just over half. Of those 42, <strong>17 never produced a single scored result</strong>. That&#8217;s 17 datasets that opened, ran, and returned nothing actionable.</p><p>Effective coverage: roughly <strong>31% of the catalog</strong> produced data you could learn from. You cannot benchmark across a catalog when two-thirds of it is structurally inoperable.</p><div><hr></div><h2>The Hello-World Test</h2><p>The clearest evidence is the simplest task.</p><p>Harbor&#8217;s hello-world benchmark asks the agent to create a file containing <code>"Hello, world!"</code> That&#8217;s it. No reasoning. No retrieval. No multi-step planning. Just: write a file.</p><p>We ran it across <strong>1,526 agent+model combinations</strong>. <strong>645 of them &#8212; 42% &#8212; never completed it cleanly even once.</strong></p><p>In one case, GPT-5.4 wrote the file perfectly. The run still errored.</p><p>The intelligence was never the bottleneck. The harness was.</p><div><hr></div><h2>What This Means for AI Benchmarking</h2><p>Benchmark infrastructure is load-bearing. It doesn&#8217;t just measure performance &#8212; it <em>defines</em> what counts as performance. When the harness fails silently, misclassifies correct answers, and drops the majority of its own datasets, the output isn&#8217;t a measurement. It&#8217;s a corrupted signal that looks like data.</p><p>The field has spent enormous energy debating which benchmarks best capture model capability. That&#8217;s the right debate to have &#8212; once the execution layer can be trusted. Our results suggest the execution layer deserves a lot more scrutiny than it typically gets.</p><p>A model that writes the correct answer deserves to have that answer counted. That&#8217;s table stakes for any evaluation system. When it isn&#8217;t met, the benchmark isn&#8217;t measuring models. It&#8217;s measuring the framework.</p>]]></content:encoded></item><item><title><![CDATA[How to Run OpenClaw Without the Frontier Tax]]></title><description><![CDATA[How inference routing cuts 60&#8211;90% of frontier model calls without changing your OpenClaw setup]]></description><link>https://blog.neurometric.ai/p/how-to-run-openclaw-without-the-frontier</link><guid isPermaLink="false">https://blog.neurometric.ai/p/how-to-run-openclaw-without-the-frontier</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Wed, 03 Jun 2026 22:12:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YSzS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YSzS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YSzS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YSzS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png" width="724" height="407.25" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:724,&quot;bytes&quot;:1189569,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/200471143?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YSzS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The OpenClaw cost problem is not a secret. Users are posting about it openly. One tech blogger documented $3,600 in a single month. Others report $200 days from runaway automation loops. Multi-agent setups on premium models routinely hit $600/month before anyone audits the config. The culprit is almost always the same: every sub-task &#8212; a heartbeat check, a JSON format, a ticket classification &#8212; is hitting a frontier model at full price.</p><p>The problem got more complicated in April. Anthropic blocked Claude Pro and Max subscribers from using their flat-rate plans with third-party tools like OpenClaw. OpenAI went the other direction. Sam Altman posted at 2am on May 2: &#8220;you can sign in to openclaw with your chatgpt account now and use your subscription there.&#8221; ChatGPT Plus and Pro now cover OpenClaw usage at a flat monthly rate via Codex OAuth &#8212; no per-token billing.</p><p>That is good news for a lot of users. But subscriptions have limits. ChatGPT Plus carries a 5-hour weekly usage quota. Pro users hit ceilings too on heavy workflows. The moment you are running serious automation &#8212; multi-agent pipelines, scheduled tasks, anything that fires dozens of requests an hour &#8212; you are back against a wall regardless of which subscription you are on.</p><p>The underlying issue is not which provider you use. It is that the wrong model is handling the wrong work.</p><p>wrong work.</p><div><hr></div><h2><strong>Why the Bill Gets Out of Control</strong></h2><p>OpenClaw routes every task to whatever model you have set as default. That model does not know the difference between a request that requires genuine reasoning and one that is just formatting a JSON object or summarizing a paragraph. It handles both the same way: full context load, full inference, full cost.</p><p>Most OpenClaw workflows are 70-80% routine. Heartbeats, memory housekeeping, classification, extraction, formatting, cron jobs. None of it requires a frontier model. But if your default is Claude Sonnet or GPT-5.4, every one of those calls is priced like it does.</p><p>That is the frontier tax. You are paying for capability you are not using.</p><div><hr></div><h2><strong>What We Built and Why It Works</strong></h2><p>Neurometric builds task-specific Small Language Models &#8212; purpose-built for narrow jobs. A model fine-tuned to classify support tickets does not need 1.7 trillion parameters to do that job well. A 7B model trained on thousands of legal extraction examples outperforms a frontier model on that specific task and runs at a fraction of the cost.</p><p>This is not a discount. It is architecture. Small models doing specific things are cheaper to run because they are smaller and more efficient at their job. The economics are structural, not promotional. That is why we can offer 100 million tokens per month for free and sustain unlimited token plans at prices that make sense.</p><p>The frontier model handles what it is actually good at: complex reasoning, multi-step planning, creative synthesis. Everything else routes to a specialist.</p><p>ClawPack is how this plugs into OpenClaw. It sits alongside your existing model as a standard provider. One model ID. Automatic routing. Your frontier model &#8212; or your ChatGPT subscription &#8212; stays in place for the work that needs it. ClawPack handles the rest.</p><p>The result is the same OpenClaw experience, with 60-90% fewer frontier model calls. If you are on a ChatGPT subscription, your quota goes further. If you are on API billing, your bill drops. Either way, you are not burning Opus-level compute to check whether your inbox has anything urgent.</p><div><hr></div><h2><strong>The Stack We Recommend</strong></h2><p>We are partnering with hosting providers like LumaDock to make this easy for OpenClaw users to set up end to end.</p><p><a href="https://lumadock.com/openclaw-vps-hosting">LumaDock</a> offers a purpose-built OpenClaw VPS template &#8212; you pick a plan, deploy a server, and OpenClaw is already installed and running when you SSH in. Starts at $1.99/month. Their<a href="https://lumadock.com/tutorials/openclaw-complete-guide"> complete OpenClaw guide</a> and<a href="https://lumadock.com/faq"> FAQ</a> cover everything from first setup to production configuration.</p><p>Once OpenClaw is running, adding ClawPack takes two steps. Go to<a href="https://marketplace.neurometric.ai/clawpack"> marketplace.neurometric.ai/clawpack</a>, get your free API key, and copy the pre-populated install command the dashboard generates. Paste it into your server terminal. Then:</p><p>openclaw models set neurometric/clawpack</p><p>That&#8217;s it. ClawPack is live. For complex reasoning tasks, configure a fallback to your frontier model or ChatGPT subscription and OpenClaw escalates automatically when the task warrants it.</p><p>Free tier covers 100M tokens/month. No credit card. Unlimited plans available through your Neurometric account.</p><div><hr></div><h2><strong>The Bigger Picture</strong></h2><p>The era of subsidized frontier inference for agentic workflows is coming to an end. Anthropic made that clear in April. OpenAI&#8217;s subscription path is a better deal for many users, but it is still a ceiling, not a solution.</p><p>The solution is not finding a cheaper frontier model. It is using the right model for the right task. Specialized intelligence delivered at the compute cost it actually requires. That is what we are building at Neurometric, and it is why the economics hold regardless of what the big providers do next.</p><div><hr></div><p><em>Get started:<a href="https://lumadock.com/tutorials/cut-openclaw-costs-clawpack/?utm_source=neurometriclumadock&amp;utm_medium=substack"> LumaDock OpenClaw x ClawPack</a></em></p>]]></content:encoded></item><item><title><![CDATA[Lean Inference Workflows: Applying "Lean" Concepts To Building AI Agents]]></title><description><![CDATA[Making inference scale in a cost effective way]]></description><link>https://blog.neurometric.ai/p/lean-inference-workflows-applying</link><guid isPermaLink="false">https://blog.neurometric.ai/p/lean-inference-workflows-applying</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 03 Jun 2026 17:39:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s a production scenario that should feel familiar: your agent hits a simple routing decision&#8212;does this user query need a database lookup or a calculator?&#8212;and it fires off a GPT-4o call with a 12,000-token context window stuffed with documentation it will never read, waits 4 seconds for a response, gets back malformed JSON, retries twice, and burns $0.40 to answer a question that a regex could have handled.</p><p>Multiply that across 10,000 daily requests. Congratulations&#8212;you&#8217;ve built an inference money pit.</p><p>The AI engineering community collectively discovered that &#8220;just throw it at a frontier model&#8221; works great in demos and collapses in production. Agents enter retry death spirals. Context windows bloat with irrelevant RAG results. Sequential LLM calls stack latency until users abandon the workflow. The tools are extraordinarily powerful, and we are using them with the efficiency of a factory floor that nobody has ever walked with a stopwatch.</p><p>Lean Manufacturing fixed this problem for physical production 40 years ago. It&#8217;s time to apply the same discipline to inference.</p><p><strong>Lean Inference Workflows</strong> are the systematic application of Lean/TPS (Toyota Production System) principles to the design of LLM-powered agent architectures. Not as metaphor&#8212;as engineering discipline.</p><div><hr></div><h2>The 7 Wastes of LLM Inference</h2><p>Taiichi Ohno&#8217;s <em>muda</em> framework identified seven categories of waste in manufacturing. Each maps cleanly onto the failure modes we build into agents every day.</p><h3>1. Overproduction &#8212; The Frontier Model Default</h3><p>The most expensive waste is calling a 70B+ frontier model for tasks that don&#8217;t need it. Routing a support ticket to the right queue? That&#8217;s an 8B classification task. Extracting structured fields from a form submission? That&#8217;s a fine-tuned 3B model with a JSON schema. Summarizing a 500-word support thread? You don&#8217;t need GPT-4o.</p><p><strong>The cost asymmetry is staggering.</strong> claude-sonnet runs ~3x the cost of haiku per token. GPT-4o runs ~10x the cost of GPT-4o-mini. When you reflexively reach for the frontier model on every step of a 15-step agent loop, you&#8217;re not just overspending&#8212;you&#8217;re adding latency at every node.  If your task is a common one, you can even move to SLMs which are faster and two orders of magnitude cheaper.</p><p>Treat your agent&#8217;s model selection the same way a traffic engineer treats routing decisions&#8212;based on payload size, complexity score, and confidence threshold, not habit.</p><h3>2. Inventory &#8212; RAG Bloat</h3><p>Your vector database returns the top-20 chunks, and you shove all 20 into the context window &#8220;just in case.&#8221; That&#8217;s inventory waste: stockpiling inputs you probably won&#8217;t use, forcing the model to process them, inflating your input token count, and degrading retrieval precision in the process. More context isn&#8217;t better&#8212;it&#8217;s a longer assembly line with more defect opportunities.</p><p><strong>Controlled inventory</strong> means retrieving fewer, better chunks via re-ranking (a cross-encoder pass over your top-k candidates), then truncating aggressively before injection.</p><h3>3. Waiting &#8212; Sequential Blocking</h3><p>Tool calls that could run in parallel are running in series. You need to fetch a user&#8217;s account history, check their subscription tier, and retrieve their recent support tickets. Instead of three parallel async calls, you have three sequential blocking calls: 300ms + 280ms + 310ms = 890ms of pure waiting.</p><p><strong>async/await + parallel execution</strong> is the <code>asyncio.gather</code> or <code>Promise.all</code> call you should have made. In a multi-step agent DAG, every synchronous bottleneck is a latency tax.</p><h3>4. Defects &#8212; Malformed Outputs and Retry Loops</h3><p>An agent asks for a JSON tool call. The model returns Markdown-wrapped JSON with an extra trailing comma. Your parser throws. The orchestrator retries. The model hallucinates a different schema on the retry. You&#8217;re now three LLM calls deep on a task that should have been one.</p><p><strong>Defects in inference are uniquely expensive</strong> because retries aren&#8217;t cheap reruns&#8212;they&#8217;re full-price LLM calls on an already-failed path. Structured outputs (OpenAI&#8217;s <code>response_format</code>, Anthropic&#8217;s tool use schemas, the <code>instructor</code> library for Python) eliminate this entirely by constraining output at the token-probability level.</p><h3>5. Over-Processing &#8212; Unnecessary Chain-of-Thought</h3><p>CoT is a forcing function for reasoning. It is not a default that belongs in every prompt. A routing classifier does not need to explain its reasoning to itself before assigning a ticket category. A field extractor does not need <code>&lt;thinking&gt;</code> tokens. Stripping CoT from non-reasoning tasks can cut your output token count by 40&#8211;60% on those steps&#8212;with zero quality loss.</p><div><hr></div><h2>Core Principles of Lean Inference</h2><h3>Just-In-Time Context: The Pull System</h3><p>In Lean manufacturing, a pull system means downstream demand triggers upstream production&#8212;nothing gets built until it&#8217;s needed. <strong>JIT Context</strong> means your agent fetches context exactly when a step requires it, scoped precisely to what that step needs.</p><p>The anti-pattern is the &#8220;God Context&#8221;: a single massive system prompt that pre-loads everything the agent <em>might</em> need across all possible execution paths. You pay the full token tax on every call, even when 80% of that context is never accessed.</p><p>The Lean pattern:</p><ul><li><p><strong>Semantic caching</strong> at the retrieval layer: if a semantically similar query was answered 30 seconds ago, return the cached embedding result, not a fresh DB round-trip.</p></li><li><p><strong>Re-ranking before injection</strong>: Run a cross-encoder (a fast, cheap model like <code>ms-marco-MiniLM-L-6-v2</code>) over your retrieved chunks <em>before</em> injecting them into the LLM context. Top-3 precision beats top-20 recall for most tasks.</p></li><li><p><strong>Step-scoped context</strong>: Each node in your agent DAG gets only the context its specific tool call requires. The summarization node doesn&#8217;t need the tool definitions. The routing node doesn&#8217;t need the document corpus.</p></li></ul><h3>Standardized Work: Deterministic Guardrails</h3><p>Lean&#8217;s <em>standardized work</em> principle says that defined, repeatable processes reduce variation and defects. In agent architecture, this translates to: <strong>make your LLM do as little undirected reasoning as possible.</strong></p><p>Logic that can be encoded deterministically should be. Your state machine transitions, routing rules, retry budgets, and tool call sequencing should live in code&#8212;not in a prompt asking the model to figure it out.</p><p>Tools like <strong>LangGraph</strong> let you encode agent control flow as an explicit graph: nodes are LLM calls or tool invocations, edges are conditional transitions, and the state machine is a first-class object you can inspect, test, and version-control. This is categorically different from a single ReAct loop where the model decides everything.</p><p><strong>Structured outputs</strong> (via <code>instructor</code>, OpenAI&#8217;s <code>strict: true</code> JSON mode, or Anthropic&#8217;s tool schemas) are the manufacturing equivalent of a jig: they physically constrain the output to the valid shape, making defects structurally impossible rather than probabilistically unlikely.</p><h3>Takt Time: The Latency Budget</h3><p>Takt time in manufacturing is the maximum allowable time per unit to meet customer demand. In agent design, every workflow should have an explicit <strong>latency budget</strong> per step and per full execution path.</p><p>Define your takt time first. If your end-to-end SLA is 2 seconds and you have 6 agent steps, your average per-step budget is ~333ms. That budget forces architectural decisions:</p><ul><li><p>Can this step use a smaller model to hit the latency target?</p></li><li><p>Should this step be parallelized?</p></li><li><p>Does this step even need an LLM, or is a heuristic or cached result sufficient?</p></li></ul><p><strong>DAG decomposition</strong> is your primary tool here. A complex task that looks like a single LLM call is often a DAG of 4&#8211;6 smaller model calls that can execute in parallel, each with faster TTFT on a smaller model, combining to lower overall latency than the single big call.</p><h3>Prompt Caching as Kanban</h3><p>Anthropic&#8217;s prompt caching and OpenAI&#8217;s equivalent cache system are <strong>Kanban cards for inference</strong>: reusable, pre-positioned work items that don&#8217;t need to be re-manufactured from scratch.</p><p>Your system prompt, tool definitions, and static knowledge base content are the same across thousands of requests. Cache them. On Anthropic&#8217;s API, a cache hit on a 10,000-token system prompt costs 10% of the base input token price. Over millions of calls, this is not a micro-optimization&#8212;it&#8217;s a cost structure change.</p><p>Design your prompts with <strong>cache-friendly prefix ordering</strong>: static system prompt first, static tool definitions second, dynamic context last. Anything that changes per-request must come after anything that doesn&#8217;t.</p><div><hr></div><h2>Before and After: Repo Analysis Agent</h2><p><strong>Before (Naive Architecture)</strong></p><p>A single ReAct loop. One GPT-4o call per step. Full repository context dumped into the window on every iteration. Tool definitions re-sent each time. Sequential file reads. No output validation. Average: <strong>14 seconds, ~85,000 tokens, ~$1.20 per run.</strong></p><p><strong>After (Lean Architecture)</strong></p><ul><li><p>A small <strong>router model</strong> (8B, fine-tuned) classifies the task type and selects the appropriate specialist pipeline &#8212; adds 80ms, saves 60% of downstream model costs</p></li><li><p><strong>Prompt caching</strong> on tool definitions and system context &#8212; 90% cache hit rate after warmup</p></li><li><p><strong>Parallel tool execution</strong> for file reads &#8212; 4 simultaneous reads instead of sequential</p></li><li><p><strong>Structured output enforcement</strong> via <code>instructor</code> &#8212; zero retry loops in 500-run benchmark</p></li><li><p><strong>Strict step budget</strong>: 6 steps max, with a fallback to human handoff at budget exhaustion</p></li></ul><p>Result: <strong>4.2 seconds average, ~18,000 tokens, ~$0.09 per run.</strong> Same output quality score on eval suite. 13x cost reduction. 3.3x latency improvement.</p><div><hr></div><h2>Conclusion</h2><p>The next frontier of AI engineering isn&#8217;t a bigger context window or a more capable base model. It&#8217;s the discipline to use what we already have without waste.</p><p>Every unnecessary frontier model call, every bloated RAG context, every sequential blocking operation, every retry loop from malformed output&#8212;these are engineering failures, not model failures. We built them. We can fix them.</p><p>Lean Inference isn&#8217;t a philosophy&#8212;it&#8217;s a set of concrete architectural decisions you can make this sprint. Audit your agent&#8217;s token burn by step. Map your sequential calls. Add structured outputs. Right-size your models. Cache your static prompts.</p><p>Build leaner. Run faster. Spend less. Ship better agents.</p>]]></content:encoded></item><item><title><![CDATA[Introducing OwlPack: SLMs That Scan And Fix Your Codebase While You Sleep]]></title><description><![CDATA[Designed for a vibecoding world]]></description><link>https://blog.neurometric.ai/p/introducing-owlpack-slms-that-scan</link><guid isPermaLink="false">https://blog.neurometric.ai/p/introducing-owlpack-slms-that-scan</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Thu, 21 May 2026 19:10:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We have over 150 fine tuned models in our <a href="http://marketplace.neurometric.ai">SLM Marketplace</a>, and over the last 6 weeks we noticed that 2 of our code testing SLMs were the most popular downloads.   We took that as a sign and decided to bundle a few coding SLMs together in a product we are announcing today -  <a href="https://marketplace.neurometric.ai/owlpack">Owlpack </a>&#8212; <em><strong>a GitHub App that runs five specialist agents against your repos while you sleep</strong></em>. </p><h2>Built on 5 Small Models</h2><p>When we say Small Language Model, we mean models under 20B parameters. Some are under 3B. They&#8217;re small enough to run on a single GPU, fast enough to return structured output in seconds, and cheap enough that running five of them in parallel against thousands of files is a rounding error, not a line item.</p><p>The catch &#8212; and it&#8217;s a real one &#8212; is that they don&#8217;t know everything. A 3B model is not going to write your novel, debate Kant, or replace Opus. But it doesn&#8217;t need to. It needs to find SQL injections. Or flag a deprecated dependency. Or notice that a function has crept past 400 lines.</p><p>Narrow the task, and small wins.</p><h2>The SLMs That Compose Owlpack</h2><p>Owlpack runs five agents every night, each focused on a different domain:</p><ul><li><p><strong>Hunter</strong> scans for security vulnerabilities &#8212; SQLi, SSRF, auth bypasses, leaked secrets, cross-referenced against live CVE feeds.</p></li><li><p><strong>Tracker</strong> hunts bugs &#8212; null references, race conditions, off-by-one errors, suspicious test coverage gaps.</p></li><li><p><strong>Keeper</strong> audits dependencies &#8212; outdated packages, breaking changes, deprecation notices, abandoned libraries.</p></li><li><p><strong>Mason</strong> identifies refactoring opportunities &#8212; duplication, coupling hotspots, long functions, outdated patterns.</p></li><li><p><strong>Scribe</strong> analyzes the codebase itself &#8212; churn, review latency, complexity drift, module ownership.</p></li></ul><p>Each one is a specialist. Each one runs in parallel. Each one returns structured findings that get deduplicated, diffed against history, and ranked by severity before they land in your inbox by morning.</p><p>Try doing that with a single frontier model. The math falls apart. A nightly full-repo scan, across thousands of users, calling GPT-class inference five times per repo, would cost more than most teams pay for their entire dev tooling stack. That&#8217;s why no one was offering this service. The unit economics didn&#8217;t work.</p><p>With SLMs, they do.</p><h2>The case for specialization</h2><p>There&#8217;s a deeper reason we built it this way, and it goes beyond cost.</p><p>When you fine-tune a small model on a specific task &#8212; say, identifying CVE patterns in JavaScript &#8212; it gets better at that task than a frontier model trained to do everything. Specialists beat generalists when the task has a defined shape. Code review is a defined shape. Dependency auditing is a defined shape. Detecting a leaked AWS key is a <em>very</em> defined shape.</p><p>A frontier model brings a trillion-plus parameters of knowledge about Roman history and SQL injection patterns. Hunter brings the SQL injection patterns. For this job, that&#8217;s the better tool.</p><p>It also means the agents don&#8217;t drift. They don&#8217;t get creative. They don&#8217;t hallucinate a function name that doesn&#8217;t exist in your repo because they read about it in a blog post once. Narrowness is a feature.</p><h2>Privacy as a byproduct</h2><p>There&#8217;s a third benefit that falls out of using SLMs: a smaller blast radius. Owlpack clones your repository at scan time, runs the agents, and deletes the clone. We don&#8217;t need a 1.5T-parameter foundation model to read your code. We need five tight, task-specific models that do their job and forget. That architecture is easier to audit, easier to contain, and easier to deploy in environments where data exfiltration risk actually matters.</p><h2>What this looks like in practice</h2><p>You install the GitHub App. You pick which repos to scan. We clone them at your chosen time, run all five agents in parallel, delete the clone, and deliver a structured briefing by morning. If there&#8217;s nothing new, we don&#8217;t email you. If there&#8217;s a CVE in a transitive dependency, you&#8217;ll know before standup.</p><p>The whole pipeline costs a fraction of what a single frontier call would. That&#8217;s what passes through to pricing. That&#8217;s what makes the service viable.</p><p>The big-model era taught the industry to reach for the biggest hammer in the room. The next era is about picking the right one. Five small, sharp tools, running every night &#8212; that&#8217;s what Owlpack is. SLMs are why it works.</p><p><a href="https://marketplace.neurometric.ai/owlpack">Try Owlpack Free For 7 Days</a>.  Or email us sales@neurometric.ai if you want to try a team plan.</p>]]></content:encoded></item><item><title><![CDATA[Beyond The Hype: What 3,000 Users Taught Us About Small Language Models In The Real World]]></title><description><![CDATA[We hear about SLMs - what are people doing with them?]]></description><link>https://blog.neurometric.ai/p/beyond-the-hype-what-3000-users-taught</link><guid isPermaLink="false">https://blog.neurometric.ai/p/beyond-the-hype-what-3000-users-taught</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 13 May 2026 20:04:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For the last two years, the AI conversation has been dominated by one question: <em>how big can we go?</em> Trillion-parameter models, frontier benchmarks, $100B compute commitments. Bigger, smarter, more general.</p><p>But somewhere along the way, a quieter revolution started inside the enterprise. Companies stopped asking &#8220;what&#8217;s the most capable model?&#8221; and started asking &#8220;what&#8217;s the most <em>useful</em> one?&#8221;</p><p>Small Language Models &#8212; generally defined as models with fewer than 10 billion parameters &#8212; have emerged as the answer. They&#8217;re cheaper to run, fast enough for real-time workflows, and small enough to deploy on-prem or even on a laptop. That changes the economics and the privacy posture of AI in ways the headline-grabbing frontier models simply can&#8217;t match.</p><p>At Neurometric, we just crossed <strong>3,000 active SLM users</strong>.  You can download a fine tuned SLM from us without becoming a customer so, only a little over 2,200 of those users have applied for a key and are using Neurometric to host the SLM. But we&#8217;ve seen which pre-fine-tuned models are most popular from our <a href="http://marketplace.neurometric.ai">SLM Marketplace</a>, and we&#8217;ve talked to some enterprises directly who have asked for our help with larger scale SLM deployments. </p><p>If you&#8217;ve ever wondered what people use SLMs for in the real world - here are the top 5 things we&#8217;ve seen.</p><h2>1. Summarization: Gist Generation at Scale</h2><p>The single biggest workload across our user base is document summarization. Legal briefs, customer transcripts, research reports, internal wikis.</p><p>Why an SLM wins here: most summarization doesn&#8217;t need creative prose. It needs the <em>gist</em> &#8212; accurate, fast, and cheap enough to run across thousands of documents a day. Pushing a 50-page PDF through a frontier model with a giant context window costs real money. An SLM does the same job at a fraction of the latency and a tiny fraction of the cost. When you&#8217;re processing 10,000 documents a week, that math becomes existential.</p><h2>2. Resume Screening: Extraction Without Hallucination</h2><p>HR teams were one of the fastest verticals to adopt. The job isn&#8217;t to &#8220;write a beautiful candidate evaluation&#8221; &#8212; it&#8217;s to pull skills, years of experience, certifications, and seniority into a structured format.</p><p>That&#8217;s an extraction task, not a creativity task. And ironically, smaller, fine-tuned models are <em>less</em> prone to embellishment than their bigger cousins. Fewer parameters means fewer paths for the model to &#8220;fill in the blank&#8221; with something that wasn&#8217;t on the resume. For HR &#8212; where a hallucinated qualification is a compliance problem &#8212; the precision of an SLM is a feature, not a limitation.</p><h2>3. Code Refactoring: Local, Low-Latency Suggestions</h2><p>Developers using SLMs for code refactoring are the most popular group who download and use it themselves, rather than have us host it. The two reasons they seem to like the SLM approach:</p><ul><li><p><strong>Latency.</strong> An autocomplete suggestion that arrives 800ms after you stop typing is useless. Local SLMs respond in tens of milliseconds.</p></li><li><p><strong>Security.</strong> Proprietary code never leaves the machine. For regulated industries and any company with a defensible codebase, that&#8217;s non-negotiable.</p></li></ul><p>The frontier model can still review the architecture. The SLM handles the thousand small edits in between.</p><h2>4. CRM Summary: Killing the Sales Busy Work</h2><p>Sales reps don&#8217;t write call notes; they scribble them. The result is a CRM full of half-finished entries that nobody trusts.</p><p>Our users are using SLMs to convert messy voice memos and chat transcripts into structured CRM fields &#8212; next steps, objections, deal stage, sentiment. This is a textbook SLM workload: repetitive, schema-bound, and high-volume. A rep doing eight calls a day generates 40 summaries a week. Multiply by a 200-person sales org and you can see why running this on a frontier model is a non-starter.</p><h2>5. Meeting Prep: Briefing Sheets Before the Call</h2><p>The fifth pattern is what users are calling &#8220;pre-meeting intelligence.&#8221; Before a customer call, the SLM ingests prior emails, past call notes, account history, and recent product activity &#8212; then generates a one-page briefing sheet.</p><p>Speed is the metric here. The briefing needs to be ready in the 90 seconds between the previous meeting ending and the next one starting. Frontier models can&#8217;t hit that latency reliably. SLMs can.</p><h2>So What Are People Actually Using SLMs For?</h2><p>They&#8217;re not writing novels. They&#8217;re not passing the bar exam. They&#8217;re not solving open-ended research problems.</p><p>They&#8217;re doing the most common work tasks.</p><p>SLMs are workhorse models. Across our thousands of users, the pattern is consistent: repetitive, high-volume, structured tasks where you need roughly <strong>90% accuracy at roughly 5% of the cost</strong> of a frontier LLM. That&#8217;s not a consolation prize &#8212; that&#8217;s the entire enterprise AI opportunity. Most business value isn&#8217;t locked behind PhD-level reasoning. It&#8217;s locked behind tasks too small and too numerous to justify a $0.10-per-call inference bill.</p><h2>The Shift: From General Purpose to Specialized Agents</h2><p>What our users are really previewing is the next phase of enterprise AI: a shift away from one giant general-purpose model handling everything, toward fleets of specialized agents each doing one thing exceptionally well.</p><p>The frontier model becomes the orchestrator. The SLMs do the work.</p><p>If you want to find where SLMs fit in your stack, don&#8217;t start with your hardest problems. Start with your most repetitive ones. Find the tasks your team does a thousand times a week that need to be <em>correct</em>, not <em>clever</em>. That&#8217;s the frontier now.</p><p><strong>Bigger isn&#8217;t better. Specialized is.</strong></p>]]></content:encoded></item><item><title><![CDATA[The CFO's Guide to Lower Inference Costs]]></title><description><![CDATA[How to Stop Paying the "AI Tax" Before It Eats Your Margins]]></description><link>https://blog.neurometric.ai/p/the-cfos-guide-to-lower-inference</link><guid isPermaLink="false">https://blog.neurometric.ai/p/the-cfos-guide-to-lower-inference</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Tue, 12 May 2026 20:02:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Introduction: The &#8220;AI Tax&#8221; and Why It&#8217;s Increasing</h2><p>Every boardroom is talking about AI. Your competitors are deploying it. Your engineers are building with it. Your investors are asking about it. What almost no one is talking about &#8212; until the invoice arrives &#8212; is the ongoing cost of keeping it running.</p><p>That cost has a name: <strong>inference</strong>. Every time your product calls an AI model to generate a response, summarize a document, or answer a customer question, you&#8217;re paying for it. And if your product is gaining traction, you&#8217;re paying more every single month.</p><p>Your engineers might describe the problem like this: <em>&#8220;We need to optimize our inference overhead to prevent compute-spend from cannibalizing our gross margins.&#8221;</em></p><p>In plain English: <strong>we need to stop spending so much money on the electricity and &#8216;brain power&#8217; it costs to run our AI features.</strong></p><p>Frontier models like GPT-5 or Claude are genuinely impressive. But they are also priced like a premium product, and using them as the default solution for every AI task in your stack is the equivalent of hiring a neurosurgeon to take your blood pressure. The capabilities are real. The expense is real. The waste is often real too.</p><p>This guide is for finance leaders, operators, and executives who want to move from &#8220;AI at any cost&#8221; to &#8220;AI at a sustainable cost.&#8221; Each strategy below includes the technical jargon your engineers might use, a plain-English translation, and &#8212; most importantly &#8212; what it means for your bottom line.</p><div><hr></div><h2>Strategy 1: Size Matters &#8212; The Rise of Small Language Models (SLMs)</h2><h3>What Your Engineers Will Say</h3><p><em>&#8220;We should migrate from monolithic LLMs to domain-specific Small Language Models (SLMs) for task-oriented workflows.&#8221;</em></p><h3>What That Actually Means</h3><p>Using a giant, expensive AI brain for simple tasks is wasteful. A smaller, cheaper model that specializes in one specific job will do that job just as well &#8212; at a fraction of the cost.</p><p>Not every AI task requires genius-level reasoning. Summarizing a short email, classifying a support ticket, extracting a date from a form, or checking whether a sentence is positive or negative &#8212; these are not hard problems. They don&#8217;t need a model trained on the entirety of human knowledge. They need a focused, efficient specialist.</p><h3>The Bottom Line</h3><p>The cost difference is not marginal &#8212; it is <strong>staggering</strong>.</p><p>If your application processes 100 million tokens per month &#8212; a realistic volume for a product with moderate usage &#8212; the difference between a frontier model and a small specialized model is <strong>$14,950 per month</strong>, or nearly <strong>$180,000 per year</strong>. For tasks where the smaller model performs equally well, that is pure margin destruction.</p><p><strong>The CFO question to ask:</strong> <em>&#8220;Which of our AI features actually require our most expensive model, and which ones are just using it because it was the default?&#8221;</em></p><div><hr></div><h2>Strategy 2: Hardware &#8212; Newer Isn&#8217;t Always Better</h2><h3>What Your Engineers Will Say</h3><p><em>&#8220;We can achieve better TCO by utilizing legacy GPU clusters or N-1 generation hardware for non-latency-critical inference.&#8221;</em></p><h3>What That Actually Means</h3><p>We don&#8217;t need the world&#8217;s fastest, brand-new computer for every task. Last year&#8217;s chips are still very fast &#8212; and significantly cheaper to rent.</p><p>The AI hardware market is driven by hype and scarcity. NVIDIA&#8217;s latest H100 and H200 GPUs are in extraordinarily high demand, which means cloud providers can charge a premium for them. For real-time, latency-sensitive tasks &#8212; think a live customer chatbot &#8212; that premium may be justified. For everything else, it usually isn&#8217;t.</p><p>Batch processing, overnight report generation, data enrichment pipelines, and other non-time-critical workloads can run on older-generation hardware with no meaningful impact on output quality.</p><h3>Comparing Your Options</h3><p>The practical play is a <strong>tiered hardware strategy</strong>: pay for premium compute only where your users will actually feel the difference. Route everything else to reserved or spot capacity. This is standard practice in mature cloud cost management &#8212; it&#8217;s time to apply the same logic to AI.</p><div><hr></div><h2>Strategy 3: The &#8220;Squish&#8221; Factor &#8212; Quantization</h2><h3>What Your Engineers Will Say</h3><p><em>&#8220;Applying 4-bit or 8-bit quantization to the model weights to reduce VRAM requirements.&#8221;</em></p><h3>What That Actually Means</h3><p>Shrinking the AI model&#8217;s file size so it fits on cheaper hardware. It&#8217;s like converting a massive 4K video file into a standard-definition version &#8212; it plays fine, takes up far less space, and most people can&#8217;t tell the difference for everyday use.</p><p>AI models are, at their core, enormous files made up of billions of numerical values called &#8220;weights.&#8221; Quantization reduces the precision of those numbers &#8212; from 32-bit floating point down to 8-bit or even 4-bit integers. The model becomes physically smaller, requires less memory to run, and can be deployed on less expensive hardware.</p><h3>The Tradeoff &#8212; And Why It&#8217;s Usually Worth It</h3><p>Quantization is not free. There is a quality tradeoff. A fully quantized model will be marginally less &#8220;intelligent&#8221; than its full-precision counterpart &#8212; typically in the range of 1&#8211;3% degradation on benchmark tests.</p><p>In exchange, you can expect:</p><ul><li><p><strong>50&#8211;75% reduction in memory requirements</strong>, enabling deployment on cheaper GPUs</p></li><li><p><strong>Significant reduction in hardware costs</strong>, often 60&#8211;80%</p></li><li><p><strong>Faster inference</strong> in many cases, due to smaller data transfers</p></li></ul><p>For most business applications &#8212; customer support, internal search, document processing, data extraction &#8212; a 1&#8211;2% quality reduction is entirely imperceptible. The math is straightforward: accepting a marginal quality trade in exchange for 70% cost savings is almost always the right financial decision.</p><div><hr></div><h2>Strategy 4: Finding the Right &#8220;Landlord&#8221; &#8212; Inference Hosting Providers</h2><h3>What Your Engineers Will Say</h3><p><em>&#8220;Moving from a general-purpose CSP to a specialized serverless inference provider to minimize cold-start latency and egress fees.&#8221;</em></p><h3>What That Actually Means</h3><p>Instead of running AI through a big, expensive general-purpose cloud platform that charges a premium for convenience, move to a specialized provider built specifically for AI inference. It&#8217;s cheaper, faster to set up, and purpose-built for the job.</p><p>AWS, Azure, and Google Cloud are exceptional general-purpose platforms. They are also priced accordingly, and they layer fees on top of fees &#8212; data egress charges, API gateway fees, storage costs, and support tiers that add up quickly. More importantly, they were not designed from the ground up for AI inference workloads.</p><p>A new generation of specialized inference providers &#8212; companies like Together AI, Fireworks AI, Replicate, and others &#8212; offer the same underlying models at lower prices because their entire infrastructure is optimized for one thing: running AI models efficiently at scale.</p><h3>What to Evaluate</h3><p>When assessing inference providers, the key financial metrics to request from your engineering team are:</p><ul><li><p><strong>Cost per million tokens</strong> for your specific models and use cases</p></li><li><p><strong>Cold-start latency</strong> &#8212; the delay when a model hasn&#8217;t been called recently</p></li><li><p><strong>Egress fees</strong> &#8212; charges for data leaving the provider&#8217;s network</p></li><li><p><strong>SLA and uptime guarantees</strong> relative to your product requirements</p></li></ul><p>The switching cost is typically low. For many teams, a migration to a specialized inference provider is a one-to-two week engineering project that delivers permanent cost reductions of 30&#8211;50%.</p><div><hr></div><h2>Strategy 5: Don&#8217;t Pay to Think Twice &#8212; Caching and Batching</h2><h3>What Your Engineers Will Say</h3><p><em>&#8220;Implementing semantic caching and request batching to improve throughput and reduce redundant compute.&#8221;</em></p><h3>What That Actually Means</h3><p><strong>Caching:</strong> If the AI already answered a question, save that answer and reuse it instead of paying the AI to think through the same problem again.</p><p><strong>Batching:</strong> Instead of sending the AI one piece of work at a time, queue up a pile of tasks and send them all at once. It&#8217;s more efficient, like doing one large grocery run instead of ten small trips.</p><p>These two techniques attack waste from different angles.</p><p><strong>Caching</strong> is particularly powerful for applications where users ask similar or identical questions &#8212; internal knowledge bases, customer FAQ bots, product recommendation engines. A semantic cache stores previous AI responses and recognizes when a new question is close enough in meaning to warrant returning the saved answer rather than generating a new one. Depending on the application, cache hit rates of 30&#8211;60% are achievable, meaning nearly half of all AI calls are resolved for free.</p><p><strong>Batching</strong> is most valuable for background processing workloads &#8212; nightly data enrichment, bulk document analysis, report generation. Running 1,000 tasks in a batch is substantially cheaper per task than running 1,000 individual requests, because the fixed overhead of spinning up compute is amortized across the entire batch.</p><p>Together, these optimizations can reduce effective inference costs by 20&#8211;50% without any change to model selection or hardware.</p><div><hr></div><h2>The CFO&#8217;s Checklist for AI Cost Optimization</h2><p>Before your next quarterly business review, ask your engineering lead to walk through these questions:</p><p><strong>On model selection:</strong></p><ul><li><p>Are we using a frontier model (the &#8220;sledgehammer&#8221;) for tasks that a smaller specialist model (the &#8220;scalpel&#8221;) could handle equally well?</p></li><li><p>Have we audited our AI feature set to categorize tasks by complexity and matched each to the appropriate model tier?</p></li><li><p>Have we run a side-by-side cost comparison for our top three highest-volume AI workflows?</p></li></ul><p><strong>On infrastructure:</strong></p><ul><li><p>Are we running all inference on on-demand, latest-generation hardware, or have we right-sized to reserved and spot capacity for non-critical workloads?</p></li><li><p>Are we paying general-purpose cloud &#8220;convenience fees&#8221; for inference that a specialized provider could handle more cheaply?</p></li></ul><p><strong>On efficiency:</strong></p><ul><li><p>Are we using quantized models where quality requirements allow?</p></li><li><p>Have we implemented caching for high-repetition query patterns?</p></li><li><p>Are background processing workloads running in batches?</p></li></ul><p><strong>On measurement:</strong></p><ul><li><p>Do we track <strong>cost per inference</strong> as a formal metric?</p></li><li><p>Do we have an alerting threshold for when compute spend exceeds a defined percentage of gross margin?</p></li></ul><p>If the answers to more than half of these questions are &#8220;no&#8221; or &#8220;we&#8217;re not sure,&#8221; there is almost certainly significant, recoverable margin sitting in your AI infrastructure.</p><div><hr></div><h2>Conclusion: Efficiency Is a Competitive Moat</h2><p>The companies that will win the AI era are not necessarily the ones using the most powerful models. They are the ones that have learned to deploy AI intelligently &#8212; matching the right tool to the right task, on the right hardware, with the right efficiency optimizations layered on top.</p><p>The company running a comparable AI product for $1,000 per month will systematically outcompete the company running it for $10,000 per month. Lower costs mean higher margins, more room to price competitively, and more budget to reinvest in product development.</p><p>None of the strategies in this guide require cutting corners on quality. They require applying the same financial discipline to AI infrastructure that good operators apply everywhere else in the business.</p><p><strong>Your call to action is simple.</strong> Schedule thirty minutes with your lead engineer and ask two questions:</p><ol><li><p><em>&#8220;What is our current cost per inference?&#8221;</em></p></li><li><p><em>&#8220;Have we tested a smaller, specialized model for any of our high-volume workflows?&#8221;</em></p></li></ol><p>The answers will tell you everything you need to know about where your next margin improvement is hiding.</p>]]></content:encoded></item><item><title><![CDATA[Auto-SLM Creator Updates]]></title><description><![CDATA[describe your task, get a model]]></description><link>https://blog.neurometric.ai/p/auto-slm-creator-updates</link><guid isPermaLink="false">https://blog.neurometric.ai/p/auto-slm-creator-updates</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Thu, 07 May 2026 11:05:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today we launched a new homepage for Neurometric.ai and right on the homepage we highlighted our Auto SLM creator.  Describe your task and in minutes get an email with a fine tuned SLM customized for that task.</p><p>This technology is made possible by our <a href="https://www.harborframework.com/">Harbor integration</a>, which allows us to test lots of models on lots of agentic tasks.  We have a good understanding of what various small base models are capable of.</p><p>The technology is still new, and we do have a brief human review of each SLM but, it happens quickly.  Let us know if you have any questions or issues when you try it out.</p>]]></content:encoded></item><item><title><![CDATA[Harbormaster: A Kubernetes-Native Web Platform for Running Harbor at Scale]]></title><description><![CDATA[A note for the Harbor open source community on what we built on top of the framework, what we learned running ~30 frontier agent benchmark runs through it, and the issues we surfaced along the way.]]></description><link>https://blog.neurometric.ai/p/harbormaster-a-kubernetes-native</link><guid isPermaLink="false">https://blog.neurometric.ai/p/harbormaster-a-kubernetes-native</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Fri, 01 May 2026 14:32:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oaWu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When <a href="https://www.harborframework.com">Harbor</a> shipped alongside Terminal-Bench 2.0, it solved the right problem at the right time: a clean, container-first abstraction for trials and jobs, sandbox adapters for Daytona, Modal, E2B, Runloop, and GKE, and an honest CLI that lets anyone run <code>harbor run --dataset ... --agent ... -n 100</code> and get useful results. For a researcher with an API key and a laptop, that&#8217;s everything you need.</p><p>But the moment you try to run Harbor as a <em>team</em> &#8212; engineers, research scientists, infra folks who don&#8217;t all live in the terminal &#8212; the rough edges show up. Who launched run 1037? What model did it use? Where are the logs? Did the run that died at 03:14 because of a Daytona auth blip ever come back? The CLI doesn&#8217;t answer those questions, and we don&#8217;t think it should &#8212; that&#8217;s not what Harbor is for.</p><p>So we built <strong>Harbormaster</strong>: a web platform that wraps Harbor and runs it on our own Kubernetes cluster. It is decisively <em>not</em> a replacement for Harbor. It is a thin shell around <code>harbor run</code>, the SQLAlchemy job tracking, and the artifact upload conventions, designed for the case where your benchmark suite is shared infrastructure rather than a personal script. Everything Harbor knows how to do, Harbormaster does &#8212; it just gives you a submit form, a dashboard, an audit trail, and a leaderboard at the end of it.</p><p>This post explains how it composes with Harbor, what it taught us about Harbor&#8217;s behavior under sustained load, and the specific bugs and edge cases we&#8217;d like to push back into the project.</p><h2>What it is, concretely</h2><p>The user-facing surface is small on purpose. You pick an agent (Claude Code, Gemini CLI, Codex), a dataset, a model, and a concurrency level, and you hit submit. From there:</p><ol><li><p>The frontend &#8212; a React SPA served by nginx &#8212; POSTs to a FastAPI backend that writes a <code>BenchmarkRun</code> row to Postgres.</p></li><li><p>The backend creates a Kubernetes <code>Job</code> whose pod is a &#8220;runner&#8221; image that has Harbor installed.</p></li><li><p>That runner builds a <code>JobConfig</code> programmatically and calls Harbor&#8217;s Python API.</p></li><li><p>Harbor does what Harbor does: it fans out trial pods at the requested concurrency, talks to the agent, runs the verifier, and writes the trajectory and reward to the configured artifact root.</p></li><li><p>We watch the artifact root &#8212; backed by S3 in production, MinIO locally &#8212; and stream per-trial state back into Postgres so the dashboard can show live progress, logs, and a running leaderboard.</p></li></ol><p>The infrastructure is intentionally boring: EKS with Karpenter for node autoscaling, an ECR pull-through cache so trial pods don&#8217;t hammer Docker Hub, Traefik IngressRoute with OIDC for auth, and k3d + MinIO for local development so contributors can iterate without a cloud account. None of this is novel by itself. The point is that <em>gluing</em> it to Harbor turned out to be straightforward, which is a real compliment to Harbor&#8217;s design &#8212; the <code>BaseEnvironment</code> interface and the Trial/Job model gave us clean seams to build against.</p><h2>The runner-as-Job pattern</h2><p>The single most useful pattern we found is treating a Harbor invocation as a Kubernetes <code>Job</code>, not as a long-lived service. Each benchmark run is its own Job, with its own pod, its own resource envelope, and its own lifecycle. Harbor&#8217;s per-trial concurrency lives one level below: the runner pod calls into Harbor, and Harbor &#8212; via whichever environment adapter you&#8217;ve configured &#8212; fans out to N parallel trial pods on the same cluster.</p><p>Two things fall out of this naturally that we wanted:</p><p><strong>Resilience to runner restarts.</strong> Harbor already uploads trial artifacts trial-by-trial as they finish. We lean on this. If a runner pod gets evicted mid-run (Karpenter consolidation, spot reclaim, OOM), the next reconciliation reads the artifact root, sees which trials have completed, and the dashboard reflects the partial state correctly. The run isn&#8217;t &#8220;lost&#8221; just because the orchestrator went away. This is mostly Harbor&#8217;s design doing the work &#8212; we just don&#8217;t get in its way.</p><p><strong>Resource isolation per run.</strong> A misbehaving Replicationbench run that wants to allocate 32 GB of RAM doesn&#8217;t take down the dashboard or the API, because the runner is in its own pod with its own limits. The Kubernetes Job boundary turns out to be a much better blast radius than &#8220;long-running daemon that orchestrates everything.&#8221;</p><p>This isn&#8217;t a new idea &#8212; it&#8217;s basically how Argo Workflows and Tekton model things &#8212; but we want to call it out for the Harbor community because it&#8217;s a pattern that maps cleanly onto Harbor&#8217;s existing <code>JobConfig</code>/<code>TrialConfig</code> abstractions without requiring any changes to the framework. If you already run Kubernetes, this is a low-friction way to operationalize Harbor for a team.</p><h2>What we ran</h2><p>We used Harbormaster to run three frontier coding agents &#8212; Claude Code (claude-opus-4-6), Gemini CLI (gemini-3.1-pro-preview), and Codex (gpt-5.4) &#8212; across roughly a dozen public datasets each. The full numbers are in the appendix, but the headline results, on the datasets where all three completed:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oaWu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oaWu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 424w, https://substackcdn.com/image/fetch/$s_!oaWu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 848w, https://substackcdn.com/image/fetch/$s_!oaWu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 1272w, https://substackcdn.com/image/fetch/$s_!oaWu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oaWu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png" width="1456" height="795" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:795,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:349330,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/196065098?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oaWu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 424w, https://substackcdn.com/image/fetch/$s_!oaWu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 848w, https://substackcdn.com/image/fetch/$s_!oaWu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 1272w, https://substackcdn.com/image/fetch/$s_!oaWu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd45e3e27-9840-4873-a74d-d4dd2afd0bd2_2872x1568.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Those numbers aren&#8217;t the point of this post &#8212; they&#8217;re a snapshot, and we&#8217;ll publish a proper write-up once SWE-Bench Verified completes for all three agents. What is the point: running this volume of trials end-to-end through Harbor, on our own cluster, on our own dime, gave us a much richer view of where Harbor is solid and where it has sharp edges.</p><h2>What we&#8217;d like to push back upstream</h2><p>Three findings are worth surfacing for the project, and we&#8217;ll be opening issues on each.</p><p><code>/logs/artifacts</code><strong> directory missing in trial environments.</strong> Several of our runs on SWE-Bench Multilingual hit a 53&#8211;86% verifier failure rate that resolved to a missing artifacts directory in the trial container. The verifier expected <code>/logs/artifacts</code> to exist; in some task environments it didn&#8217;t, and the verifier failed fast rather than creating the directory. Both Claude Code and Gemini CLI runs were affected, which suggests it&#8217;s a task-environment issue, not an agent issue. We&#8217;ve worked around it locally by ensuring the directory exists in our runner image preflight, but a fix in the relevant adapter would be cleaner.</p><p><strong>Aider Polyglot verifier errors at scale.</strong> On Aider Polyglot at <code>n_concurrent=4</code>, we saw 88 verifier errors out of 225 trials for Claude Code (94 of 225 for Gemini CLI). The pattern looked race-y &#8212; re-running affected trials individually passed. We suspect contention on a shared verifier resource at high concurrency, but we haven&#8217;t fully isolated it. We&#8217;ll attach our trajectory bundles to the issue.</p><p><strong>Pod OOM on Replicationbench HPC tasks.</strong> A handful of Replicationbench tasks run physics or ML simulations that allocate gigabytes of memory before Harbor finishes attaching. The pod gets killed before the trial can even start, which means the trial doesn&#8217;t show up as a failure in the usual sense &#8212; it just disappears. Two changes would help: a configurable per-task memory ceiling (so we can opt these tasks into a higher limit) and clearer surfacing of &#8220;pod OOMKilled before trial start&#8221; as a distinct failure class in the trial output.</p><p>None of these are dealbreakers. They&#8217;re the kind of thing you only find by running the framework hard, in an environment that isn&#8217;t the reference Daytona setup, against the full heterogeneity of the public benchmark catalog. We&#8217;d rather find them and report them than work around them silently.</p><h2>What&#8217;s next</h2><p>Three things on our roadmap that we think are interesting for the community:</p><p>We&#8217;re publishing the runner image and the K8s Job templates. The web UI itself has too much of our internal auth model baked in to open-source as-is, but the runner &#8212; the part that turns &#8220;Harbor on a laptop&#8221; into &#8220;Harbor on a cluster&#8221; &#8212; is general, and we&#8217;ll have it on GitHub by month&#8217;s end. If you have an EKS or GKE cluster and want a starting point that isn&#8217;t &#8220;stand up Daytona,&#8221; that&#8217;s the artifact to grab.</p><p>We want to contribute a richer EKS-flavored environment adapter alongside the existing GKE one. The two clouds aren&#8217;t identical (IRSA vs. Workload Identity, ECR vs. Artifact Registry), and the differences matter for image pull behavior at high concurrency. We&#8217;ve worked through them; we&#8217;d like that work to live in <code>harbor-framework/harbor</code> rather than only in our fork.</p><p>And we&#8217;d like to keep running the public catalog. The numbers above represent perhaps two weeks of cluster time. There are clear gaps &#8212; Codex hasn&#8217;t finished SWE-Bench Verified, Gemini CLI never completed it across three attempts, and Replicationbench is blocked on the OOM issue above. We&#8217;ll push through those and publish complete leaderboards as runs land.</p><p>If you&#8217;re building something similar &#8212; we&#8217;d love to compare notes. Harbor&#8217;s whole bet is that a shared standard makes the ecosystem move faster, and operationalizing that standard for teams is exactly the kind of thing the community can do better together than any one of us can do alone.</p><p>Over time we will be using Harbormaster to run as much of the Harbor suite as possible.  Our work at Neurometric is heavily tied to task-specific AI and so, running Harbor on lots of models to understand the smallest, the fastest, and the most cost effective models for each task will help us move our product forward.  As we go through this, we will make sure we publish the results.  Stay tuned.</p><p></p>]]></content:encoded></item><item><title><![CDATA[HX Is The New UX: What You Need To Know About Harness Experience.]]></title><description><![CDATA[The interface is dead. Long live the harness.]]></description><link>https://blog.neurometric.ai/p/hx-is-the-new-ux-what-you-need-to</link><guid isPermaLink="false">https://blog.neurometric.ai/p/hx-is-the-new-ux-what-you-need-to</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Thu, 23 Apr 2026 23:38:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yF0M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For thirty years, the central obsession of product design has been a single question: <em>how do we make it easier for a human to click the right button?</em> We built funnels. We A/B tested button colors. We agonized over empty states and loading spinners.</p><p>That era is ending &#8212; not gradually, but structurally.</p><div><hr></div><h2>The Death of the Funnel</h2><p>Consider a traditional travel booking site. Its UX is a masterpiece of guided constraint: search bar, filter panel, calendar picker, seat map, confirmation modal. Every screen funnels the human toward a single, monetizable action. The design <em>is</em> the product.</p><p>Now consider what happens when an AI agent books your travel. It doesn&#8217;t load the homepage. It doesn&#8217;t hover over the &#8220;flexible dates&#8221; toggle. It hits an API, cross-references your calendar, checks your preference history, and surfaces a ranked shortlist &#8212; bypassing every carefully crafted screen entirely.</p><p>The funnel isn&#8217;t just broken. It&#8217;s irrelevant.</p><p>Agents don&#8217;t navigate UIs. They negotiate with systems. And when the agent is the primary &#8220;user&#8221; of software, the human behind it occupies an entirely different role &#8212; one for which we have almost no design vocabulary. Until now.</p><div><hr></div><h2>Defining HX: The Harness Experience</h2><p><strong>HX &#8212; Harness Experience &#8212; is the design discipline governing the interface between a human and their agentic fleet.</strong></p><p>Where UX asks &#8220;how do I get the human from A to B?&#8221;, HX asks something more complex: how do I let a human <em>steer, trust, and audit</em> a system acting on their behalf, at speed, across dozens of simultaneous tasks?</p><p>The word &#8220;harness&#8221; is deliberate. A harness isn&#8217;t a cage &#8212; it doesn&#8217;t constrain. It channels energy, distributes load, and <strong>keeps you connected to something far more powerful than yourself</strong>. A well-designed harness means you stay in control without doing all the work. A poorly designed one means you get dragged.</p><p>That is the failure mode of most AI-native products being built today</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yF0M!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yF0M!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!yF0M!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!yF0M!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!yF0M!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yF0M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png" width="1456" height="794" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:794,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5106022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/195296009?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yF0M!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 424w, https://substackcdn.com/image/fetch/$s_!yF0M!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 848w, https://substackcdn.com/image/fetch/$s_!yF0M!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!yF0M!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6150c2bc-6049-4e9e-9051-463805873a75_2816x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>.</p><h3>From User to Director</h3><p>There is a psychological shift happening that designers are not yet taking seriously.</p><p>A <em>user</em> takes action. They click, type, swipe. Engagement is transactional and immediate &#8212; the button changes color; the form submits.</p><p>A <em>director</em> evaluates outcomes. They set intent, observe execution, and make judgment calls about whether the agent&#8217;s behavior actually reflects their goals. It requires a fundamentally different kind of trust.</p><p>Think about the difference between driving a car and managing a driver. As a driver, you feel every turn. As a manager of a driver, you need a map, an ETA, and a reason to believe they know the route.</p><p>HX must be built for directors, not drivers.</p><div><hr></div><h2>The Three Pillars of HX</h2><p><strong>Steerability</strong> is the first. Beyond a simple prompt, how do you express nuance? A great travel agent doesn&#8217;t just hear &#8220;I want a beach vacation.&#8221; They remember you hate crowds and ask whether this trip is romantic or family. Steerability means designing systems where human intent can be expressed in layers &#8212; with context, constraints, and corrections &#8212; not just a one-line instruction.</p><p><strong>Transparency and Auditability</strong> is the second. An agent that acts like a black box will eventually fail in ways that feel inexplicable, and therefore unforgivable. But flooding the director with raw logs is its own design failure. The challenge is <em>legibility at the right altitude</em> &#8212; letting a human understand <em>why</em> a decision was made without requiring them to become an engineer. Think of it like a flight manifest versus cockpit telemetry &#8212; same journey, entirely different information needs.</p><p><strong>Intervention Points</strong> are the third and most underrated. The moment a human most needs control is often the moment automation is least prepared to hand it back. Designing intervention points means thinking carefully about <em>when</em> to surface a pause, <em>how</em> to hand off context cleanly, and <em>how</em> to let a human step in &#8212; and back out &#8212; without shattering what the agent was doing. This is not a button. It is an architecture.</p><div><hr></div><h2>The New Design Toolkit</h2><p>So what replaces Figma?</p><p>Not a single tool, but a new set of design primitives. <strong>Designing for latency</strong> &#8212; because agents operate over seconds or minutes, not milliseconds, and unexplained silence breeds distrust faster than failure does. <strong>Designing for uncertainty</strong> &#8212; because an agent operating at 78% confidence needs to surface that honestly, not paper over it with false precision. <strong>Designing for verification</strong> &#8212; because the human&#8217;s core job is now quality control, and the interface must make that fast, clear, and low-friction.</p><p>The new craft lives in feedback loops, not flows. In state communication, not screen transitions. In systems honest about the edges of their own competence.</p><div><hr></div><h2>Stop Building Cages. Start Building Harnesses.</h2><p>The best UX designers of the last decade built elegant cages &#8212; beautiful, invisible paths that guided users exactly where the product needed them to go. It was a form of benevolent choreography.</p><p>That skill, applied to agentic systems, produces something genuinely dangerous: automation that feels smooth right up until it catastrophically diverges from what you actually wanted.</p><p>The call to action is urgent. <strong>Stop designing for obedience. Start designing for collaboration.</strong> The agent is your most powerful colleague &#8212; capable, fast, and tireless. HX is how you ensure it is actually working <em>for</em> you, not just working.</p><p>The funnel is over. Build the harness.</p>]]></content:encoded></item></channel></rss>