<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Neurometric Blog]]></title><description><![CDATA[A substack for Neurometric - AI system orchestration and workflow SLMs ]]></description><link>https://blog.neurometric.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png</url><title>Neurometric Blog</title><link>https://blog.neurometric.ai</link></image><generator>Substack</generator><lastBuildDate>Thu, 08 Oct 2026 00:31:56 GMT</lastBuildDate><atom:link href="https://blog.neurometric.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[neurometric]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[neurometric@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[neurometric@substack.com]]></itunes:email><itunes:name><![CDATA[neurometric]]></itunes:name></itunes:owner><itunes:author><![CDATA[neurometric]]></itunes:author><googleplay:owner><![CDATA[neurometric@substack.com]]></googleplay:owner><googleplay:email><![CDATA[neurometric@substack.com]]></googleplay:email><googleplay:author><![CDATA[neurometric]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Introducing Taskrouter: The AI Router API That Learns To Optimize For Your Workloads]]></title><description><![CDATA[All the major models plus task specific ones]]></description><link>https://blog.neurometric.ai/p/introducing-taskrouter-the-ai-router</link><guid isPermaLink="false">https://blog.neurometric.ai/p/introducing-taskrouter-the-ai-router</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 07 Oct 2026 21:27:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We have a big announcement today - Taskrouter.com is officially out of beta.  </p><p>AI has moved from the playground to production, and the bills moved with it. Teams now spend thousands of dollars sending every query to a massive frontier model, including the queries that are simple.</p><p>Most developers know this is wasteful. You want frontier-level reasoning for the complex edge cases. For the repetitive tasks that make up most of your traffic, you want the speed and cost of a specialized small language model (SLM).</p><p>The hard part is knowing which is which. Deciding when a smaller model is good enough takes constant benchmarking, testing and custom routing logic. That is a full-time job, and it isn&#8217;t the product you set out to build.</p><p>Taskrouter does that job for you. It isn&#8217;t another API aggregator. It&#8217;s a router that learns your specific tasks and optimizes them for quality, speed and cost.</p><h2>How Taskrouter works: from frontier to specialized</h2><p>Integration is one API key. You point your application at Taskrouter the same way you would at OpenRouter or OpenAI.</p><p>From there, four things happen:</p><ol><li><p><strong>Start strong.</strong> Your initial workloads go to top-tier frontier models. That sets a quality baseline and captures gold-standard outputs for each task.</p></li><li><p><strong>Test continuously.</strong> Behind the scenes, Taskrouter runs those same tasks against a fleet of smaller, specialized and open-source models.</p></li><li><p><strong>Hand off when the data says so.</strong> Once a cheaper, faster SLM is shown to match or beat the frontier model on your specific task, Taskrouter routes that traffic to it automatically.</p></li><li><p><strong>Keep the safety net.</strong> Automatic failover to frontier models stays in place for tasks that require deep reasoning.</p></li></ol><p>The result is lower cost and lower latency at the same quality. <strong>Our first Fortune 500 customer to deploy Taskrouter in production saved millions of dollars, over 90% of their inference costs, while seeing their latency drop by 200ms on average and accuracy improving 5 percentage points.</strong></p><h2>The elephant in the room: privacy, IP and your data</h2><p>When enterprise developers hear &#8220;learns from your data,&#8221; alarm bells ring. You picture your proprietary data being fed into a massive public model. That is a reasonable fear, and we built Taskrouter around it.</p><p>We come from enterprise. Taskrouter was built by the team that built my first company.  We ran backup for Fortune 500 firms and had to pass rigorous security audits.  In our current form we have been building SLMs for enterprise customers for over a year, and have active Fortune 500 customers on those SLMs. We know what it takes to pass a corporate infosec review.</p><p>Here is the Taskrouter privacy guarantee:</p><ul><li><p><strong>It&#8217;s your IP.</strong> We don&#8217;t train public models on your data. Period.  We only use it to route to the right models, and to fine tune models that are just yours, specialized to your unique workflows.  You own those models.</p></li><li><p><strong>Optimization is isolated.</strong> The routing intelligence and the specialized models we spin up for your workloads are entirely isolated. They learn for you, not for us, and certainly not for your competitors.</p></li><li><p><strong>You control the data.</strong> We act as a strict pass-through with enterprise-grade data retention policies. When we use your data to benchmark a smaller model, it stays inside a secure, tenant-isolated environment.</p></li></ul><h2>Measured models, clearer choices</h2><p>Routing decisions shouldn&#8217;t be a black box, so we show our work.</p><p><strong>See the comparisons.</strong> Our Task Explorer page shows hundreds of comparisons of various models on common enterprise workloads. You can check how each model performs before you route a single request.</p><p><strong>Keep control.</strong> Taskrouter can handle routing automatically, but you keep ultimate control over which models are used and where your data goes.</p><p><strong>Prove the ROI.</strong> It&#8217;s not enough to be doing AI; you have to show the business value. Taskrouter&#8217;s analytics show exactly how much money and time efficient routing saves you.</p><p>That is the whole idea: real tasks, measured models, and clearer choices.</p><h2>Try it, and tell us where it breaks</h2><p>Sign up by October 10th and get $25 in free inference credits, good on any model. <strong>taskrouter.com</strong></p><p>We also want your feedback. Throw your hardest tasks at Taskrouter and tell us where it breaks.</p><p>Stop paying frontier prices for everyday tasks. Let Taskrouter learn your workloads and optimize your AI infrastructure automatically.</p>]]></content:encoded></item><item><title><![CDATA[Compounding Inference Is As Powerful As Compounding Interest]]></title><description><![CDATA[The best idea from finance applies to AI as well]]></description><link>https://blog.neurometric.ai/p/compounding-inference-is-as-powerful</link><guid isPermaLink="false">https://blog.neurometric.ai/p/compounding-inference-is-as-powerful</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Tue, 15 Sep 2026 18:01:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Z6RJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Compound interest works because of one decision: you reinvest the return instead of spending it. The rate matters less than most people think. A modest rate reinvested for long enough beats a great rate paid out as income. Linear versus exponential is entirely about whether the output gets fed back into the base.</p><p>Most AI teams run inference as income. A request comes in, you call a model, you get an answer, the answer is consumed. The next request costs the same, uses the same model, and produces the same quality. You&#8217;ve spent the token. Nothing accrued.</p><p>Tobi L&#252;tke posted a chart on September 1 that shows what reinvestment looks like. Shopify&#8217;s ML team fine-tuned a 0.8B-parameter model for one job, building buyer profiles, and it now beats GPT-5.6-sol at its highest reasoning setting on that task: 84.6 versus 83.0 on their judge. Tobi&#8217;s take was that tiny models for special-purpose tasks work &#8220;incredibly well&#8221; when you have &#8220;a great self improving recursive flywheel.&#8221; Read that carefully. The result isn&#8217;t the small model. The result is the machine that produced it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Z6RJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 424w, https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 848w, https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 1272w, https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png" width="974" height="978" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:978,&quot;width&quot;:974,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:406441,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.neurometric.ai/i/215866532?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 424w, https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 848w, https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 1272w, https://substackcdn.com/image/fetch/$s_!Z6RJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b08f48e-ebe4-420b-b5fd-48a9289bdc06_974x978.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What the chart actually shows</h2><p>Three training runs in one week. July 23, 29k samples, score 75.3. July 27, 42k samples, 78.1. July 30, 54k samples, 84.6. The gains got bigger with each run, not smaller. That&#8217;s the shape of compounding, and it came from three loops running simultaneously.</p><p>The quality loop. The frontier model is the teacher. Its outputs get graded, and the ones that pass become training data for the student. Every expensive API call is doing two jobs: serving the request and depositing a labeled example into your dataset.</p><p>The cost loop. The system prompt went from 9.1K tokens written out in full to 1.1K tokens &#8220;gisted,&#8221; an 8x cut. Once the model has learned the instructions through training, you stop paying to re-explain them on every call. Smaller model, shorter prompt, cheaper request.</p><p>The volume loop. Throughput went from 2M profiles per day on the prior 2B production model to 72M per day on the 0.8B model across 100 H100s. A 36x increase. More volume means more graded samples, which means better training data, which means the next run is cheaper and better still. The savings from cheap inference fund the inference that generates the next round of training data.</p><p>Each loop feeds the other two. That&#8217;s the reinvestment.</p><h2>What it takes to build this</h2><p>Anyone can fine-tune Qwen. That isn&#8217;t the hard part and it isn&#8217;t the moat. The Shopify team went from behind the frontier to ahead of it in seven days because the surrounding infrastructure already existed. Four pieces matter.</p><p>A task-level definition of quality. Not &#8220;is the model good&#8221; but &#8220;is this specific output a good buyer profile.&#8221; You need a judge that scores individual outputs on this task, calibrated well enough that you&#8217;d trust it to decide what goes into the training set. Without this, you have no way to know whether run three is better than run two, and no way to filter the teacher&#8217;s outputs into clean data. Everything downstream depends on it.</p><p>A pipeline from production traffic to training samples. The data isn&#8217;t in a lab. It&#8217;s in the requests you&#8217;re already serving. If your inference logs are write-only, you&#8217;re paying the teacher and throwing the lesson away. The plumbing that turns yesterday&#8217;s traffic into today&#8217;s graded dataset is the actual asset.</p><p>A cadence. Three runs in a week. Not a quarterly retraining project with a Jira epic. The flywheel compounds at the frequency you turn it, and a team that ships a new checkpoint every few days will lap a team that does it twice a year, even if the second team&#8217;s individual runs are better.</p><p>Clear routing between teacher and student. Early on, the frontier model handles everything and the student learns. As the student closes the gap, traffic shifts. You need to know, per task, which model is winning right now, and the answer changes weekly. Guessing here either wastes money on the teacher or ships bad outputs from the student.</p><h2>The mindset shift</h2><p>Stop thinking about your API bill as rent. Think of it as tuition. The frontier model&#8217;s job on any narrow, high-volume task is to make itself unnecessary for that task. If you&#8217;re six months into production and still calling the same frontier model with the same 9K-token prompt for the same job, you&#8217;ve been paying tuition and skipping class.</p><p>The teams that internalize this won&#8217;t look dramatically different next quarter. Compounding never does at first. They&#8217;ll look untouchable in two years, with dozens of tasks running on models that cost a fraction of frontier pricing and outperform it, and it&#8217;ll seem sudden to everyone who was paying simple interest the whole time.</p><p>If you want a system that makes it easy to compound your IP into AI models, that&#8217;s what we do here at Neurometric.  Contact us if you would like to chat.</p>]]></content:encoded></item><item><title><![CDATA[Four New Task Specific Models Now On Trusted Router]]></title><description><![CDATA[Lighting fast, ultra cheap single task model]]></description><link>https://blog.neurometric.ai/p/four-new-task-specific-models-now</link><guid isPermaLink="false">https://blog.neurometric.ai/p/four-new-task-specific-models-now</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Fri, 04 Sep 2026 15:56:04 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We just launched four new task specific models, available on our <a href="https://trustedrouter.com/providers/neurometric">Trustedrouter page</a>.  They are fine tuned to be very fast and very cheap at individual tasks, including:  document extraction, grounded document Q&amp;A, conversation summary, and classification routing.  Below are instructions for using them.</p><p><strong>1. Document Structured Extraction --Published<br></strong><br>Endpoint: neurometric/document-structured-extraction<br><br>Use this to extract selected fields from document text. The caller supplies the document and an output schema defining the requested fields and types. Missing fields are returned as<br>null.<br><br>Example input:<br><br>{&#8221;document&#8221;:&#8221;Invoice 1042. Vendor: Acme Corporation. Date: 2026-09-01. Total: USD 250.00.&#8221;,&#8221;output_schema&#8221;:<br>{&#8221;invoice_id&#8221;:&#8221;string&#8221;,&#8221;vendor&#8221;:&#8221;string&#8221;,&#8221;date&#8221;:&#8221;string&#8221;,&#8221;total&#8221;:&#8221;number&#8221;,&#8221;currency&#8221;:&#8221;string&#8221;,&#8221;purchase_order&#8221;:&#8221;string|null&#8221;}}<br><br>Expected output:<br><br>{&#8221;invoice_id&#8221;:&#8221;1042&#8221;,&#8221;vendor&#8221;:&#8221;Acme Corporation&#8221;,&#8221;date&#8221;:&#8221;2026-09-01&#8221;,&#8221;total&#8221;:250,&#8221;currency&#8221;:&#8221;USD&#8221;,&#8221;purchase_order&#8221;:null}<br><br>It currently expects extracted text rather than a raw PDF.<br><br><strong> 2. Grounded Document QA</strong><br></p><p>Endpoint: neurometric/grounded-document-qa<br><br>Use this to answer questions using only supplied document chunks. The response includes the answer and supporting source IDs. If the documents do not contain the answer, the endpoint abstains.<br><br>Example input:<br><br>{&#8221;chunks&#8221;:[{&#8221;id&#8221;:&#8221;POLICY-1&#8221;,&#8221;status&#8221;:&#8221;authoritative&#8221;,&#8221;text&#8221;:&#8221;Customers may return unused products within 30 days of purchase with a receipt.&#8221;},{&#8221;id&#8221;:&#8221;POLICY-<br>2&#8221;,&#8221;status&#8221;:&#8221;authoritative&#8221;,&#8221;text&#8221;:&#8221;Refunds are issued to the original payment method.&#8221;}],&#8221;question&#8221;:&#8221;How long do customers have to return an unused product?&#8221;}<br><br>Expected output:<br><br>{&#8221;answer&#8221;:&#8221;30 days&#8221;,&#8221;citations&#8221;:[&#8221;POLICY-1&#8221;]}<br><br>Example when the answer is unavailable:<br><br>{&#8221;answer&#8221;:null,&#8221;citations&#8221;:[]}<br><br><strong>3. Conversation Summary<br></strong><br>Endpoint: neurometric/conversation-summary<br><br>Use this to extract the latest decision, current status, outstanding actions, owners, deadlines, and material risks from a business conversation. Completed and superseded information<br>is excluded.<br><br>Example input:<br><br>{&#8221;conversation_id&#8221;:&#8221;THREAD-001&#8221;,&#8221;messages&#8221;:[{&#8221;id&#8221;:&#8221;MSG-001&#8221;,&#8221;body&#8221;:&#8221;FINAL DECISION: Launch the new customer portal&#8221;},{&#8221;id&#8221;:&#8221;MSG-002&#8221;,&#8221;body&#8221;:&#8221;CURRENT STATUS: Waiting for security<br>approval&#8221;},{&#8221;id&#8221;:&#8221;MSG-003&#8221;,&#8221;body&#8221;:&#8221;OPEN ACTION: Send the approval packet. OWNER: Amina. DUE: 2026-09-10.&#8221;},{&#8221;id&#8221;:&#8221;MSG-004&#8221;,&#8221;body&#8221;:&#8221;MATERIAL RISK: Security review may delay the<br>launch&#8221;}]}<br><br>Expected output:<br><br>{&#8221;decision&#8221;:&#8221;Launch the new customer portal&#8221;,&#8221;current_status&#8221;:&#8221;Waiting for security approval&#8221;,&#8221;open_items&#8221;:[{&#8221;action&#8221;:&#8221;Send the approval packet&#8221;,&#8221;owner&#8221;:&#8221;Amina&#8221;,&#8221;due_date&#8221;:&#8221;2026-09-<br>10&#8221;}],&#8221;risks&#8221;:[&#8221;Security review may delay the launch&#8221;]}<br><br><strong>4. Classification Router<br>Endpoint: neurometric/classification-router<br></strong><br>Use this to classify requests using a taxonomy supplied at request time. The caller provides labels, their meanings, a fallback label, and the requests to classify.<br><br>Example input:<br><br>{&#8221;labels&#8221;:{&#8221;Billing&#8221;:&#8221;charges or refunds&#8221;,&#8221;Technical&#8221;:&#8221;errors or outages&#8221;,&#8221;Other&#8221;:&#8221;anything else&#8221;},&#8221;fallback_label&#8221;:&#8221;Other&#8221;,&#8221;requests&#8221;:[{&#8221;id&#8221;:&#8221;REQ-1&#8221;,&#8221;text&#8221;:&#8221;Please refund this<br>duplicate charge.&#8221;},{&#8221;id&#8221;:&#8221;REQ-2&#8221;,&#8221;text&#8221;:&#8221;The API returns a 502 error.&#8221;},{&#8221;id&#8221;:&#8221;REQ-3&#8221;,&#8221;text&#8221;:&#8221;What is your office address?&#8221;}]}<br><br>Expected output:<br><br>{&#8221;REQ-1&#8221;:&#8221;Billing&#8221;,&#8221;REQ-2&#8221;:&#8221;Technical&#8221;,&#8221;REQ-3&#8221;:&#8221;Other&#8221;}</p>]]></content:encoded></item><item><title><![CDATA[Introducing A Task Specific Tool Calling Model - Available on TrustedRouter]]></title><description><![CDATA[When you need fast efficient tool calling]]></description><link>https://blog.neurometric.ai/p/introducing-a-task-specific-tool</link><guid isPermaLink="false">https://blog.neurometric.ai/p/introducing-a-task-specific-tool</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Sat, 29 Aug 2026 20:23:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Today we&#8217;re making our Neurometric tool calling SLM <a href="https://trustedrouter.com/models/neurometric/tool-choice">available on TrustedRouter</a>. It does one thing: it turns intent into valid, schema-bound tool calls. It has its own pipeline and harness tuned for that job, and nothing else. Pricing is $0.01 per million input tokens and $0.10 per million output tokens.</p><p>Most teams building agents are paying frontier-model prices for a task that does not need a frontier model. Tool selection is a narrow, highly structured problem. Treating it as one changes the economics of the whole system.</p><h2>What you get</h2><p><strong>Cost reduction on the largest line item.</strong> In a running agent, the expensive part is not reasoning. It&#8217;s context accumulation: re-sending tool definitions, past tool outputs, and environment state on every single turn. That accounts for 80&#8211;90% of agent spend in most architectures. Moving tool selection onto a small, cheap model drops cost per turn by 70&#8211;90%, and the effect compounds with every additional turn in a loop.</p><p><strong>Lower latency end to end.</strong> Small models deliver much faster time to first token and higher generation throughput. In multi-step agents, 30&#8211;40% of wall-clock time disappears into inter-call orchestration overhead rather than useful work. Shortening each hop shortens the entire chain. Users experience this as an agent that feels responsive instead of one that appears to stall between steps.</p><p><strong>Higher schema reliability.</strong> A model fine-tuned exclusively on JSON schema adherence produces fewer syntax errors, fewer missing required arguments, and fewer hallucinated parameters than a general-purpose model juggling reasoning, tone, and format compliance in a single pass. Specialization beats breadth here. Fewer malformed calls also means fewer retries, which is a second-order cost and latency win.</p><p><strong>Tighter control over what leaves your perimeter.</strong> Tool schemas are a description of your internal systems: database fields, endpoint names, argument semantics. Isolating schema binding in a dedicated small model lets you keep those definitions out of the context you send to a general-purpose provider, and gives you a deployment story that fits inside a VPC or on-prem environment when that matters.</p><h2>How to use it</h2><p>Four patterns cover most of what we&#8217;ve seen work.</p><p><strong>Intent router (fast path).</strong> Put the model at the entry point of your pipeline. Standard single-turn requests such as &#8220;look up user ID 1234&#8221; or &#8220;fetch local weather&#8221; go straight to it, which generates the API payload and executes. Your reasoning model never sees the request. In most production agents, a large majority of turns look like this.</p><p><strong>Planner/executor split.</strong> Use a large reasoning model strictly to decompose a complex problem into natural-language sub-goals. Hand each sub-goal to the small model, which converts the instruction into valid arguments and calls the tool. The expensive model does planning, which is what it&#8217;s good at. The cheap model does binding, which is what it&#8217;s good at.</p><p><strong>Structured output compiler.</strong> Let your reasoning model emit its decision as lightweight text rather than strict JSON. Feed that text to the small model to map intent onto schema-bound parameters. This removes a real failure mode, where a reasoning model degrades its own reasoning because it is simultaneously trying to satisfy a format constraint.</p><p><strong>Speculative tool generation.</strong> Fire the small model in parallel to predict likely tool calls from the raw user input while your primary model is still drafting its strategy. When the prediction matches, you&#8217;ve already paid the latency. At these prices, speculating and discarding is cheap enough to be a rounding error.</p><h2>Getting started</h2><p>The model is live on TrustedRouter now. If you&#8217;re already routing through TrustedRouter, you can point tool-calling traffic at it and compare against your current path directly, on your own workloads.</p>]]></content:encoded></item><item><title><![CDATA[Gemma 4b vs Gemini Flash: You Don't Need Frontier Models For Tool Calling Workflows]]></title><description><![CDATA[What our Acebench analysis showed about the models]]></description><link>https://blog.neurometric.ai/p/gemma-4b-vs-gemini-flash-you-dont</link><guid isPermaLink="false">https://blog.neurometric.ai/p/gemma-4b-vs-gemini-flash-you-dont</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Tue, 11 Aug 2026 21:19:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!twzy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Most of what makes an AI assistant useful isn&#8217;t prose &#8212; it&#8217;s calling tools: booking the meeting, adding the cart item, pulling the shipping estimate. When these products fail, it&#8217;s rarely bad writing. It&#8217;s the wrong function, a mangled date, or three calls when one would do.</p><p>AceBench measures exactly that: ~2,000 hand-annotated tasks drawn from ~4,500 synthetic APIs across eight domains (finance, health, travel, tech, entertainment, and more), published January 2025 and later accepted at EMNLP.</p><p>Three things make it worth a practitioner&#8217;s attention:</p><p><strong>Grading is mechanical and brutal.</strong> Right function, right call count, right arguments &#8212; or it&#8217;s wrong. No partial credit, no AI judge. One bad field voids the form. Harsher than real life, but reproducible.</p><p><strong>It comes in tiers.</strong> Normal (clear requests), Special (vague), Agent (multi-turn). We ran Normal, English only &#8212; this says nothing about ambiguity or agent loops.</p><p><strong>It&#8217;s designed to be taken apart.</strong> Tasks isolate specific failure modes: value types, near-duplicate functions, mid-conversation drops. A leaderboard number says which model wins overall. A benchmark that comes apart says whether a cheap model is good enough for <em>your</em> work &#8212; usually the real question.</p><h2>What we ran</h2><p>Five models, 772 tasks each, 3,860 attempts, English Normal split:</p><ul><li><p><code>gemini-3.6-flash</code>, <code>gemini-3.1-flash-lite</code> &#8212; Google&#8217;s hosted models</p></li><li><p><code>gemma-4-12B-it</code>, <code>gemma-4-E4B-it</code>, <code>gemma-4-E2B-it</code> &#8212; open-weight, self-served (big/small/very small)</p></li></ul><p>Same harness and tasks for everyone. Tool calls return a bare acknowledgement &#8212; no feedback, one shot per call.</p><p><strong>Key point:  the open model on hardware we control tied Google&#8217;s hosted model, using about an eighth of the words.</strong></p><h2>The scoreboard</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!twzy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!twzy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 424w, https://substackcdn.com/image/fetch/$s_!twzy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 848w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1272w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!twzy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png" width="1292" height="460" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:460,&quot;width&quot;:1292,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:66008,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!twzy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 424w, https://substackcdn.com/image/fetch/$s_!twzy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 848w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1272w, https://substackcdn.com/image/fetch/$s_!twzy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F981ba10e-920a-4875-a4f5-c0f564ec8438_1292x460.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The top two are a real tie: 45 tasks the hosted model got and the open one missed, 35 the other way. That&#8217;s the signature of two equally capable models, not a better and a worse one. Run it again and the order could flip. Every other gap in the table is real &#8212; but the tie at the top is between a hosted frontier model and a 12B open model on a single GPU.</p><h2>What it costs</h2><p>Scoring the same isn&#8217;t interesting. Scoring the same <em>this cheaply</em> is.</p><p>ModelWords per correct answer*:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8OpP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8OpP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 424w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 848w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1272w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8OpP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp" width="1306" height="476" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:476,&quot;width&quot;:1306,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:16632,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8OpP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 424w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 848w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1272w, https://substackcdn.com/image/fetch/$s_!8OpP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff885b587-2bb7-48f9-a4c5-a215db2bae14_1306x476.webp 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>*tokens (~&#190; word each), total output</p><p>The hosted model generates ~8x more text per correct answer than the 12B model it&#8217;s tied with &#8212; almost all invisible &#8220;thinking&#8221; (655,000 tokens across the run). The other four models did none of that: read, call, stop.</p><p>And the thinking isn&#8217;t buying anything. When the hosted model got a task wrong, it generated <em>more than twice</em> as much text as when it got one right. Extra effort here is a sign of being stuck, not a way out. On a benchmark where the winning move is two steps, there&#8217;s not much to think about.</p><p>Generated tokens are the expensive half of any pricing page, and what determines latency. So: same accuracy, a fraction of the tokens, and a deployment you own instead of rent.</p><h2>The catch</h2><p>Cheap on tokens &#8800; cheap in practice:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lNL9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lNL9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 424w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 848w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1272w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lNL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png" width="1334" height="562" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:562,&quot;width&quot;:1334,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:73710,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lNL9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 424w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 848w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1272w, https://substackcdn.com/image/fetch/$s_!lNL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fdfc63d-69ec-454a-bcc9-55ad0f8057d4_1334x562.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The 12B model that ties Google&#8217;s best was 4x slower per answer than the API it ties, despite generating far less text. The smaller E4B, on the same GPU, hit 5.6s at ~91% of the top score. Note: these numbers are default vLLM, full BF16, on a single L40S &#8212; plenty of room to push both further with better hardware.</p><h2>The smallest model is better than its score</h2><p><code>gemma-4-E2B-it</code> came last at 67%. Look at <em>how</em> it failed and half the gap disappears.</p><p>Its top mistake wasn&#8217;t picking the wrong tool &#8212; it was making the right call plus extra, already-completed ones from earlier in the conversation:</p><blockquote><p><strong>Called:</strong> ask about gift etiquette &#8594; send the gift &#8594; schedule the meeting <strong>Wanted:</strong> schedule the meeting</p></blockquote><p>It knew the current turn. It just also re-did two things already done. Graded: zero.</p><p>72 of its 255 failures end with exactly the right call, buried under repeated history. Score only the current turn and it jumps from 67% to 76% &#8212; close to Google&#8217;s smaller model. That&#8217;s a prompt fix, not a bigger model &#8212; and it barely moves the other four (0-3 failures each of this kind).</p><h2>Where bigger models still earn their keep</h2><p>Three places:</p><p><strong>Nested arguments.</strong> A tool wanting <code>{"journey": {"from": "Shanghai", "to": "Hangzhou", "times": {...}}}</code> trips up everyone &#8212; best models barely clear 60%, versus 85-96% for flatter arguments.</p><p><strong>Several calls at once.</strong> Same tool, three times, different details: the hosted model pulls ahead ~5 points. Coordinating calls is harder than making one.</p><p><strong>Long conversations</strong> &#8212; the sharpest split. Four of five models <em>improve</em> turn over turn as context narrows the options. The smallest model collapses:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4T2I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4T2I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 424w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 848w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1272w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4T2I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png" width="1290" height="594" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:594,&quot;width&quot;:1290,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77120,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4T2I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 424w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 848w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1272w, https://substackcdn.com/image/fetch/$s_!4T2I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc05c5ce7-068d-4bb2-8baf-2ec5d079c18f_1290x594.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>That last column is what matters in production: a conversation only works if every turn lands. 87% per-turn becomes 74% overall.</p><h2>Two models beat one</h2><p>Since every model saw the same tasks, we can check what a pair covers:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QlGi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QlGi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 424w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 848w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1272w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QlGi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png" width="1300" height="376" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:376,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45669,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/210816210?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QlGi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 424w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 848w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1272w, https://substackcdn.com/image/fetch/$s_!QlGi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ddeab2-af96-421f-ab2e-06fad9536ee5_1300x376.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Pairing the hosted model with the open one adds 4.5 points; pairing it with Google&#8217;s other model adds just 2.1. The two hosted models fail on the same tasks &#8212; the open one fails on different ones. If you can check an answer and retry, the self-served model is the better (and cheaper) second opinion.</p><h2>So what do you do with this</h2><p>If the task is &#8220;here&#8217;s what I want, here are the tools, go&#8221; &#8212; use a small self-hosted model. There&#8217;s no ambiguity for extra reasoning to resolve. A 12B model matches the frontier here; even 4B gets within ~90%.</p><p>If the work needs complex structures, coordinated multi-calls, or long conversations, the gap reopens &#8212; fastest for the smallest models.</p><p>Check what a scoreboard actually measures before trusting it: one category here was scoring a missing input, not the model, so we dropped it. Another score was understated nine points over a technicality about which calls &#8220;count&#8221; for a turn. Neither shows up in a leaderboard number &#8212; and both change what you&#8217;d buy.</p>]]></content:encoded></item><item><title><![CDATA[New Podcast Episode - TMax: Closing the Frontier Gap With Open Data]]></title><description><![CDATA[A conversation with Yash Sharma, Director of AI Research at Neurometric AI]]></description><link>https://blog.neurometric.ai/p/new-podcast-episode-tmax-closing</link><guid isPermaLink="false">https://blog.neurometric.ai/p/new-podcast-episode-tmax-closing</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Fri, 07 Aug 2026 18:24:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This week, Cooper and Yash (Neurometric AI) unpack TMax, a paper out of AI2 (the Allen Institute for AI) that takes a genuinely different approach to closing the gap between small open-weight models and frontier labs &#8212; not by training a bigger model, but by open-sourcing the entire recipe: data, code, and technique.</p><p>The paper focuses on terminal agents &#8212; models given nothing but command-line access as a tool, which turns out to be enough to tackle an enormous range of tasks, from software engineering to security work to scheduling. Using this recipe, AI2 boosted a Qwen model&#8217;s Terminal-Bench score from roughly 25% to 31%, closing meaningful ground on Claude Haiku&#8217;s roughly 33%.</p><p>What made this one worth an episode wasn&#8217;t just the score bump. It was the design choices behind it.</p><p>Watch on YouTube: <a href="https://youtu.be/4U0gqv9sSn4?si=ejv04gktGbTOMJM8">https://youtu.be/4U0gqv9sSn4?si=ejv04gktGbTOMJM8</a></p><p>Listen:<a href="https://tokenengineering.podbean.com/"> https://tokenengineering.podbean.com</a></p><p>Paper: TMax: A recipe for terminal agents &#8212;<a href="https://arxiv.org/abs/2606.23321"> https://arxiv.org/abs/2606.23321</a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Glass Is Half... Correct? Half Our SLM Benchmark 'Failures' Contained The Right Answer]]></title><description><![CDATA[An analysis of small model failures on CRM arena]]></description><link>https://blog.neurometric.ai/p/the-glass-is-half-correct-half-our</link><guid isPermaLink="false">https://blog.neurometric.ai/p/the-glass-is-half-correct-half-our</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 05 Aug 2026 18:55:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!DhE3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What we ran</h2><p>Three models, two tool interfaces, 340 CRM tasks each &#8212; 2,040 agent rollouts in total.</p><p>The benchmark is CRMArena, run through the Harbor/dockworker harness. Every task asks a question about a read-only Salesforce org (&#8221;in May 2021, which state had the quickest case closures?&#8221;, &#8220;which knowledge article does this quote violate?&#8221;) and grades the submitted answer by exact match. Seventeen task categories, split evenly across a B2B and a B2C org.</p><p>The interesting design choice is that the same 340 questions are posed twice, behind two different tool surfaces:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DhE3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DhE3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 424w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 848w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1272w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DhE3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png" width="1324" height="434" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:434,&quot;width&quot;:1324,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:78759,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DhE3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 424w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 848w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1272w, https://substackcdn.com/image/fetch/$s_!DhE3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4607764b-86ee-4c66-8a06-85ed20059266_1324x434.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Identical questions, identical ground truth &#8212; we verified this holds on every task where both suites produced a graded answer. So the two suites isolate one variable: does the model do better writing SQL, or picking the right pre-built tool and filling in its arguments?</p><p>The three models: <strong>gemini-3.6-flash</strong> and <strong>gemini-3.1-flash-lite</strong> (hosted), and <strong>gemma-4-E4B-it</strong> (open weights, served locally on vLLM). One is a small local model; two are hosted. Every trial used the same <code>noshell</code> agent scaffold, whose only way to answer is to call a terminal <code>submit_answer</code> tool.</p><h2>The scoreboard</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!37I2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!37I2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 424w, https://substackcdn.com/image/fetch/$s_!37I2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 848w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1272w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!37I2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png" width="1312" height="348" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:348,&quot;width&quot;:1312,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53643,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!37I2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 424w, https://substackcdn.com/image/fetch/$s_!37I2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 848w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1272w, https://substackcdn.com/image/fetch/$s_!37I2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a6f7edb-d2a6-4243-ad59-d46641a431e1_1312x348.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Read it and you&#8217;d conclude the ordering is obvious and the small model isn&#8217;t close: 39.7% against 67.6% is a 28-point gap. Every pairwise difference here is statistically significant (p &#8804; 0.009).</p><p>Then you look at how the failures happen, and the picture inverts.</p><h2>Half of the small model&#8217;s &#8220;failures&#8221; contain the right answer</h2><p>Not every zero is a wrong answer. A trial scores zero if the answer file was never written at all &#8212; and the scaffold only writes it when the model calls <code>submit_answer</code>. Classifying all 2,040 trials by how they failed:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AbqI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AbqI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 424w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 848w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1272w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AbqI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png" width="1348" height="778" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:778,&quot;width&quot;:1348,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:98995,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AbqI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 424w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 848w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1272w, https://substackcdn.com/image/fetch/$s_!AbqI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F895a63f0-cd78-40b2-9eb1-07263bef7084_1348x778.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>More than half of gemma&#8217;s SQL trials never submitted anything. And they didn&#8217;t crash or time out &#8212; in 171 of them the model finished its work, wrote the answer out in prose, and simply never called the tool. Something like:</p><blockquote><p>The <code>case_metrics</code> call returned an average closure time of 4.2 days for CA, the lowest of any state. The answer is CA.</p></blockquote><p>Graded: zero.</p><p>So we tested the obvious question &#8212; were those prose answers right? Ground truth isn&#8217;t recorded for ungraded trials, but because the same 340 tasks appear in both suites, we could recover the expected answer for every one of them and grep the final message for it.</p><p>Of gemma&#8217;s 250 prose non-submissions, <strong>44&#8211;52% contained the correct answer</strong>. The range is the strict and loose reading of the same check: the loose count is any trial whose prose contains the expected value; the strict count additionally requires that the model wasn&#8217;t hedging across a list of candidates (&#8804;1 other record ID mentioned). Both bound the same conclusion.</p><p>That reframes the scoreboard as a lower bound. Fixing one scaffold behaviour &#8212; get the model to call the tool &#8212; moves gemma to:</p><div class="captioned-image-container"><figure><div class="image-link image2" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!w90h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 424w, https://substackcdn.com/image/fetch/$s_!w90h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 848w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1272w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!w90h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png" width="1326" height="282" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:282,&quot;width&quot;:1326,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!w90h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 424w, https://substackcdn.com/image/fetch/$s_!w90h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 848w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1272w, https://substackcdn.com/image/fetch/$s_!w90h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff847b9d2-7dc1-4a77-9b7b-cd7e30aa784f_1326x282.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></div></figure></div><p>A model that looked like it scored 26% on SQL was doing work worth about 50%. Its measured number was roughly half its actual competence, and every point of that gap is a formatting bug.</p><p>The same correction barely moves the hosted models &#8212; gemini-3.6-flash has exactly zero prose non-submissions, and gemini-3.1-flash-lite has one. They always call the tool. What we were measuring, for a third of the benchmark, was instruction-following on the harness contract, not CRM reasoning.</p><p>There&#8217;s a cleaner way to see it. Restrict to trials that submitted anything, and the reasoning quality behind the scoreboard separates from the plumbing:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RiZl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RiZl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 424w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 848w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1272w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RiZl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png" width="1330" height="346" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4436c62-6974-44e8-a378-003463c52213_1330x346.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:346,&quot;width&quot;:1330,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47261,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RiZl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 424w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 848w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1272w, https://substackcdn.com/image/fetch/$s_!RiZl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4436c62-6974-44e8-a378-003463c52213_1330x346.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On answers it actually submits, the small local model is within two points of the hosted flash-lite model &#8212; a difference well inside the noise at this sample size. The 11-point headline gap between them is almost entirely tool-calling discipline.</p><h2>What it costs</h2><p>This is where the small model stops being a curiosity. Token totals are the whole run; the per-win column divides by correct answers, which is the number that matters if you&#8217;re paying for throughput.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nOs-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nOs-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 424w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 848w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1272w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nOs-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png" width="1322" height="850" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33631410-c178-489a-a32c-2baa2db91b67_1322x850.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:850,&quot;width&quot;:1322,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:122754,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nOs-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 424w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 848w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1272w, https://substackcdn.com/image/fetch/$s_!nOs-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33631410-c178-489a-a32c-2baa2db91b67_1322x850.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Gemma on the tool API buys a correct answer for 45.4k tokens. gemini-3.6-flash needs 196k for the same thing &#8212; <strong>4.3&#215; more</strong>. On SQL it needs 598k, or <strong>13&#215; gemma&#8217;s best configuration</strong>.</p><p>The driver is visible in the reasoning-token column. gemini-3.6-flash spent 1.36M reasoning tokens on the API suite and 2.59M on SQL. The other two models spent none. That&#8217;s what the extra accuracy is bought with: 3.6-flash&#8217;s median trial emits 3,978 output tokens on the API suite against gemma&#8217;s 501 &#8212; an 8&#215; difference in generated text per attempt.</p><p>And time to completion &#8212; median wall time of the agent execution phase alone, excluding container build and verification:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LtyB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LtyB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 424w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 848w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1272w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LtyB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png" width="1350" height="616" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:616,&quot;width&quot;:1350,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:83023,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LtyB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 424w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 848w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1272w, https://substackcdn.com/image/fetch/$s_!LtyB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F876d5bc3-d17d-4580-b96d-1bab9c004e3e_1350x616.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Here the small model does not win, and it&#8217;s worth being precise about why. gemma is 12.6s per attempt against 3.6-flash&#8217;s 30.0s &#8212; faster per attempt &#8212; but because it converts fewer attempts into correct answers, it lands at 73.8s per win versus 51.8s. And this column mixes model speed with serving infrastructure: a local vLLM instance against Google&#8217;s production endpoints. It is not a property of the models, and it&#8217;s the one metric here we&#8217;d throw out of a purchasing decision. Tokens are the honest efficiency measure; wall-clock is an artifact of where each model happened to be running.</p><p>One more pattern worth noting: every model spends more tokens on the trials it gets wrong &#8212; 1.5&#215; to 2.1&#215; its passing median. Failure is not cheap. Effort is a symptom of being lost, not a route out of it, which means the token cost of a wrong answer exceeds the token cost of a right one across the board.</p><h2>Where each model actually fails</h2><p>The three models have almost nothing in common in their failure profiles.</p><p><strong>gemma-4-E4B-it</strong> &#8212; a formatting problem wearing a capability problem&#8217;s clothes. 250 of its failures are prose-instead-of-tool-call; roughly half contain the right answer. Its second tendency is over-caution: 50 trials answered <code>None</code> (&#8221;nothing matches&#8221;) when a real record existed. It abstains too readily and it won&#8217;t call the terminal tool. Both are addressable without touching the model.</p><p><strong>gemini-3.1-flash-lite</strong> &#8212; quietly broken generations. 34 API trials (10%) ended with an empty completion: a single EOS token, no text and no tool call. Not prose, not a wrong answer &#8212; nothing at all. That signature points at serving or sampling rather than the prompt, and it&#8217;s the cheapest 10% anyone in this comparison could recover. Beyond that its losses are ordinary wrong answers (21&#8211;23%), the highest wrong-answer rate of the three.</p><p><strong>gemini-3.6-flash</strong> &#8212; runs out of budget, and won&#8217;t say &#8220;none.&#8221; It never fails to submit when it finishes, but it frequently doesn&#8217;t finish: 60 SQL trials (18%) and 17 API trials (5%) hit the 25-step cap mid-work. Its failing trials burn a median 9,747 output tokens against 6,452 when passing &#8212; it iterates on queries that never land. Its other weakness is the mirror image of gemma&#8217;s: 63 trials invented a value where <code>None</code> was correct. The strongest model is the one most likely to manufacture an answer rather than concede there isn&#8217;t one.</p><p>That last contrast is the most useful thing here for anyone building on these models. The two abstention errors point in opposite directions &#8212; gemma says &#8220;none&#8221; when an answer exists, 3.6-flash asserts an answer when none does &#8212; so there is no single prompt that fixes both. It&#8217;s a calibration problem per model, not a benchmark-wide one.</p><h2>The interface matters, and not the way you&#8217;d guess</h2><p>Because the same questions appear behind both tool surfaces, we can ask whether SQL or semantic tools suit each model better as a paired comparison &#8212; the same task, two interfaces &#8212; using an exact McNemar test on the tasks where the two disagree.</p><p>For gemini-3.6-flash, the only model whose two runs had matched step budgets, the tool API wins clearly: +8.5 points, p &lt; 0.001, with 47 tasks solved only through the API against 18 solved only through SQL.</p><p>But the aggregate hides something better. Across the models, between 65 and 133 of the 340 tasks flip outcome between the two interfaces. For gemini-3.1-flash-lite the headline rates are nearly identical (49.1% vs 50.9%) while 94 tasks flip &#8212; 44 solved only via SQL, 50 only via the API, with just 123 solved by both. <strong>The interface doesn&#8217;t change how many tasks it gets right; it changes which third of the benchmark it gets right.</strong> A model that looks interface-indifferent on the scoreboard is nothing of the kind, and running both interfaces and taking either success would score far above either alone.</p><p>Category detail shows where this bites. gemini-3.6-flash gets <code>sales-amount-understanding</code> right 60% of the time through semantic tools and 5% through SQL &#8212; but that collapse is not an inability to write the query. 18 of its 20 SQL attempts hit the step cap mid-work. Where a semantic tool answers &#8220;total order amount by owner in this window&#8221; in one call, SQL requires discovering the schema, joining line items to orders, and windowing on the right date column &#8212; and it ran out of budget getting there. <code>lead-routing</code>, by contrast, it solves 100% either way.</p><h2>Not all 17 categories are the same benchmark</h2><p>Treating CRMArena as one number hides the most actionable result in the run. Here is every category, worst-first, across all six runs (<code>36</code> = gemini-3.6-flash, <code>31</code> = gemini-3.1-flash-lite, <code>gm</code> = gemma-4-E4B-it):</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sMca!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sMca!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 424w, https://substackcdn.com/image/fetch/$s_!sMca!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 848w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1272w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sMca!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png" width="1334" height="1474" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1474,&quot;width&quot;:1334,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:207655,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sMca!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 424w, https://substackcdn.com/image/fetch/$s_!sMca!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 848w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1272w, https://substackcdn.com/image/fetch/$s_!sMca!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F06331dfa-e101-4579-baba-4654896bc4db_1334x1474.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7rD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7rD8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 424w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 848w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7rD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png" width="1350" height="1038" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1038,&quot;width&quot;:1350,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:142926,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7rD8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 424w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 848w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1272w, https://substackcdn.com/image/fetch/$s_!7rD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ebf1ddd-0383-4b34-829f-b021da887f93_1350x1038.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Group the 17 categories by what they actually ask for, and a pattern appears that the headline rates completely obscure:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!llyz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!llyz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 424w, https://substackcdn.com/image/fetch/$s_!llyz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 848w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1272w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!llyz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png" width="1330" height="944" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:944,&quot;width&quot;:1330,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:126035,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/209966752?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!llyz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 424w, https://substackcdn.com/image/fetch/$s_!llyz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 848w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1272w, https://substackcdn.com/image/fetch/$s_!llyz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea2012b7-237e-4c2b-b5aa-8cb61b39951a_1330x944.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On policy judgement, model scale buys almost nothing. Reading a chat transcript and deciding whether an agent breached policy, or whether a lead is qualified: gemini-3.6-flash manages 37.5%, gemma manages 32.5%. The small model is at 87% of the best model&#8217;s performance. Compare that to aggregate metrics, where the same small model reaches only 57% of the frontier score and the spread across runs hits 85 points on <code>monthly-trend-analysis</code>.</p><p>The reason is visible in the failure modes. On <code>policy-violation-identification</code>, half of gemini-3.6-flash&#8217;s API failures are missed_abstention &#8212; 8 trials asserting a violation where the correct answer was &#8220;none&#8221; &#8212; against 8 passes. gemma&#8217;s failures on the same category split 4 missed_abstention and 4 false_abstention. This category isn&#8217;t primarily testing reasoning; it&#8217;s testing whether a model will decline to answer. That&#8217;s a calibration property, and calibration doesn&#8217;t scale with capability the way multi-step numeric work does.</p><p>So the practical read is: the categories where a small model is nearly as good are the judgement ones, and the categories where it falls off a cliff are the multi-step aggregations. If your workload is &#8220;read this conversation and classify it,&#8221; a 4B local model is a serious candidate. If it&#8217;s &#8220;compute this metric across three joins and a date window,&#8221; it is not.</p><p>Two more things the table says:</p><p><strong>invalid-config is the one genuine capability cliff.</strong> gemini-3.6-flash gets 50% on both interfaces; every other run scores 5&#8211;10%. A 45-point gap that survives both tool surfaces is the clearest evidence in this run of something the smaller models simply cannot do &#8212; matching a quote&#8217;s configuration against the knowledge article it violates. Notably it&#8217;s not a formatting artifact: gemma&#8217;s failures here are 8 wrong answers alongside 7 prose non-submissions, so even the recovered ceiling stays low.</p><p><strong>quote-approval defeats everything.</strong> 8% pooled, topping out at 20%, and gemini-3.6-flash scores 5% and 0%. No model, no interface, no step budget helps. When the strongest model in a comparison scores 5% on a category, that&#8217;s usually a signal about the task or the grader rather than the models &#8212; and it&#8217;s the first thing we&#8217;d re-examine before drawing conclusions about the remaining headroom.</p><p>And where does the small model actually beat a hosted model head-to-head? On the tool API, gemma matches or beats gemini-3.1-flash-lite in 5 of 17 categories: <code>policy-violation-identification</code> (35% vs 15%, +20 points), <code>wrong-stage-rectification</code> (40% vs 35%), <code>invalid-config</code> (10% vs 5%), and ties on <code>case-routing</code> and <code>lead-routing</code>. Four of those five are judgement or routing tasks, not aggregations &#8212; the same pattern again.</p><h2>What we&#8217;d take away</h2><p>The headline number on an agent benchmark is a joint measurement of the model and the scaffold around it, and for small models the scaffold dominates. gemma-4-E4B-it looked like a 26% model on SQL. It was doing work worth about 50%, and the missing half was one unmade tool call. Anyone comparing a small open model against a hosted frontier model on a leaderboard number, without looking at the failure modes underneath it, is partly measuring which model was better at following the harness&#8217;s calling convention.</p><p>Small models win on the axis that gets left off the leaderboard. At 45.4k tokens per correct answer against 196k, gemma is 4.3&#215; cheaper per unit of useful output than the model that beats it by 28 points &#8212; and it gets there with zero reasoning tokens against 1.36M. If your workload tolerates 50% task accuracy, or you can put a verifier behind it and retry, the small model is the better economic choice by a wide margin.</p><p>But the honest version of &#8220;small models can win&#8221; is narrower than the slogan. gemma did not beat gemini-3.6-flash in a single one of the 17 categories, on either interface. What it did was reach parity with a hosted flash-lite model on reasoning quality &#8212; 56.0% against 57.9% on submitted answers &#8212; at a fraction of the token cost, while running locally. That&#8217;s the claim the data supports: not that small models beat frontier models, but that the gap to the tier below frontier is mostly plumbing, and the plumbing is cheap to fix.</p><p>The category breakdown sharpens that into a selection rule. The gap is not uniform across work: on policy judgement the small model is at 87% of the frontier model, on multi-step aggregation it is at 57%, and on one category (<code>invalid-config</code>) it is nowhere near. &#8220;Can a small model do this?&#8221; has no general answer, but it has a reliable one per task family &#8212; and the families where the answer is yes are the judgement-shaped ones, which is the opposite of where most people assume a small model will struggle.</p><p>The fixes are unglamorous and they&#8217;re all in the harness: make the terminal tool call unmissable for models that write prose, investigate the empty completions, raise the step cap for models that iterate, and calibrate abstention per model rather than globally. None of them require a bigger model.</p>]]></content:encoded></item><item><title><![CDATA[Neurometric Is Now On TrustedRouter]]></title><description><![CDATA[Hosted generic and task specific SLMs are available.]]></description><link>https://blog.neurometric.ai/p/neurometric-is-now-on-trustedrouter</link><guid isPermaLink="false">https://blog.neurometric.ai/p/neurometric-is-now-on-trustedrouter</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 29 Jul 2026 19:28:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Neurometric is now a <a href="https://trustedrouter.com/providers/neurometric">listed provider on TrustedRouter</a>. If you already route inference through TrustedRouter, our models are available to you today &#8212; no new account, no new SDK, no new base URL.</p><h2>What Is TrustedRouter?</h2><p>TrustedRouter is an inference gateway in the same family as OpenRouter: one OpenAI-compatible API, one base URL, and hundreds of models and routes behind it. You point your existing client at the gateway and swap model strings instead of rewriting integrations every time you want to try a different provider.</p><p>What makes it interesting is the trust posture. TrustedRouter keeps no record of what you route &#8212; no prompt or output logs, by design, and the gateway is attested and fails closed rather than silently degrading. It publishes a live status page and a continuously sampled performance leaderboard, so latency and uptime claims are measured rather than asserted. There&#8217;s also a documented migration path from OpenRouter for teams that want to move without a rewrite.</p><p>For anyone who has spent time explaining to a security team why prompts containing customer data are being logged by a third party, that combination is worth something.</p><h2>What We&#8217;re Serving Today</h2><p>We&#8217;re live with three small models on prepaid routes:</p><ul><li><p><code>ibm-granite/granite-4.1-8b</code> &#8212; 32K context, $0.0525/1M prompt and $0.105/1M completion</p></li><li><p><code>qwen/qwen3-vl-8b-instruct</code> &#8212; 262K context, vision-capable</p></li><li><p><code>qwen/qwen3-vl-8b-thinking</code> &#8212; 32K context, reasoning traces</p></li></ul><p>Early measured numbers on our routes: 863 ms p50 time-to-first-token on Granite and 100% uptime across the sampling window. Our listed policy note is accurate about what we do and don&#8217;t guarantee &#8212; upstream trace logging is disabled, prompts and completions are not written to observability or object storage, and we retain only aggregate request counts and token totals. That&#8217;s a no-store posture, not contractual ZDR, and we&#8217;d rather say so plainly than let people assume more than we&#8217;ve committed to.</p><h2>Task-Specific Models Are Coming Next</h2><p>Small general-purpose models are the entry point, not the thesis. The reason we build the way we do is that most production workloads aren&#8217;t &#8220;chat&#8221; &#8212; they&#8217;re a narrow, repetitive task where a tuned 8B model matches or beats a frontier model at a fraction of the cost.</p><p>Three task-specific models from our model marketplace we expect to list on TrustedRouter in the coming weeks:</p><ol><li><p><strong>Text-to-SQL</strong> &#8212; natural language to correct, schema-aware queries against a known database, tuned for join accuracy rather than conversational fluency.</p></li><li><p><strong>Structured extraction</strong> &#8212; pulling typed JSON out of invoices, contracts, and claims documents against a supplied schema, with predictable failure behavior on missing fields.</p></li><li><p><strong>Support intent classification</strong> &#8212; routing inbound tickets to the right queue and priority, the kind of high-volume, low-token task where per-call cost dominates everything else.</p></li></ol><p>Each is graded against our proprietary task benchmark corpus, so you can see how a specialist performs on your task class before you route production traffic to it.</p><p>If you want to try the current models, grab a key at TrustedRouter and use <code>neurometric</code> as the provider. If you have a task class you&#8217;d like us to build for, tell us &#8212; that&#8217;s how the roadmap gets set.</p>]]></content:encoded></item><item><title><![CDATA[Navigating the Enterprise AI Labyrinth: Build, Buy, Route, or Wait]]></title><description><![CDATA[How to operate when the world is changing so fast.]]></description><link>https://blog.neurometric.ai/p/navigating-the-enterprise-ai-labyrinth</link><guid isPermaLink="false">https://blog.neurometric.ai/p/navigating-the-enterprise-ai-labyrinth</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 29 Jul 2026 00:20:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>At Neurometric, we spend our days in the trenches with enterprise technology leaders helping them save inference costs and optimize model usage. In most companies the pressure to &#8220;do something with AI&#8221; is intense, but the rush to production often leads to bloated budgets, architectural dead ends, and a staggering accumulation of technical debt. When the dust settles, the organizations that succeed are not the ones that deploy the most models; they are the ones that apply a rigorous, intellectually honest framework to how they adopt them.</p><p>To cut through the hype, organizations must fundamentally rethink their deployment strategies. In our experience, the most critical decision a technical leader makes isn&#8217;t which specific foundational model to choose, but rather the strategic posture they take toward the underlying business task. We have observed that enterprise AI strategy ultimately boils down to a fundamental set of decisions. There are usually four options on the table, and they require a level of candor that is often missing from vendor pitches and internal strategy meetings.</p><h3>The Four Options </h3><p>When evaluating a new AI capability or use case, technology leaders must choose between four distinct paths: Buy, Build, Route, or Wait.</p><p><strong>Buy</strong> is the right answer much more often than technical leaders want it to be. Engineers are naturally wired to create, and the allure of constructing a bespoke AI system is incredibly strong. However, if a vendor has already solved a generic enterprise problem&#8212;like drafting marketing copy, summarizing meeting notes, or triaging customer support tickets&#8212;purchasing that solution is almost always the superior economic choice.</p><p><strong>Build</strong> is the right answer only <strong>when the task at hand </strong><em><strong>is</strong></em><strong> the core business</strong>. If the AI system is going to directly drive your competitive advantage in the marketplace, you cannot outsource it.</p><p><strong>Route</strong> is the correct posture when the task is high-volume and the models are fungible. If you are processing millions of identical queries where the subtle nuances of a massive frontier model are unnecessary, routing queries to the most efficient model available is the only way to scale without destroying your margins.</p><p>Finally, there is <strong>Wait</strong>. Wait is the honest default. It is the right decision more often than anyone wants to admit, and yet it is almost never proposed in a strategy meeting. There is immense career risk in telling a CEO to wait on AI. But for highly volatile use cases, or problems where the underlying foundational models are currently struggling but rapidly improving, waiting six months for the ecosystem to mature is often vastly superior to burning capital on a brittle, premature V1.</p><h3>The &#8220;Build Test&#8221;</h3><p>If you are leaning toward building a custom AI solution, you must subject your proposal to the &#8220;Build Test.&#8221; Building bespoke AI is an expensive, resource-intensive endeavor that requires long-term commitment. At Neurometric, we advise clients to build only if they can definitively answer &#8220;yes&#8221; to all three of the following criteria:</p><p>First, is the task core to your differentiation? The AI must do something that separates you from your competitors. If it is merely an operational efficiency that every other company in your sector will eventually adopt, it fails this test.</p><p>Second, do you possess proprietary data or a unique workflow that a vendor structurally cannot access? You need an unfair advantage. If you are building a model using the exact same public datasets and standard enterprise tools as the major SaaS vendors, they will eventually commoditize your creation. You must have a moat built on data or processes that are uniquely yours.</p><p>Third, can you staff the maintenance of this system for the next three years? Building the model is merely the starting line. Models drift, APIs change, underlying data distributions shift, and security vulnerabilities emerge. You are not just funding a build phase; you are funding a permanent product team.</p><p>If you meet two out of these three criteria, it is not a &#8220;maybe.&#8221; Two out of three means you default to Buy. The economics of maintaining a sub-scale, non-differentiated AI system will slowly drain your engineering resources.</p><h3>Model Selection: The Frontier Premium</h3><p>Once you have decided how to acquire the capability, you must choose the right engine. The market is currently bifurcated between massive &#8220;frontier&#8221; models and smaller, specialized, or distilled models.</p><p>Frontier models are the bleeding-edge giants of the industry. They possess incredible reasoning capabilities and vast world knowledge. You should reserve these models strictly for tasks requiring deep judgment, handling high ambiguity, and managing low-volume, high-stakes scenarios. If an AI is reviewing a complex legal contract for a multi-million dollar merger, you want the frontier model.</p><p>Conversely, small, specialized, or distilled models are the workhorses of the modern enterprise. These should be deployed for high-volume, well-specified, and latency-sensitive tasks. When you are processing tens of thousands of basic data extraction requests per hour, you do not need an AI that can write a sonnet or pass the bar exam.</p><p>The divergence in economics here is huge, often separated by one to two orders of magnitude. A distilled model can easily be 10x to 100x cheaper per token than a frontier model, while returning responses in a fraction of the time. Yet, we routinely see organizations using frontier models for absolutely everything simply because it requires only one API integration. They are paying a 30&#215; premium for the privilege of architectural laziness.</p><h3>Routing as an Operating Discipline</h3><p>To capture the economic benefits of smaller models without sacrificing quality, organizations must adopt routing as a core operating discipline.</p><p>The biggest mistake teams make is routing at the application level&#8212;deciding that an entire application will use just one model. Instead, you must route per task. A single customer service application might use a cheap, fast model to identify the language of an incoming ticket, a specialized model to extract the customer&#8217;s account number, and only invoke a frontier model if the ticket requires a complex, nuanced apology for a service failure.</p><p>Implementing this requires establishing a strict quality floor for every specific task. You determine the minimum acceptable accuracy, and then you dynamically route the workload to the absolute cheapest model that clears that floor. Because the open-source and proprietary model landscapes are evolving at a breakneck pace, this is not a set-it-and-forget-it architecture. You must re-benchmark your routing logic monthly. The model that was the most cost-effective in January might be entirely obsolete by April.</p><h3>Vendor Risk and AI Sovereignty </h3><p>Finally, organizations must wake up to the reality of vendor risk in the AI space. It is real, it is severe, and it is currently vastly underweighted in enterprise risk assessments.</p><p>When you build a product highly dependent on a third-party API, you are at the mercy of their roadmap. Deprecation schedules are often brutally short, forcing expensive emergency migrations when a vendor decides to sunset a specific model version. Price changes can occur overnight, destroying the unit economics of your application. Rate limits can be abruptly tightened, throttling your application&#8217;s ability to scale during a surge in user demand.</p><p>Furthermore, data privacy terms are a constantly shifting target. An update to a vendor&#8217;s terms of service could suddenly allow them to use your enterprise data to train their next generation of models, violating your internal compliance policies.</p><p>Most insidiously, a model can change its behavior right underneath you without any version bump. Vendors continuously tweak, align, and &#8220;improve&#8221; their models behind the scenes. A prompt that reliably output perfect JSON formatting on Tuesday might suddenly start wrapping its output in conversational pleasantries on Thursday, breaking your entire data pipeline.</p><p>To mitigate this, technology leaders must push back during procurement. Do not accept standard click-wrap agreements for critical AI infrastructure. You must explicitly contract for extended notice periods regarding deprecations, strict rate limit guarantees, immutable data terms, and rigid version control to ensure the model you test is the exact model you run in production.</p><p>The AI ecosystem is moving incredibly fast, but the fundamental laws of enterprise software engineering and business economics have not been suspended. By rigorously applying the Build Test, rightsizing your model selection, establishing a disciplined routing architecture, and aggressively managing vendor risk, you can navigate this landscape successfully. The goal isn&#8217;t to deploy AI the fastest; the goal is to deploy it in a way that actually works for your business.</p><p>If you want to shield yourself from the chaos of the model selection and routing ecosystem, Neurometric&#8217;s core platform evaluates, chooses, and routes models automatically.  <a href="https://studio.neurometric.ai/">Try it</a> for free.</p>]]></content:encoded></item><item><title><![CDATA[What Does A Token Engineering Platform Do?]]></title><description><![CDATA[The new tool for the most important job in AI]]></description><link>https://blog.neurometric.ai/p/what-does-a-token-engineering-platform</link><guid isPermaLink="false">https://blog.neurometric.ai/p/what-does-a-token-engineering-platform</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Tue, 14 Jul 2026 09:36:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>If your company runs AI in production, you&#8217;re buying intelligence by the token. And if you&#8217;re like most enterprises, you have almost no tooling to manage that spend the way you manage your cloud spend.</p><p>This should feel familiar. When companies moved to the cloud, an entire layer of cost management and optimization tooling emerged around it &#8212; because once compute became a metered utility, optimizing that meter became a discipline with real dollars attached. AI inference is following the same path, except the stakes are arriving faster. A production AI workload at enterprise scale &#8212; say, a million model calls a day &#8212; can swing by millions of dollars a year based on decisions most teams made once, early, and never revisited.</p><p>Token engineering is the discipline that fixes this. A token engineering platform is the system that operationalizes it.</p><h2>What Is Token Engineering?</h2><p>Token engineering is the practice of treating tokens as an engineered resource: measured, benchmarked, routed, and continuously optimized. It&#8217;s a systems discipline, not a procurement exercise.</p><p>The common misconception is that token engineering means &#8220;use a cheaper model.&#8221; It doesn&#8217;t. It means optimizing every AI workload across three dimensions simultaneously: <strong>cost, speed, and reliability</strong>. Sometimes the right answer is a smaller, cheaper model. Sometimes it&#8217;s a faster one. Sometimes it&#8217;s the frontier model, but with a compressed prompt and an aggressive caching layer in front of it. The point is that the answer is different for every task, and it changes constantly.</p><p>Three forces make this urgent right now. First, model proliferation: frontier LLMs, open-weight models, and small language models (SLMs) now number in the hundreds, with meaningful new releases every month. Second, price variance: the cost of completing the same task can vary by 100x or more depending on which model, technique, and hardware you choose. Third, the capability crossover: for a growing share of enterprise tasks, purpose-built SLMs now match or beat frontier models at a fraction of the cost.</p><p>Meanwhile, most teams pick a model at the start of a project, hardcode it, and move on. Every month that decision goes unexamined, the gap between what they pay and what they should pay gets wider.</p><h2>The Core Components of a Token Engineering Platform</h2><p>Managing this problem manually doesn&#8217;t scale. A token engineering platform automates it. Here&#8217;s what a complete platform looks like &#8212; and how the Neurometric platform implements each piece.</p><h3>1. Workload Evaluation and Benchmarking</h3><p>You cannot optimize what you cannot measure, and public leaderboards won&#8217;t save you. Generic benchmarks tell you how models perform on academic tasks, not on <em>your</em> workloads &#8212; your customer service intents, your document extraction formats, your code review standards.</p><p>A token engineering platform continuously evaluates models and token engineering techniques against your actual tasks. When a new model ships, you shouldn&#8217;t have to wonder whether it&#8217;s better for your use case; the platform should tell you, with graded evidence. Neurometric&#8217;s Harbor evaluation infrastructure does exactly this, with more than 15,000 graded tasks and thousands of benchmark runs powering every recommendation the platform makes.</p><h3>2. Task-Level Routing</h3><p>Evaluation tells you which model is best for which task. Routing acts on it &#8212; automatically, per request, in production.</p><p>This is the decision engine at the heart of the platform. Most applications don&#8217;t have one workload; they have dozens of distinct tasks hiding inside a single product, each with different cost, latency, and quality requirements. Routing at the application level means paying frontier prices for tasks a model one-tenth the cost handles perfectly. Routing at the task level means every request goes to the cheapest model that meets its quality bar. Our Task Level Router keeps this workload-to-model mapping current as models, prices, and your traffic all change &#8212; so the routing decision you&#8217;d make today doesn&#8217;t quietly decay into the wrong decision six months from now.</p><h3>3. Automated SLM Creation and Fine-Tuning</h3><p>Sometimes no existing model sits at the right point on the cost/quality frontier for your task. The frontier model is overkill and overpriced; the small models miss your quality bar. Historically, the answer was a fine-tuning project: hire ML engineers, build a data pipeline, spend a quarter.</p><p>A token engineering platform makes this a feature, not a project. When evaluation data shows that a task is a candidate for a purpose-built model, the platform can distill and fine-tune an SLM automatically, validate it against your benchmarks, and slot it into the routing layer. The economics are hard to ignore: a task-specific SLM can run at 1/50th the cost of a frontier model while matching its accuracy on that narrow task. Our Auto-SLM Creator turns what used to be a specialized data science effort into a platform capability.</p><h3>4. Hardware and Deployment Optimization</h3><p>Which model you run is only half the equation. Where you run it matters just as much.</p><p>The same open-weight model has wildly different economics depending on whether you consume it through an API, self-host it on dedicated GPUs, run it quantized on cheaper hardware, or batch it for throughput over latency. At scale, the API-versus-self-hosting decision alone can be worth seven figures annually &#8212; and the right answer flips as your volume grows and hardware prices move. A token engineering platform models these deployment economics continuously and recommends the optimal placement for each model in your stack, so the decision gets revisited by software instead of forgotten by people.</p><h3>5. Prompt Compression, Rewriting, and Caching</h3><p>The cheapest token is the one you never send.</p><p>Before a request ever reaches a model, there are three opportunities to shrink it: compress the prompt to strip redundancy, rewrite it for token efficiency without losing intent, and cache aggressively &#8212; both exact-match and semantic &#8212; so repeated or near-repeated requests never hit the model at all. Individually these look like small percentage gains. At enterprise volume they compound into real money: at a million calls a day, a 20% reduction in tokens per call is not a rounding error, it&#8217;s a budget line. The platform applies these techniques automatically and only where evaluation shows they don&#8217;t degrade quality.</p><h3>6. Cost Attribution</h3><p>None of the above works as a one-time exercise. Optimization needs a feedback loop, and that loop starts with knowing exactly where your tokens go.</p><p>A token engineering platform attributes token spend at the level your business actually thinks in: per task, per team, per product feature, per customer. This is the FinOps layer for AI. It&#8217;s what lets you answer questions like &#8220;what does our IVR summarization actually cost per call?&#8221; or &#8220;which feature&#8217;s AI spend grew 40% last quarter, and was that traffic or inefficiency?&#8221; Without attribution, every optimization is a guess. With it, the platform can show you &#8212; in dollars &#8212; what each routing decision, SLM deployment, and caching policy is saving.</p><h3>7. Governance, Budgets, and Reliability Controls</h3><p>Finally, the layer that makes all of this safe to run in production: spend limits per team or application, fallback chains when a provider degrades, SLA enforcement on latency and quality, and defined degradation policies for when things go wrong.</p><p>This is the difference between a developer tool and an enterprise platform. Your finance team gets budget enforcement. Your platform team gets reliability guarantees. Your compliance team gets an audit trail of which model handled which request and why.</p><h2>How It All Fits Together</h2><p>These aren&#8217;t seven point solutions bolted together. They&#8217;re a flywheel: <strong>benchmark &#8594; route &#8594; optimize &#8594; attribute &#8594; re-benchmark.</strong></p><p>Evaluation data drives routing decisions. Routing data reveals which tasks are candidates for purpose-built SLMs. Deployment optimization changes the economics that feed back into routing. Cost attribution measures the impact of all of it and surfaces the next opportunity. Every component makes the others smarter, and the loop runs continuously &#8212; which matters, because the model landscape changes monthly and your traffic changes daily.</p><p>This is also the honest answer to the build-versus-buy question. Any strong engineering team can build one of these components. Very few can build all seven, and almost none can afford to <em>maintain</em> all seven against a model market that reprices and re-ranks itself every few weeks. Nobody builds their own cloud cost management platform anymore. The same logic is arriving for tokens.</p><h2>In Summary</h2><p>AI inference is becoming one of the largest new line items in enterprise technology budgets, and it&#8217;s currently one of the least managed. Token engineering is the discipline that closes that gap. A token engineering platform &#8212; evaluation, routing, automated SLM creation, deployment optimization, prompt optimization, cost attribution, and governance &#8212; is how you run it at scale.</p><p>If you&#8217;re spending real money on model calls and can&#8217;t say with confidence that each task is running on the right model, at the right price, on the right hardware, that&#8217;s the gap Neurometric was built to close. [Get in touch to benchmark your workloads -sales@neurometric.ai]</p>]]></content:encoded></item><item><title><![CDATA[Tokenmaxxing to Tokenminning: The Case for Token Engineering]]></title><description><![CDATA[In this episode, Rob May and Calvin Cooper unpack why "token engineering" is on track to become a discipline every AI company needs, on the same trajectory as DevOps or UX before it.]]></description><link>https://blog.neurometric.ai/p/tokenmaxxing-to-tokenminning-the</link><guid isPermaLink="false">https://blog.neurometric.ai/p/tokenmaxxing-to-tokenminning-the</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Sat, 11 Jul 2026 18:04:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>What's in this episode:</p><ul><li><p><strong>The tokenminning.com origin story &#8212; (</strong><a href="https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html">As seen in the NYT</a><strong>)</strong> The Tokenminning Manifesto by Neurometric is gaining traction, with coverage in the NYT. As enterprises prove out AI use cases while blowing through their token budgets, token engineering is becoming a first-order concern.</p></li></ul><ul><li><p>Specialized models vs. frontier models &#8212; Small, task-specific models can almost always beat a frontier model on a given task, even at a fraction of the parameter count. The catch: that specialized model can only do that one thing well. This is a feature, not a bug.</p><p></p></li><li><p>The new Token Engineering Platform &#8212; Neurometric's answer to a fragmented tooling landscape. It brings SLM fine-tuning, distillation, and token caching into a single system so teams get one view across their entire model stack.</p></li></ul><ul><li><p>&nbsp;COGS vs. OPEX &#8212; how Neurometric segments its customer base, and why companies whose AI spend hits gross margin (not just headcount-adjacent budgets) are the ones moving fastest toward token efficiency.</p></li></ul><p></p><p>Listen to the full episode:<strong> <a href="https://tokenengineering.podbean.com/">https://tokenengineering.podbean.com/</a></strong></p><p><strong>Watch on YouTube: <a href="https://youtu.be/JHFeraXq3RU?si=fzzRcEBuPo3vbF6A">https://youtu.be/JHFeraXq3RU?si=fzzRcEBuPo3vbF6A</a></strong></p><p></p><p>**Resources mentioned:**</p><ul><li><p>Tokenminning Manifesto:<a href="http://tokenminning.com/"> tokenminning.com</a></p></li><li><p> Neurometric AI: <a href="http://neurometric.ai/">neurometric.ai</a></p></li><li><p>NYT Article: <a href="https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html">https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html</a></p></li></ul><p></p>]]></content:encoded></item><item><title><![CDATA[Introducing The Neurometric Token Engineering Platform]]></title><description><![CDATA[optimize your inference system]]></description><link>https://blog.neurometric.ai/p/introducing-the-neurometric-token</link><guid isPermaLink="false">https://blog.neurometric.ai/p/introducing-the-neurometric-token</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Thu, 25 Jun 2026 11:08:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2oWO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When we started Neurometric we spoke to many large enterprises about what they were doing with AI and how they were building out their systems.  Most were not very far along but the few that were had all independently come to the same architecture.</p><p>They all started with one large frontier model, and as their inference costs grew, they started to peel off their high volume simpler workloads and set up &#8220;task specific endpoints.&#8221;  Workloads like customer sentiment analysis of support tickets, named entity extraction from documents, email summarization, these don&#8217;t need frontier intelligence.  By setting up an endpoint with a small model for just those workflows, they saw lower costs and faster latency.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2oWO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2oWO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 424w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 848w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1272w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2oWO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp" width="1456" height="1167" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/beb30d01-738b-442d-b053-332956b4519e_1916x1536.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1167,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:71524,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/203535954?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2oWO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 424w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 848w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1272w, https://substackcdn.com/image/fetch/$s_!2oWO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbeb30d01-738b-442d-b053-332956b4519e_1916x1536.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The problem is, there are thousands of models out there, and hundreds of techniques to optimize them - from where you run the model to prompting and token management techniques.  And that is all on top of the complicated issue of which model works best for your use case in the first place.</p><p>Today we are announcing that Neurometric has pulled all of these tools together in a single platform that makes it easy to make token engineering decisions.  With our platform you can:</p><ul><li><p>Monitor and measure AI workloads</p></li><li><p>Optimize workloads for cost or latency</p></li><li><p>Build custom SLMs for workloads that benefit from that approach</p></li><li><p>Evaluate and test many models, techniques, and hosting platforms</p></li></ul><p>Our customers typically see an 80% drop in inference charges and a 4x improvement in latency using the Neurometric platform.  </p><p>If you want take your AI optimization to the next level, hire a token engineer, and use the Neurometric platform.  Reach out to us at sales@neurometric.ai if you want to learn more.</p>]]></content:encoded></item><item><title><![CDATA[Tokenmaxxing is out, tokenminning is in]]></title><description><![CDATA[The New York Times (Eli Tan) reported on a shift companies are now making after a year of unchecked AI spending.]]></description><link>https://blog.neurometric.ai/p/tokenmaxxing-is-out-tokenminning</link><guid isPermaLink="false">https://blog.neurometric.ai/p/tokenmaxxing-is-out-tokenminning</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Fri, 19 Jun 2026 15:15:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The New York Times (Eli Tan) reported on a shift companies are now making after a year of unchecked AI spending.</p><p></p><p>Rob May, our CEO, told the Times: CEOs who couldn&#8217;t measure AI savviness defaulted to &#8220;who&#8217;s using the most tokens.&#8221; Volume over efficiency was never going to hold once the bills landed.</p><p></p><p>Uber blew through its full-year AI budget in four months. Meta is capping usage after an &#8220;exponential increase&#8221; in costs. AT&amp;T&#8217;s shows companies can save up to 90% by using less powerful models for most tasks.</p><p></p><p>This is the insight we built Neurometric on.&nbsp;</p><p></p><p>Read the full article: <a href="https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html">https://www.nytimes.com/2026/06/18/technology/ai-token-minimizing.html</a></p>]]></content:encoded></item><item><title><![CDATA[Automating the SLM Development Loop — Pioneer Agent Paper Breakdown]]></title><description><![CDATA[The bottleneck in deploying small language models isn&#8217;t training.]]></description><link>https://blog.neurometric.ai/p/automating-the-slm-development-loop</link><guid isPermaLink="false">https://blog.neurometric.ai/p/automating-the-slm-development-loop</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Tue, 16 Jun 2026 22:05:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The bottleneck in deploying small language models isn&#8217;t training. It&#8217;s data curation, failure diagnosis, regression avoidance, and iteration control.</p><p>That&#8217;s the central argument of the Pioneer Agent paper from Fastino Labs &#8212; and it maps pretty cleanly onto what we&#8217;ve been building at Neurometric.</p><p>In this episode of Inference Time Tactics, Director of AI Research Yash Sharma breaks it down with co-founder Calvin Cooper: what Pioneer Agent actually is, how it uses Claude Sonnet as an ML engineer in a box, and what the results tell us about where SLMs go next.</p><p></p><p>Key topics:</p><p>&#8212; Cold start data curation with zero customer traces</p><p>&#8212; Why naive retraining fails with noisy production data (and how an agentic loop fixes it)</p><p>&#8212; 83-point benchmark improvements &#8212; and why the 1.6-point cases matter just as much</p><p>&#8212; Regression rollback, hyperparameter search, and the heuristics baked into the system</p><p>&#8212; Where we think SLMs can go beyond &#8220;simple tasks&#8221;</p><p></p><p>Watch the full episode on YouTube: <a href="https://youtu.be/PiIrywMGsAA?si=gs_3nDwgS7QbvwQG">https://youtu.be/PiIrywMGsAA?si=gs_3nDwgS7QbvwQG</a></p><p></p><p>Or listen in on any platform: <a href="https://inferencetimetactics.podstream.com/">https://inferencetimetactics.podstream.com</a></p><p></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://blog.neurometric.ai/subscribe?utm_source=email&r=&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://blog.neurometric.ai/subscribe?utm_source=email&r="><span>Subscribe</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Discipline of Token Engineering: Why Your AI Infrastructure Is Bleeding Cash]]></title><description><![CDATA[And how to fix it.]]></description><link>https://blog.neurometric.ai/p/the-discipline-of-token-engineering</link><guid isPermaLink="false">https://blog.neurometric.ai/p/the-discipline-of-token-engineering</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Sat, 13 Jun 2026 17:54:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Remember when building an LLM application simply meant wrapping a basic system prompt around a single frontier API key? Those days are officially over. In 2026, we find ourselves in the middle of a massive architectural shift. We are no longer just building chatbots; we are deploying complex, multi-step agent fleets. Yet this evolution has exposed a glaring operational vulnerability: every company is now a token company, but almost none of them have a token engineer.</p><p>Currently, enterprise infrastructure spend is run with zero discipline. We regularly see frontier models costing $5 to $25 per million tokens assigned to commodity work that $0.10 to $0.50 small models handle easily on benchmarks. This creates a 50x to 250x price spread, paid on every request, every single day. To survive this efficiency gap, a new engineering discipline has emerged: <strong>Token Engineering</strong>. It is the systematic optimization of which model runs which task&#8212;at what size, with what prompt structure, and at what cost.</p><h2>The Agentic Scale Problem</h2><p>The root of this cost crisis is that our foundational software design patterns have completely changed. Over 70% of routed inference traffic now comes from autonomous agents and CLI tools&#8212;not human chat interfaces. While a standard chat turn consumes just a few thousand tokens, a single agentic task regularly chews through 100K to 1M tokens as it loops, reasons, and self-corrects.</p><p>This volume shift has caused platform-wide token volume to explode by more than 10x in roughly a year, with 16 to 18 trillion tokens per week routed on OpenRouter alone. When workloads scale to this magnitude, token waste becomes an engineering failure, not a model failure.</p><h2>The Five Pillars of Token Waste</h2><p>When auditing modern agent pipelines, token waste typically boils down to five core engineering oversights:</p><ul><li><p><strong>Oversized Models</strong>: Allocating expensive frontier models to simple extraction, classification, and formatting tasks that small models win on benchmarks.</p></li><li><p><strong>Prompt Bloat</strong>: Deploying unversioned, unmeasured prompts that carry thousands of redundant tokens into every single call.</p></li><li><p><strong>No Caching Strategy</strong>: Re-sending massive chunks of static context on every request instead of caching it at 10% to 20% of the standard price.</p></li><li><p><strong>Sequential Sprawl</strong>: Running agent steps serially with full context when steps could be decomposed, parallelized, and right-sized across lean endpoints.</p></li><li><p><strong>Blind Retries</strong>: Retrying unexpected failures on the same expensive model with no confidence scoring or cheaper fallback path.</p></li></ul><h2>Why Human Optimization Fails</h2><p>Fixing these leaks manually is a noble goal, but token engineering simply does not scale as a human job. First, we face continuous <strong>Model Churn</strong>. Major model releases land weekly. Look at Gemma 4: it went from non-existent to routing 240 billion tokens per week&#8212;roughly one-third of all small-model traffic on OpenRouter&#8212;in a mere 70 days. No human team can re-benchmark thousands of model-by-task combinations on that clock.</p><p>Second, we hit <strong>Price Churn</strong>. Providers reprice continuously; caching, batching, and changing infrastructure economics shift the optimal choice even when models remain static. The half-life of an optimization is measured in weeks; a hand-tuned pipeline is stale before the next sprint ends. The engineer has to be automated.</p><h2>The FinOps Parallel</h2><p>Every era of computing waste eventually forces the transition from a heroic manual task into an automated platform practice. When web applications struggled with uptime, we turned manual on-call duties into Site Reliability Engineering (SRE). When cloud spend spiraled out of control, cloud waste created FinOps platforms that natively paid for themselves. Today, token spend is the fastest-growing cost line in software, and it demands its own dedicated engineering platform.</p><p>As LLM infrastructure commoditizes, value is rapidly migrating away from foundational providers and straight to the orchestration layer&#8212;the core intelligence about which intelligence to use. By decoupling your software workflows from rigid API keys and transitioning to dynamic task endpoints, you can let automation drive delivery costs down via leaner prompts, continuous model re-matching, and purpose-built small models. In this new landscape, remember: every standard inference vendor makes more money when your system wastes tokens. True operational maturity belongs to the engineering teams who build systems that actively eliminate the bleed.</p><p>If you want help, Neurometric offers a platform ideal for token engineering.  Contact us to chat more about it.</p>]]></content:encoded></item><item><title><![CDATA[What We Learned About The Harbor Framework From More than 5,700 Benchmark Runs.]]></title><description><![CDATA[The infrastructure you run on matters a lot.]]></description><link>https://blog.neurometric.ai/p/what-we-learned-about-the-harbor</link><guid isPermaLink="false">https://blog.neurometric.ai/p/what-we-learned-about-the-harbor</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 10 Jun 2026 18:13:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We ran a large-scale benchmark study to evaluate AI agent performance across a broad model catalog. What we expected to produce was a leaderboard. What we got instead was a lesson in how evaluation infrastructure can quietly invalidate your results before you ever read them.</p><div><hr></div><h2>The Numbers</h2><p><strong>5,791 benchmark runs. 15,750 individual graded tasks.</strong></p><p>That&#8217;s a meaningful dataset. But the headline isn&#8217;t a model score &#8212; it&#8217;s that the framework couldn&#8217;t reliably finish.</p><p><strong>53% of runs errored out.</strong> An errored run produces nothing: zero tasks completed, no usable result. More than half our compute returned empty-handed, not because the models failed, but because the harness crashed.</p><p>This matters beyond the obvious waste. When your error rate crosses 50%, you&#8217;re no longer sampling performance &#8212; you&#8217;re sampling infrastructure luck.</p><div><hr></div><h2>The Grading Problem</h2><p>The errors weren&#8217;t just crashes. They were silent distortions inside the results that <em>did</em> come back.</p><p>In <strong>765 cases</strong>, the agent produced the exactly correct answer &#8212; and Harbor logged it as an error anyway. That&#8217;s <strong>one in five of all correct results</strong> misclassified as failures. A grading system that can&#8217;t recognize its own correct answers isn&#8217;t grading. It&#8217;s noise with a spreadsheet attached.</p><p>The implication: any model rankings produced under these conditions would systematically understate performance, with the degree of understatement varying arbitrarily across runs. You can&#8217;t normalize your way out of that.</p><div><hr></div><h2>Dataset Coverage</h2><p>Harbor ships with 80 datasets. We got <strong>42 of them to run at all</strong> &#8212; just over half. Of those 42, <strong>17 never produced a single scored result</strong>. That&#8217;s 17 datasets that opened, ran, and returned nothing actionable.</p><p>Effective coverage: roughly <strong>31% of the catalog</strong> produced data you could learn from. You cannot benchmark across a catalog when two-thirds of it is structurally inoperable.</p><div><hr></div><h2>The Hello-World Test</h2><p>The clearest evidence is the simplest task.</p><p>Harbor&#8217;s hello-world benchmark asks the agent to create a file containing <code>"Hello, world!"</code> That&#8217;s it. No reasoning. No retrieval. No multi-step planning. Just: write a file.</p><p>We ran it across <strong>1,526 agent+model combinations</strong>. <strong>645 of them &#8212; 42% &#8212; never completed it cleanly even once.</strong></p><p>In one case, GPT-5.4 wrote the file perfectly. The run still errored.</p><p>The intelligence was never the bottleneck. The harness was.</p><div><hr></div><h2>What This Means for AI Benchmarking</h2><p>Benchmark infrastructure is load-bearing. It doesn&#8217;t just measure performance &#8212; it <em>defines</em> what counts as performance. When the harness fails silently, misclassifies correct answers, and drops the majority of its own datasets, the output isn&#8217;t a measurement. It&#8217;s a corrupted signal that looks like data.</p><p>The field has spent enormous energy debating which benchmarks best capture model capability. That&#8217;s the right debate to have &#8212; once the execution layer can be trusted. Our results suggest the execution layer deserves a lot more scrutiny than it typically gets.</p><p>A model that writes the correct answer deserves to have that answer counted. That&#8217;s table stakes for any evaluation system. When it isn&#8217;t met, the benchmark isn&#8217;t measuring models. It&#8217;s measuring the framework.</p>]]></content:encoded></item><item><title><![CDATA[How to Run OpenClaw Without the Frontier Tax]]></title><description><![CDATA[How inference routing cuts 60&#8211;90% of frontier model calls without changing your OpenClaw setup]]></description><link>https://blog.neurometric.ai/p/how-to-run-openclaw-without-the-frontier</link><guid isPermaLink="false">https://blog.neurometric.ai/p/how-to-run-openclaw-without-the-frontier</guid><dc:creator><![CDATA[neurometric]]></dc:creator><pubDate>Wed, 03 Jun 2026 22:12:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YSzS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YSzS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YSzS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YSzS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png" width="724" height="407.25" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:724,&quot;bytes&quot;:1189569,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://neurometric.substack.com/i/200471143?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YSzS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!YSzS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26f33274-03fb-42e3-8df4-19455a9d3ef2_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The OpenClaw cost problem is not a secret. Users are posting about it openly. One tech blogger documented $3,600 in a single month. Others report $200 days from runaway automation loops. Multi-agent setups on premium models routinely hit $600/month before anyone audits the config. The culprit is almost always the same: every sub-task &#8212; a heartbeat check, a JSON format, a ticket classification &#8212; is hitting a frontier model at full price.</p><p>The problem got more complicated in April. Anthropic blocked Claude Pro and Max subscribers from using their flat-rate plans with third-party tools like OpenClaw. OpenAI went the other direction. Sam Altman posted at 2am on May 2: &#8220;you can sign in to openclaw with your chatgpt account now and use your subscription there.&#8221; ChatGPT Plus and Pro now cover OpenClaw usage at a flat monthly rate via Codex OAuth &#8212; no per-token billing.</p><p>That is good news for a lot of users. But subscriptions have limits. ChatGPT Plus carries a 5-hour weekly usage quota. Pro users hit ceilings too on heavy workflows. The moment you are running serious automation &#8212; multi-agent pipelines, scheduled tasks, anything that fires dozens of requests an hour &#8212; you are back against a wall regardless of which subscription you are on.</p><p>The underlying issue is not which provider you use. It is that the wrong model is handling the wrong work.</p><p>wrong work.</p><div><hr></div><h2><strong>Why the Bill Gets Out of Control</strong></h2><p>OpenClaw routes every task to whatever model you have set as default. That model does not know the difference between a request that requires genuine reasoning and one that is just formatting a JSON object or summarizing a paragraph. It handles both the same way: full context load, full inference, full cost.</p><p>Most OpenClaw workflows are 70-80% routine. Heartbeats, memory housekeeping, classification, extraction, formatting, cron jobs. None of it requires a frontier model. But if your default is Claude Sonnet or GPT-5.4, every one of those calls is priced like it does.</p><p>That is the frontier tax. You are paying for capability you are not using.</p><div><hr></div><h2><strong>What We Built and Why It Works</strong></h2><p>Neurometric builds task-specific Small Language Models &#8212; purpose-built for narrow jobs. A model fine-tuned to classify support tickets does not need 1.7 trillion parameters to do that job well. A 7B model trained on thousands of legal extraction examples outperforms a frontier model on that specific task and runs at a fraction of the cost.</p><p>This is not a discount. It is architecture. Small models doing specific things are cheaper to run because they are smaller and more efficient at their job. The economics are structural, not promotional. That is why we can offer 100 million tokens per month for free and sustain unlimited token plans at prices that make sense.</p><p>The frontier model handles what it is actually good at: complex reasoning, multi-step planning, creative synthesis. Everything else routes to a specialist.</p><p>ClawPack is how this plugs into OpenClaw. It sits alongside your existing model as a standard provider. One model ID. Automatic routing. Your frontier model &#8212; or your ChatGPT subscription &#8212; stays in place for the work that needs it. ClawPack handles the rest.</p><p>The result is the same OpenClaw experience, with 60-90% fewer frontier model calls. If you are on a ChatGPT subscription, your quota goes further. If you are on API billing, your bill drops. Either way, you are not burning Opus-level compute to check whether your inbox has anything urgent.</p><div><hr></div><h2><strong>The Stack We Recommend</strong></h2><p>We are partnering with hosting providers like LumaDock to make this easy for OpenClaw users to set up end to end.</p><p><a href="https://lumadock.com/openclaw-vps-hosting">LumaDock</a> offers a purpose-built OpenClaw VPS template &#8212; you pick a plan, deploy a server, and OpenClaw is already installed and running when you SSH in. Starts at $1.99/month. Their<a href="https://lumadock.com/tutorials/openclaw-complete-guide"> complete OpenClaw guide</a> and<a href="https://lumadock.com/faq"> FAQ</a> cover everything from first setup to production configuration.</p><p>Once OpenClaw is running, adding ClawPack takes two steps. Go to<a href="https://marketplace.neurometric.ai/clawpack"> marketplace.neurometric.ai/clawpack</a>, get your free API key, and copy the pre-populated install command the dashboard generates. Paste it into your server terminal. Then:</p><p>openclaw models set neurometric/clawpack</p><p>That&#8217;s it. ClawPack is live. For complex reasoning tasks, configure a fallback to your frontier model or ChatGPT subscription and OpenClaw escalates automatically when the task warrants it.</p><p>Free tier covers 100M tokens/month. No credit card. Unlimited plans available through your Neurometric account.</p><div><hr></div><h2><strong>The Bigger Picture</strong></h2><p>The era of subsidized frontier inference for agentic workflows is coming to an end. Anthropic made that clear in April. OpenAI&#8217;s subscription path is a better deal for many users, but it is still a ceiling, not a solution.</p><p>The solution is not finding a cheaper frontier model. It is using the right model for the right task. Specialized intelligence delivered at the compute cost it actually requires. That is what we are building at Neurometric, and it is why the economics hold regardless of what the big providers do next.</p><div><hr></div><p><em>Get started:<a href="https://lumadock.com/tutorials/cut-openclaw-costs-clawpack/?utm_source=neurometriclumadock&amp;utm_medium=substack"> LumaDock OpenClaw x ClawPack</a></em></p>]]></content:encoded></item><item><title><![CDATA[Lean Inference Workflows: Applying "Lean" Concepts To Building AI Agents]]></title><description><![CDATA[Making inference scale in a cost effective way]]></description><link>https://blog.neurometric.ai/p/lean-inference-workflows-applying</link><guid isPermaLink="false">https://blog.neurometric.ai/p/lean-inference-workflows-applying</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 03 Jun 2026 17:39:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Here&#8217;s a production scenario that should feel familiar: your agent hits a simple routing decision&#8212;does this user query need a database lookup or a calculator?&#8212;and it fires off a GPT-4o call with a 12,000-token context window stuffed with documentation it will never read, waits 4 seconds for a response, gets back malformed JSON, retries twice, and burns $0.40 to answer a question that a regex could have handled.</p><p>Multiply that across 10,000 daily requests. Congratulations&#8212;you&#8217;ve built an inference money pit.</p><p>The AI engineering community collectively discovered that &#8220;just throw it at a frontier model&#8221; works great in demos and collapses in production. Agents enter retry death spirals. Context windows bloat with irrelevant RAG results. Sequential LLM calls stack latency until users abandon the workflow. The tools are extraordinarily powerful, and we are using them with the efficiency of a factory floor that nobody has ever walked with a stopwatch.</p><p>Lean Manufacturing fixed this problem for physical production 40 years ago. It&#8217;s time to apply the same discipline to inference.</p><p><strong>Lean Inference Workflows</strong> are the systematic application of Lean/TPS (Toyota Production System) principles to the design of LLM-powered agent architectures. Not as metaphor&#8212;as engineering discipline.</p><div><hr></div><h2>The 7 Wastes of LLM Inference</h2><p>Taiichi Ohno&#8217;s <em>muda</em> framework identified seven categories of waste in manufacturing. Each maps cleanly onto the failure modes we build into agents every day.</p><h3>1. Overproduction &#8212; The Frontier Model Default</h3><p>The most expensive waste is calling a 70B+ frontier model for tasks that don&#8217;t need it. Routing a support ticket to the right queue? That&#8217;s an 8B classification task. Extracting structured fields from a form submission? That&#8217;s a fine-tuned 3B model with a JSON schema. Summarizing a 500-word support thread? You don&#8217;t need GPT-4o.</p><p><strong>The cost asymmetry is staggering.</strong> claude-sonnet runs ~3x the cost of haiku per token. GPT-4o runs ~10x the cost of GPT-4o-mini. When you reflexively reach for the frontier model on every step of a 15-step agent loop, you&#8217;re not just overspending&#8212;you&#8217;re adding latency at every node.  If your task is a common one, you can even move to SLMs which are faster and two orders of magnitude cheaper.</p><p>Treat your agent&#8217;s model selection the same way a traffic engineer treats routing decisions&#8212;based on payload size, complexity score, and confidence threshold, not habit.</p><h3>2. Inventory &#8212; RAG Bloat</h3><p>Your vector database returns the top-20 chunks, and you shove all 20 into the context window &#8220;just in case.&#8221; That&#8217;s inventory waste: stockpiling inputs you probably won&#8217;t use, forcing the model to process them, inflating your input token count, and degrading retrieval precision in the process. More context isn&#8217;t better&#8212;it&#8217;s a longer assembly line with more defect opportunities.</p><p><strong>Controlled inventory</strong> means retrieving fewer, better chunks via re-ranking (a cross-encoder pass over your top-k candidates), then truncating aggressively before injection.</p><h3>3. Waiting &#8212; Sequential Blocking</h3><p>Tool calls that could run in parallel are running in series. You need to fetch a user&#8217;s account history, check their subscription tier, and retrieve their recent support tickets. Instead of three parallel async calls, you have three sequential blocking calls: 300ms + 280ms + 310ms = 890ms of pure waiting.</p><p><strong>async/await + parallel execution</strong> is the <code>asyncio.gather</code> or <code>Promise.all</code> call you should have made. In a multi-step agent DAG, every synchronous bottleneck is a latency tax.</p><h3>4. Defects &#8212; Malformed Outputs and Retry Loops</h3><p>An agent asks for a JSON tool call. The model returns Markdown-wrapped JSON with an extra trailing comma. Your parser throws. The orchestrator retries. The model hallucinates a different schema on the retry. You&#8217;re now three LLM calls deep on a task that should have been one.</p><p><strong>Defects in inference are uniquely expensive</strong> because retries aren&#8217;t cheap reruns&#8212;they&#8217;re full-price LLM calls on an already-failed path. Structured outputs (OpenAI&#8217;s <code>response_format</code>, Anthropic&#8217;s tool use schemas, the <code>instructor</code> library for Python) eliminate this entirely by constraining output at the token-probability level.</p><h3>5. Over-Processing &#8212; Unnecessary Chain-of-Thought</h3><p>CoT is a forcing function for reasoning. It is not a default that belongs in every prompt. A routing classifier does not need to explain its reasoning to itself before assigning a ticket category. A field extractor does not need <code>&lt;thinking&gt;</code> tokens. Stripping CoT from non-reasoning tasks can cut your output token count by 40&#8211;60% on those steps&#8212;with zero quality loss.</p><div><hr></div><h2>Core Principles of Lean Inference</h2><h3>Just-In-Time Context: The Pull System</h3><p>In Lean manufacturing, a pull system means downstream demand triggers upstream production&#8212;nothing gets built until it&#8217;s needed. <strong>JIT Context</strong> means your agent fetches context exactly when a step requires it, scoped precisely to what that step needs.</p><p>The anti-pattern is the &#8220;God Context&#8221;: a single massive system prompt that pre-loads everything the agent <em>might</em> need across all possible execution paths. You pay the full token tax on every call, even when 80% of that context is never accessed.</p><p>The Lean pattern:</p><ul><li><p><strong>Semantic caching</strong> at the retrieval layer: if a semantically similar query was answered 30 seconds ago, return the cached embedding result, not a fresh DB round-trip.</p></li><li><p><strong>Re-ranking before injection</strong>: Run a cross-encoder (a fast, cheap model like <code>ms-marco-MiniLM-L-6-v2</code>) over your retrieved chunks <em>before</em> injecting them into the LLM context. Top-3 precision beats top-20 recall for most tasks.</p></li><li><p><strong>Step-scoped context</strong>: Each node in your agent DAG gets only the context its specific tool call requires. The summarization node doesn&#8217;t need the tool definitions. The routing node doesn&#8217;t need the document corpus.</p></li></ul><h3>Standardized Work: Deterministic Guardrails</h3><p>Lean&#8217;s <em>standardized work</em> principle says that defined, repeatable processes reduce variation and defects. In agent architecture, this translates to: <strong>make your LLM do as little undirected reasoning as possible.</strong></p><p>Logic that can be encoded deterministically should be. Your state machine transitions, routing rules, retry budgets, and tool call sequencing should live in code&#8212;not in a prompt asking the model to figure it out.</p><p>Tools like <strong>LangGraph</strong> let you encode agent control flow as an explicit graph: nodes are LLM calls or tool invocations, edges are conditional transitions, and the state machine is a first-class object you can inspect, test, and version-control. This is categorically different from a single ReAct loop where the model decides everything.</p><p><strong>Structured outputs</strong> (via <code>instructor</code>, OpenAI&#8217;s <code>strict: true</code> JSON mode, or Anthropic&#8217;s tool schemas) are the manufacturing equivalent of a jig: they physically constrain the output to the valid shape, making defects structurally impossible rather than probabilistically unlikely.</p><h3>Takt Time: The Latency Budget</h3><p>Takt time in manufacturing is the maximum allowable time per unit to meet customer demand. In agent design, every workflow should have an explicit <strong>latency budget</strong> per step and per full execution path.</p><p>Define your takt time first. If your end-to-end SLA is 2 seconds and you have 6 agent steps, your average per-step budget is ~333ms. That budget forces architectural decisions:</p><ul><li><p>Can this step use a smaller model to hit the latency target?</p></li><li><p>Should this step be parallelized?</p></li><li><p>Does this step even need an LLM, or is a heuristic or cached result sufficient?</p></li></ul><p><strong>DAG decomposition</strong> is your primary tool here. A complex task that looks like a single LLM call is often a DAG of 4&#8211;6 smaller model calls that can execute in parallel, each with faster TTFT on a smaller model, combining to lower overall latency than the single big call.</p><h3>Prompt Caching as Kanban</h3><p>Anthropic&#8217;s prompt caching and OpenAI&#8217;s equivalent cache system are <strong>Kanban cards for inference</strong>: reusable, pre-positioned work items that don&#8217;t need to be re-manufactured from scratch.</p><p>Your system prompt, tool definitions, and static knowledge base content are the same across thousands of requests. Cache them. On Anthropic&#8217;s API, a cache hit on a 10,000-token system prompt costs 10% of the base input token price. Over millions of calls, this is not a micro-optimization&#8212;it&#8217;s a cost structure change.</p><p>Design your prompts with <strong>cache-friendly prefix ordering</strong>: static system prompt first, static tool definitions second, dynamic context last. Anything that changes per-request must come after anything that doesn&#8217;t.</p><div><hr></div><h2>Before and After: Repo Analysis Agent</h2><p><strong>Before (Naive Architecture)</strong></p><p>A single ReAct loop. One GPT-4o call per step. Full repository context dumped into the window on every iteration. Tool definitions re-sent each time. Sequential file reads. No output validation. Average: <strong>14 seconds, ~85,000 tokens, ~$1.20 per run.</strong></p><p><strong>After (Lean Architecture)</strong></p><ul><li><p>A small <strong>router model</strong> (8B, fine-tuned) classifies the task type and selects the appropriate specialist pipeline &#8212; adds 80ms, saves 60% of downstream model costs</p></li><li><p><strong>Prompt caching</strong> on tool definitions and system context &#8212; 90% cache hit rate after warmup</p></li><li><p><strong>Parallel tool execution</strong> for file reads &#8212; 4 simultaneous reads instead of sequential</p></li><li><p><strong>Structured output enforcement</strong> via <code>instructor</code> &#8212; zero retry loops in 500-run benchmark</p></li><li><p><strong>Strict step budget</strong>: 6 steps max, with a fallback to human handoff at budget exhaustion</p></li></ul><p>Result: <strong>4.2 seconds average, ~18,000 tokens, ~$0.09 per run.</strong> Same output quality score on eval suite. 13x cost reduction. 3.3x latency improvement.</p><div><hr></div><h2>Conclusion</h2><p>The next frontier of AI engineering isn&#8217;t a bigger context window or a more capable base model. It&#8217;s the discipline to use what we already have without waste.</p><p>Every unnecessary frontier model call, every bloated RAG context, every sequential blocking operation, every retry loop from malformed output&#8212;these are engineering failures, not model failures. We built them. We can fix them.</p><p>Lean Inference isn&#8217;t a philosophy&#8212;it&#8217;s a set of concrete architectural decisions you can make this sprint. Audit your agent&#8217;s token burn by step. Map your sequential calls. Add structured outputs. Right-size your models. Cache your static prompts.</p><p>Build leaner. Run faster. Spend less. Ship better agents.</p>]]></content:encoded></item><item><title><![CDATA[Introducing OwlPack: SLMs That Scan And Fix Your Codebase While You Sleep]]></title><description><![CDATA[Designed for a vibecoding world]]></description><link>https://blog.neurometric.ai/p/introducing-owlpack-slms-that-scan</link><guid isPermaLink="false">https://blog.neurometric.ai/p/introducing-owlpack-slms-that-scan</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Thu, 21 May 2026 19:10:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We have over 150 fine tuned models in our <a href="http://marketplace.neurometric.ai">SLM Marketplace</a>, and over the last 6 weeks we noticed that 2 of our code testing SLMs were the most popular downloads.   We took that as a sign and decided to bundle a few coding SLMs together in a product we are announcing today -  <a href="https://marketplace.neurometric.ai/owlpack">Owlpack </a>&#8212; <em><strong>a GitHub App that runs five specialist agents against your repos while you sleep</strong></em>. </p><h2>Built on 5 Small Models</h2><p>When we say Small Language Model, we mean models under 20B parameters. Some are under 3B. They&#8217;re small enough to run on a single GPU, fast enough to return structured output in seconds, and cheap enough that running five of them in parallel against thousands of files is a rounding error, not a line item.</p><p>The catch &#8212; and it&#8217;s a real one &#8212; is that they don&#8217;t know everything. A 3B model is not going to write your novel, debate Kant, or replace Opus. But it doesn&#8217;t need to. It needs to find SQL injections. Or flag a deprecated dependency. Or notice that a function has crept past 400 lines.</p><p>Narrow the task, and small wins.</p><h2>The SLMs That Compose Owlpack</h2><p>Owlpack runs five agents every night, each focused on a different domain:</p><ul><li><p><strong>Hunter</strong> scans for security vulnerabilities &#8212; SQLi, SSRF, auth bypasses, leaked secrets, cross-referenced against live CVE feeds.</p></li><li><p><strong>Tracker</strong> hunts bugs &#8212; null references, race conditions, off-by-one errors, suspicious test coverage gaps.</p></li><li><p><strong>Keeper</strong> audits dependencies &#8212; outdated packages, breaking changes, deprecation notices, abandoned libraries.</p></li><li><p><strong>Mason</strong> identifies refactoring opportunities &#8212; duplication, coupling hotspots, long functions, outdated patterns.</p></li><li><p><strong>Scribe</strong> analyzes the codebase itself &#8212; churn, review latency, complexity drift, module ownership.</p></li></ul><p>Each one is a specialist. Each one runs in parallel. Each one returns structured findings that get deduplicated, diffed against history, and ranked by severity before they land in your inbox by morning.</p><p>Try doing that with a single frontier model. The math falls apart. A nightly full-repo scan, across thousands of users, calling GPT-class inference five times per repo, would cost more than most teams pay for their entire dev tooling stack. That&#8217;s why no one was offering this service. The unit economics didn&#8217;t work.</p><p>With SLMs, they do.</p><h2>The case for specialization</h2><p>There&#8217;s a deeper reason we built it this way, and it goes beyond cost.</p><p>When you fine-tune a small model on a specific task &#8212; say, identifying CVE patterns in JavaScript &#8212; it gets better at that task than a frontier model trained to do everything. Specialists beat generalists when the task has a defined shape. Code review is a defined shape. Dependency auditing is a defined shape. Detecting a leaked AWS key is a <em>very</em> defined shape.</p><p>A frontier model brings a trillion-plus parameters of knowledge about Roman history and SQL injection patterns. Hunter brings the SQL injection patterns. For this job, that&#8217;s the better tool.</p><p>It also means the agents don&#8217;t drift. They don&#8217;t get creative. They don&#8217;t hallucinate a function name that doesn&#8217;t exist in your repo because they read about it in a blog post once. Narrowness is a feature.</p><h2>Privacy as a byproduct</h2><p>There&#8217;s a third benefit that falls out of using SLMs: a smaller blast radius. Owlpack clones your repository at scan time, runs the agents, and deletes the clone. We don&#8217;t need a 1.5T-parameter foundation model to read your code. We need five tight, task-specific models that do their job and forget. That architecture is easier to audit, easier to contain, and easier to deploy in environments where data exfiltration risk actually matters.</p><h2>What this looks like in practice</h2><p>You install the GitHub App. You pick which repos to scan. We clone them at your chosen time, run all five agents in parallel, delete the clone, and deliver a structured briefing by morning. If there&#8217;s nothing new, we don&#8217;t email you. If there&#8217;s a CVE in a transitive dependency, you&#8217;ll know before standup.</p><p>The whole pipeline costs a fraction of what a single frontier call would. That&#8217;s what passes through to pricing. That&#8217;s what makes the service viable.</p><p>The big-model era taught the industry to reach for the biggest hammer in the room. The next era is about picking the right one. Five small, sharp tools, running every night &#8212; that&#8217;s what Owlpack is. SLMs are why it works.</p><p><a href="https://marketplace.neurometric.ai/owlpack">Try Owlpack Free For 7 Days</a>.  Or email us sales@neurometric.ai if you want to try a team plan.</p>]]></content:encoded></item><item><title><![CDATA[Beyond The Hype: What 3,000 Users Taught Us About Small Language Models In The Real World]]></title><description><![CDATA[We hear about SLMs - what are people doing with them?]]></description><link>https://blog.neurometric.ai/p/beyond-the-hype-what-3000-users-taught</link><guid isPermaLink="false">https://blog.neurometric.ai/p/beyond-the-hype-what-3000-users-taught</guid><dc:creator><![CDATA[Rob May]]></dc:creator><pubDate>Wed, 13 May 2026 20:04:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hnne!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ac5327f-75ee-4122-90c8-29331a4392a0_800x800.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>For the last two years, the AI conversation has been dominated by one question: <em>how big can we go?</em> Trillion-parameter models, frontier benchmarks, $100B compute commitments. Bigger, smarter, more general.</p><p>But somewhere along the way, a quieter revolution started inside the enterprise. Companies stopped asking &#8220;what&#8217;s the most capable model?&#8221; and started asking &#8220;what&#8217;s the most <em>useful</em> one?&#8221;</p><p>Small Language Models &#8212; generally defined as models with fewer than 10 billion parameters &#8212; have emerged as the answer. They&#8217;re cheaper to run, fast enough for real-time workflows, and small enough to deploy on-prem or even on a laptop. That changes the economics and the privacy posture of AI in ways the headline-grabbing frontier models simply can&#8217;t match.</p><p>At Neurometric, we just crossed <strong>3,000 active SLM users</strong>.  You can download a fine tuned SLM from us without becoming a customer so, only a little over 2,200 of those users have applied for a key and are using Neurometric to host the SLM. But we&#8217;ve seen which pre-fine-tuned models are most popular from our <a href="http://marketplace.neurometric.ai">SLM Marketplace</a>, and we&#8217;ve talked to some enterprises directly who have asked for our help with larger scale SLM deployments. </p><p>If you&#8217;ve ever wondered what people use SLMs for in the real world - here are the top 5 things we&#8217;ve seen.</p><h2>1. Summarization: Gist Generation at Scale</h2><p>The single biggest workload across our user base is document summarization. Legal briefs, customer transcripts, research reports, internal wikis.</p><p>Why an SLM wins here: most summarization doesn&#8217;t need creative prose. It needs the <em>gist</em> &#8212; accurate, fast, and cheap enough to run across thousands of documents a day. Pushing a 50-page PDF through a frontier model with a giant context window costs real money. An SLM does the same job at a fraction of the latency and a tiny fraction of the cost. When you&#8217;re processing 10,000 documents a week, that math becomes existential.</p><h2>2. Resume Screening: Extraction Without Hallucination</h2><p>HR teams were one of the fastest verticals to adopt. The job isn&#8217;t to &#8220;write a beautiful candidate evaluation&#8221; &#8212; it&#8217;s to pull skills, years of experience, certifications, and seniority into a structured format.</p><p>That&#8217;s an extraction task, not a creativity task. And ironically, smaller, fine-tuned models are <em>less</em> prone to embellishment than their bigger cousins. Fewer parameters means fewer paths for the model to &#8220;fill in the blank&#8221; with something that wasn&#8217;t on the resume. For HR &#8212; where a hallucinated qualification is a compliance problem &#8212; the precision of an SLM is a feature, not a limitation.</p><h2>3. Code Refactoring: Local, Low-Latency Suggestions</h2><p>Developers using SLMs for code refactoring are the most popular group who download and use it themselves, rather than have us host it. The two reasons they seem to like the SLM approach:</p><ul><li><p><strong>Latency.</strong> An autocomplete suggestion that arrives 800ms after you stop typing is useless. Local SLMs respond in tens of milliseconds.</p></li><li><p><strong>Security.</strong> Proprietary code never leaves the machine. For regulated industries and any company with a defensible codebase, that&#8217;s non-negotiable.</p></li></ul><p>The frontier model can still review the architecture. The SLM handles the thousand small edits in between.</p><h2>4. CRM Summary: Killing the Sales Busy Work</h2><p>Sales reps don&#8217;t write call notes; they scribble them. The result is a CRM full of half-finished entries that nobody trusts.</p><p>Our users are using SLMs to convert messy voice memos and chat transcripts into structured CRM fields &#8212; next steps, objections, deal stage, sentiment. This is a textbook SLM workload: repetitive, schema-bound, and high-volume. A rep doing eight calls a day generates 40 summaries a week. Multiply by a 200-person sales org and you can see why running this on a frontier model is a non-starter.</p><h2>5. Meeting Prep: Briefing Sheets Before the Call</h2><p>The fifth pattern is what users are calling &#8220;pre-meeting intelligence.&#8221; Before a customer call, the SLM ingests prior emails, past call notes, account history, and recent product activity &#8212; then generates a one-page briefing sheet.</p><p>Speed is the metric here. The briefing needs to be ready in the 90 seconds between the previous meeting ending and the next one starting. Frontier models can&#8217;t hit that latency reliably. SLMs can.</p><h2>So What Are People Actually Using SLMs For?</h2><p>They&#8217;re not writing novels. They&#8217;re not passing the bar exam. They&#8217;re not solving open-ended research problems.</p><p>They&#8217;re doing the most common work tasks.</p><p>SLMs are workhorse models. Across our thousands of users, the pattern is consistent: repetitive, high-volume, structured tasks where you need roughly <strong>90% accuracy at roughly 5% of the cost</strong> of a frontier LLM. That&#8217;s not a consolation prize &#8212; that&#8217;s the entire enterprise AI opportunity. Most business value isn&#8217;t locked behind PhD-level reasoning. It&#8217;s locked behind tasks too small and too numerous to justify a $0.10-per-call inference bill.</p><h2>The Shift: From General Purpose to Specialized Agents</h2><p>What our users are really previewing is the next phase of enterprise AI: a shift away from one giant general-purpose model handling everything, toward fleets of specialized agents each doing one thing exceptionally well.</p><p>The frontier model becomes the orchestrator. The SLMs do the work.</p><p>If you want to find where SLMs fit in your stack, don&#8217;t start with your hardest problems. Start with your most repetitive ones. Find the tasks your team does a thousand times a week that need to be <em>correct</em>, not <em>clever</em>. That&#8217;s the frontier now.</p><p><strong>Bigger isn&#8217;t better. Specialized is.</strong></p>]]></content:encoded></item></channel></rss>