This week, Cooper and Yash (Neurometric AI) unpack TMax, a paper out of AI2 (the Allen Institute for AI) that takes a genuinely different approach to closing the gap between small open-weight models and frontier labs — not by training a bigger model, but by open-sourcing the entire recipe: data, code, and technique.
The paper focuses on terminal agents — models given nothing but command-line access as a tool, which turns out to be enough to tackle an enormous range of tasks, from software engineering to security work to scheduling. Using this recipe, AI2 boosted a Qwen model’s Terminal-Bench score from roughly 25% to 31%, closing meaningful ground on Claude Haiku’s roughly 33%.
What made this one worth an episode wasn’t just the score bump. It was the design choices behind it.
Watch on YouTube: https://youtu.be/4U0gqv9sSn4?si=ejv04gktGbTOMJM8
Listen: https://tokenengineering.podbean.com
Paper: TMax: A recipe for terminal agents — https://arxiv.org/abs/2606.23321

