Browsing: Assistant

TL;DR — I replayed 27 real tasks from my own AI agent against two local models, one hardware upgrade apart, scoring both against the same frozen Claude baseline. On a single RTX 3090 capped at a 16K context, a 30B model scored 22.8/100 to Claude’s 89.4, and leaked malformed tool-call syntax into a quarter of…