Squeezy vs external baseline (Mini)
Both on gpt-5.4-mini, median of 10 runs. Every language came in cheaper at equal-or-better recall.
Two boards pit Squeezy against an external coding-agent baseline on the same code-understanding tasks: same prompt, same model, same grader. Squeezy answers from its local knowledge of the code, and at the median it spent 60 cents for every dollar the baseline spent.
Both on gpt-5.4-mini, median of 10 runs. Every language came in cheaper at equal-or-better recall.
Both on claude-haiku-4-5, median of 10 runs. Every language came in cheaper at equal-or-better recall.
Ratio is Squeezy's cost divided by the baseline's: 0.60 means Squeezy spent 60 cents where the baseline spent a dollar. Cost only counts when the answer passes the board's recall rule, so a cheap wrong answer can't win. The 15 rows are the general-purpose programming languages in this suite; markup and documentation formats such as CSS, HTML, and Markdown are not scored. C and C++, and JavaScript and TypeScript, count as separate rows.
This is a fixed publication snapshot built from medians of 10 runs. Each baseline uses the same repository question, grader, model tier, and pricing assumptions.
Each row compares answers to the same real-world code question, scored by the same grader.
Every figure is the median of 10 runs, because single agent runs vary in both cost and recall.
These results come from this task suite on these model tiers. Your own bill depends on the models you pick, your provider's pricing, and the shape of your work.