squeezy vs external baseline (Mini)
Both sides on gpt-5.4-mini, median of 10 runs. Every language came in cheaper at the baseline's recall or better.
one recorded measurement · medians of 10 runs, same grader on both sides
show me the numbers, not the adjectives.
both boards
Both sides on gpt-5.4-mini, median of 10 runs. Every language came in cheaper at the baseline's recall or better.
Both sides on Claude Haiku 4.5, median of 10 runs. Every language came in cheaper at the baseline's recall or better.
Add up all 15 tasks on a board and squeezy's total is 43% less across the Mini board, 56% less across Haiku.
Ratio is squeezy's cost divided by the baseline's, so 0.60 means squeezy spent 60 cents where the baseline spent a dollar. A language counts as a win only if squeezy found at least as much as the baseline did and came in at least 5% cheaper, so a cheap wrong answer can't win. These 15 languages are the ones this suite scores. squeezy indexes more than 40 languages and formats; the full list lives in the language docs.
mini board · gpt-5.4-mini on both sides
haiku board · Claude Haiku 4.5 on both sides
Both agents read the same repository at the same commit, answered the same question, and were scored by the same grader. On the Mini board both sides are also pinned to the same reasoning effort -- how hard the model is allowed to think before it answers -- because that setting moves cost on its own. The baseline's cost is recomputed from its own token counts against the price list squeezy prices itself with, so the ratio compares two agents rather than two vendors' invoices. squeezy's figure includes whatever it handed to a subagent.
Each row is one hard code-navigation question, asked of a real codebase pinned to a commit -- nginx, Terraform, Laravel and Akka among them. List every subtype of a type, trace a call graph, classify how each implementation is written. The grader scores how many of the required symbols the answer actually found.
Agent runs vary in both cost and what they find, so each number is the median of 10 runs. A run that returned no cost at all, from a timeout or a failed launch, is dropped rather than scored as a free win.
This is one task suite on two model tiers, recorded once. Your own bill depends on the models you pick, your provider's rates, and the shape of your work.