benchmarks

The saving, measured.

Two boards pit Squeezy against an external coding-agent baseline on the same code-understanding tasks: same prompt, same model, same grader. Squeezy answers from its local knowledge of the code, and at the median it spent 60 cents for every dollar the baseline spent.

Mini board 15 / 15

Squeezy vs external baseline (Mini)

Both on gpt-5.4-mini, median of 10 runs. Every language came in cheaper at equal-or-better recall.

Haiku board 15 / 15

Squeezy vs external baseline (Haiku)

Both on claude-haiku-4-5, median of 10 runs. Every language came in cheaper at equal-or-better recall.

language ratios

Cost ratio by language.

Ratio is Squeezy's cost divided by the baseline's: 0.60 means Squeezy spent 60 cents where the baseline spent a dollar. Cost only counts when the answer passes the board's recall rule, so a cheap wrong answer can't win. The 15 rows are the general-purpose programming languages in this suite; markup and documentation formats such as CSS, HTML, and Markdown are not scored. C and C++, and JavaScript and TypeScript, count as separate rows.

Mini tier

15 language tasks

Haiku tier

15 language tasks

This is a fixed publication snapshot built from medians of 10 runs. Each baseline uses the same repository question, grader, model tier, and pricing assumptions.

same task

Prompt and repo are matched

Each row compares answers to the same real-world code question, scored by the same grader.

median

Multiple runs reduce noise

Every figure is the median of 10 runs, because single agent runs vary in both cost and recall.

scope

What the numbers cover

These results come from this task suite on these model tiers. Your own bill depends on the models you pick, your provider's pricing, and the shape of your work.

GitHub

Repository access is under construction.

Squeezy's repository is not public yet. The product site and documentation are available here in the meantime.