Skip to content
← Blog

Do two small models beat one big one? We measured it

· The Toskar team · 4 min read

A popular idea for getting more out of small local models: have several of them answer the same question, compare the answers, and let them check each other. If three drafts agree, you can probably trust the answer. If they disagree, something is worth a second look.

Toskar 1.8 can do this. We call it Deliberate. Before turning it on for anyone automatically, we measured whether it actually helps.

How Deliberate works

When a profile has Deliberate on:

  1. Three drafts. The model writes its answer, and two more drafts are written without seeing it, at different settings so they don't just repeat each other. With paired computers, the drafts can run on different machines at once.
  2. A quick vote. When the answer is short, such as a number or a name, Toskar compares the three. If most agree, that answer is kept, and nothing more runs.
  3. Checks and a judge. When the drafts disagree, or the answers are too long to compare, each draft is checked against the others for claims, disagreements, and likely mistakes. A judge then writes the answer, and says so where the drafts still disagree.

You can see all of it under the answer: each draft, which model and computer wrote it, its final answer, the checks, and the judge.

What we measured

We wrote 39 questions that each have one right answer:

  • Math: word problems, such as splitting a bill or how long a leaking tank takes to fill.
  • Logic: puzzles, such as who finished third in a race.
  • Two-step facts: for example, "In which century was the author of Pride and Prejudice born?"
  • Trick questions: such as the bat and the ball.

We ran each question on four models, first answering alone, then with Deliberate. Everything ran on one desktop with an AMD Radeon RX 7900 graphics card (24 GB). We counted right answers, and the time and tokens each question took.

Model Alone With Deliberate Time Tokens
Qwen 2.5 7B 37 of 39 (1.0 s) 38 of 39 (2.4 s) 2.3× 2.8×
Gemma 2 9B 35 of 39 (0.9 s) 33 of 39 (3.1 s) 3.4× 3.9×
Qwen 2.5 14B 37 of 39 (1.9 s) 37 of 39 (4.2 s) 2.3× 2.9×
Qwen 2.5 32B 37 of 39 (3.8 s) 39 of 39 (8.4 s) 2.2× 2.9×

What we found

Overall, about even, at two to four times the cost. Across all four models, 146 of 156 answers were right alone, and 147 with Deliberate. Every question took two to four times as long and used about three times the tokens.

A small model checking itself caught up with a big one. Qwen 2.5 7B with Deliberate got 38 of 39 right in 2.4 seconds a question. Qwen 2.5 32B, answering once, got 37 right in 3.8 seconds. If a 32B model doesn't fit your computer, Deliberate can get a 7B close to it on questions like these.

But one model's drafts can share its mistakes. On Gemma 2 9B, Deliberate made things worse: two logic puzzles it got right alone came out wrong, because its other drafts agreed on the wrong answer. Three drafts from the same model are like asking one person the same question three times: they tend to make the same slip.

Logic puzzles are where it matters. The models were already nearly perfect at math, facts, and trick questions. Every change, for better or worse, was on a logic puzzle.

It also found a mistake in our own test. Qwen 2.5 32B answered "one kilogram" to "Which is heavier: a kilogram of feathers or a kilogram of steel?". That's right, but our scoring only accepted "neither" or "the same". We fixed the scoring.

What we did with it

Deliberate stays off by default, and Auto doesn't turn it on. The numbers don't show a kind of question where it reliably helps enough to be worth the extra time on most computers.

It's still there for anyone who wants it. Turn it on per profile, in Profiles & Orchestration → Strategy and computers → Compare independent drafts. It's most worth trying on a small model, for questions where getting it right matters more than getting it fast.

What's next

The design expects drafts from different model families to catch more than one model at three settings, say a Qwen, a Gemma, and a Llama answering together. A Qwen slip that Gemma doesn't make is exactly what the vote should catch. Our test machine loads one model at a time, so that's the next thing we'll measure. We're also writing harder logic questions, since the rest of the set is close to its ceiling.

We'll publish those numbers here too.

Keep reading