Do two small models beat one big one? We measured it
· The Toskar team · 4 min read
A popular idea for getting more out of small local models: have several of them answer the same question, compare the answers, and let them check each other. If three drafts agree, you can probably trust the answer. If they disagree, something is worth a second look.
Toskar 1.8 can do this. We call it Deliberate. Before turning it on for anyone automatically, we measured whether it actually helps.
How Deliberate works
When a profile has Deliberate on:
- Three drafts. The model writes its answer, and two more drafts are written without seeing it, at different settings so they don't just repeat each other. With paired computers, the drafts can run on different machines at once.
- A quick vote. When the answer is short, such as a number or a name, Toskar compares the three. If most agree, that answer is kept, and nothing more runs.
- Checks and a judge. When the drafts disagree, or the answers are too long to compare, each draft is checked against the others for claims, disagreements, and likely mistakes. A judge then writes the answer, and says so where the drafts still disagree.
You can see all of it under the answer: each draft, which model and computer wrote it, its final answer, the checks, and the judge.
What we measured
We wrote 39 questions that each have one right answer:
- Math: word problems, such as splitting a bill or how long a leaking tank takes to fill.
- Logic: puzzles, such as who finished third in a race.
- Two-step facts: for example, "In which century was the author of Pride and Prejudice born?"
- Trick questions: such as the bat and the ball.
We ran each question on four models, first answering alone, then with Deliberate. Everything ran on one desktop with an AMD Radeon RX 7900 graphics card (24 GB). We counted right answers, and the time and tokens each question took.
| Model | Alone | With Deliberate | Time | Tokens |
|---|---|---|---|---|
| Qwen 2.5 7B | 37 of 39 (1.0 s) | 38 of 39 (2.4 s) | 2.3× | 2.8× |
| Gemma 2 9B | 35 of 39 (0.9 s) | 33 of 39 (3.1 s) | 3.4× | 3.9× |
| Qwen 2.5 14B | 37 of 39 (1.9 s) | 37 of 39 (4.2 s) | 2.3× | 2.9× |
| Qwen 2.5 32B | 37 of 39 (3.8 s) | 39 of 39 (8.4 s) | 2.2× | 2.9× |
What we found
Overall, about even, at two to four times the cost. Across all four models, 146 of 156 answers were right alone, and 147 with Deliberate. Every question took two to four times as long and used about three times the tokens.
A small model checking itself caught up with a big one. Qwen 2.5 7B with Deliberate got 38 of 39 right in 2.4 seconds a question. Qwen 2.5 32B, answering once, got 37 right in 3.8 seconds. If a 32B model doesn't fit your computer, Deliberate can get a 7B close to it on questions like these.
But one model's drafts can share its mistakes. On Gemma 2 9B, Deliberate made things worse: two logic puzzles it got right alone came out wrong, because its other drafts agreed on the wrong answer. Three drafts from the same model are like asking one person the same question three times: they tend to make the same slip.
Logic puzzles are where it matters. The models were already nearly perfect at math, facts, and trick questions. Every change, for better or worse, was on a logic puzzle.
It also found a mistake in our own test. Qwen 2.5 32B answered "one kilogram" to "Which is heavier: a kilogram of feathers or a kilogram of steel?". That's right, but our scoring only accepted "neither" or "the same". We fixed the scoring.
What we did with it
Deliberate stays off by default, and Auto doesn't turn it on. The numbers don't show a kind of question where it reliably helps enough to be worth the extra time on most computers.
It's still there for anyone who wants it. Turn it on per profile, in Profiles & Orchestration → Strategy and computers → Compare independent drafts. It's most worth trying on a small model, for questions where getting it right matters more than getting it fast.
What's next
The design expects drafts from different model families to catch more than one model at three settings, say a Qwen, a Gemma, and a Llama answering together. A Qwen slip that Gemma doesn't make is exactly what the vote should catch. Our test machine loads one model at a time, so that's the next thing we'll measure. We're also writing harder logic questions, since the rest of the set is close to its ceiling.
We'll publish those numbers here too.
Keep reading
- What can my computer run? A plain guide to local AI and memoryMemory decides which AI models your computer can run. How much each model size needs, what fits in 8, 16, 32, and 64 GB, and Mac vs. PC.October 5, 2026 · 4 min read
- Toskar vs. ChatGPT for business: a private alternativeChatGPT's business plans charge per person each month and run in OpenAI's cloud. Toskar runs on computers you own. Here's an honest look at the trade-offs.October 7, 2026 · 3 min read
- Toskar vs. LM Studio: which local AI app fits you?Both run AI models on your own computer for free. LM Studio is a polished app for one machine; Toskar adds documents, automations, and your other computers.October 7, 2026 · 3 min read