Skip to content
← Blog

What can my computer run? A plain guide to local AI and memory

· The Toskar team · 4 min read

The first question about running AI on your own computer is always the same: will it run? The answer comes down to one number: memory.

A model has to fit in memory while it works, and it needs some extra room for the conversation it's having. If it fits with room to spare, it runs well. If it barely fits, it runs, but it's slow and leaves little room for anything else. If it doesn't fit, it doesn't run.

If you'd rather skip the reading, the tool on our home page answers it for your computer in two clicks, using the same math as the app.

What a model needs

Models are described by their size in parameters: a "7B" model has about 7 billion. Toskar's catalog uses compressed (quantized) versions, usually labeled Q4_K_M. Each parameter takes a little over half a byte, so the file is about a quarter of the full-size model's, at a small cost in quality.

Toskar adds up what a model needs while it runs:

  • The model itself, about the size of its download.
  • Room for the conversation, which grows with how much the model keeps in mind at once. Toskar plans for 8,192 tokens, about 6,000 words.
  • A little working space for the program that runs it: half a gigabyte.

Here's that total for models in Toskar's catalog:

Model size Examples Download Memory it needs
1–2B Llama 3.2 1B, Qwen 2.5 Coder 1.5B, Gemma 2 2B 0.7–1.6 GB about 1.3–2.2 GB
3–4B Llama 3.2 3B, Qwen 2.5 3B, Phi 3.5 Mini 2 GB about 2.5–3 GB
7B Mistral 7B, Qwen 2.5 7B, DeepSeek R1 Distill 7B 4–4.5 GB about 5–5.5 GB
9B Gemma 2 9B 5.4 GB about 6.5 GB
14B Qwen 2.5 14B, Qwen 2.5 Coder 14B 8.4 GB about 10 GB
32B Qwen 2.5 32B 18.6 GB about 21 GB

How Toskar grades the fit

Toskar compares that number with your computer's memory and gives each model a grade:

  • Excellent fit: memory is at least 1.5 times what the model needs. Plenty to spare.
  • Good fit: at least 1.2 times. It should run comfortably.
  • Tight fit: about the same as it needs. It runs, but close other big apps.
  • May run slowly: a bit less than it needs. It may lean on swap, or need a shorter conversation window.

What fits in common amounts of memory

Your memory Runs well Runs, but slowly
8 GB Everything up to 7B, and Gemma 2 9B 14B models
16 GB Everything up to 14B Qwen 2.5 32B
32 GB Everything in the catalog, including 32B —
64 GB and up Everything, with room to run more than one model —

For most people, a 7B model is where local AI starts to feel genuinely useful, and 14B is a clear step up in quality. With 32 GB or more, the 32B models come within reach.

Macs and PCs count memory differently

A Mac with Apple silicon (M1 and later) shares one pool of memory between the processor and the graphics, and the graphics side does the AI work. All of it counts. That's why a Mac with plenty of memory is one of the easiest ways to run larger models.

A PC has system memory (RAM) and, if it has a graphics card, separate memory on the card (VRAM). Toskar grades the fit against system memory. The graphics card decides the speed:

  • If the whole model fits in the card's memory, the card does the work, and it's fast. Toskar calls this a good GPU fit.
  • If the model is bigger than the card's memory, part of it runs on the card and the rest on the processor. It works, but it's slower: partial GPU offload.
  • With no graphics card, the processor does everything. Small models are fine, and larger ones are slow.

More than one computer

You don't have to fit everything on one machine. Toskar connects several computers into one system and sends each request to a computer that can run the model. An older laptop can handle quick questions while a desktop with more memory takes the big model.

Try it

Check your computer on the home page, or install Toskar. It looks at your hardware and recommends models that fit, so you never have to do this math yourself.