For your home lab
Every computer you own, one private AI
The laptop, the desktop, the old gaming PC in the closet: Toskar pairs them, puts each model where it fits, and gives you one assistant and one API. Free, open source, and nothing leaves your network unless you say so.
Start with what you have
What your hardware can run
From the same estimate the app uses, on Toskar's model catalog. Check your exact machine with What can my computer run?
A laptop with 8 GB
Small, fast models for chat, summaries, and code help — most of Toskar's catalog runs comfortably.
A Mac or PC with 16 GB
Models up to 14B parameters, like Qwen 2.5 14B and Qwen 2.5 Coder 14B, run comfortably.
A gaming PC with a 24 GB card
The largest model in the catalog, Qwen 2.5 32B, fits entirely on the graphics card.
All of them, paired
Ask from the laptop and let the GPU box answer. Toskar places each request on a computer that can run it.
What to build
Weekend projects that keep running
Ask your own files
Point it at folders, PDFs, spreadsheets, and databases; answers cite where they came from.
Automations that run at home
"Every morning, check this and tell me if it changed": schedules and alerts to your phone with ntfy, email, or a webhook.
Your local AI everywhere
Use it from scripts over an OpenAI-compatible API, and from Claude Desktop, Claude Code, Cursor, or VS Code over MCP.
Plug in tools
Add MCP servers — folders, a browser, Notion, Linear, or anything else — usually by answering a question or two.
Train your own model
Turn examples and documents into a specialized model on Apple silicon or an NVIDIA GPU, and export it as one GGUF file.
See what it's doing
Live memory, speed, and placement per computer, a health check, and a list of everything that left your network.
Set it up
Two computers in four steps
Step 1
Install on each computer
Homebrew on a Mac, apt or a package on Linux, or the archive on Windows. Each one runs in the background.
Step 2
Let Toskar recommend models
It reads each computer's memory and graphics card and suggests models that fit, then downloads them.
Step 3
Pair them
On the Computers page, the other computer shows up on your network. Pair it once; Toskar trusts it from then on.
Step 4
Ask from anywhere on your network
Chat on either computer, or call the API from a script, and Toskar places the work where it runs best.
Questions
- Which operating systems does it run on?
- macOS (Apple silicon and Intel), Linux (x86-64 and ARM64, with apt, deb, and rpm packages), and Windows. On a server it runs headless with a web interface.
- Does it work without the internet?
- Yes, once the models are downloaded. Web search and connected services are off until you turn them on, and Settings lists everything that ever left.
- How is it different from running llama.cpp or Ollama myself?
- Toskar runs llama.cpp underneath and adds what's around it: model recommendations for your hardware, pairing several computers with placement across them, an assistant with your documents, tools, memory, and automations, an OpenAI-compatible API with per-key limits, and MCP in both directions.
- Can I use a model that isn't in the catalog?
- Yes. Install any GGUF model from a URL, including from Hugging Face.
- Is it really free?
- Toskar Core is free, open-source software under AGPL-3.0. Contributions and hardware reports are welcome on GitHub.
Install it tonight
One command on a Mac or Linux box, and it's running. Found a bug or have hardware we haven't seen? Tell us on GitHub. Why local AI has the bigger picture.