Skip to content

For your home lab

Every computer you own, one private AI

The laptop, the desktop, the old gaming PC in the closet: Toskar pairs them, puts each model where it fits, and gives you one assistant and one API. Free, open source, and nothing leaves your network unless you say so.

Start with what you have

What your hardware can run

From the same estimate the app uses, on Toskar's model catalog. Check your exact machine with What can my computer run?

  • A laptop with 8 GB

    Small, fast models for chat, summaries, and code help — most of Toskar's catalog runs comfortably.

  • A Mac or PC with 16 GB

    Models up to 14B parameters, like Qwen 2.5 14B and Qwen 2.5 Coder 14B, run comfortably.

  • A gaming PC with a 24 GB card

    The largest model in the catalog, Qwen 2.5 32B, fits entirely on the graphics card.

  • All of them, paired

    Ask from the laptop and let the GPU box answer. Toskar places each request on a computer that can run it.

What to build

Weekend projects that keep running

  • Ask your own files

    Point it at folders, PDFs, spreadsheets, and databases; answers cite where they came from.

  • Automations that run at home

    "Every morning, check this and tell me if it changed": schedules and alerts to your phone with ntfy, email, or a webhook.

  • Your local AI everywhere

    Use it from scripts over an OpenAI-compatible API, and from Claude Desktop, Claude Code, Cursor, or VS Code over MCP.

  • Plug in tools

    Add MCP servers — folders, a browser, Notion, Linear, or anything else — usually by answering a question or two.

  • Train your own model

    Turn examples and documents into a specialized model on Apple silicon or an NVIDIA GPU, and export it as one GGUF file.

  • See what it's doing

    Live memory, speed, and placement per computer, a health check, and a list of everything that left your network.

Set it up

Two computers in four steps

  1. Step 1

    Install on each computer

    Homebrew on a Mac, apt or a package on Linux, or the archive on Windows. Each one runs in the background.

  2. Step 2

    Let Toskar recommend models

    It reads each computer's memory and graphics card and suggests models that fit, then downloads them.

  3. Step 3

    Pair them

    On the Computers page, the other computer shows up on your network. Pair it once; Toskar trusts it from then on.

  4. Step 4

    Ask from anywhere on your network

    Chat on either computer, or call the API from a script, and Toskar places the work where it runs best.

Questions

Which operating systems does it run on?
macOS (Apple silicon and Intel), Linux (x86-64 and ARM64, with apt, deb, and rpm packages), and Windows. On a server it runs headless with a web interface.
Does it work without the internet?
Yes, once the models are downloaded. Web search and connected services are off until you turn them on, and Settings lists everything that ever left.
How is it different from running llama.cpp or Ollama myself?
Toskar runs llama.cpp underneath and adds what's around it: model recommendations for your hardware, pairing several computers with placement across them, an assistant with your documents, tools, memory, and automations, an OpenAI-compatible API with per-key limits, and MCP in both directions.
Can I use a model that isn't in the catalog?
Yes. Install any GGUF model from a URL, including from Hugging Face.
Is it really free?
Toskar Core is free, open-source software under AGPL-3.0. Contributions and hardware reports are welcome on GitHub.

Install it tonight

One command on a Mac or Linux box, and it's running. Found a bug or have hardware we haven't seen? Tell us on GitHub. Why local AI has the bigger picture.