For developers
An OpenAI-compatible API on your own hardware
Point the OpenAI SDK at your own computers. No per-token bill, no rate limits but your hardware, and source code, customer data, and prompts that never leave the building.
What you get
Built to be called from code
An OpenAI-compatible API
Change the base URL and keep your SDK. Chat completions with streaming, JSON and JSON-schema output, and reasoning effort, answered by models on your hardware.
Your models in your editor
Claude Desktop, Claude Code, Cursor, and VS Code can ask your local models and search your connected docs over MCP. API Access gives the settings to paste.
A key per developer
Each key has its own limits on memory, knowledge, and tools, and can be rotated or revoked. Nothing needs a key on the computer itself.
Coding models and tools
Qwen 2.5 Coder and other coding models from the catalog, and a Programming profile that can read files, run commands, and use git when you allow it.
Several machines, one endpoint
Pair a GPU box with your laptop. Toskar places each request on the computer that can run it, behind the same address.
Open source, and yours to change
Toskar Core is AGPL-3.0 Go and TypeScript, with an OpenAPI spec. Read it, run it in CI, or build your own client against it.
Try it
Call it like you call OpenAI
Your local AI inside the apps you use
These apps can ask your local AI and search your connected knowledge over MCP. Toskar gives you the settings to paste.
- Claude Desktop
- Claude Code
- Cursor
- VS Code
- any MCP app
Anything that speaks the OpenAI API
Point any app or SDK that lets you set an OpenAI base URL at Toskar. Use auto as the model and Toskar picks one for each request. No key is needed on this computer.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:7331/v1", api_key="local")
reply = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Hello!"}],
)
print(reply.choices[0].message.content)With Toskar running, this prints a reply from a model on your computer.
For a team
One box for the whole team
Step 1
Put Toskar on a machine with memory
A Mac with 64 GB of unified memory, or a PC with 32 GB of memory and a 24 GB graphics card, runs every model in the catalog, coding models included.
Step 2
Open it to the network
Turn on local network access and give each developer a key, limited to what their tools need.
Step 3
Use it from everything
SDKs and scripts over the OpenAI API, editors over MCP, and the Toskar chat itself on that machine.
Good to know
Where it differs from OpenAI's API
- No function-calling passthrough yet. Tool definitions you send are not passed to the model; Toskar's own tools are configured per profile.
- Chat completions and models only. There is no embeddings, images, or audio endpoint.
- Some parameters are ignored.
temperatureandmax_tokensare accepted and not applied, and a model must be installed before a request can use it. - HTTP on the local network. Put a TLS proxy in front of it before reaching it from anywhere you don't trust.
The full list is in the API docs, and the code is on GitHub.
Questions
- Which OpenAI endpoints does Toskar support?
- GET /v1/models and POST /v1/chat/completions, with streaming, response_format (json_object and json_schema), and reasoning_effort. There is no embeddings endpoint yet; document search runs inside Toskar.
- Can I pass my own tools (function calling)?
- Not yet. Tool definitions in a request are not passed through; Toskar's own tools, set per profile, are what the assistant can use. Apps that depend on calling their own functions through the API won't work against it today.
- Can Cursor or Claude Code use Toskar as their main model?
- They use Toskar over MCP: as tools to ask your local models and search your docs from inside the editor. Running an editor's agent on top of a local model depends on the editor; without function-calling passthrough, agents that need it won't work yet.
- What hardware do coding models need?
- Qwen 2.5 Coder 14B runs comfortably on a Mac with 16 GB of unified memory or a PC with 16 GB of memory; the 7B and 1.5B versions run on less. The home page's "What can my computer run?" shows your computer.
- Is temperature or max_tokens honored?
- They are accepted and ignored for now. The API docs list every known difference from OpenAI's API.
Run it for your team
Install it yourself in a few minutes, or have Yeix size the machine, set up keys and a TLS proxy, and keep it running. See also why local AI.