AI providers
Connect Farabi to Ollama, LM Studio, llama.cpp, Claude, OpenAI, Gemini, Claude Code or Codex, keep several providers, and set model and effort.
Pro
Farabi and every AI feature are part of Gatesys Pro — free for 3 months, then $20 a year. See plans
Farabi runs on a model you choose. Eight runtimes share one code path: local model servers, cloud APIs, and agent CLIs you are already signed in to. Pro unlocks the features; the model is yours.
Choose a runtime
Open Settings › Farabi › Model › Inference provider and pick one. This is the default provider, which answers everything that does not pick its own model:
| Runtime | Kind | Default endpoint | You need | Where prompts go |
|---|---|---|---|---|
| Ollama | Local | http://127.0.0.1:11434 | Ollama running, with a chat model pulled | This computer |
| LM Studio | Local | http://127.0.0.1:1234 | LM Studio’s local server, with a model loaded | This computer |
| llama.cpp | Local | http://127.0.0.1:8080 | llama.cpp’s server, started with a model file | This computer |
| Claude | Cloud | https://api.anthropic.com | An API key from console.anthropic.com | Anthropic |
| OpenAI | Cloud | https://api.openai.com | An API key from platform.openai.com | OpenAI |
| Gemini | Cloud | Google’s API | An API key from aistudio.google.com | |
| Claude Code | Agent CLI | None | The claude CLI installed and signed in | Anthropic, under your sign-in |
| Codex | Agent CLI | None | The codex CLI installed and signed in | OpenAI, under your sign-in |
The status pill beside Inference provider reads Connected when the runtime answers. Otherwise it says what is missing: Key needed, Not installed, Sign in needed or Offline. Test connection checks again and shows the round-trip time.
Local runtimes
A local runtime keeps every prompt on this computer, at the endpoint you set. The defaults are loopback addresses.
Ollama
Install Ollama and pull a chat model:
ollama pull qwen3:8bIn Gatesys SSH, pick Ollama. Leave the endpoint as it is unless Ollama listens elsewhere.
Click Test connection, then choose the model under Installed models.
Gatesys SSH asks Ollama for a 4,096-token context window only when Ollama’s own could be smaller (a Modelfile below 4,096, or a server older than 0.6.7), so a prompt carrying host facts and the terminal tail is not quietly cut. A larger context you set is never overridden.
Ollama cloud models leave this computer
A model such as gpt-oss:120b-cloud, or any model the server lists with a remote host, is marked Runs on Ollama’s cloud — prompts leave this machine. Your local Ollama forwards its prompts to ollama.com, so Gatesys SSH treats it like a cloud provider, and footers say sent to Ollama’s cloud.
LM Studio
- In LM Studio, load a model and start its local server.
- In Gatesys SSH, pick LM Studio, click Test connection and choose the model.
llama.cpp
Start llama.cpp’s server with a model file, for example:
llama-server -m ./model.gguf --port 8080In Gatesys SSH, pick llama.cpp, click Test connection and choose the model.
A model server on another machine
You can point a local runtime at a GPU box on your network, or at a Gatesys SSH tunnel that leads to one. Settings then says the prompts go to that host, and Gatesys SSH treats them as leaving this computer, with the stricter masking described in What runs on the server.
Cloud APIs
- Pick Claude, OpenAI or Gemini. The endpoint is set to that vendor’s own API over HTTPS.
- Paste your API key and click Save. It is encrypted in the local vault and never written to the config file. Forget removes it.
- Choose a model under Available models.
Usage is billed to your own account with that vendor.
Gateways and proxies
The endpoint stays editable, for a gateway or proxy in front of the vendor. Whatever you type there receives the prompt and your key instead. Settings, footers and audit entries then name that host rather than the vendor, and add over plain HTTP when it is an http:// address on another machine.
Agent CLIs
Claude Code and Codex run as child processes under your own sign-in. The app holds no key, and the prompt goes to the CLI on standard input, never on its command line.
Install the CLI and sign in to it once:
npm install -g @anthropic-ai/claude-code npm install -g @openai/codexIn Gatesys SSH, pick Claude Code or Codex. The dot on the button shows whether the CLI was found. Detect again looks for it after you install it.
Choose a model. default uses whatever the CLI is set to. Claude Code also offers
sonnet,opusandhaiku; Codex lists the models from its own picker.
How the CLIs are kept from acting
Each CLI is made unable to act, not just asked not to:
- Claude Code runs with every tool off, no MCP servers, one turn and no saved session. It still loads your user settings, because its model and sign-in live there, and with them any hooks you configured there, which will see the prompt.
- Codex is checked before its first run. Gatesys SSH asks it what it has, switches off by name its shell tools, browser and computer use, apps, plugins, sub-agents, image generation, memories, tool suggestions, hooks and every enabled MCP server, starts it with no inherited environment, then asks again. If anything still reads as on, or a listing cannot be read, Codex is not run, and the message names what could not be switched off. It then runs in a read-only sandbox in an empty scratch folder.
- A tripwire watches both CLIs. The first tool call of any kind stops the CLI and everything it started, fails the request with a reason, and is written to the audit log.
More than one provider
Settings › Farabi › Model › Providers keeps a list of providers, up to twelve. Farabi’s panel switches between them per conversation; the default answers everything else: Explain, watch and setup drafts, Safe Change drafts and Test this model.
- Click Add a provider.
- Choose the runtime and, for a local server or a cloud API, its endpoint. A cloud API with no key saved yet asks for one; it goes to the local vault, never to the config file. An agent CLI needs nothing more: the app finds the CLI itself.
- Click Add provider.
Each row says what the app found there: how many models, Key needed, Not installed, Sign in needed or Offline, and where its prompts go. Make default makes a row the default provider, and the bin removes one; a saved API key stays in the vault. The last provider cannot be removed. The refresh button checks every provider again.
A cloud endpoint on another machine that is not the vendor’s own API, set by an edit of settings.json, reads Needs approval. Nothing is sent to it, key included, until you click Approve.
More than one provider is part of Pro
Without Pro, the list keeps every provider already in it, the default one answers, and Add a provider shows an upgrade card instead.
Model and effort
Settings › Farabi › Model sets the default provider’s model, and under it the Effort: how hard the model thinks before it answers. Default leaves it to the provider. Only the levels the chosen model takes are offered:
| Runtime | Effort levels |
|---|---|
| Claude | The levels the model lists, from low to max, with adaptive thinking where the model takes it |
| OpenAI | Minimal to high on GPT-5 models, low to high on the o series; none on other models |
| Gemini | Low to high on 2.5, minimal to high on 3 |
| Ollama | Think off and think on for a model that can think; low to high for gpt-oss |
| LM Studio, llama.cpp | None |
| Claude Code | Low to max; none for Haiku |
| Codex | The levels Codex lists for the model |
In Farabi’s panel, each conversation can pick its own model from any provider in the list, with its own effort. See Pick a model and effort.
Choosing a model
- A small local model is enough. Every feature is built to be useful with a model of about 8B parameters on a laptop, such as
qwen3:8b, because the model’s job is kept small: explain, rank and phrase, never decide. - No model still works. Without one, features fall back to their rules and deterministic halves, and say so.
- Measure rather than guess. Test this model scores your model against the rules on the app’s own test material.
- Embedding models are listed but cannot be picked for chat.
Where your prompt goes
Only the main process sends prompts: to the provider a Farabi conversation picked, and otherwise to the default provider. It dials only providers saved in the list, by their id, and refuses any other. The window itself never reaches the network.
| Provider | Leaves this computer? |
|---|---|
| Ollama, LM Studio, llama.cpp on a loopback endpoint | No |
| A local runtime on another machine, or through a Gatesys SSH tunnel | Yes |
| An Ollama cloud model | Yes, to ollama.com |
| Claude, OpenAI, Gemini | Yes, to the vendor or the endpoint you set |
| Claude Code, Codex | Yes, to Anthropic or OpenAI under your sign-in |
Everything that leaves gets stricter masking. See What runs on the server.
Something unclear or wrong? Tell us.