Provider Setup
Last updated Reviewed against Prism v0.3.26
Which providers can Prism use?
| Provider | Category | How Prism reaches it |
|---|---|---|
| Ollama Cloud | Built in, API key | OpenAI-compatible Chat Completions at https://ollama.com/v1/chat/completions |
| OpenCode Go | Built in, API key | OpenAI-compatible Chat Completions at https://opencode.ai/zen/go |
| Custom providers | Unlimited, API key | Any OpenAI-compatible base URL: OpenRouter, Groq, DeepSeek, Together AI, Cerebras, Vercel AI Gateway, and local servers such as Ollama, LM Studio, llama.cpp or vLLM |
| Codex / ChatGPT | OAuth, no API key | https://chatgpt.com/backend-api/codex/responses |
/api/chat surface. That path is where the cloud honours graded reasoning effort per model, where tool calls carry real ids, and where errors come back in the OpenAI shape. The native Ollama API is still supported as an upstream when you add a local Ollama server as a custom provider.Configure a provider in the Admin UI
Open the Admin UI
http://127.0.0.1:8765/admin and select the Provider tab. The Connect tab lists the live service URLs and the shared token if you need them.Pick a default provider
Enter the API key
Add custom providers
custom_openrouter_ab12cd that you can reference from a model.Add a Codex account instead (optional)
Configure a provider in config.json
{
"default_provider": "ollama_cloud",
"ollama_cloud": {
"id": "ollama_cloud",
"name": "Ollama Cloud",
"base_url": "https://ollama.com",
"api_key": "your-ollama-cloud-key"
},
"opencode_go": {
"id": "opencode_go",
"name": "OpenCode Go",
"base_url": "https://opencode.ai/zen/go",
"api_key": "your-opencode-go-key"
},
"custom_providers": [
{
"id": "custom_groq_9f2c11",
"name": "Groq",
"base_url": "https://api.groq.com/openai/v1",
"api_key": "gsk_..."
},
{
"id": "custom_local_ollama",
"name": "Local Ollama",
"base_url": "http://localhost:11434",
"api_key": "ollama"
}
]
}config.json takes priority. If it is empty, Prism falls back to OLLAMA_API_KEY or OPENCODE_GO_API_KEY from the environment.Provider-per-model routing
Routing is decided per model, not per session. Each entry in known_models names the provider that serves it, so Claude Code can run its main thread on one upstream and its sub-agents on another in the same session.
{
"default_model": "glm-5.1:cloud",
"known_models": [
{
"id": "glm-5.1:cloud",
"provider": "ollama_cloud",
"reasoning": true,
"context_length": 128000,
"max_output_tokens": 16384,
"capabilities": { "tool_calling": true, "vision": true }
},
{
"id": "openrouter/anthropic/claude-3.5-sonnet",
"provider": "custom_openrouter_ab12cd",
"context_length": 200000
},
{
"id": "gpt-5-codex",
"provider": "codex_ab12cd",
"api": "responses"
}
]
}When a request arrives, Prism resolves the model name — aliases first, then known models, then the default model — and forwards to that model's provider. Nothing in the agent config changes.
Model capabilities and the API protocol
| Field | Purpose |
|---|---|
| provider | Which upstream serves the model. |
| reasoning | Whether the model accepts thinking / reasoning parameters. |
| reasoning_effort | Allowed effort values. Prism normalizes invalid values and strips the parameter from non-reasoning models instead of returning an error. |
| context_length | Fills the model metadata agents display (ZCode, ZCode-style limits, Codex catalog). |
| max_output_tokens | Maximum output tokens, also surfaced to agents. |
| capabilities | tool_calling, structured_outputs and vision, so agents only advertise what the model supports. |
| api | chat_completions (default) or responses. Forces which protocol Prism speaks to that upstream, model by model. |
models.dev. For Ollama Cloud it asks the provider directly first, and the provider's own values win.Reasoning effort is validated, not rejected
Clients send reasoning parameters in several dialects. Prism normalizes them: enabled, on and true become medium; disabled, off, false and none are dropped; an Anthropic thinking block is translated to reasoning_effort=medium. An effort value a model does not accept is clamped to its first allowed value, and the parameter is stripped entirely from non-reasoning models.
Related
Model Remapping
Aliases, per-tier mappings and how routing resolves.
OAuth Accounts
Add a Codex account and route it without an API key.
Config Reference
Every config.json key, environment variable and flag.
API Formats
The endpoints agents call and how translation works.
Guide: Prism vs LiteLLM
When a local agent proxy beats a shared gateway, and how to run both.
Guide: Groq
Add Groq as a custom provider and route Claude Code to it.
Guide: DeepSeek
Cheap, capable models for everyday agent work.
Guide: OpenRouter
Hundreds of models behind one key.