Providers & Search

Provider Setup

Quick answer
Prism ships with two built-in providers — Ollama Cloud and OpenCode Go — accepts any number of custom OpenAI-compatible providers, and can authenticate to OpenAI/Codex through OAuth with no API key. Each model is assigned a provider, so a single agent session can mix models from different upstreams.

Last updated Reviewed against Prism v0.3.26

Which providers can Prism use?

ProviderCategoryHow Prism reaches it
Ollama CloudBuilt in, API keyOpenAI-compatible Chat Completions at https://ollama.com/v1/chat/completions
OpenCode GoBuilt in, API keyOpenAI-compatible Chat Completions at https://opencode.ai/zen/go
Custom providersUnlimited, API keyAny OpenAI-compatible base URL: OpenRouter, Groq, DeepSeek, Together AI, Cerebras, Vercel AI Gateway, and local servers such as Ollama, LM Studio, llama.cpp or vLLM
Codex / ChatGPTOAuth, no API keyhttps://chatgpt.com/backend-api/codex/responses
Ollama Cloud is not the native Ollama API
Ollama Cloud is reached over its OpenAI-compatible Chat Completions endpoint, not the native /api/chat surface. That path is where the cloud honours graded reasoning effort per model, where tool calls carry real ids, and where errors come back in the OpenAI shape. The native Ollama API is still supported as an upstream when you add a local Ollama server as a custom provider.

Configure a provider in the Admin UI

1

Open the Admin UI

Go to http://127.0.0.1:8765/admin and select the Provider tab. The Connect tab lists the live service URLs and the shared token if you need them.
2

Pick a default provider

Choose Ollama Cloud or OpenCode Go, or leave the default empty if every model has an explicit provider. The default is only a fallback for models with no provider of their own.
3

Enter the API key

Paste the key and click Save. Prism restarts the proxy automatically and the change is live within seconds.
4

Add custom providers

Click Add Provider for each additional upstream. Give it a name, the base URL, and the key. Prism creates an id such as custom_openrouter_ab12cd that you can reference from a model.
5

Add a Codex account instead (optional)

For OpenAI/Codex, go to the OAuth tab and click Add Codex Account. Your browser opens, you sign in, and Prism stores and refreshes the token itself. See OAuth Accounts.

Configure a provider in config.json

{
  "default_provider": "ollama_cloud",
  "ollama_cloud": {
    "id": "ollama_cloud",
    "name": "Ollama Cloud",
    "base_url": "https://ollama.com",
    "api_key": "your-ollama-cloud-key"
  },
  "opencode_go": {
    "id": "opencode_go",
    "name": "OpenCode Go",
    "base_url": "https://opencode.ai/zen/go",
    "api_key": "your-opencode-go-key"
  },
  "custom_providers": [
    {
      "id": "custom_groq_9f2c11",
      "name": "Groq",
      "base_url": "https://api.groq.com/openai/v1",
      "api_key": "gsk_..."
    },
    {
      "id": "custom_local_ollama",
      "name": "Local Ollama",
      "base_url": "http://localhost:11434",
      "api_key": "ollama"
    }
  ]
}
Key precedence
A key in config.json takes priority. If it is empty, Prism falls back to OLLAMA_API_KEY or OPENCODE_GO_API_KEY from the environment.

Provider-per-model routing

Routing is decided per model, not per session. Each entry in known_models names the provider that serves it, so Claude Code can run its main thread on one upstream and its sub-agents on another in the same session.

{
  "default_model": "glm-5.1:cloud",
  "known_models": [
    {
      "id": "glm-5.1:cloud",
      "provider": "ollama_cloud",
      "reasoning": true,
      "context_length": 128000,
      "max_output_tokens": 16384,
      "capabilities": { "tool_calling": true, "vision": true }
    },
    {
      "id": "openrouter/anthropic/claude-3.5-sonnet",
      "provider": "custom_openrouter_ab12cd",
      "context_length": 200000
    },
    {
      "id": "gpt-5-codex",
      "provider": "codex_ab12cd",
      "api": "responses"
    }
  ]
}

When a request arrives, Prism resolves the model name — aliases first, then known models, then the default model — and forwards to that model's provider. Nothing in the agent config changes.

Model capabilities and the API protocol

FieldPurpose
providerWhich upstream serves the model.
reasoningWhether the model accepts thinking / reasoning parameters.
reasoning_effortAllowed effort values. Prism normalizes invalid values and strips the parameter from non-reasoning models instead of returning an error.
context_lengthFills the model metadata agents display (ZCode, ZCode-style limits, Codex catalog).
max_output_tokensMaximum output tokens, also surfaced to agents.
capabilitiestool_calling, structured_outputs and vision, so agents only advertise what the model supports.
apichat_completions (default) or responses. Forces which protocol Prism speaks to that upstream, model by model.
Auto-fill from models.dev
In the Models tab, type a model name and click Search. Prism pulls context length, output limits, reasoning support, tool calling, vision and structured outputs from models.dev. For Ollama Cloud it asks the provider directly first, and the provider's own values win.

Reasoning effort is validated, not rejected

Clients send reasoning parameters in several dialects. Prism normalizes them: enabled, on and true become medium; disabled, off, false and none are dropped; an Anthropic thinking block is translated to reasoning_effort=medium. An effort value a model does not accept is clamped to its first allowed value, and the parameter is stripped entirely from non-reasoning models.