Claude Code has a hard limitation: it only knows how to talk to the Anthropic Messages API. That means you are locked to Anthropic's models for the main thread, for quick completions and for every background sub-agent it spawns. If you want a cheaper daily driver, a faster-inference model, or a provider with different rate limits, Claude Code cannot do it on its own.
Prism solves this by acting as a local proxy between Claude Code and your upstream provider. When Claude Code sends an Anthropic Messages request, Prism translates it to the protocol that provider actually speaks — OpenAI Chat Completions or OpenAI Responses — and streams the response back in Anthropic's shape. Claude Code never knows the difference.
How it works
Claude Code
│ Anthropic Messages → http://127.0.0.1:11434
▼
┌──────────────────────┐
│ Prism proxy │ Anthropic ⇄ Chat Completions / Responses
│ /v1/messages │ per-model provider routing
└──────────────────────┘
│ forwards in the provider's native protocol
▼
Ollama Cloud · OpenCode Go · OpenRouter · Groq · DeepSeek · any OpenAI-compatible APIClaude Code sends all requests to one base URL. Point that at Prism instead of api.anthropic.com and every request flows through the proxy. Prism resolves the requested model to a provider, translates, and forwards. Streaming, thinking blocks, tool calls and system prompts all pass through.
Step by step
Install and run Prism
127.0.0.1:11434.Configure a provider
http://127.0.0.1:8765/admin and go to the Provider tab. Pick a built-in provider such as Ollama Cloud or OpenCode Go, or add a custom OpenAI-compatible provider with its base URL and API key.Add models
Connect Claude Code
~/.claude/settings.json and writes the environment keys for you, including a model per tier.Restart Claude Code
ANTHROPIC_BASE_URL at startup, so restart it before testing a prompt.What Prism writes
{
"env": {
"ANTHROPIC_BASE_URL": "http://127.0.0.1:11434",
"ANTHROPIC_AUTH_TOKEN": "prism",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.1:cloud",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "deepseek-v4-flash:cloud",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "deepseek-v4-flash:cloud",
"CLAUDE_CODE_SUBAGENT_MODEL": "deepseek-v4-flash:cloud"
}
}ANTHROPIC_AUTH_TOKEN. The value is the fixed literal prism— your upstream provider's real key goes in Prism's admin UI, never in Claude Code's config.Claude Code watches and rewrites ~/.claude/settings.json itself, so those keys can disappear. Prism re-writes them on every startup, which means restarting Prism recovers the integration.
Per-tier model mapping
Claude Code uses different models for different jobs: a strong model for the main thread, a fast one for quick completions, and a lightweight one for the sub-agents it spawns in the background. Prism exposes each as its own environment variable, so all four can point at different upstream models.
ANTHROPIC_DEFAULT_OPUS_MODEL → main thread (most capable) ANTHROPIC_DEFAULT_SONNET_MODEL → balanced default ANTHROPIC_DEFAULT_HAIKU_MODEL → fast background completions CLAUDE_CODE_SUBAGENT_MODEL → parallel sub-agents
This is the single biggest cost lever in a Claude Code setup, because background tasks usually outnumber the messages you type. Point haiku and subagent at a small fast model and the main thread at your strongest one.
Aliasing model names
If a client sends an Anthropic model name that is not in your catalog, Prism serves your default model. To control that mapping explicitly, add an alias under the aliases key in model_remapping.json:
{
"default_model": "glm-5.1:cloud",
"known_models": [
{ "id": "glm-5.1:cloud", "provider": "ollama_cloud" }
],
"aliases": {
"claude-3-5-sonnet-20241022": "glm-5.1:cloud",
"claude-3-5-haiku-20241022": "deepseek-v4-flash:cloud"
}
}When Claude Code asks for claude-3-5-sonnet-20241022, Prism swaps the name before forwarding and the client still sees its original model name in the response. See Model Remapping for the resolution order.
Supported providers
Anything OpenAI-compatible works, plus the two built-ins and Codex via OAuth:
Ollama Cloud
Built in. Managed Ollama models over an OpenAI-compatible endpoint, with an API key.
OpenCode Go
Built in. An API-key provider with its own model catalog.
OpenRouter, Groq, DeepSeek and anything else
Unlimited custom providers with a base URL and API key.
Codex via OAuth
Sign in with a ChatGPT account instead of pasting an API key. Routed over the Responses API.
/api/chat route; a local Ollama server can be used as an upstream provider instead.Get started
Install Prism, add a provider, add your models, and click Setup for Claude Code. In a few minutes you can be running it on models from Ollama Cloud, OpenCode Go, OpenRouter, Groq or DeepSeek — with no patches, no forks and no changes to Claude Code itself.