Claude Code

How to Use Any LLM Provider with Claude Code and Prism

By Kavin M KPublished Updated 9 min read

Quick answer
Set ANTHROPIC_BASE_URL to http://127.0.0.1:11434 and ANTHROPIC_AUTH_TOKEN to prism in ~/.claude/settings.json, then restart Claude Code. Prism receives Anthropic Messages requests and forwards them to whichever provider and protocol your chosen model uses.

Claude Code has a hard limitation: it only knows how to talk to the Anthropic Messages API. That means you are locked to Anthropic's models for the main thread, for quick completions and for every background sub-agent it spawns. If you want a cheaper daily driver, a faster-inference model, or a provider with different rate limits, Claude Code cannot do it on its own.

Prism solves this by acting as a local proxy between Claude Code and your upstream provider. When Claude Code sends an Anthropic Messages request, Prism translates it to the protocol that provider actually speaks — OpenAI Chat Completions or OpenAI Responses — and streams the response back in Anthropic's shape. Claude Code never knows the difference.

How it works

  Claude Code
       │  Anthropic Messages → http://127.0.0.1:11434
       ▼
  ┌──────────────────────┐
  │  Prism proxy          │  Anthropic ⇄ Chat Completions / Responses
  │  /v1/messages         │  per-model provider routing
  └──────────────────────┘
       │  forwards in the provider's native protocol
       ▼
  Ollama Cloud · OpenCode Go · OpenRouter · Groq · DeepSeek · any OpenAI-compatible API

Claude Code sends all requests to one base URL. Point that at Prism instead of api.anthropic.com and every request flows through the proxy. Prism resolves the requested model to a provider, translates, and forwards. Streaming, thinking blocks, tool calls and system prompts all pass through.

Step by step

1

Install and run Prism

Grab the installer from the latest release — a per-user MSI on Windows, a DMG on macOS, an AppImage on Linux — or run the portable binary. A tray icon appears and the proxy starts listening on 127.0.0.1:11434.
2

Configure a provider

Open the admin UI at http://127.0.0.1:8765/admin and go to the Provider tab. Pick a built-in provider such as Ollama Cloud or OpenCode Go, or add a custom OpenAI-compatible provider with its base URL and API key.
3

Add models

In the Models tab, add the models you want. Prism can auto-fill limits and capabilities from models.dev for many providers, so you rarely type model names by hand.
4

Connect Claude Code

In the Agents tab, click Setup next to Claude Code. Prism backs up ~/.claude/settings.json and writes the environment keys for you, including a model per tier.
5

Restart Claude Code

It reads ANTHROPIC_BASE_URL at startup, so restart it before testing a prompt.

What Prism writes

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://127.0.0.1:11434",
    "ANTHROPIC_AUTH_TOKEN": "prism",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.1:cloud",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "deepseek-v4-flash:cloud",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "deepseek-v4-flash:cloud",
    "CLAUDE_CODE_SUBAGENT_MODEL": "deepseek-v4-flash:cloud"
  }
}
Use ANTHROPIC_AUTH_TOKEN, not ANTHROPIC_API_KEY
Prism expects ANTHROPIC_AUTH_TOKEN. The value is the fixed literal prism— your upstream provider's real key goes in Prism's admin UI, never in Claude Code's config.

Claude Code watches and rewrites ~/.claude/settings.json itself, so those keys can disappear. Prism re-writes them on every startup, which means restarting Prism recovers the integration.

Per-tier model mapping

Claude Code uses different models for different jobs: a strong model for the main thread, a fast one for quick completions, and a lightweight one for the sub-agents it spawns in the background. Prism exposes each as its own environment variable, so all four can point at different upstream models.

ANTHROPIC_DEFAULT_OPUS_MODEL   → main thread (most capable)
ANTHROPIC_DEFAULT_SONNET_MODEL → balanced default
ANTHROPIC_DEFAULT_HAIKU_MODEL  → fast background completions
CLAUDE_CODE_SUBAGENT_MODEL     → parallel sub-agents

This is the single biggest cost lever in a Claude Code setup, because background tasks usually outnumber the messages you type. Point haiku and subagent at a small fast model and the main thread at your strongest one.

Aliasing model names

If a client sends an Anthropic model name that is not in your catalog, Prism serves your default model. To control that mapping explicitly, add an alias under the aliases key in model_remapping.json:

{
  "default_model": "glm-5.1:cloud",
  "known_models": [
    { "id": "glm-5.1:cloud", "provider": "ollama_cloud" }
  ],
  "aliases": {
    "claude-3-5-sonnet-20241022": "glm-5.1:cloud",
    "claude-3-5-haiku-20241022": "deepseek-v4-flash:cloud"
  }
}

When Claude Code asks for claude-3-5-sonnet-20241022, Prism swaps the name before forwarding and the client still sees its original model name in the response. See Model Remapping for the resolution order.

Supported providers

Anything OpenAI-compatible works, plus the two built-ins and Codex via OAuth:

Ollama Cloud

Built in. Managed Ollama models over an OpenAI-compatible endpoint, with an API key.

OpenCode Go

Built in. An API-key provider with its own model catalog.

OpenRouter, Groq, DeepSeek and anything else

Unlimited custom providers with a base URL and API key.

Codex via OAuth

Sign in with a ChatGPT account instead of pasting an API key. Routed over the Responses API.

No /api/chat involved
Prism's inbound API is Anthropic Messages plus OpenAI Chat Completions and Responses. It does not serve Ollama's native /api/chat route; a local Ollama server can be used as an upstream provider instead.

Get started

Install Prism, add a provider, add your models, and click Setup for Claude Code. In a few minutes you can be running it on models from Ollama Cloud, OpenCode Go, OpenRouter, Groq or DeepSeek — with no patches, no forks and no changes to Claude Code itself.

Download Prism

Reference

Every claim on this page is checked against the Prism source, and the reference documentation is where those details live in full.

Try it yourself

Prism is a free, MIT-licensed local proxy for AI coding agents. Install it, point one agent at http://127.0.0.1:11434, and the rest of this post applies as written.