Reference

API Formats

Quick answer
Prism exposes three inference protocols — Anthropic Messages at /v1/messages, OpenAI Chat Completions at /v1/chat/completions, and OpenAI Responses at /v1/responses — plus an MCP JSON-RPC gateway at /mcp. It detects the format from the path, resolves the model to a provider, and translates to whatever that upstream expects.

Last updated Reviewed against Prism v0.3.26

Inbound endpoints

EndpointMethodAuthProtocol
/v1/messagesPOSTx-api-key or BearerAnthropic Messages
/v1/messages/count_tokensPOSTrequiredAnthropic-shaped 404 (counting is not supported)
/v1/chat/completionsPOSTBearerOpenAI Chat Completions
/v1/responsesPOSTBearerOpenAI Responses
/mcpPOSTBearer or x-api-keyMCP JSON-RPC — every enabled server
/mcp/<agent>POSTBearer or x-api-keyMCP JSON-RPC — one agent's allowlist
There is no inbound Ollama route
Prism does not serve /api/chat. The Ollama-native API is an upstream format only: if you add a local Ollama server as a custom provider, Prism can talk to it in its native shape. Clients that speak only Ollama should be pointed at an OpenAI- or Anthropic-compatible endpoint instead, or at real Ollama.

Utility endpoints

EndpointAuthWhat it returns
GET /v1/modelsnoneThe model list built from your config. Intentionally unauthenticated so OpenAI-compatible clients that send no key during discovery still work.
GET /healthnone{ "status": "ok" }
GET /noneService name and version.
GET /api/model-info?id=&provider=noneModel metadata looked up from models.dev, or from Ollama Cloud when provider=ollama_cloud.
GET /v1/statsnoneLive stats as JSON.

Authentication

Agents authenticate to Prism with a fixed shared token, prism. It is not a secret; it only proves the request is coming from an agent Prism is configured to serve. Upstream credentials never leave Prism.

ProtocolHeader
OpenAI / MCPAuthorization: Bearer prism
Anthropicx-api-key: prism
curl http://127.0.0.1:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer prism" \
  -d '{
    "model": "glm-5.1:cloud",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

How routing resolves a request

Client request
     │
     ▼  path + User-Agent identify the protocol
Detect client format
     │
     ▼
Resolve model: alias → known model → default model
     │
     ▼  each model names its provider
Pick upstream provider and its wire format
     │
     ▼
Translate request  →  forward  →  translate response
     │
     ▼
Stream events back in the client's protocol

Translation support

Tool calls, thinking / reasoning blocks, images and structured outputs are translated on every path, not just text.

From (inbound)To (upstream)Notes
Anthropic MessagesOpenAI Chat CompletionsThe Claude Code path to non-Anthropic providers.
Anthropic MessagesOllama nativeFor local Ollama servers added as custom providers.
OpenAI Chat CompletionsOllama nativeLocal model routing.
OpenAI Chat CompletionsOpenAI Chat CompletionsPass-through to any OpenAI-compatible upstream.
OpenAI ResponsesOpenAI Chat CompletionsLets Responses-only clients reach Chat-Completions providers.
AnyOpenAI ResponsesUsed for Codex OAuth models and any model with api: "responses".

Client auto-detection

You do not declare which client you are. Prism reads the request path and the User-Agent header. Recognised clients include Claude Code, Codex, Cursor, GitHub Copilot, Factory Droid, OpenCode, Empryo, Hermes, DeepSeek Harness, Aider, Continue, Supermaven, Windsurf and Trae. Anything else shows up as its raw User-Agent, or as "Unknown".

Override the client name
Send an X-Client-Name header to label a request yourself. It takes priority over User-Agent detection and appears as-is in the Stats Dashboard.

Streaming

All three inference protocols stream over Server-Sent Events, including the full Responses API event sequence. Ollama uses NDJSON upstream, which Prism converts to SSE for the client. See Streaming for the event details.