API Formats
/v1/messages, OpenAI Chat Completions at /v1/chat/completions, and OpenAI Responses at /v1/responses — plus an MCP JSON-RPC gateway at /mcp. It detects the format from the path, resolves the model to a provider, and translates to whatever that upstream expects.Last updated Reviewed against Prism v0.3.26
Inbound endpoints
| Endpoint | Method | Auth | Protocol |
|---|---|---|---|
| /v1/messages | POST | x-api-key or Bearer | Anthropic Messages |
| /v1/messages/count_tokens | POST | required | Anthropic-shaped 404 (counting is not supported) |
| /v1/chat/completions | POST | Bearer | OpenAI Chat Completions |
| /v1/responses | POST | Bearer | OpenAI Responses |
| /mcp | POST | Bearer or x-api-key | MCP JSON-RPC — every enabled server |
| /mcp/<agent> | POST | Bearer or x-api-key | MCP JSON-RPC — one agent's allowlist |
/api/chat. The Ollama-native API is an upstream format only: if you add a local Ollama server as a custom provider, Prism can talk to it in its native shape. Clients that speak only Ollama should be pointed at an OpenAI- or Anthropic-compatible endpoint instead, or at real Ollama.Utility endpoints
| Endpoint | Auth | What it returns |
|---|---|---|
| GET /v1/models | none | The model list built from your config. Intentionally unauthenticated so OpenAI-compatible clients that send no key during discovery still work. |
| GET /health | none | { "status": "ok" } |
| GET / | none | Service name and version. |
| GET /api/model-info?id=&provider= | none | Model metadata looked up from models.dev, or from Ollama Cloud when provider=ollama_cloud. |
| GET /v1/stats | none | Live stats as JSON. |
Authentication
Agents authenticate to Prism with a fixed shared token, prism. It is not a secret; it only proves the request is coming from an agent Prism is configured to serve. Upstream credentials never leave Prism.
| Protocol | Header |
|---|---|
| OpenAI / MCP | Authorization: Bearer prism |
| Anthropic | x-api-key: prism |
curl http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer prism" \
-d '{
"model": "glm-5.1:cloud",
"messages": [{"role": "user", "content": "Hello"}]
}'How routing resolves a request
Client request
│
▼ path + User-Agent identify the protocol
Detect client format
│
▼
Resolve model: alias → known model → default model
│
▼ each model names its provider
Pick upstream provider and its wire format
│
▼
Translate request → forward → translate response
│
▼
Stream events back in the client's protocolTranslation support
Tool calls, thinking / reasoning blocks, images and structured outputs are translated on every path, not just text.
| From (inbound) | To (upstream) | Notes |
|---|---|---|
| Anthropic Messages | OpenAI Chat Completions | The Claude Code path to non-Anthropic providers. |
| Anthropic Messages | Ollama native | For local Ollama servers added as custom providers. |
| OpenAI Chat Completions | Ollama native | Local model routing. |
| OpenAI Chat Completions | OpenAI Chat Completions | Pass-through to any OpenAI-compatible upstream. |
| OpenAI Responses | OpenAI Chat Completions | Lets Responses-only clients reach Chat-Completions providers. |
| Any | OpenAI Responses | Used for Codex OAuth models and any model with api: "responses". |
Client auto-detection
You do not declare which client you are. Prism reads the request path and the User-Agent header. Recognised clients include Claude Code, Codex, Cursor, GitHub Copilot, Factory Droid, OpenCode, Empryo, Hermes, DeepSeek Harness, Aider, Continue, Supermaven, Windsurf and Trae. Anything else shows up as its raw User-Agent, or as "Unknown".
X-Client-Name header to label a request yourself. It takes priority over User-Agent detection and appears as-is in the Stats Dashboard.Streaming
All three inference protocols stream over Server-Sent Events, including the full Responses API event sequence. Ollama uses NDJSON upstream, which Prism converts to SSE for the client. See Streaming for the event details.