Reference

FAQ

Quick answer
Prism is a single-binary local proxy that routes AI coding agents to any LLM provider. It translates Anthropic, OpenAI and Ollama protocols, brokers an MCP gateway, and includes free web search. These are the questions people ask most often after installing it.

Last updated Reviewed against Prism v0.3.26

Frequently asked questions

If your question is not here, the Troubleshooting page covers failure modes, and the rest of the docs cover each feature in depth.

What API formats does Prism support?

Prism accepts Anthropic Messages at /v1/messages, OpenAI Chat Completions at /v1/chat/completions, OpenAI Responses at /v1/responses, and MCP JSON-RPC at /mcp and /mcp/<agent>. It translates between them and each upstream's native format in real time. There is no inbound Ollama /api/chat route; Ollama's native API is only used upstream.

Which upstream providers can Prism use?

Prism ships with two built-in providers, Ollama Cloud and OpenCode Go, and accepts any number of custom OpenAI-compatible providers such as OpenRouter, Groq, DeepSeek, Together AI, Cerebras, or a local Ollama, LM Studio, llama.cpp or vLLM server. You can also sign in with an OpenAI/ChatGPT account via OAuth and use Codex with no API key.

Is Ollama Cloud reached through the native Ollama API?

No. Prism reaches Ollama Cloud over its OpenAI-compatible Chat Completions endpoint, not the native /api/chat surface. That path honours graded reasoning effort per model, returns tool calls with real ids, and reports errors in the OpenAI shape. The native Ollama API is supported upstream when you add a local Ollama server as a custom provider.

What is the MCP gateway?

Prism re-exposes Model Context Protocol servers to your agents behind one Prism-authenticated endpoint. /mcp aggregates every enabled server and /mcp/<agent> serves a single agent's allowlist. Prism supports stdio, Streamable HTTP and HTTP+SSE transports, holds the upstream credentials, and can broker OAuth by opening your browser the first time an agent uses a server that is not authorised yet.

Which web search providers does Prism support?

Prism ships a managed local SearXNG instance that needs no API key, and also supports Exa, Tavily, Brave and Serper with an API key, plus user-defined declarative REST providers. You can configure a fallback chain so a failing provider is skipped automatically.

Do I need Python or Docker to run Prism?

No. Prism is a single binary with no runtime dependencies. Python is only relevant to the optional managed SearXNG instance, which downloads its own isolated interpreter if your machine has no Python 3.11 or newer.

What API key do clients use to reach Prism?

Use "prism" as the API key for every client. Send it as Authorization: Bearer prism for OpenAI-style and MCP requests, or x-api-key: prism for Anthropic-style requests. It is a fixed local token, not a secret. Your real upstream keys are stored in Prism.

How do I add an upstream provider?

Open the Admin UI at http://127.0.0.1:8765/admin, go to the Provider tab, and add the provider with its base URL and API key. You can also edit config.json directly. Prism hot-reloads the file within seconds.

Does Prism support streaming responses?

Yes. Every inference path streams over Server-Sent Events, including the full OpenAI Responses event sequence. Thinking and reasoning blocks, tool calls with incremental arguments, and image content all stream correctly.

Can I use Prism with my ChatGPT account?

Yes. Add a Codex account from the Admin UI OAuth tab. Prism stores the tokens, refreshes them automatically, and shows session and weekly usage. Requests route to chatgpt.com/backend-api/codex rather than api.openai.com, which avoids Cloudflare restrictions.

How does model routing decide which provider to use?

Prism resolves the requested model name in order: aliases first, then known models, then the default model. Each known model names its provider, so routing is decided per model and a single session can mix providers. A model can also override which protocol Prism speaks upstream with an api field of chat_completions or responses.

What is the default model?

The default model is whatever you set as default_model in model_remapping.json. A fresh install points at an Ollama Cloud model. If a client requests a model that matches no alias or known model, Prism routes it to the default.

Which AI coding agents does Prism integrate with?

Prism has one-click setup for 14 agents: Claude Code, Codex Desktop and CLI, Cursor, Factory Droid, OpenCode, ZCode, Grok Build, Zed, Pi, Oh My Pi, Kimi Code, Prime Agent, Empryo, Hermes and DeepSeek Harness. Claude Desktop, the OpenAI SDK and generic Ollama-compatible clients work with manual configuration.

Is request content stored on disk?

No. Only metadata such as token counts, model names, timing and the detected client name is written to a local SQLite database. Prompts and responses are never persisted.

Does Prism send telemetry?

Prism sends one anonymous heartbeat per calendar day so the maintainer can count active installs. It carries six fields: a random id, version, OS, architecture, whether Prism was used that day, and a coarse request bucket. It never includes prompts, responses, file paths or keys. You can disable it in the Admin UI Proxy tab or with PRISM_ANALYTICS_DISABLED=1.

Can Prism start automatically when I log in?

Yes, on Windows, macOS and Linux. Enable Start at Login in the Admin UI Proxy tab or the system tray menu. Prism uses a per-user mechanism on each platform, so it needs no administrator rights or system service.

Does Prism work on Windows, macOS and Linux?

Yes. Windows ships a per-user MSI installer and a portable executable, macOS ships a notarised DMG with an app bundle, and Linux ships a portable AppImage and a tarball. The native system tray, auto-start and one-click agent setup all work on every platform.

What happens if Prism crashes?

Prism runs as two processes: the tray hosts the Admin UI and can restart the proxy if it exits, and the headless proxy serves requests. Stats use SQLite in WAL mode, so they survive a crash. If the tray process itself exits, run the binary again.