"OpenAI-compatible" is a promise that covers two different APIs. Almost every provider implements /v1/chat/completions. Some implement /v1/responses instead, usually because that is the endpoint their newest reasoning models were built against. A few implement both, with different features on each.
Prism decides which one to call per model, not per provider and not globally. One config entry switches a single model onto the Responses API while its neighbours stay on Chat Completions.
Two directions, four shapes
It helps to separate what Prism accepts from what it sends, because the two are independent.
Inbound (what agents send) Outbound (what Prism calls)
────────────────────────── ──────────────────────────
POST /v1/messages Anthropic → chat_completions (default)
POST /v1/chat/completions OpenAI → responses (per-model opt-in)
POST /v1/responses OpenAI → provider-specific handling
when the provider is Codex OAuthClaude Code sends Anthropic Messages. Codex sends Responses. Cursor, OpenCode, Zed and most others send Chat Completions. Each inbound shape can land on either outbound API, which is the point: your agent's protocol does not have to match your provider's.
Setting it
The field is api on a known model. In the admin UI it is a control in the Models tab; in the file it looks like this:
// %APPDATA%\prism\model_remapping.json
{
"default_model": "z-ai/glm-4.6",
"known_models": [
{
"id": "z-ai/glm-4.6",
"provider": "custom_openrouter_7g8h9i",
"reasoning": true,
"reasoning_effort": ["low", "medium", "high"],
"context_length": 200000,
"max_output_tokens": 128000,
"capabilities": { "tool_calling": true, "vision": false }
},
{
"id": "muse-spark",
"provider": "custom_local_2b3c4d",
"api": "responses",
"capabilities": { "tool_calling": true }
}
]
}Leaving api out means chat_completions. It is the safe default because it is the shape nearly every OpenAI-compatible endpoint implements.
api: "responses" for their models automatically. Setting it by hand to chat_completions on a Codex model is not a valid configuration — the upstream does not serve that route.How to tell which one a provider wants
Try Responses first if the model is a reasoner
/v1/responses and hide it on Chat Completions.Ask the endpoint
/v1/responses route. A 404 means it is not implemented; a 400 about an unsupported field means it is.Set api and watch the log
curl -i https://api.provider.example/v1/responses \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"model":"their-model","input":"hi"}'What actually differs
| Behaviour | chat_completions | responses |
|---|---|---|
| Message array | messages[] | input[] with typed items |
| Streaming | SSE deltas per choice | SSE with typed event names |
| Reasoning output | reasoning_content or none | reasoning items |
| Tool calls | tool_calls in a choice delta | function_call items |
| Statefulness | Stateless | A response id can be referenced |
Prism translates both directions, event by event, without buffering — a streamed response stays streamed. Anthropic thinking blocks map to the provider's reasoning fields and back, and tool-call arguments are relayed as fragments so the client can assemble them exactly as it does from a native provider.
The search interaction worth remembering
Prism's web-search interception applies to Anthropic web_search and web_fetch calls, and to Responses web_search and x_search calls — except on models configured with api: "responses", and except on Codex OAuth accounts. On those routes the provider's own search runs and your configured providers are not consulted.
So if you switched a model to Responses to get its reasoning output and your searches suddenly cost money or disappear, that is why. The two features are independent, and the protocol choice decides whether local search can answer at all.
Aliases are a separate layer
The api field sits on a known model. Aliases map a name an agent sends onto one of those models, and resolution runs aliases → known models → default model. An alias does not carry protocol information itself; it points at a model that has it.
{
"default_model": "z-ai/glm-4.6",
"known_models": [{ "id": "muse-spark", "provider": "custom_local_2b3c4d", "api": "responses" }],
"aliases": {
"claude-sonnet-4-5-20250929": "z-ai/glm-4.6",
"claude-opus-4-1-20250805": "muse-spark"
}
}With that config, a Claude Code request for claude-opus-4-1-20250805 is translated from Anthropic Messages into the Responses API, while claude-sonnet-4-5-20250929 goes out as Chat Completions. Two tiers, two protocols, one agent, one config file.
reasoning, reasoning_effort, context_length, max_output_tokens and capabilities describe what a model can do; api describes how to ask it. Setting reasoning: true on a model whose route never returns reasoning does not create the behaviour — it only tells Prism to expect it.