Internals

Chat Completions vs Responses: Setting Protocol per Model

By Kavin M KPublished Updated 8 min read

Quick answer
Give a model the field api: "responses" in your Prism catalog and it is served over /v1/responses upstream; leave it unset or set chat_completions for everything else. Agents that bind one protocol per provider get two provider entries written for them, so both kinds of model stay selectable.

"OpenAI-compatible" is a promise that covers two different APIs. Almost every provider implements /v1/chat/completions. Some implement /v1/responses instead, usually because that is the endpoint their newest reasoning models were built against. A few implement both, with different features on each.

Prism decides which one to call per model, not per provider and not globally. One config entry switches a single model onto the Responses API while its neighbours stay on Chat Completions.

Two directions, four shapes

It helps to separate what Prism accepts from what it sends, because the two are independent.

Inbound (what agents send)          Outbound (what Prism calls)
──────────────────────────          ──────────────────────────
POST /v1/messages        Anthropic  →  chat_completions   (default)
POST /v1/chat/completions OpenAI    →  responses         (per-model opt-in)
POST /v1/responses       OpenAI      →  provider-specific handling
                                       when the provider is Codex OAuth

Claude Code sends Anthropic Messages. Codex sends Responses. Cursor, OpenCode, Zed and most others send Chat Completions. Each inbound shape can land on either outbound API, which is the point: your agent's protocol does not have to match your provider's.

Setting it

The field is api on a known model. In the admin UI it is a control in the Models tab; in the file it looks like this:

// %APPDATA%\prism\model_remapping.json
{
  "default_model": "z-ai/glm-4.6",
  "known_models": [
    {
      "id": "z-ai/glm-4.6",
      "provider": "custom_openrouter_7g8h9i",
      "reasoning": true,
      "reasoning_effort": ["low", "medium", "high"],
      "context_length": 200000,
      "max_output_tokens": 128000,
      "capabilities": { "tool_calling": true, "vision": false }
    },
    {
      "id": "muse-spark",
      "provider": "custom_local_2b3c4d",
      "api": "responses",
      "capabilities": { "tool_calling": true }
    }
  ]
}

Leaving api out means chat_completions. It is the safe default because it is the shape nearly every OpenAI-compatible endpoint implements.

Codex providers are always Responses
Accounts added through OAuth under Codex are Responses-only, so Prism sets api: "responses" for their models automatically. Setting it by hand to chat_completions on a Codex model is not a valid configuration — the upstream does not serve that route.

How to tell which one a provider wants

1

Try Responses first if the model is a reasoner

Newer reasoning models from providers that have moved to the Responses API usually expose their thinking on /v1/responses and hide it on Chat Completions.
2

Ask the endpoint

Send a minimal request to the provider's /v1/responses route. A 404 means it is not implemented; a 400 about an unsupported field means it is.
3

Set api and watch the log

Turn on Debug Logs, make one request, and read which endpoint Prism called and what the upstream answered. This is the fastest way to settle the question.
curl -i https://api.provider.example/v1/responses \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"their-model","input":"hi"}'

What actually differs

Behaviourchat_completionsresponses
Message arraymessages[]input[] with typed items
StreamingSSE deltas per choiceSSE with typed event names
Reasoning outputreasoning_content or nonereasoning items
Tool callstool_calls in a choice deltafunction_call items
StatefulnessStatelessA response id can be referenced

Prism translates both directions, event by event, without buffering — a streamed response stays streamed. Anthropic thinking blocks map to the provider's reasoning fields and back, and tool-call arguments are relayed as fragments so the client can assemble them exactly as it does from a native provider.

The search interaction worth remembering

Prism's web-search interception applies to Anthropic web_search and web_fetch calls, and to Responses web_search and x_search calls — except on models configured with api: "responses", and except on Codex OAuth accounts. On those routes the provider's own search runs and your configured providers are not consulted.

So if you switched a model to Responses to get its reasoning output and your searches suddenly cost money or disappear, that is why. The two features are independent, and the protocol choice decides whether local search can answer at all.

Aliases are a separate layer

The api field sits on a known model. Aliases map a name an agent sends onto one of those models, and resolution runs aliases → known models → default model. An alias does not carry protocol information itself; it points at a model that has it.

{
  "default_model": "z-ai/glm-4.6",
  "known_models": [{ "id": "muse-spark", "provider": "custom_local_2b3c4d", "api": "responses" }],
  "aliases": {
    "claude-sonnet-4-5-20250929": "z-ai/glm-4.6",
    "claude-opus-4-1-20250805":   "muse-spark"
  }
}

With that config, a Claude Code request for claude-opus-4-1-20250805 is translated from Anthropic Messages into the Responses API, while claude-sonnet-4-5-20250929 goes out as Chat Completions. Two tiers, two protocols, one agent, one config file.

Protocol is not the same as capability
reasoning, reasoning_effort, context_length, max_output_tokens and capabilities describe what a model can do; api describes how to ask it. Setting reasoning: true on a model whose route never returns reasoning does not create the behaviour — it only tells Prism to expect it.

Reference

Every claim on this page is checked against the Prism source, and the reference documentation is where those details live in full.

Try it yourself

Prism is a free, MIT-licensed local proxy for AI coding agents. Install it, point one agent at http://127.0.0.1:11434, and the rest of this post applies as written.