Model Remapping
Last updated Reviewed against Prism v0.3.26
How a model name is resolved
When a request arrives, Prism resolves the model in three ordered steps:
| Step | Source | Result |
|---|---|---|
| 1 | aliases | The requested name is swapped for its target and resolution continues. |
| 2 | known_models | The model is found in the catalog, which supplies its provider, limits, capabilities and API protocol. |
| 3 | default_model | Nothing matched, so the request goes to the default model and its provider. |
Client asks for "claude-3-5-haiku-20241022"
│
▼ aliases lookup
"deepseek-v4-flash:cloud"
│
▼ known_models lookup
provider = ollama_cloud, capabilities, limits
│
▼
Forward to Ollama Cloud over Chat Completions
│
▼ response
Translated back; the client still sees its original model nameThe model_remapping.json file
{
"default_model": "glm-5.1:cloud",
"known_models": [
{
"id": "glm-5.1:cloud",
"provider": "ollama_cloud",
"reasoning": true,
"reasoning_effort": ["low", "medium", "high"],
"context_length": 128000,
"max_output_tokens": 16384,
"capabilities": {
"tool_calling": true,
"structured_outputs": true,
"vision": true
},
"api": "chat_completions"
}
],
"aliases": {
"claude-3-5-sonnet-20241022": "glm-5.1:cloud",
"claude-3-5-haiku-20241022": "deepseek-v4-flash:cloud",
"gpt-4o": "glm-5.1:cloud"
}
}aliases, alongside default_model and known_models.Per-tier mapping for Claude Code
Claude Code uses a different model for the main conversation, for quick completions, and for background sub-agents. Prism exposes those as four tiers so each can point at a different upstream model:
| Tier | Used for | Typical choice |
|---|---|---|
| opus | The main thread — the most capable work | A strong reasoning model |
| sonnet | The balanced daily driver | A mid-tier model |
| haiku | Fast, cheap completions | A small fast model |
| subagent | Background and parallel tasks | A cheap model to keep cost down |
Tiers are configured once in the Admin UI's Agents tab when you set up Claude Code. The chosen models are written into its config as the tier variables.
{
"agent_integrations": {
"claude_code_tiers": {
"opus": "glm-5.1:cloud",
"sonnet": "deepseek-v4-flash:cloud",
"haiku": "deepseek-v4-flash:cloud",
"subagent": "deepseek-v4-flash:cloud"
}
}
}Provider-qualified model ids
A model's provider field is what decides routing. Provider ids are the built-in ollama_cloud and opencode_go, a custom id such as custom_groq_9f2c11, or a Codex account id such as codex_ab12cd. Two models with the same name can coexist as long as their provider differs.
Choosing the upstream protocol per model
The api field overrides which protocol Prism uses to talk to the upstream for that model: chat_completions (the default) or responses. This is what lets a Responses-only model such as a Codex model sit alongside Chat-Completions models in one catalog.
Edit remappings in the Admin UI
Open the Models tab
http://127.0.0.1:8765/admin and select Models.Search and add a model
Assign a provider and protocol
Add an alias
Common uses
| Goal | How |
|---|---|
| Cut cost without touching agent config | Alias an expensive model name to a cheaper one. |
| Keep Claude Code unchanged | Alias its Anthropic model names to your upstream models. |
| Split spend across tiers | Map opus, sonnet, haiku and subagent to different models. |
| Test a provider | Repoint a single alias and compare results. |
| Serve Responses-only models | Add the model with api: "responses" and a Codex provider. |
Related
Provider Setup
Add the providers these models route to.
Claude Code
Per-tier mappings in the one-click setup flow.
Config Reference
Every field in model_remapping.json.
API Formats
What chat_completions and responses mean upstream.
Guide: per-model API protocol
Why one provider can need Responses for one model and Chat Completions for another.