Groq runs inference on custom LPU hardware, delivering some of the fastest token generation in the industry. With throughput often exceeding 500 tokens per second on models like Llama 3.3 70B, Groq makes coding agents feel instant. The problem is that Claude Code only speaks the Anthropic Messages API, and Groq speaks OpenAI Chat Completions. They cannot talk directly.
Prism bridges that gap. By running Prism as a local proxy, you can route Claude Code's Anthropic-format requests to Groq's OpenAI-compatible endpoint. Prism translates the request, forwards it to Groq, and translates the streaming response back — all in real time.
Why Groq
Speed is the headline benefit. At 500+ tokens per second, Groq's LPU-based inference eliminates the waiting that makes slower providers feel sluggish. For coding agents that make many small requests, the latency difference is dramatic. Groq also offers a generous free tier, making it a great choice for everyday coding work.
Step by step
Get a Groq API key
gsk_.Add Groq as a custom provider in Prism
http://127.0.0.1:8765/admin, go to the Provider tab, and click Add custom provider. Set the base URL to https://api.groq.com/openai/v1 and paste your API key.Add models
models.dev. Click Auto-fill from models.dev and select the models you want, such as llama-3.3-70b-versatile.Set up Claude Code
Manual environment variables
If you prefer to configure Claude Code yourself:
$env:ANTHROPIC_BASE_URL = "http://127.0.0.1:11434"
$env:ANTHROPIC_AUTH_TOKEN = "prism"Model remapping
Claude Code requests Anthropic model names like claude-3-5-sonnet-20241022. Since Groq does not host Claude models, you need to remap those names to Groq models. In the Models tab or in model_remapping.json:
// %APPDATA%\prism\model_remapping.json
{
"default_model": "llama-3.3-70b-versatile",
"known_models": [
{ "id": "llama-3.3-70b-versatile", "provider": "custom_groq" },
{ "id": "llama-3.1-8b-instant", "provider": "custom_groq" }
],
"aliases": {
"claude-3-5-sonnet-20241022": "llama-3.3-70b-versatile",
"claude-3-5-haiku-20241022": "llama-3.1-8b-instant"
}
}Now when Claude Code asks for Sonnet, Prism sends the request to Groq with llama-3.3-70b-versatile instead. The fast Haiku tier maps to llama-3.1-8b-instantfor even lower latency on quick completions. To control the tiers directly rather than by alias, point Claude Code's ANTHROPIC_DEFAULT_*_MODEL variables at Groq model names — see the Claude Code docs.
Config example
The custom provider entry in Prism's config file looks like this:
// %APPDATA%\prism\config.json
{
"custom_providers": [
{
"id": "custom_groq_1a2b3c",
"name": "Groq",
"base_url": "https://api.groq.com/openai/v1",
"api_key": "gsk_your-key-here"
}
]
}model_remapping.json, under known_models, where each model names its provider. The Models tab writes both for you.Get started
Download Prism, add Groq as a custom provider, remap a model, and point Claude Code at it. In under two minutes you can be coding with Groq's ultra-fast inference through Claude Code.