Access the Vercel AI Gateway with managed API key authentication. Browse the model catalog, run chat, responses, messages, and embeddings requests across providers, and check credits and usage.
Reference
This app is Vercel's inference gateway. For projects, deployments, and domains, use Vercel. Every path needs the /v1 prefix, and model IDs are always {creator}/{model}, for example anthropic/claude-haiku-4.5.
Inference requests are billed to the connected Vercel account. Check a model's pricing before sending many requests.
List Models
GET /vercel-ai-gateway/v1/modelsReturns the whole catalog. limit and type query parameters are ignored, so filter the results yourself.
Get Model
GET /vercel-ai-gateway/v1/models/{creator}/{model}List Model Endpoints
GET /vercel-ai-gateway/v1/models/{creator}/{model}/endpointsLists every provider that serves the model, with per-provider pricing, limits, and live uptime and latency.
Chat Completions
POST /vercel-ai-gateway/v1/chat/completions
Content-Type: application/json
{
"model": "anthropic/claude-haiku-4.5",
"messages": [
{ "role": "user", "content": "Say hello in five words." }
],
"max_tokens": 100
}Add "stream": true for server-sent events ending in data: [DONE]. provider_metadata.gateway.routing.finalProvider names the provider that served the request.
Responses
POST /vercel-ai-gateway/v1/responses
Content-Type: application/json
{
"model": "openai/gpt-4o-mini",
"input": "Say hello in five words."
}output is an array of typed items. Filter by type rather than reading output[0], since reasoning models emit a reasoning item first.
Messages
POST /vercel-ai-gateway/v1/messages
Content-Type: application/json
{
"model": "anthropic/claude-haiku-4.5",
"max_tokens": 100,
"messages": [
{ "role": "user", "content": "Say hello in five words." }
]
}Anthropic's Messages API shape. max_tokens is required here.
Embeddings
POST /vercel-ai-gateway/v1/embeddings
Content-Type: application/json
{
"model": "openai/text-embedding-3-small",
"input": ["first string", "second string"]
}Only models whose type is embedding work here.
Get Credit Balance
GET /vercel-ai-gateway/v1/creditsbalance and total_used are decimal strings in USD.
Get Generation Usage
GET /vercel-ai-gateway/v1/generation?id={generationId}Cost and token usage for one completed request, using the gen_... ID from an inference response. Usage is recorded a few seconds after the request, so an immediate lookup can return 404; retry after a short wait.
Notes:
- Inference depends on the Vercel account's state:
| Account state | Status | type |
|---|---|---|
| No card on file | 403 | customer_verification_required |
| Card on file, free credits | 429 when throttled | rate_limit_exceeded |
| Paid credits | 200 | — |
- The free-tier
429is account-wide and comes from Vercel, without aRetry-Afterheader. Back off for a few minutes. - An unknown model, or one without the
{creator}/prefix, returns403. Check model IDs against/v1/modelsfirst. - Token usage, reasoning output, and error shapes differ between
/chat/completions,/responses,/messages, and/embeddings, so parse each route separately. /v1/report(spend reports) needs a paid Vercel plan and returns403otherwise. Use/v1/generationor/v1/creditsinstead.- None of the endpoints are paginated.