Maton

Vercel AI Gateway

Access the Vercel AI Gateway with managed API key authentication. Browse the model catalog, run chat, responses, messages, and embeddings requests across providers, and check credits and usage.

Reference

This app is Vercel's inference gateway. For projects, deployments, and domains, use Vercel. Every path needs the /v1 prefix, and model IDs are always {creator}/{model}, for example anthropic/claude-haiku-4.5.

Inference requests are billed to the connected Vercel account. Check a model's pricing before sending many requests.

List Models

GET /vercel-ai-gateway/v1/models

Returns the whole catalog. limit and type query parameters are ignored, so filter the results yourself.

Get Model

GET /vercel-ai-gateway/v1/models/{creator}/{model}

List Model Endpoints

GET /vercel-ai-gateway/v1/models/{creator}/{model}/endpoints

Lists every provider that serves the model, with per-provider pricing, limits, and live uptime and latency.

Chat Completions

POST /vercel-ai-gateway/v1/chat/completions
Content-Type: application/json

{
  "model": "anthropic/claude-haiku-4.5",
  "messages": [
    { "role": "user", "content": "Say hello in five words." }
  ],
  "max_tokens": 100
}

Add "stream": true for server-sent events ending in data: [DONE]. provider_metadata.gateway.routing.finalProvider names the provider that served the request.

Responses

POST /vercel-ai-gateway/v1/responses
Content-Type: application/json

{
  "model": "openai/gpt-4o-mini",
  "input": "Say hello in five words."
}

output is an array of typed items. Filter by type rather than reading output[0], since reasoning models emit a reasoning item first.

Messages

POST /vercel-ai-gateway/v1/messages
Content-Type: application/json

{
  "model": "anthropic/claude-haiku-4.5",
  "max_tokens": 100,
  "messages": [
    { "role": "user", "content": "Say hello in five words." }
  ]
}

Anthropic's Messages API shape. max_tokens is required here.

Embeddings

POST /vercel-ai-gateway/v1/embeddings
Content-Type: application/json

{
  "model": "openai/text-embedding-3-small",
  "input": ["first string", "second string"]
}

Only models whose type is embedding work here.

Get Credit Balance

GET /vercel-ai-gateway/v1/credits

balance and total_used are decimal strings in USD.

Get Generation Usage

GET /vercel-ai-gateway/v1/generation?id={generationId}

Cost and token usage for one completed request, using the gen_... ID from an inference response. Usage is recorded a few seconds after the request, so an immediate lookup can return 404; retry after a short wait.

Notes:

  • Inference depends on the Vercel account's state:
Account stateStatustype
No card on file403customer_verification_required
Card on file, free credits429 when throttledrate_limit_exceeded
Paid credits200—
  • The free-tier 429 is account-wide and comes from Vercel, without a Retry-After header. Back off for a few minutes.
  • An unknown model, or one without the {creator}/ prefix, returns 403. Check model IDs against /v1/models first.
  • Token usage, reasoning output, and error shapes differ between /chat/completions, /responses, /messages, and /embeddings, so parse each route separately.
  • /v1/report (spend reports) needs a paid Vercel plan and returns 403 otherwise. Use /v1/generation or /v1/credits instead.
  • None of the endpoints are paginated.

Resources

On this page