BUILDINGtool

AI Gateway probe

Call every model an account can reach through Vercel AI Gateway on OIDC alone, and find where the free tier stops you.

CREATED AUG 24, 2026 · UPDATED AUG 24, 2026

PROBLEM

Every AI Gateway quickstart ends the same way: install the SDK, paste a model slug, run it. None of them say what happens when the account behind the token is not entitled to the model in the snippet, which is the case most people are actually in. The refusal arrives as a 403 buried six frames down a stack trace, and it reads like a broken setup rather than a billing tier.

HYPOTHESIS

That authentication would be the hard part and model access would be automatic. Exactly backwards.

ARCHITECTURE

A single TypeScript file run through Node's type stripping, no build step. It asks the gateway for the credit balance and the model catalogue to prove the token authenticated, then walks a candidate list smallest model first and stops at the first one that answers. Refusals are classified rather than counted: 403 and 429 mean opposite things and only one of them is worth waiting out. maxRetries is 0 on purpose, because the SDK's default backoff wraps a 429 in a RetryError and hides the status code that decides whether to wait.

STACK

TypeScriptAI SDK 7Vercel AI GatewayNode

RESULT

OIDC alone works. vercel env pull mints a VERCEL_OIDC_TOKEN and the SDK authenticates on it with no API key present anywhere in the project, for both catalogue reads and generation. 327 models are reachable. Entitlement, not auth, is the wall: on the free monthly credit grant gpt-5-nano answers immediately, gpt-5.6-luna returns 429 (permitted but throttled hard), and gpt-5.5 returns 403 (not permitted at all). A Vercel Pro plan does not change this, because gateway credits bill separately from the plan.

MEASURED

metricvaluebaseline
cost of one streamed answer$0.00077 · gpt-5-nano, roughly 6,500 calls inside the grant$5.00 free monthly grant
models reachable through the gateway327 · catalogue call, authenticated on OIDC alone—
API keys required01 per provider without the gateway
largest OpenAI model the free grant servesgpt-5-nano · gpt-5.6-luna is permitted but throttled to unusablegpt-5.5 refused with 403

LEARNINGS

  • 403 and 429 are not the same answer. A 403 will never clear by waiting; a 429 might. Collapsing both into "it failed" is what turns a five-minute setup into an afternoon.
  • Probe smallest first. The cheapest model that answers proves the whole pipeline, and finding it costs a rounding error.
  • Space the probes. Firing them back to back trips the free-tier limiter, and then you are measuring your own impatience instead of the account. The first version of this spent four minutes convincing itself a working model was broken.
  • A stale AI_GATEWAY_API_KEY takes priority over the OIDC token and shadows it silently, so an unset variable is part of the setup.
  • The plan tier and the gateway tier are different things, and the error message says "free tier" while the dashboard says Pro.

PROVENANCE

Built with Claude Code in a terminal. The 403 vs 429 distinction came out of running the probe badly first: fifteen calls in a burst exhausted the limiter, which made a working model look restricted.

TAGS

ai-gatewayoidcauthcost