Skip to content

API reference

Endpoints, errors and limits

All endpoints live under https://infermesh.dev/v1. Authenticate with Authorization: Bearer <key> or x-api-key: <key>. Keys carry scopes; an endpoint rejects keys without its scope.

Endpoints

Method & pathScopeDescription
POST /v1/chat/completionsinferenceOpenAI-compatible chat completions for every model, streaming or not.
POST /v1/messagesinferenceAnthropic Messages API for Claude models, including prompt caching.
POST /v1/messages/count_tokensinferenceAnthropic token counting (free).
GET /v1/models—Model catalog with list prices (no key needed).
GET /v1/models/{id}—A single model by id or alias.
GET /v1/account/creditsaccountBalance, reserved amount, per-source buckets, tier and this key's limits.
GET /v1/usage?window=24haccountRequests, tokens, spend and latency for 1h, 24h, 7d or 30d, by model.
GET /v1/keys/nonce?wallet=…—Message to sign for wallet-authenticated key creation.
POST /v1/keys/generatekeys or wallet signatureCreate an API key programmatically (see below).
POST /v1/mcpmcpMCP server (Streamable HTTP, JSON-RPC).
GET /v1/mcp/statusmcpBalance, agent policy, recent top-ups and gateway health.
GET /v1/mcp/balancemcpCompact balance check for agents.
GET /v1/health—Liveness.

Creating keys programmatically

With an existing key that has the keys scope, child keys inherit its account; they can't gain scopes, exceed the parent's spend cap or its rate limit. Without a key, sign the message from /v1/keys/nonce with your wallet.

curl https://infermesh.dev/v1/keys/generate \
  -H "Authorization: Bearer sk-infer-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "name": "ci-runner", "scopes": ["inference"], "spend_cap": "25", "spend_cap_period": "daily", "rate_limit_rpm": 120 }'

The response contains key exactly once. Scopes: inference, account, keys, mcp.

Errors

OpenAI-style endpoints return { "error": { "message", "type", "param", "code" } }; /v1/messages returns Anthropic's { "type": "error", "error": { "type", "message" } }. Both include the InferMesh code and machine-readable extras.

HTTPcodeMeaning & what to do
401invalid_api_keyMissing, unknown, paused or revoked key. The body includes replenish_url where you can create a key and add credits.
402insufficient_creditsBalance can't cover the prompt. Includes balance_micros, required_micros and replenish_url.
402spend_cap_exceededThis key hit its daily, monthly or lifetime ceiling. Includes remaining_micros.
403scope_deniedThe key lacks the scope this endpoint needs.
404model_not_foundUnknown model id. List models with GET /v1/models.
429rate_limitedPer-key rate limit. Retry after retry_after_ms (also sent as retry-after).
429ip_blockedToo many failed authentications from your IP; temporarily blocked.
4xxprovider errorThe provider rejected the request (for example an invalid parameter). Returned exactly as the provider sent it; nothing is charged.
502upstream_errorThe provider's stream was interrupted or empty; retry.
503upstream_unavailableEvery route for the model failed or is unavailable; nothing is charged. Retry with backoff.

When credits run out mid-stream

A streaming request that exhausts your balance is cut off cleanly: the gateway stops reading from the provider and sends a final error event, then ends the stream.

OpenAI-style stream
data: {"error":{"message":"Credits ran out mid-stream; the response was truncated. Add credits at https://infermesh.dev/compute","type":"insufficient_credits","param":null,"code":"insufficient_credits","replenish_url":"https://infermesh.dev/compute"}}

data: [DONE]

Anthropic-style streams receive the same error as an event: error. Clients should treat it like any stream error; the partial output already received is valid.

Rate limits and spend caps

  • Each key has a requests-per-minute limit (default 600, up to 10,000) enforced as a token bucket that allows bursts of a quarter of the per-minute limit (at least 5 requests).
  • Spend caps apply per UTC day, per UTC month or for the key's lifetime, and count in-flight reservations so concurrent requests can't overshoot.
  • Agent keys are capped at the agent's daily spend limit.

Failover

If a provider returns an error before streaming starts, the gateway retries the next route for the model. A route that keeps failing is skipped for a short cool-down (circuit breaker) and probed again afterwards.