API Reference

The complete contract for the AltLLM inference API: every public endpoint, every request field, the exact response and streaming schemas, the model catalog, and what “OpenAI-compatible” does and does not cover.

Overview

AltLLM serves an OpenAI-compatible inference API over HTTPS. Requests are JSON, responses are JSON or server-sent events, and every billed call is metered against your account balance.
Base URL
https://api.altllm.ai/v1
PropertyValue
Content-Typeapplication/json on every request with a body.
AuthorizationBearer sk-alt-… — required on everything except model discovery, provider metadata, and health.
X-Request-IDOptional. Any correlation ID you send is echoed back on the response and recorded in AltLLM’s logs.
Streamingtext/event-stream when stream: true.
Rate limitsPer account, per model, over a 60-second sliding window.

One reference, one set of facts

Model IDs, context windows, output ceilings, prices, and per-model parameter behavior on this page are generated from the gateway’s own catalog — the same data GET /v1/models serves. A CI check fails the build if this page and the gateway disagree.

Authentication

Every authenticated request carries a Portal API key as a bearer token. Keys are issued in the Portal and map to exactly one account.
Authorization: Bearer sk-alt-your_key_here
PropertyDetail
Formatsk-alt- followed by 48 URL-safe characters (55 characters total).
CreationPortal → API Keys. The full key is shown once, at creation, and is never retrievable again — only its 12-character prefix.
RotationCreate the replacement, deploy it, then delete the old key. Keys have no expiry, so rotation is always an explicit action.
Disable vs. deleteDisabling suspends a key while keeping its usage history; deleting revokes it permanently. Both take effect on the next request.
ScopeA key inherits its account’s plan, balance, and model entitlements. Keys are not individually scoped to models.

Never call AltLLM from a browser or mobile app

There is no CORS-safe, origin-restricted key type. A key in client-side code is a key anyone can extract and spend your balance with. Call AltLLM from your own backend and expose your own authenticated endpoint to clients.

Authentication vs. entitlement failures

These fail for different reasons and need different fixes. A retry never helps with either.

401 — the key itself is the problem
{
  "error": {
    "message": "Invalid API key.",
    "type": "authentication_error",
    "code": "invalid_api_key"
  }
}
403 — the key is valid, the plan is not
{
  "error": {
    "message": "Model 'altllm-flex-gpt-5.6' requires Flex tier or higher. Your current tier: Free. Upgrade at https://platform.altllm.ai/billing/upgrade",
    "type": "tier_upgrade_required",
    "code": "forbidden",
    "current_tier": "free",
    "required_tier": "flex",
    "model": "altllm-flex-gpt-5.6",
    "upgrade_url": "https://platform.altllm.ai/billing/upgrade"
  }
}

Business Flex SKUs are track-locked. A Personal account does not merely lack permission — the models are filtered out of GET /v1/models entirely, and GET /v1/models/{model_id} returns 404 rather than revealing that they exist.

OpenAI compatibility

AltLLM implements a defined subset of the OpenAI API — enough that the official clients work unchanged against chat completions, models, and responses. It is not a reimplementation of the whole OpenAI surface. This table is the contract.
AreaStatusDetail
Client librariessupportedThe official openai clients work by pointing them at the AltLLM base URL. The examples here use the modern client constructor, which needs openai-python 1.0+ or openai-node 4.0+.
chat.completions.createsupportedStreaming and non-streaming, with tools and structured output.
responses.createpartialImplemented via translation onto the chat pipeline. Text in, text out.
models.list / models.retrievesupportedReturns AltLLM's catalog with extra pricing and capability fields alongside the OpenAI ones.
Multimodal inputunsupportedImage content parts are accepted by the API but dropped before the model sees them.
Sampling parameterspartialParameters a provider rejects are dropped rather than erroring, so the same request body works across models. See the per-model table.
Error shapesupportedErrors use OpenAI's { "error": { message, type, code } } envelope, with extra fields on billing and entitlement errors.
finish_reasonsupportedstop, length, tool_calls, and content_filter, normalized across providers.
Response model fieldsupportedAlways echoes the model ID you requested, never the underlying provider model.
Embeddings, images, audio, files, batches, fine-tuning, assistantsunsupportedNot served. Authenticated requests return 404 route_not_found.

Endpoints AltLLM does not implement

An authenticated request to any of these returns 404 with type: "not_found" and code: "route_not_found". They fail immediately rather than timing out or half-working.

EndpointWhy
POST /v1/embeddingsNo embedding models are served.
POST /v1/images/generationsNo image generation.
POST /v1/audio/speechNo text-to-speech.
POST /v1/audio/transcriptionsNo speech-to-text.
POST /v1/moderationsNo standalone moderation endpoint. Provider safety filters still apply.
POST /v1/filesNo file storage. Send content inline in messages.
POST /v1/batchesNo Batch API.
POST /v1/fine_tuning/jobsNo fine-tuning.
GET /v1/assistantsNo Assistants API. Use /v1/responses for multi-turn state.
POST /v1/completionsLegacy text completions are not served. Use chat completions.

Endpoints

The complete public surface. Operator endpoints — metrics, routing statistics, analytics, and rate-limit administration — are excluded: they are not part of the public contract and may change without notice.
POST/v1/chat/completionsAPI keyrate limited

Create a chat completion. Supports streaming, custom tools, structured output, and AltLLM's server-side crypto tools.

Entitlement: Model must be in your plan; Flex SKUs need the Business track on the Flex tier.

Status codes: 200 Completion returned, or SSE stream opened. · 400 Malformed JSON, or a value the gateway rejects outright. · 401 Missing or invalid API key. · 402 No balance, expired credits, or no Portal account. · 403 Your plan does not include the requested model. · 404 Unknown or suppressed model ID. · 429 Per-model RPM or TPM limit exceeded. · 503 Rate-limit, credit-reservation, or admission state temporarily unavailable. · 500 Gateway error. · 502 Upstream provider error passed through.
POST/v1/responsesAPI keyrate limited

OpenAI Responses API. Translated onto the same routing, billing, and tool pipeline as chat completions.

Entitlement: Same model entitlements as /v1/chat/completions.

Status codes: 200 Response returned, or SSE stream opened. · 400 Malformed JSON body. · 401 Missing or invalid API key. · 402 No balance, expired credits, or no Portal account. · 403 Your plan does not include the requested model. · 404 Unknown or suppressed model ID. · 429 Per-model RPM or TPM limit exceeded. · 503 Rate-limit, credit-reservation, or admission state temporarily unavailable. · 500 Gateway error.
  • Conversation state is keyed by previous_response_id; store defaults to true. Stored state is owner-scoped, retained for up to 30 days, limited to 256 KiB per response, 100 responses per account, and 1,000 responses globally. Storage is best-effort, so persist important application state in your own system.
  • Responses input keeps text parts and drops unsupported non-text parts. Chat Completions forwards structured content to the selected provider, but the public catalogue guarantees text input only.
GET/v1/modelspublic

List available models with pricing, limits, modalities, and supported features.

Entitlement: Optional. Send your API key to see the models your account can actually call, including Flex SKUs.

Query parameters:
  • supported_parameters (string) — Comma-separated sampling parameters. Returns only models supporting all of them.
  • category (string) — coding or general. Filters on the model's instruct type.
Status codes: 200 Model list returned.
  • Unauthenticated callers are treated as Personal/free, so Flex SKUs are omitted.
  • Models an admin has disabled are omitted for everyone.
GET/v1/models/{model_id}public

Retrieve one model's full metadata.

Entitlement: Optional. Same per-caller filtering as the list endpoint.

Path parameters: model_id (string) — Exact model ID, e.g. altllm-basic.

Status codes: 200 Model metadata returned. · 404 Unknown model, or a model your account cannot access.
  • A model you lack access to returns 404, not 403 — restricted SKUs stay invisible rather than advertising themselves.
GET/v1/providerpublic

Provider metadata: supported features, model count, indicative tier pricing, and specializations.

Status codes: 200 Provider metadata returned.
  • Intended for aggregator integrations. Its model count and prices are the published Personal-catalog upper bound, not account-specific availability or an SLA. Use /v1/models for the models and prices available to the current caller.
GET/healthpublic

Readiness check for the gateway and its required rate-limit Redis dependency.

Status codes: 200 Gateway is ready to serve inference traffic. · 503 A required dependency is unavailable; retry later.
GET/health/livepublic

Process-only liveness check used by Kubernetes; it does not assert dependency readiness.

Status codes: 200 The gateway process is alive.

Chat Completions

POST/v1/chat/completions

Request fields

Every field the gateway reads, plus the OpenAI fields it forwards untouched. The Support column is the one that matters: AltLLM never rejects a request because a provider dislikes a sampling parameter — it drops the parameter instead, so one request body works across the whole catalog.

FieldTypeRequiredDefaultSupportBehavior
modelstringYes—supportedModel ID from GET /v1/models. Unknown or suppressed IDs return 404; models outside your plan return 403.
messagesMessage[]Yes—supportedConversation so far. See the message schema below for roles and content parts.
streambooleanNofalsesupportedReturn text/event-stream server-sent events instead of a single JSON body.
max_tokensintegerNo—model-dependentCap on generated tokens. Several models enforce a minimum — sending less returns 400, and omitting it makes the gateway apply that minimum for you.
temperaturenumberNo1strippedSampling temperature, 0–2. Dropped for models whose provider rejects it, rather than failing the request.
top_pnumberNo1strippedNucleus sampling, 0–1. Dropped for the same models as temperature.
top_kintegerNo—strippedTop-k sampling. Supported by most catalog models; dropped on Claude and Gemini Flex SKUs.
stopstring | string[]No—model-dependentUp to 4 sequences that end generation. finish_reason becomes stop.
toolsTool[]No—supportedYour own function definitions. Sending any tool disables AltLLM's built-in crypto tools for that request.
tool_choice"none" | "auto" | "required" | objectNo"auto"model-dependentConstrain tool selection. Supported by every model that supports tools.
response_formatobjectNo—model-dependent{"type": "json_object"} or {"type": "json_schema", ...}. Available on models listing json_mode or structured_outputs.
seedintegerNo—strippedBest-effort determinism. Dropped on GPT Flex, which rejects it.
frequency_penaltynumberNo0strippedPenalize token frequency, -2 to 2. Dropped on GPT and Gemini Flex SKUs.
presence_penaltynumberNo0strippedPenalize repeated tokens, -2 to 2. Dropped on GPT and Gemini Flex SKUs.
logit_biasRecord<string, number>No—strippedPer-token bias. Dropped on GPT Flex.
logprobsbooleanNo—passthroughForwarded to the routing layer. Support depends on the upstream provider — verify against your model before relying on it.
top_logprobsintegerNo—passthroughForwarded with logprobs. Same caveat.
nintegerNo1passthroughForwarded unchanged. AltLLM meters and bills every generated completion, and tool execution assumes choices[0].
parallel_tool_callsbooleanNo—passthroughForwarded to the provider. AltLLM executes returned tool calls sequentially either way.
userstringNo—supportedOverwritten by the gateway with the account ID resolved from your API key, so usage always attributes to you.
reasoning_effort"low" | "medium" | "high"No—passthroughForwarded to reasoning-capable providers that accept it.
thinkingobjectNo—passthroughAnthropic extended-thinking block, forwarded to Claude-family models.
web_search_optionsobjectNo—AltLLM extension{"search_context_size": "low" | "medium" | "high"} requests provider-specific hosted-search effort. Current environments return 400 until per-search fees are metered; future Gemini enablement remains mutually exclusive with function tools.
tool_execution"execute" | "passthrough"No"execute"AltLLM extensionpassthrough skips built-in tool injection and returns raw tool_calls for you to execute.

Parameters dropped per model

These models reject some sampling controls upstream, so the gateway removes them before the provider call. Your request still succeeds; the parameter simply has no effect.

  • altllm-flex-gemini-3.6 — drops frequency_penalty, presence_penalty, temperature, top_k, top_p
  • altllm-flex-gpt-5.6 — drops frequency_penalty, logit_bias, presence_penalty, seed, temperature, top_p

Models with a minimum output budget reject an explicit max_tokens below their floor with a 400, and apply the floor for you when you omit it: altllm-flex-gemini-3.6 (128), altllm-light-coding (4,096), altllm-native-flash (4,096), altllm-native-standard (4,096), altllm-pro (128), altllm-pro-coding (128), altllm-standard (128).

Routing and credential fields are operator-owned and removed for every model: api_key, api_base, fallbacks, extra_body, and related overrides cannot be set by a caller.

Minimal request

cURL
curl https://api.altllm.ai/v1/chat/completions \
  -H "Authorization: Bearer $ALTLLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "altllm-standard",
    "messages": [
      {"role": "system", "content": "You are a concise crypto research assistant."},
      {"role": "user", "content": "Summarize what EIP-4844 changed for rollup costs."}
    ],
    "max_tokens": 500
  }'

Messages and content

A message is a role, its content, and — for tool traffic — the identifiers that tie a result back to the call that requested it.

Roles

RoleRequired fieldsMeaning
systemcontentInstructions for the model. On catalog models AltLLM prepends its own identity prompt; on altllm-native-* and altllm-flex-* models yours is sent unmodified.
usercontentInput from the end user.
assistantcontent or tool_callsA previous model turn. Carries tool_calls when the model asked for a tool; content may then be null.
toolcontent, tool_call_idThe result of one tool call. Must directly follow the assistant message that requested it.

Message schema

{
  "role": "system" | "user" | "assistant" | "tool",

  // Required unless the assistant message carries tool_calls.
  // String, or an array of content parts (see below).
  "content": string | ContentPart[] | null,

  // Assistant messages only. Present when the model requested a tool.
  "tool_calls": [
    {
      "id": "call_abc123",
      "type": "function",
      "function": { "name": "get_token_price", "arguments": "{\"symbol\":\"ETH\"}" }
    }
  ],

  // Tool messages only. Must match the id of the call being answered.
  "tool_call_id": "call_abc123",

  // Optional label. Not interpreted by the gateway.
  "name": "string"
}

Ordering rules

  • A system message, if present, comes first. On catalog models AltLLM prepends its own identity prompt to yours; on altllm-native-* and altllm-flex-* models your system prompt is sent unmodified.
  • Every tool message must follow the assistant message whose tool_calls contains its tool_call_id.
  • An assistant message with tool_calls needs one tool message per call before the next completion, or the provider rejects the conversation.
  • Total input must fit the model’s context window. AltLLM does not truncate for you — an oversized request fails upstream.

Content parts

PartStatusBehavior
string contentsupportedThe simplest form. "content": "Explain EIP-4844".
{ "type": "text", "text": "..." }supportedText parts are concatenated in order, joined by newlines.
{ "type": "image_url", "image_url": { ... } }flattenedAccepted by the API but dropped during flattening — the model never sees the image. Do not send image parts today, and do not read image in a model's input_modalities as a promise that this endpoint accepts them.

Image input is not delivered to the model

Array content is flattened to a single string before routing: each part’s text is joined with newlines, and a part with no text field contributes nothing. An image_url part is therefore accepted and then silently discarded.

0 models list image in input_modalities because their underlying providers support vision. That capability is not reachable through this API today. Treat AltLLM as text-in, text-out.

Multi-turn with text parts
{
  "model": "altllm-standard",
  "messages": [
    {"role": "system", "content": "You are a concise crypto research assistant."},
    {"role": "user", "content": [
      {"type": "text", "text": "Here is the contract ABI:"},
      {"type": "text", "text": "[{\"name\":\"transfer\",\"type\":\"function\"}]"}
    ]},
    {"role": "assistant", "content": "This ABI exposes a single transfer function."},
    {"role": "user", "content": "What access control should it have?"}
  ]
}

Response schema

A non-streaming completion returns one JSON object. Every field below is always present unless marked optional.
200 — complete response
{
  "id": "chatcmpl-9f2b1c4a8e7d",
  "object": "chat.completion",
  "created": 1755500000,
  "model": "altllm-standard",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "EIP-4844 introduced blob-carrying transactions...",
        "tool_calls": null
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 42,
    "completion_tokens": 128,
    "total_tokens": 170
  }
}
FieldTypeMeaning
idstringCompletion identifier. Not the same as X-Request-ID — quote the header when contacting support.
objectstringchat.completion, or chat.completion.chunk on stream events.
createdintegerUnix timestamp in seconds.
modelstringAlways the model ID you requested. The gateway rewrites this field, so the underlying provider model is never disclosed.
choices[].indexintegerPosition of the choice. Tool execution operates on index 0.
choices[].message.contentstring | nullGenerated text. null when the model returned only tool calls.
choices[].message.tool_callsarray | nullTools the model wants you to run. Present only for tools you supplied — AltLLM’s built-in crypto tools are executed server-side and never surface here.
choices[].finish_reasonstringstop (completed or hit a stop sequence), length (hit max_tokens), tool_calls (awaiting your tool results), or content_filter (provider safety refusal).
usage.prompt_tokensintegerTotal input tokens, including any cached ones. Tokens consumed by server-side tool execution are included.
usage.completion_tokensintegerGenerated tokens, including hidden reasoning tokens on reasoning models. Those are billed at the output rate.
usage.prompt_tokens_details.cached_tokensintegerOptional. Present on models with prompt_caching. Billed at the cached input rate, not the standard one.

Streaming

Set stream: true to receive text/event-stream chunks. Each event is a data: line holding one JSON object, terminated by a blank line.
Request
curl https://api.altllm.ai/v1/chat/completions \
  -H "Authorization: Bearer $ALTLLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "altllm-standard",
    "messages": [{"role": "user", "content": "Name three L2 rollups."}],
    "stream": true
  }'
Response — complete SSE sequence
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":"Arbitrum"},"finish_reason":null}]}

data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":", Optimism"},"finish_reason":null}]}

data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":", and Base."},"finish_reason":null}]}

data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Event ordering

  • The first chunk carries delta.role. Later chunks carry delta.content fragments.
  • Exactly one chunk carries a non-null finish_reason with an empty delta.
  • The stream always ends with the literal line data: [DONE]. Stop reading there; do not attempt to parse it as JSON.
  • id, created, and model stay constant across every chunk of one completion.

Errors after the stream starts

Once headers are sent the status is already 200 and cannot be changed. If the provider connection drops mid-generation, AltLLM emits an error event in the stream and closes it. Treat any stream that ends without [DONE] as incomplete.

Mid-stream failure
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":"Arbitrum"},"finish_reason":null}]}

data: {"error":{"message":"The model stream ended before completion. Please retry the request.","type":"upstream_stream_error","code":"stream_interrupted"}}

This is retryable. A Chat stream that uses AltLLM server-side tools is charged only after a valid terminal event and final usage block, so an interrupted or malformed tool stream creates no charge. Other streaming paths settle only from provider-reported usage; use the Portal transaction record as the billing source of truth rather than partial output.

Streaming with server-side tools

When AltLLM’s built-in crypto tools are active, the gateway resolves tool calls before opening the stream to you. Time to first token is therefore longer on tool-using turns, and you never see the intermediate tool traffic — only the final answer streams.

Structured output

Models advertising json_mode accept response_format. Models advertising structured_outputs additionally accept a JSON Schema and conform to it.
JSON Schema output
curl https://api.altllm.ai/v1/chat/completions \
  -H "Authorization: Bearer $ALTLLM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "altllm-standard",
    "messages": [{"role": "user", "content": "Classify the risk of a 95% LTV lending position."}],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "risk_assessment",
        "strict": true,
        "schema": {
          "type": "object",
          "properties": {
            "level": {"type": "string", "enum": ["low", "medium", "high"]},
            "rationale": {"type": "string"}
          },
          "required": ["level", "rationale"],
          "additionalProperties": false
        }
      }
    }
  }'

The JSON arrives as a string in choices[0].message.content and still needs parsing. Check the model’s supported_features before relying on strict schema conformance — 17 of 17 catalog models support it.

Tool calling

AltLLM has two distinct tool systems. Knowing which one is active for a request is the difference between getting tool calls back and never seeing them.
Built-in crypto toolsYour own tools
Defined byAltLLMThe tools array in your request
Executed byThe gateway, server-sideYour application
Visible in the responseNo — only the final answerYes — as tool_calls
Active whenYou send no tools, on a catalog model that supports functionsYou send tools

The two systems are mutually exclusive

Sending any tools array turns off built-in crypto tool injection for that request. This is deliberate: injecting dozens of AltLLM tools alongside yours drowns them out and the model stops calling yours. You cannot combine both in one request.

Built-in tools are also never injected for:

  • altllm-native-* and altllm-flex-* models — these reach their provider without AltLLM prompt or tool additions
  • requests sending "tool_execution": "passthrough"

Full custom-tool lifecycle

Five steps. The conversation grows by two messages per tool round trip, and you resend the whole thing each time — the API is stateless.

1. Define the tool and ask
{
  "model": "altllm-standard",
  "messages": [
    {"role": "user", "content": "What is the portfolio value of 0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045?"}
  ],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_portfolio",
      "description": "Return the total USD value of a wallet's holdings.",
      "parameters": {
        "type": "object",
        "properties": {
          "wallet_address": {"type": "string", "description": "0x-prefixed EVM address"}
        },
        "required": ["wallet_address"],
        "additionalProperties": false
      }
    }
  }],
  "tool_choice": "auto"
}
2. The model asks you to run it
{
  "id": "chatcmpl-9f2b1c4a8e7d",
  "object": "chat.completion",
  "created": 1755500000,
  "model": "altllm-standard",
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": null,
      "tool_calls": [{
        "id": "call_7d3f9a1b",
        "type": "function",
        "function": {
          "name": "get_portfolio",
          "arguments": "{\"wallet_address\":\"0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045\"}"
        }
      }]
    },
    "finish_reason": "tool_calls"
  }],
  "usage": {"prompt_tokens": 91, "completion_tokens": 33, "total_tokens": 124}
}

3. Your application parses function.arguments — always a JSON string, never an object — and runs the tool.

4. Send the result back with the full history
{
  "model": "altllm-standard",
  "messages": [
    {"role": "user", "content": "What is the portfolio value of 0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045?"},
    {
      "role": "assistant",
      "content": null,
      "tool_calls": [{
        "id": "call_7d3f9a1b",
        "type": "function",
        "function": {
          "name": "get_portfolio",
          "arguments": "{\"wallet_address\":\"0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045\"}"
        }
      }]
    },
    {
      "role": "tool",
      "tool_call_id": "call_7d3f9a1b",
      "content": "{\"total_usd\": 1284300.55, \"chain\": \"ethereum\"}"
    }
  ],
  "tools": [{"type": "function", "function": {"name": "get_portfolio", "description": "Return the total USD value of a wallet's holdings.", "parameters": {"type": "object", "properties": {"wallet_address": {"type": "string"}}, "required": ["wallet_address"]}}}]
}
5. The model answers
{
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "That wallet holds about $1,284,300 on Ethereum mainnet.",
      "tool_calls": null
    },
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens": 148, "completion_tokens": 21, "total_tokens": 169}
}

Edge cases

  • Keep resending tools on every follow-up request. Omitting it after a tool result leaves the model unable to reference the definition.
  • Parallel calls: a single assistant message can contain several entries in tool_calls. Answer every one with its own tool message before the next request, matching each tool_call_id.
  • Tool failures: return the error as the tool message content — for example {"error": "address not found"}. Do not omit the message; the conversation becomes invalid without it.
  • Streaming tool calls: arguments arrive fragmented across delta.tool_calls[].function.arguments chunks. Concatenate by index before parsing — an individual fragment is not valid JSON.
  • Name collisions: your tool names take precedence, because built-in tools are not injected once you send any tool.
  • Tool budgets: providers cap how many tool definitions they accept per request — roughly 50 for Claude models and around 100 for others. Keep your set small.

Models

GET /v1/models is authoritative. Send your API key and it returns exactly the models your account can call; call it unauthenticated and you get the Personal free-tier view. The tables below are generated from the same catalog.
Discovery
# Everything your account can call
curl https://api.altllm.ai/v1/models -H "Authorization: Bearer $ALTLLM_API_KEY"

# One model
curl https://api.altllm.ai/v1/models/altllm-standard -H "Authorization: Bearer $ALTLLM_API_KEY"

# Only models supporting specific sampling parameters
curl "https://api.altllm.ai/v1/models?supported_parameters=temperature,top_p"
Model object
{
  "id": "altllm-standard",
  "object": "model",
  "name": "AltLLM Standard",
  "created": 1704067200,
  "description": "AltLLM's recommended tier for daily crypto and DeFi tasks...",
  "pricing": {
    "prompt": "0.000002",
    "completion": "0.000009",
    "image": "0",
    "request": "0",
    "input_cache_read": "0.0000002"
  },
  "context_length": 1048576,
  "max_output_length": 65536,
  "input_modalities": ["text"],
  "output_modalities": ["text"],
  "quantization": "bf16",
  "architecture": {"tokenizer": "AltLLM", "instruct_type": "chat"},
  "top_provider": {"is_moderated": false},
  "supported_sampling_parameters": ["temperature", "top_p", "top_k", "stop", "seed", "..."],
  "supported_features": ["tools", "json_mode", "structured_outputs", "streaming", "reasoning", "prompt_caching"]
}

Catalog models (15)

Personal-track models, included in your subscription and gated by tier. Prices are USD per million tokens.

ModelContextMax outputInput / OutputFeaturesNotes
altllm-standard
AltLLM Standard
1.05M66K$2.00 / $9.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 128; cached input $0.20/M
altllm-native-fast
AltLLM Native Fast
1.05M66K$0.20 / $0.80
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
no built-in crypto tools; cached input $0.02/M
altllm-native-flash
AltLLM Native Flash
1.05M66K$0.14 / $0.80
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 4,096; no built-in crypto tools; cached input $0.01/M
altllm-native-light
AltLLM Native Light
1.05M66K$0.20 / $0.80
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
no built-in crypto tools; cached input $0.02/M
altllm-native-standard
AltLLM Native Standard
1.05M66K$1.20 / $4.40
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 4,096; no built-in crypto tools; cached input $0.12/M
altllm-native-promax
AltLLM Native Pro Max
1.05M66K$4.80 / $28.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
no built-in crypto tools; cached input $0.48/M
altllm-basic
AltLLM Basic
1.05M66K$5.00 / $30.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
cached input $0.50/M
altllm-pro
AltLLM Pro
1.05M66K$0.80 / $3.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 128; cached input $0.08/M
altllm-light-coding
AltLLM Light Coding
1.05M66K$0.20 / $0.76
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 4,096; cached input $0.02/M
altllm-pro-coding
AltLLM Pro Coding
1.05M66K$0.80 / $3.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 128; cached input $0.08/M
altllm-max-coding
AltLLM Max Coding
1.05M66K$2.00 / $10.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
provider safety filters; cached input $0.20/M
altllm-hybrid-semantic
AltLLM Hybrid Semantic
1.05M66KSelected tier
toolsjson_modestructured_outputsstreamingreasoningrouting
billed at selected target tier
altllm-hybrid-cost
AltLLM Hybrid Cost
1.05M66KSelected tier
toolsjson_modestructured_outputsstreamingreasoningrouting
billed at selected target tier
altllm-hybrid-tiered
AltLLM Hybrid Tiered
1.05M66KSelected tier
toolsjson_modestructured_outputsstreamingreasoningrouting
billed at selected target tier
altllm-hybrid-turn
AltLLM Hybrid Turn
1.05M66KSelected tier
toolsjson_modestructured_outputsstreamingreasoningrouting
billed at selected target tier

Business Flex models (2)

Provider-native frontier models, billed per token rather than against a subscription. They require the Business track on the Flex tier and are invisible to every other account — filtered out of GET /v1/models, and 404 on direct lookup.

ModelContextMax outputInput / OutputFeaturesNotes
altllm-flex-gpt-5.6
Flex: GPT-5.6 Sol
1.05M128K$6.00 / $36.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
drops frequency_penalty, logit_bias, presence_penalty, seed, temperature, top_p; no built-in crypto tools; cached input $0.60/M
altllm-flex-gemini-3.6
Flex: Gemini 3.6 Flash Native
1.05M66K$1.80 / $9.00
toolsjson_modestructured_outputsstreamingreasoningprompt_caching
max_tokens ≥ 128; drops frequency_penalty, presence_penalty, temperature, top_k, top_p; no built-in crypto tools; cached input $0.18/M

Provider data handling on Flex models

Flex requests reach the provider named by the selected SKU. Currently advertised Flex models use OpenAI or Google infrastructure. Data handling follows that provider and the approved AltLLM account controls; AltLLM does not promise Zero Data Retention unless a customer agreement explicitly says so. See the pricing page for current provider-specific notes.

Access and entitlement

  • Each model has a minimum plan tier, configured in the Portal. Requesting one above your plan returns 403 tier_upgrade_required, naming the required_tier.
  • Flex SKUs are track-locked and fail closed: an account without the Business track cannot reach them under any tier.
  • Models an administrator has disabled disappear from discovery for everyone.
  • Because entitlement is per account, do not hard-code a model list. Call GET /v1/models at startup and select from what comes back.

Capabilities at a glance

  • Streaming: every published model.
  • Function calling: every published model.
  • Prompt caching: altllm-standard, altllm-native-fast, altllm-native-flash, altllm-native-light, altllm-native-standard, altllm-native-promax, altllm-basic, altllm-pro, altllm-light-coding, altllm-pro-coding, altllm-max-coding, altllm-flex-gpt-5.6, altllm-flex-gemini-3.6. Cached input bills at the reduced rate shown in each model’s pricing.input_cache_read.
  • Reasoning: models listing reasoning in supported_features generate hidden reasoning tokens that count as output tokens and are billed as such.
  • Vision: not reachable — see content parts.

Usage and rate limits

Limits are enforced per account, per model, over a 60-second sliding window — not per API key. Multiple keys on one account share one budget.
LimitUnitScopeDefault
Requests per minuterequestsaccount × model60
Tokens per minutereserved input + output tokensaccount × model100,000
Windowsecondssliding60
Context windowtokensper request, per modelSee the model tables
Max outputtokensper request, per modelSee the model tables

Counters are isolated per account and model; ceilings are configured per model. The values above are the fallback when a model has no explicit configuration. Before generation, the TPM check atomically reserves estimated message and multimodal input plus prompt-bearing tool/structured-output schemas plus the largest valid explicit max_output_tokens, max_completion_tokens, or max_tokens budget. If none is set, it reserves a conservative 4,096-token output budget. Set an explicit realistic maximum to avoid consuming unused concurrent capacity. Completed usage replaces that reservation when the provider reports final usage. If the atomic reservation service is unavailable, AltLLM rejects the request before provider invocation with a retryable 503 rate_limit_unavailable instead of silently bypassing your account limits.

Context capacity and throughput are separate limits

A model’s context window is its technical input-plus-output capacity, not a guaranteed per-request allowance for every account. Your configured TPM ceiling can be lower, so a large single request needs enough TPM for its estimated input plus output reservation. Contact support before relying on near-context-window requests in production.

Rate-limit headers

Rate-limit headers appear only on 429 responses

A successful response carries X-Request-ID and, when applicable, X-Credits-Expire-At — but no X-RateLimit-* headers. You cannot track remaining quota by reading successful responses. Handle 429s instead.
HeaderPresent onMeaning
X-Request-IDEvery chat completion and response, success or failureCorrelation ID for the request. Echoes your own X-Request-ID if you send one. Include it in support reports.
X-Credits-Expire-AtWhen your credits have an expiry dateISO-8601 timestamp at which the current credit grant expires.
X-RateLimit-Limit429 responses onlyRequests-per-minute ceiling for this model.
X-RateLimit-Remaining429 responses onlyRequests left in the current 60-second window.
X-RateLimit-Reset429 responses onlyUnix timestamp (seconds) when the window resets.
X-RateLimit-Limit-Tokens429 responses onlyTokens-per-minute ceiling for this model.
X-RateLimit-Remaining-Tokens429 responses onlyTokens left in the current 60-second window.
Retry-After429 responses, and retryable rate-limit 503 responsesSeconds to wait before retrying.

How usage is billed

  • Cost is prompt_tokens × input rate + completion_tokens × output rate, at the per-model rates in the tables above, deducted from your credit balance.
  • Cached input is billed at input_cache_read instead of the standard input rate, and is subtracted from the fresh-input count rather than charged twice.
  • Reasoning tokens are part of completion_tokens and are billed at the output rate even though you never see them.
  • Server-side tool execution adds tokens to prompt_tokens: tool schemas and results re-enter the model as input. A tool-using turn costs more than the same question without tools.
  • Interrupted Chat tool streams create no charge unless the gateway receives both a valid terminal event and final usage. Other stream modes settle from completed provider-reported usage.
  • Credits expire. Check X-Credits-Expire-At, and see the pricing page for cycle behavior.

Errors and retries

Errors use the OpenAI envelope. Billing and entitlement errors add fields so you can act on them without parsing the message string. See Error Handling for worked examples.
Error envelope
{
  "error": {
    "message": "Human-readable description",
    "type": "insufficient_credits",
    "code": "payment_required"
  }
}
StatustypecodeMeaningRetryExtra fields
401authentication_errormissing_api_keyNo Authorization: Bearer header was sent.fix, then retry—
401authentication_errorinvalid_api_keyThe key is unknown, disabled, or revoked.fix, then retry—
402insufficient_creditspayment_requiredBalance is too low to start the request.fix, then retrybalance, billing_url
402credits_expiredpayment_requiredYour credits passed their expiry date.fix, then retrybalance, expired_at, billing_url
402no_accountpayment_requiredThe identity behind the request has no Portal account.fix, then retrybalance, billing_url
403tier_upgrade_requiredforbiddenYour plan does not include this model.fix, then retrycurrent_tier, required_tier, model, upgrade_url
404not_foundroute_not_foundAuthenticated request to a path AltLLM does not serve.not retryable—
404invalid_request_errorprevious_response_not_foundThe requested conversation state is missing, expired, evicted, or belongs to a different account.fix, then retry—
400invalid_request_error—The request violates a gateway rule, such as max_tokens below a model's documented minimum.fix, then retry—
429rate_limit_exceededrpm_exceededToo many requests for this model in the last 60 seconds.retryable—
503service_unavailablerate_limit_unavailableAtomic RPM/TPM reservation is temporarily unavailable, so the request was not sent upstream.retryable—
503server_errorcredit_reservation_unavailableThe prepaid-credit hold or its durable settlement state is temporarily unavailable, so the request cannot continue safely.retryable—
503server_erroradmission_state_unavailableA multi-step request lost the internal admission state required before its next provider call.retryable—
429rate_limit_exceededtpm_exceededToo many tokens for this model in the last 60 seconds.retryable—
200upstream_stream_errorstream_interruptedThe provider stream ended before completion. Delivered as an SSE error event after the 200 headers were already sent.retryable—
500gateway_error—Gateway processing failed after request validation. Quote the X-Request-ID when contacting support.retryable—
500server_errorinternal_errorUnhandled gateway failure. Quote the X-Request-ID when contacting support.retryable—
502upstream_error—The provider failed or returned an unparseable body. code carries the upstream status when there is one.retryable—
503service_unavailable—Response storage is unavailable, so previous_response_id cannot be resolved on /v1/responses.retryable—

Retry policy

  • 429: honor Retry-After when present; otherwise back off exponentially from one second with jitter.
  • 500, 502, 503: retry up to three times with exponential backoff; honor Retry-After when present.
  • Interrupted streams: retry the whole request. Partial output is not resumable.
  • 400, 401, 402, 403, 404: never retry unchanged. Fix the request, the key, the balance, or the plan first.

Reporting a problem

Include the X-Request-ID response header, the model ID, the UTC timestamp, and the full error body. That header is the correlation key in AltLLM’s logs — a report without it is much slower to investigate.

SDK examples

The official OpenAI clients work unchanged — point them at the AltLLM base URL. See the Python and TypeScript SDK guides for fuller walkthroughs.
Python — non-streaming
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ALTLLM_API_KEY"],
    base_url="https://api.altllm.ai/v1",
)

response = client.chat.completions.create(
    model="altllm-standard",
    messages=[{"role": "user", "content": "Explain rollup data availability."}],
    max_tokens=500,
)
print(response.choices[0].message.content)
print(response.usage.total_tokens)
Python — streaming
stream = client.chat.completions.create(
    model="altllm-standard",
    messages=[{"role": "user", "content": "Name three L2 rollups."}],
    stream=True,
)

for chunk in stream:
    delta = chunk.choices[0].delta.content
    if delta:
        print(delta, end="", flush=True)
TypeScript — non-streaming
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.ALTLLM_API_KEY,
  baseURL: "https://api.altllm.ai/v1",
});

const response = await client.chat.completions.create({
  model: "altllm-standard",
  messages: [{ role: "user", content: "Explain rollup data availability." }],
  max_tokens: 500,
});

console.log(response.choices[0].message.content);
TypeScript — streaming
const stream = await client.chat.completions.create({
  model: "altllm-standard",
  messages: [{ role: "user", content: "Name three L2 rollups." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
Selecting a model at runtime
# Never hard-code a model list — entitlements are per account.
available = {m.id for m in client.models.list().data}

for candidate in ("altllm-flex-gpt-5.6", "altllm-basic", "altllm-standard"):
    if candidate in available:
        model = candidate
        break
else:
    raise RuntimeError("no usable model for this account")
Handling errors and retries
import time
import openai

def complete_with_retry(**kwargs):
    for attempt in range(4):
        try:
            return client.chat.completions.create(**kwargs)
        except openai.RateLimitError as exc:
            wait = float(exc.response.headers.get("Retry-After", 2 ** attempt))
            time.sleep(wait)
        except openai.InternalServerError:
            time.sleep(2 ** attempt)
        except (openai.AuthenticationError, openai.PermissionDeniedError, openai.BadRequestError):
            # Key, plan, balance, or request shape — retrying changes nothing.
            raise
    raise RuntimeError("exhausted retries")

Versioning and lifecycle

The API is versioned by path prefix. There is no version header and no date-pinned API version.
TopicPolicy
API version/v1. A breaking change to request or response shapes would ship under a new prefix, not by mutating /v1.
Additive changesNew endpoints, new optional request fields, and new response fields can appear in /v1 at any time. Ignore unknown response fields rather than failing on them.
Model IDsCatalog IDs such as altllm-standard are stable aliases. The provider model behind an alias can change; the ID, its pricing, and its advertised capabilities are the contract.
Pinned versionsFlex SKUs carry the provider version in the ID — altllm-flex-gpt-5.6, altllm-flex-gemini-3.6. A new provider version arrives as a new ID rather than changing an existing one.
DeprecationA superseded model can remain callable as an unlisted compatibility alias for pinned integrations while disappearing from discovery, detail, provider-pricing, SDK, and current guidance surfaces. Removal of that compatibility path is a separate breaking lifecycle decision; do not infer it merely from absence in GET /v1/models.
MigrationDiscover models at startup and fall back across a preference list, as in the SDK example above. An application that hard-codes one ID breaks on retirement.
ChangelogModel and pricing changes are reflected in GET /v1/models and on the Models and Pricing pages as they ship.