API Reference
The complete contract for the AltLLM inference API: every public endpoint, every request field, the exact response and streaming schemas, the model catalog, and what “OpenAI-compatible” does and does not cover.
Overview
https://api.altllm.ai/v1| Property | Value |
|---|---|
Content-Type | application/json on every request with a body. |
Authorization | Bearer sk-alt-… — required on everything except model discovery, provider metadata, and health. |
X-Request-ID | Optional. Any correlation ID you send is echoed back on the response and recorded in AltLLM’s logs. |
| Streaming | text/event-stream when stream: true. |
| Rate limits | Per account, per model, over a 60-second sliding window. |
One reference, one set of facts
GET /v1/models serves. A CI check fails the build if this page and the gateway disagree.Authentication
Authorization: Bearer sk-alt-your_key_here| Property | Detail |
|---|---|
| Format | sk-alt- followed by 48 URL-safe characters (55 characters total). |
| Creation | Portal → API Keys. The full key is shown once, at creation, and is never retrievable again — only its 12-character prefix. |
| Rotation | Create the replacement, deploy it, then delete the old key. Keys have no expiry, so rotation is always an explicit action. |
| Disable vs. delete | Disabling suspends a key while keeping its usage history; deleting revokes it permanently. Both take effect on the next request. |
| Scope | A key inherits its account’s plan, balance, and model entitlements. Keys are not individually scoped to models. |
Never call AltLLM from a browser or mobile app
Authentication vs. entitlement failures
These fail for different reasons and need different fixes. A retry never helps with either.
{
"error": {
"message": "Invalid API key.",
"type": "authentication_error",
"code": "invalid_api_key"
}
}{
"error": {
"message": "Model 'altllm-flex-gpt-5.6' requires Flex tier or higher. Your current tier: Free. Upgrade at https://platform.altllm.ai/billing/upgrade",
"type": "tier_upgrade_required",
"code": "forbidden",
"current_tier": "free",
"required_tier": "flex",
"model": "altllm-flex-gpt-5.6",
"upgrade_url": "https://platform.altllm.ai/billing/upgrade"
}
}Business Flex SKUs are track-locked. A Personal account does not merely lack permission — the models are filtered out of GET /v1/models entirely, and GET /v1/models/{model_id} returns 404 rather than revealing that they exist.
OpenAI compatibility
| Area | Status | Detail |
|---|---|---|
| Client libraries | supported | The official openai clients work by pointing them at the AltLLM base URL. The examples here use the modern client constructor, which needs openai-python 1.0+ or openai-node 4.0+. |
| chat.completions.create | supported | Streaming and non-streaming, with tools and structured output. |
| responses.create | partial | Implemented via translation onto the chat pipeline. Text in, text out. |
| models.list / models.retrieve | supported | Returns AltLLM's catalog with extra pricing and capability fields alongside the OpenAI ones. |
| Multimodal input | unsupported | Image content parts are accepted by the API but dropped before the model sees them. |
| Sampling parameters | partial | Parameters a provider rejects are dropped rather than erroring, so the same request body works across models. See the per-model table. |
| Error shape | supported | Errors use OpenAI's { "error": { message, type, code } } envelope, with extra fields on billing and entitlement errors. |
| finish_reason | supported | stop, length, tool_calls, and content_filter, normalized across providers. |
Response model field | supported | Always echoes the model ID you requested, never the underlying provider model. |
| Embeddings, images, audio, files, batches, fine-tuning, assistants | unsupported | Not served. Authenticated requests return 404 route_not_found. |
Endpoints AltLLM does not implement
An authenticated request to any of these returns 404 with type: "not_found" and code: "route_not_found". They fail immediately rather than timing out or half-working.
| Endpoint | Why |
|---|---|
POST /v1/embeddings | No embedding models are served. |
POST /v1/images/generations | No image generation. |
POST /v1/audio/speech | No text-to-speech. |
POST /v1/audio/transcriptions | No speech-to-text. |
POST /v1/moderations | No standalone moderation endpoint. Provider safety filters still apply. |
POST /v1/files | No file storage. Send content inline in messages. |
POST /v1/batches | No Batch API. |
POST /v1/fine_tuning/jobs | No fine-tuning. |
GET /v1/assistants | No Assistants API. Use /v1/responses for multi-turn state. |
POST /v1/completions | Legacy text completions are not served. Use chat completions. |
Endpoints
/v1/chat/completionsAPI keyrate limitedCreate a chat completion. Supports streaming, custom tools, structured output, and AltLLM's server-side crypto tools.
Entitlement: Model must be in your plan; Flex SKUs need the Business track on the Flex tier.
/v1/responsesAPI keyrate limitedOpenAI Responses API. Translated onto the same routing, billing, and tool pipeline as chat completions.
Entitlement: Same model entitlements as /v1/chat/completions.
- Conversation state is keyed by
previous_response_id;storedefaults totrue. Stored state is owner-scoped, retained for up to 30 days, limited to 256 KiB per response, 100 responses per account, and 1,000 responses globally. Storage is best-effort, so persist important application state in your own system. - Responses input keeps text parts and drops unsupported non-text parts. Chat Completions forwards structured content to the selected provider, but the public catalogue guarantees text input only.
/v1/modelspublicList available models with pricing, limits, modalities, and supported features.
Entitlement: Optional. Send your API key to see the models your account can actually call, including Flex SKUs.
supported_parameters(string) — Comma-separated sampling parameters. Returns only models supporting all of them.category(string) —codingorgeneral. Filters on the model's instruct type.
- Unauthenticated callers are treated as Personal/free, so Flex SKUs are omitted.
- Models an admin has disabled are omitted for everyone.
/v1/models/{model_id}publicRetrieve one model's full metadata.
Entitlement: Optional. Same per-caller filtering as the list endpoint.
Path parameters: model_id (string) — Exact model ID, e.g. altllm-basic.
- A model you lack access to returns 404, not 403 — restricted SKUs stay invisible rather than advertising themselves.
/v1/providerpublicProvider metadata: supported features, model count, indicative tier pricing, and specializations.
- Intended for aggregator integrations. Its model count and prices are the published Personal-catalog upper bound, not account-specific availability or an SLA. Use
/v1/modelsfor the models and prices available to the current caller.
/healthpublicReadiness check for the gateway and its required rate-limit Redis dependency.
/health/livepublicProcess-only liveness check used by Kubernetes; it does not assert dependency readiness.
Chat Completions
/v1/chat/completionsRequest fields
Every field the gateway reads, plus the OpenAI fields it forwards untouched. The Support column is the one that matters: AltLLM never rejects a request because a provider dislikes a sampling parameter — it drops the parameter instead, so one request body works across the whole catalog.
| Field | Type | Required | Default | Support | Behavior |
|---|---|---|---|---|---|
model | string | Yes | — | supported | Model ID from GET /v1/models. Unknown or suppressed IDs return 404; models outside your plan return 403. |
messages | Message[] | Yes | — | supported | Conversation so far. See the message schema below for roles and content parts. |
stream | boolean | No | false | supported | Return text/event-stream server-sent events instead of a single JSON body. |
max_tokens | integer | No | — | model-dependent | Cap on generated tokens. Several models enforce a minimum — sending less returns 400, and omitting it makes the gateway apply that minimum for you. |
temperature | number | No | 1 | stripped | Sampling temperature, 0–2. Dropped for models whose provider rejects it, rather than failing the request. |
top_p | number | No | 1 | stripped | Nucleus sampling, 0–1. Dropped for the same models as temperature. |
top_k | integer | No | — | stripped | Top-k sampling. Supported by most catalog models; dropped on Claude and Gemini Flex SKUs. |
stop | string | string[] | No | — | model-dependent | Up to 4 sequences that end generation. finish_reason becomes stop. |
tools | Tool[] | No | — | supported | Your own function definitions. Sending any tool disables AltLLM's built-in crypto tools for that request. |
tool_choice | "none" | "auto" | "required" | object | No | "auto" | model-dependent | Constrain tool selection. Supported by every model that supports tools. |
response_format | object | No | — | model-dependent | {"type": "json_object"} or {"type": "json_schema", ...}. Available on models listing json_mode or structured_outputs. |
seed | integer | No | — | stripped | Best-effort determinism. Dropped on GPT Flex, which rejects it. |
frequency_penalty | number | No | 0 | stripped | Penalize token frequency, -2 to 2. Dropped on GPT and Gemini Flex SKUs. |
presence_penalty | number | No | 0 | stripped | Penalize repeated tokens, -2 to 2. Dropped on GPT and Gemini Flex SKUs. |
logit_bias | Record<string, number> | No | — | stripped | Per-token bias. Dropped on GPT Flex. |
logprobs | boolean | No | — | passthrough | Forwarded to the routing layer. Support depends on the upstream provider — verify against your model before relying on it. |
top_logprobs | integer | No | — | passthrough | Forwarded with logprobs. Same caveat. |
n | integer | No | 1 | passthrough | Forwarded unchanged. AltLLM meters and bills every generated completion, and tool execution assumes choices[0]. |
parallel_tool_calls | boolean | No | — | passthrough | Forwarded to the provider. AltLLM executes returned tool calls sequentially either way. |
user | string | No | — | supported | Overwritten by the gateway with the account ID resolved from your API key, so usage always attributes to you. |
reasoning_effort | "low" | "medium" | "high" | No | — | passthrough | Forwarded to reasoning-capable providers that accept it. |
thinking | object | No | — | passthrough | Anthropic extended-thinking block, forwarded to Claude-family models. |
web_search_options | object | No | — | AltLLM extension | {"search_context_size": "low" | "medium" | "high"} requests provider-specific hosted-search effort. Current environments return 400 until per-search fees are metered; future Gemini enablement remains mutually exclusive with function tools. |
tool_execution | "execute" | "passthrough" | No | "execute" | AltLLM extension | passthrough skips built-in tool injection and returns raw tool_calls for you to execute. |
Parameters dropped per model
These models reject some sampling controls upstream, so the gateway removes them before the provider call. Your request still succeeds; the parameter simply has no effect.
altllm-flex-gemini-3.6— dropsfrequency_penalty,presence_penalty,temperature,top_k,top_paltllm-flex-gpt-5.6— dropsfrequency_penalty,logit_bias,presence_penalty,seed,temperature,top_p
Models with a minimum output budget reject an explicit max_tokens below their floor with a 400, and apply the floor for you when you omit it: altllm-flex-gemini-3.6 (128), altllm-light-coding (4,096), altllm-native-flash (4,096), altllm-native-standard (4,096), altllm-pro (128), altllm-pro-coding (128), altllm-standard (128).
Routing and credential fields are operator-owned and removed for every model: api_key, api_base, fallbacks, extra_body, and related overrides cannot be set by a caller.
Minimal request
curl https://api.altllm.ai/v1/chat/completions \
-H "Authorization: Bearer $ALTLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "altllm-standard",
"messages": [
{"role": "system", "content": "You are a concise crypto research assistant."},
{"role": "user", "content": "Summarize what EIP-4844 changed for rollup costs."}
],
"max_tokens": 500
}'Messages and content
Roles
| Role | Required fields | Meaning |
|---|---|---|
system | content | Instructions for the model. On catalog models AltLLM prepends its own identity prompt; on altllm-native-* and altllm-flex-* models yours is sent unmodified. |
user | content | Input from the end user. |
assistant | content or tool_calls | A previous model turn. Carries tool_calls when the model asked for a tool; content may then be null. |
tool | content, tool_call_id | The result of one tool call. Must directly follow the assistant message that requested it. |
Message schema
{
"role": "system" | "user" | "assistant" | "tool",
// Required unless the assistant message carries tool_calls.
// String, or an array of content parts (see below).
"content": string | ContentPart[] | null,
// Assistant messages only. Present when the model requested a tool.
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": { "name": "get_token_price", "arguments": "{\"symbol\":\"ETH\"}" }
}
],
// Tool messages only. Must match the id of the call being answered.
"tool_call_id": "call_abc123",
// Optional label. Not interpreted by the gateway.
"name": "string"
}Ordering rules
- A
systemmessage, if present, comes first. On catalog models AltLLM prepends its own identity prompt to yours; onaltllm-native-*andaltllm-flex-*models your system prompt is sent unmodified. - Every
toolmessage must follow the assistant message whosetool_callscontains itstool_call_id. - An assistant message with
tool_callsneeds onetoolmessage per call before the next completion, or the provider rejects the conversation. - Total input must fit the model’s context window. AltLLM does not truncate for you — an oversized request fails upstream.
Content parts
| Part | Status | Behavior |
|---|---|---|
string content | supported | The simplest form. "content": "Explain EIP-4844". |
{ "type": "text", "text": "..." } | supported | Text parts are concatenated in order, joined by newlines. |
{ "type": "image_url", "image_url": { ... } } | flattened | Accepted by the API but dropped during flattening — the model never sees the image. Do not send image parts today, and do not read image in a model's input_modalities as a promise that this endpoint accepts them. |
Image input is not delivered to the model
Array content is flattened to a single string before routing: each part’s text is joined with newlines, and a part with no text field contributes nothing. An image_url part is therefore accepted and then silently discarded.
0 models list image in input_modalities because their underlying providers support vision. That capability is not reachable through this API today. Treat AltLLM as text-in, text-out.
{
"model": "altllm-standard",
"messages": [
{"role": "system", "content": "You are a concise crypto research assistant."},
{"role": "user", "content": [
{"type": "text", "text": "Here is the contract ABI:"},
{"type": "text", "text": "[{\"name\":\"transfer\",\"type\":\"function\"}]"}
]},
{"role": "assistant", "content": "This ABI exposes a single transfer function."},
{"role": "user", "content": "What access control should it have?"}
]
}Response schema
{
"id": "chatcmpl-9f2b1c4a8e7d",
"object": "chat.completion",
"created": 1755500000,
"model": "altllm-standard",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "EIP-4844 introduced blob-carrying transactions...",
"tool_calls": null
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 42,
"completion_tokens": 128,
"total_tokens": 170
}
}| Field | Type | Meaning |
|---|---|---|
id | string | Completion identifier. Not the same as X-Request-ID — quote the header when contacting support. |
object | string | chat.completion, or chat.completion.chunk on stream events. |
created | integer | Unix timestamp in seconds. |
model | string | Always the model ID you requested. The gateway rewrites this field, so the underlying provider model is never disclosed. |
choices[].index | integer | Position of the choice. Tool execution operates on index 0. |
choices[].message.content | string | null | Generated text. null when the model returned only tool calls. |
choices[].message.tool_calls | array | null | Tools the model wants you to run. Present only for tools you supplied — AltLLM’s built-in crypto tools are executed server-side and never surface here. |
choices[].finish_reason | string | stop (completed or hit a stop sequence), length (hit max_tokens), tool_calls (awaiting your tool results), or content_filter (provider safety refusal). |
usage.prompt_tokens | integer | Total input tokens, including any cached ones. Tokens consumed by server-side tool execution are included. |
usage.completion_tokens | integer | Generated tokens, including hidden reasoning tokens on reasoning models. Those are billed at the output rate. |
usage.prompt_tokens_details.cached_tokens | integer | Optional. Present on models with prompt_caching. Billed at the cached input rate, not the standard one. |
Streaming
stream: true to receive text/event-stream chunks. Each event is a data: line holding one JSON object, terminated by a blank line.curl https://api.altllm.ai/v1/chat/completions \
-H "Authorization: Bearer $ALTLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "altllm-standard",
"messages": [{"role": "user", "content": "Name three L2 rollups."}],
"stream": true
}'data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":"Arbitrum"},"finish_reason":null}]}
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":", Optimism"},"finish_reason":null}]}
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":", and Base."},"finish_reason":null}]}
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]Event ordering
- The first chunk carries
delta.role. Later chunks carrydelta.contentfragments. - Exactly one chunk carries a non-null
finish_reasonwith an emptydelta. - The stream always ends with the literal line
data: [DONE]. Stop reading there; do not attempt to parse it as JSON. id,created, andmodelstay constant across every chunk of one completion.
Errors after the stream starts
Once headers are sent the status is already 200 and cannot be changed. If the provider connection drops mid-generation, AltLLM emits an error event in the stream and closes it. Treat any stream that ends without [DONE] as incomplete.
data: {"id":"chatcmpl-9f2b1c4a8e7d","object":"chat.completion.chunk","created":1755500000,"model":"altllm-standard","choices":[{"index":0,"delta":{"content":"Arbitrum"},"finish_reason":null}]}
data: {"error":{"message":"The model stream ended before completion. Please retry the request.","type":"upstream_stream_error","code":"stream_interrupted"}}This is retryable. A Chat stream that uses AltLLM server-side tools is charged only after a valid terminal event and final usage block, so an interrupted or malformed tool stream creates no charge. Other streaming paths settle only from provider-reported usage; use the Portal transaction record as the billing source of truth rather than partial output.
Streaming with server-side tools
When AltLLM’s built-in crypto tools are active, the gateway resolves tool calls before opening the stream to you. Time to first token is therefore longer on tool-using turns, and you never see the intermediate tool traffic — only the final answer streams.
Structured output
json_mode accept response_format. Models advertising structured_outputs additionally accept a JSON Schema and conform to it.curl https://api.altllm.ai/v1/chat/completions \
-H "Authorization: Bearer $ALTLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "altllm-standard",
"messages": [{"role": "user", "content": "Classify the risk of a 95% LTV lending position."}],
"response_format": {
"type": "json_schema",
"json_schema": {
"name": "risk_assessment",
"strict": true,
"schema": {
"type": "object",
"properties": {
"level": {"type": "string", "enum": ["low", "medium", "high"]},
"rationale": {"type": "string"}
},
"required": ["level", "rationale"],
"additionalProperties": false
}
}
}
}'The JSON arrives as a string in choices[0].message.content and still needs parsing. Check the model’s supported_features before relying on strict schema conformance — 17 of 17 catalog models support it.
Tool calling
| Built-in crypto tools | Your own tools | |
|---|---|---|
| Defined by | AltLLM | The tools array in your request |
| Executed by | The gateway, server-side | Your application |
| Visible in the response | No — only the final answer | Yes — as tool_calls |
| Active when | You send no tools, on a catalog model that supports functions | You send tools |
The two systems are mutually exclusive
Sending any tools array turns off built-in crypto tool injection for that request. This is deliberate: injecting dozens of AltLLM tools alongside yours drowns them out and the model stops calling yours. You cannot combine both in one request.
Built-in tools are also never injected for:
altllm-native-*andaltllm-flex-*models — these reach their provider without AltLLM prompt or tool additions- requests sending
"tool_execution": "passthrough"
Full custom-tool lifecycle
Five steps. The conversation grows by two messages per tool round trip, and you resend the whole thing each time — the API is stateless.
{
"model": "altllm-standard",
"messages": [
{"role": "user", "content": "What is the portfolio value of 0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045?"}
],
"tools": [{
"type": "function",
"function": {
"name": "get_portfolio",
"description": "Return the total USD value of a wallet's holdings.",
"parameters": {
"type": "object",
"properties": {
"wallet_address": {"type": "string", "description": "0x-prefixed EVM address"}
},
"required": ["wallet_address"],
"additionalProperties": false
}
}
}],
"tool_choice": "auto"
}{
"id": "chatcmpl-9f2b1c4a8e7d",
"object": "chat.completion",
"created": 1755500000,
"model": "altllm-standard",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_7d3f9a1b",
"type": "function",
"function": {
"name": "get_portfolio",
"arguments": "{\"wallet_address\":\"0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045\"}"
}
}]
},
"finish_reason": "tool_calls"
}],
"usage": {"prompt_tokens": 91, "completion_tokens": 33, "total_tokens": 124}
}3. Your application parses function.arguments — always a JSON string, never an object — and runs the tool.
{
"model": "altllm-standard",
"messages": [
{"role": "user", "content": "What is the portfolio value of 0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045?"},
{
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_7d3f9a1b",
"type": "function",
"function": {
"name": "get_portfolio",
"arguments": "{\"wallet_address\":\"0xd8dA6BF26964aF9D7eEd9e03E53415D37aA96045\"}"
}
}]
},
{
"role": "tool",
"tool_call_id": "call_7d3f9a1b",
"content": "{\"total_usd\": 1284300.55, \"chain\": \"ethereum\"}"
}
],
"tools": [{"type": "function", "function": {"name": "get_portfolio", "description": "Return the total USD value of a wallet's holdings.", "parameters": {"type": "object", "properties": {"wallet_address": {"type": "string"}}, "required": ["wallet_address"]}}}]
}{
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "That wallet holds about $1,284,300 on Ethereum mainnet.",
"tool_calls": null
},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 148, "completion_tokens": 21, "total_tokens": 169}
}Edge cases
- Keep resending
toolson every follow-up request. Omitting it after a tool result leaves the model unable to reference the definition. - Parallel calls: a single assistant message can contain several entries in
tool_calls. Answer every one with its owntoolmessage before the next request, matching eachtool_call_id. - Tool failures: return the error as the
toolmessage content — for example{"error": "address not found"}. Do not omit the message; the conversation becomes invalid without it. - Streaming tool calls: arguments arrive fragmented across
delta.tool_calls[].function.argumentschunks. Concatenate byindexbefore parsing — an individual fragment is not valid JSON. - Name collisions: your tool names take precedence, because built-in tools are not injected once you send any tool.
- Tool budgets: providers cap how many tool definitions they accept per request — roughly 50 for Claude models and around 100 for others. Keep your set small.
Models
GET /v1/models is authoritative. Send your API key and it returns exactly the models your account can call; call it unauthenticated and you get the Personal free-tier view. The tables below are generated from the same catalog.# Everything your account can call
curl https://api.altllm.ai/v1/models -H "Authorization: Bearer $ALTLLM_API_KEY"
# One model
curl https://api.altllm.ai/v1/models/altllm-standard -H "Authorization: Bearer $ALTLLM_API_KEY"
# Only models supporting specific sampling parameters
curl "https://api.altllm.ai/v1/models?supported_parameters=temperature,top_p"{
"id": "altllm-standard",
"object": "model",
"name": "AltLLM Standard",
"created": 1704067200,
"description": "AltLLM's recommended tier for daily crypto and DeFi tasks...",
"pricing": {
"prompt": "0.000002",
"completion": "0.000009",
"image": "0",
"request": "0",
"input_cache_read": "0.0000002"
},
"context_length": 1048576,
"max_output_length": 65536,
"input_modalities": ["text"],
"output_modalities": ["text"],
"quantization": "bf16",
"architecture": {"tokenizer": "AltLLM", "instruct_type": "chat"},
"top_provider": {"is_moderated": false},
"supported_sampling_parameters": ["temperature", "top_p", "top_k", "stop", "seed", "..."],
"supported_features": ["tools", "json_mode", "structured_outputs", "streaming", "reasoning", "prompt_caching"]
}Catalog models (15)
Personal-track models, included in your subscription and gated by tier. Prices are USD per million tokens.
| Model | Context | Max output | Input / Output | Features | Notes |
|---|---|---|---|---|---|
altllm-standardAltLLM Standard | 1.05M | 66K | $2.00 / $9.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 128; cached input $0.20/M |
altllm-native-fastAltLLM Native Fast | 1.05M | 66K | $0.20 / $0.80 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | no built-in crypto tools; cached input $0.02/M |
altllm-native-flashAltLLM Native Flash | 1.05M | 66K | $0.14 / $0.80 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 4,096; no built-in crypto tools; cached input $0.01/M |
altllm-native-lightAltLLM Native Light | 1.05M | 66K | $0.20 / $0.80 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | no built-in crypto tools; cached input $0.02/M |
altllm-native-standardAltLLM Native Standard | 1.05M | 66K | $1.20 / $4.40 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 4,096; no built-in crypto tools; cached input $0.12/M |
altllm-native-promaxAltLLM Native Pro Max | 1.05M | 66K | $4.80 / $28.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | no built-in crypto tools; cached input $0.48/M |
altllm-basicAltLLM Basic | 1.05M | 66K | $5.00 / $30.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | cached input $0.50/M |
altllm-proAltLLM Pro | 1.05M | 66K | $0.80 / $3.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 128; cached input $0.08/M |
altllm-light-codingAltLLM Light Coding | 1.05M | 66K | $0.20 / $0.76 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 4,096; cached input $0.02/M |
altllm-pro-codingAltLLM Pro Coding | 1.05M | 66K | $0.80 / $3.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 128; cached input $0.08/M |
altllm-max-codingAltLLM Max Coding | 1.05M | 66K | $2.00 / $10.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | provider safety filters; cached input $0.20/M |
altllm-hybrid-semanticAltLLM Hybrid Semantic | 1.05M | 66K | Selected tier | toolsjson_modestructured_outputsstreamingreasoningrouting | billed at selected target tier |
altllm-hybrid-costAltLLM Hybrid Cost | 1.05M | 66K | Selected tier | toolsjson_modestructured_outputsstreamingreasoningrouting | billed at selected target tier |
altllm-hybrid-tieredAltLLM Hybrid Tiered | 1.05M | 66K | Selected tier | toolsjson_modestructured_outputsstreamingreasoningrouting | billed at selected target tier |
altllm-hybrid-turnAltLLM Hybrid Turn | 1.05M | 66K | Selected tier | toolsjson_modestructured_outputsstreamingreasoningrouting | billed at selected target tier |
Business Flex models (2)
Provider-native frontier models, billed per token rather than against a subscription. They require the Business track on the Flex tier and are invisible to every other account — filtered out of GET /v1/models, and 404 on direct lookup.
| Model | Context | Max output | Input / Output | Features | Notes |
|---|---|---|---|---|---|
altllm-flex-gpt-5.6Flex: GPT-5.6 Sol | 1.05M | 128K | $6.00 / $36.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | drops frequency_penalty, logit_bias, presence_penalty, seed, temperature, top_p; no built-in crypto tools; cached input $0.60/M |
altllm-flex-gemini-3.6Flex: Gemini 3.6 Flash Native | 1.05M | 66K | $1.80 / $9.00 | toolsjson_modestructured_outputsstreamingreasoningprompt_caching | max_tokens ≥ 128; drops frequency_penalty, presence_penalty, temperature, top_k, top_p; no built-in crypto tools; cached input $0.18/M |
Provider data handling on Flex models
Access and entitlement
- Each model has a minimum plan tier, configured in the Portal. Requesting one above your plan returns 403
tier_upgrade_required, naming therequired_tier. - Flex SKUs are track-locked and fail closed: an account without the Business track cannot reach them under any tier.
- Models an administrator has disabled disappear from discovery for everyone.
- Because entitlement is per account, do not hard-code a model list. Call
GET /v1/modelsat startup and select from what comes back.
Capabilities at a glance
- Streaming: every published model.
- Function calling: every published model.
- Prompt caching: altllm-standard, altllm-native-fast, altllm-native-flash, altllm-native-light, altllm-native-standard, altllm-native-promax, altllm-basic, altllm-pro, altllm-light-coding, altllm-pro-coding, altllm-max-coding, altllm-flex-gpt-5.6, altllm-flex-gemini-3.6. Cached input bills at the reduced rate shown in each model’s
pricing.input_cache_read. - Reasoning: models listing
reasoninginsupported_featuresgenerate hidden reasoning tokens that count as output tokens and are billed as such. - Vision: not reachable — see content parts.
Usage and rate limits
| Limit | Unit | Scope | Default |
|---|---|---|---|
| Requests per minute | requests | account × model | 60 |
| Tokens per minute | reserved input + output tokens | account × model | 100,000 |
| Window | seconds | sliding | 60 |
| Context window | tokens | per request, per model | See the model tables |
| Max output | tokens | per request, per model | See the model tables |
Counters are isolated per account and model; ceilings are configured per model. The values above are the fallback when a model has no explicit configuration. Before generation, the TPM check atomically reserves estimated message and multimodal input plus prompt-bearing tool/structured-output schemas plus the largest valid explicit max_output_tokens, max_completion_tokens, or max_tokens budget. If none is set, it reserves a conservative 4,096-token output budget. Set an explicit realistic maximum to avoid consuming unused concurrent capacity. Completed usage replaces that reservation when the provider reports final usage. If the atomic reservation service is unavailable, AltLLM rejects the request before provider invocation with a retryable 503 rate_limit_unavailable instead of silently bypassing your account limits.
Context capacity and throughput are separate limits
Rate-limit headers
Rate-limit headers appear only on 429 responses
X-Request-ID and, when applicable, X-Credits-Expire-At — but no X-RateLimit-* headers. You cannot track remaining quota by reading successful responses. Handle 429s instead.| Header | Present on | Meaning |
|---|---|---|
X-Request-ID | Every chat completion and response, success or failure | Correlation ID for the request. Echoes your own X-Request-ID if you send one. Include it in support reports. |
X-Credits-Expire-At | When your credits have an expiry date | ISO-8601 timestamp at which the current credit grant expires. |
X-RateLimit-Limit | 429 responses only | Requests-per-minute ceiling for this model. |
X-RateLimit-Remaining | 429 responses only | Requests left in the current 60-second window. |
X-RateLimit-Reset | 429 responses only | Unix timestamp (seconds) when the window resets. |
X-RateLimit-Limit-Tokens | 429 responses only | Tokens-per-minute ceiling for this model. |
X-RateLimit-Remaining-Tokens | 429 responses only | Tokens left in the current 60-second window. |
Retry-After | 429 responses, and retryable rate-limit 503 responses | Seconds to wait before retrying. |
How usage is billed
- Cost is
prompt_tokens × input rate + completion_tokens × output rate, at the per-model rates in the tables above, deducted from your credit balance. - Cached input is billed at
input_cache_readinstead of the standard input rate, and is subtracted from the fresh-input count rather than charged twice. - Reasoning tokens are part of
completion_tokensand are billed at the output rate even though you never see them. - Server-side tool execution adds tokens to
prompt_tokens: tool schemas and results re-enter the model as input. A tool-using turn costs more than the same question without tools. - Interrupted Chat tool streams create no charge unless the gateway receives both a valid terminal event and final usage. Other stream modes settle from completed provider-reported usage.
- Credits expire. Check
X-Credits-Expire-At, and see the pricing page for cycle behavior.
Errors and retries
{
"error": {
"message": "Human-readable description",
"type": "insufficient_credits",
"code": "payment_required"
}
}| Status | type | code | Meaning | Retry | Extra fields |
|---|---|---|---|---|---|
| 401 | authentication_error | missing_api_key | No Authorization: Bearer header was sent. | fix, then retry | — |
| 401 | authentication_error | invalid_api_key | The key is unknown, disabled, or revoked. | fix, then retry | — |
| 402 | insufficient_credits | payment_required | Balance is too low to start the request. | fix, then retry | balance, billing_url |
| 402 | credits_expired | payment_required | Your credits passed their expiry date. | fix, then retry | balance, expired_at, billing_url |
| 402 | no_account | payment_required | The identity behind the request has no Portal account. | fix, then retry | balance, billing_url |
| 403 | tier_upgrade_required | forbidden | Your plan does not include this model. | fix, then retry | current_tier, required_tier, model, upgrade_url |
| 404 | not_found | route_not_found | Authenticated request to a path AltLLM does not serve. | not retryable | — |
| 404 | invalid_request_error | previous_response_not_found | The requested conversation state is missing, expired, evicted, or belongs to a different account. | fix, then retry | — |
| 400 | invalid_request_error | — | The request violates a gateway rule, such as max_tokens below a model's documented minimum. | fix, then retry | — |
| 429 | rate_limit_exceeded | rpm_exceeded | Too many requests for this model in the last 60 seconds. | retryable | — |
| 503 | service_unavailable | rate_limit_unavailable | Atomic RPM/TPM reservation is temporarily unavailable, so the request was not sent upstream. | retryable | — |
| 503 | server_error | credit_reservation_unavailable | The prepaid-credit hold or its durable settlement state is temporarily unavailable, so the request cannot continue safely. | retryable | — |
| 503 | server_error | admission_state_unavailable | A multi-step request lost the internal admission state required before its next provider call. | retryable | — |
| 429 | rate_limit_exceeded | tpm_exceeded | Too many tokens for this model in the last 60 seconds. | retryable | — |
| 200 | upstream_stream_error | stream_interrupted | The provider stream ended before completion. Delivered as an SSE error event after the 200 headers were already sent. | retryable | — |
| 500 | gateway_error | — | Gateway processing failed after request validation. Quote the X-Request-ID when contacting support. | retryable | — |
| 500 | server_error | internal_error | Unhandled gateway failure. Quote the X-Request-ID when contacting support. | retryable | — |
| 502 | upstream_error | — | The provider failed or returned an unparseable body. code carries the upstream status when there is one. | retryable | — |
| 503 | service_unavailable | — | Response storage is unavailable, so previous_response_id cannot be resolved on /v1/responses. | retryable | — |
Retry policy
- 429: honor
Retry-Afterwhen present; otherwise back off exponentially from one second with jitter. - 500, 502, 503: retry up to three times with exponential backoff; honor
Retry-Afterwhen present. - Interrupted streams: retry the whole request. Partial output is not resumable.
- 400, 401, 402, 403, 404: never retry unchanged. Fix the request, the key, the balance, or the plan first.
Reporting a problem
Include the X-Request-ID response header, the model ID, the UTC timestamp, and the full error body. That header is the correlation key in AltLLM’s logs — a report without it is much slower to investigate.
SDK examples
from openai import OpenAI
client = OpenAI(
api_key=os.environ["ALTLLM_API_KEY"],
base_url="https://api.altllm.ai/v1",
)
response = client.chat.completions.create(
model="altllm-standard",
messages=[{"role": "user", "content": "Explain rollup data availability."}],
max_tokens=500,
)
print(response.choices[0].message.content)
print(response.usage.total_tokens)stream = client.chat.completions.create(
model="altllm-standard",
messages=[{"role": "user", "content": "Name three L2 rollups."}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.ALTLLM_API_KEY,
baseURL: "https://api.altllm.ai/v1",
});
const response = await client.chat.completions.create({
model: "altllm-standard",
messages: [{ role: "user", content: "Explain rollup data availability." }],
max_tokens: 500,
});
console.log(response.choices[0].message.content);const stream = await client.chat.completions.create({
model: "altllm-standard",
messages: [{ role: "user", content: "Name three L2 rollups." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}# Never hard-code a model list — entitlements are per account.
available = {m.id for m in client.models.list().data}
for candidate in ("altllm-flex-gpt-5.6", "altllm-basic", "altllm-standard"):
if candidate in available:
model = candidate
break
else:
raise RuntimeError("no usable model for this account")import time
import openai
def complete_with_retry(**kwargs):
for attempt in range(4):
try:
return client.chat.completions.create(**kwargs)
except openai.RateLimitError as exc:
wait = float(exc.response.headers.get("Retry-After", 2 ** attempt))
time.sleep(wait)
except openai.InternalServerError:
time.sleep(2 ** attempt)
except (openai.AuthenticationError, openai.PermissionDeniedError, openai.BadRequestError):
# Key, plan, balance, or request shape — retrying changes nothing.
raise
raise RuntimeError("exhausted retries")Versioning and lifecycle
| Topic | Policy |
|---|---|
| API version | /v1. A breaking change to request or response shapes would ship under a new prefix, not by mutating /v1. |
| Additive changes | New endpoints, new optional request fields, and new response fields can appear in /v1 at any time. Ignore unknown response fields rather than failing on them. |
| Model IDs | Catalog IDs such as altllm-standard are stable aliases. The provider model behind an alias can change; the ID, its pricing, and its advertised capabilities are the contract. |
| Pinned versions | Flex SKUs carry the provider version in the ID — altllm-flex-gpt-5.6, altllm-flex-gemini-3.6. A new provider version arrives as a new ID rather than changing an existing one. |
| Deprecation | A superseded model can remain callable as an unlisted compatibility alias for pinned integrations while disappearing from discovery, detail, provider-pricing, SDK, and current guidance surfaces. Removal of that compatibility path is a separate breaking lifecycle decision; do not infer it merely from absence in GET /v1/models. |
| Migration | Discover models at startup and fall back across a preference list, as in the SDK example above. An application that hard-codes one ID breaks on retirement. |
| Changelog | Model and pricing changes are reflected in GET /v1/models and on the Models and Pricing pages as they ship. |