Error Handling

Every error AltLLM returns, what causes it, and whether retrying helps.

Error response format

Errors use the OpenAI envelope, so an OpenAI client raises its usual exception types. Billing and entitlement errors add fields alongside the standard three, letting you react without parsing the message string.

{
  "error": {
    "message": "Human-readable description",
    "type": "insufficient_credits",
    "code": "payment_required",

    // Present on 402 balance and credit errors:
    "balance": 0.0,
    "billing_url": "https://platform.altllm.ai/billing",

    // Present on 403 entitlement errors:
    "current_tier": "free",
    "required_tier": "flex",
    "model": "altllm-flex-gpt-5.6",
    "upgrade_url": "https://platform.altllm.ai/billing/upgrade"
  }
}

Errors raised by an upstream provider are passed through with the provider's own type, so treat type as an open set and branch on the HTTP status first.

HTTP status codes

200Success

Completion returned, or an SSE stream opened. A stream can still fail mid-flight — see below.

400Bad Request

Malformed JSON, or a value the gateway rejects such as an out-of-range max_tokens.

401Unauthorized

The Authorization header is missing, or the key is unknown, disabled, or revoked.

402Payment Required

The account has no balance, its credits expired, or it has no Portal account.

403Forbidden

The key is valid but the plan does not include the requested model.

404Not Found

Unknown model, a model your account cannot access, or a route AltLLM does not serve.

429Rate Limited

The per-model RPM or TPM budget for this account is exhausted. Check Retry-After.

500Internal Error

Gateway failure. Retry with exponential backoff.

502Bad Gateway

An upstream provider error passed through with the provider's own error body.

Error types

The complete set of error.type values the gateway itself emits.

StatusTypeCodeMeaningRetry
401authentication_errormissing_api_keyNo Authorization: Bearer header was sent.Fix, then retry
401authentication_errorinvalid_api_keyThe key is unknown, disabled, or revoked.Fix, then retry
402insufficient_creditspayment_requiredBalance is too low to start the request.Fix, then retry
402credits_expiredpayment_requiredYour credits passed their expiry date.Fix, then retry
402no_accountpayment_requiredThe identity behind the request has no Portal account.Fix, then retry
403tier_upgrade_requiredforbiddenYour plan does not include this model.Fix, then retry
404not_foundroute_not_foundAuthenticated request to a path AltLLM does not serve.No
404invalid_request_errorprevious_response_not_foundThe requested conversation state is missing, expired, evicted, or belongs to a different account.Fix, then retry
400invalid_request_error—The request violates a gateway rule, such as max_tokens below a model's documented minimum.Fix, then retry
429rate_limit_exceededrpm_exceededToo many requests for this model in the last 60 seconds.Yes
503service_unavailablerate_limit_unavailableAtomic RPM/TPM reservation is temporarily unavailable, so the request was not sent upstream.Yes
503server_errorcredit_reservation_unavailableThe prepaid-credit hold or its durable settlement state is temporarily unavailable, so the request cannot continue safely.Yes
503server_erroradmission_state_unavailableA multi-step request lost the internal admission state required before its next provider call.Yes
429rate_limit_exceededtpm_exceededToo many tokens for this model in the last 60 seconds.Yes
200upstream_stream_errorstream_interruptedThe provider stream ended before completion. Delivered as an SSE error event after the 200 headers were already sent.Yes
500gateway_error—Gateway processing failed after request validation. Quote the X-Request-ID when contacting support.Yes
500server_errorinternal_errorUnhandled gateway failure. Quote the X-Request-ID when contacting support.Yes
502upstream_error—The provider failed or returned an unparseable body. code carries the upstream status when there is one.Yes
503service_unavailable—Response storage is unavailable, so previous_response_id cannot be resolved on /v1/responses.Yes

Errors during a stream

Once a streaming response starts, the status is already 200 and cannot change. A provider failure mid-generation arrives as an error event inside the stream, and the connection then closes without the usual data: [DONE] terminator.

data: {"error":{"message":"The model stream ended before completion. Please retry the request.","type":"upstream_stream_error","code":"stream_interrupted"}}

Treat any stream that ends without [DONE] as incomplete. Retrying is safe; you are billed only for tokens generated before the interruption.

Handling examples

Python

import os
import time

import openai
from openai import OpenAI

client = OpenAI(
    base_url="https://api.altllm.ai/v1",
    api_key=os.environ["ALTLLM_API_KEY"],
)


def chat_with_retry(message: str, max_retries: int = 3) -> str:
    for attempt in range(max_retries):
        try:
            response = client.chat.completions.create(
                model="altllm-standard",
                messages=[{"role": "user", "content": message}],
            )
            return response.choices[0].message.content or ""

        except openai.RateLimitError as exc:
            # Retry-After is present when the gateway can compute a wait.
            wait = float(exc.response.headers.get("Retry-After", 2 ** attempt))
            if attempt == max_retries - 1:
                raise
            time.sleep(wait)

        except openai.InternalServerError:
            if attempt == max_retries - 1:
                raise
            time.sleep(2 ** attempt)

        except openai.AuthenticationError:
            # 401 — the key is wrong. Retrying cannot fix it.
            raise

        except openai.PermissionDeniedError as exc:
            # 403 — valid key, insufficient plan. error.required_tier says what is needed.
            body = exc.response.json().get("error", {})
            raise RuntimeError(
                f"{body.get('model')} needs the {body.get('required_tier')} tier"
            ) from exc

        except openai.APIStatusError as exc:
            # 402 — no balance, expired credits, or no Portal account.
            if exc.status_code == 402:
                body = exc.response.json().get("error", {})
                raise RuntimeError(f"Add credits at {body.get('billing_url')}") from exc
            raise

    raise RuntimeError("exhausted retries")

TypeScript

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.altllm.ai/v1",
  apiKey: process.env.ALTLLM_API_KEY,
});

const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

export async function chatWithRetry(message: string, maxRetries = 3): Promise<string> {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      const response = await client.chat.completions.create({
        model: "altllm-standard",
        messages: [{ role: "user", content: message }],
      });
      return response.choices[0].message.content ?? "";
    } catch (error) {
      if (!(error instanceof OpenAI.APIError)) throw error;

      // 401 / 403 / 402 / 400 — retrying changes nothing.
      if ([400, 401, 402, 403, 404].includes(error.status ?? 0)) throw error;

      if (attempt === maxRetries - 1) throw error;

      if (error.status === 429) {
        const retryAfter = Number(error.headers?.["retry-after"] ?? 2 ** attempt);
        await sleep(retryAfter * 1000);
        continue;
      }

      if ((error.status ?? 0) >= 500) {
        await sleep(2 ** attempt * 1000);
        continue;
      }

      throw error;
    }
  }
  throw new Error("exhausted retries");
}

Best practices

  • Branch on HTTP status, not the message. Message strings change; statuses and the extra error fields do not.
  • Never retry 400, 401, 402, 403, or 404. Fix the request, key, balance, or plan first.
  • Honor Retry-After on 429, and fall back to exponential backoff with jitter when it is absent.
  • Record the X-Request-ID header on failures. It is the correlation key in AltLLM's logs and the first thing support will ask for.
  • Treat a stream ending without [DONE] as failed rather than accepting the partial text.
  • Set generous timeouts — 30–60s for completions, longer for reasoning models and tool-using turns.
  • Watch X-Credits-Expire-At so credit expiry does not surface as a surprise 402.