# SarangAI — Complete Reference (LLM-optimized) This is the full plain-text documentation of the SarangAI AI gateway, its CLI, and its public API surface, written for LLM crawlers and coding agents. A shorter index is available at https://sarangai.id/llms.txt. - Product: SarangAI — unified AI gateway (https://sarangai.id) - Gateway base URL: `https://sarangai.id/api/gateway/v1` - Official CLI: `sarangai-cli` (https://github.com/sarangpenyamun/sarangai-cli) - API style: OpenAI-compatible (drop-in for OpenAI SDKs and agents) - Billing: prepaid credit wallet, per-token accounting in credits (USD-equivalent pricing + markup) - Auth: `Authorization: Bearer sk-sarang-...` API keys ## 1. Overview SarangAI exposes one OpenAI-compatible endpoint that routes to 200+ models from multiple providers (Anthropic Claude, OpenAI GPT-4o, Google Gemini, DeepSeek, Meta Llama, Qwen, Mistral, xAI Grok, Groq, MiniMax, Kimi, and more). Gateway capabilities: - Single API key for all providers. - Automatic transient-error retries with exponential backoff. - Provider fallback routing (OpenRouter `route: "fallback"` behavior). - Equivalent-tier internal model cascade: if the primary model is unavailable, up to 3 tier-equivalent models are tried transparently; billing follows the model that actually served the request. - Admin-disabled models are silently rerouted instead of failing with 403. - Per-key rate limiting: 10 requests per 10-second window per API key. - Request body limit: 1 MiB (1,048,576 bytes). - Streaming (SSE) support with forced `stream_options.include_usage: true`. - Minimum balance check: requests are rejected with HTTP 402 when credits are depleted. ## 2. Authentication All gateway endpoints require a Bearer API key. ``` Authorization: Bearer sk-sarang-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx ``` - Key format: `sk-sarang-` followed by 48 hexadecimal characters. - Keys are created in the dashboard at https://sarangai.id/user/api-keys. - The CLI can also mint a key automatically through the device-code flow (see section 6). - Invalid or inactive keys return HTTP 401 with `{"error": {"message": "Invalid API key", "type": "auth_error"}}`. Full authentication documentation: https://sarangai.id/auth.md ## 3. Endpoints ### 3.1 POST /api/gateway/v1/chat/completions OpenAI-compatible chat completions. Streams when `"stream": true`. ``` curl https://sarangai.id/api/gateway/v1/chat/completions \ -H "Authorization: Bearer sk-sarang-..." \ -H "Content-Type: application/json" \ -d '{ "model": "anthropic/claude-sonnet-4.6", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Hello SarangAI!"} ] }' ``` Request fields (OpenAI schema): - `model` (string, required): provider model ID. Common aliases are accepted and normalized, e.g. `gpt-4o` -> `openai/gpt-4o`, `claude-3.5-sonnet` -> `anthropic/claude-3.5-sonnet-20241022`, `deepseek-v3` -> `deepseek/deepseek-chat`. Unknown IDs are forwarded as-is so provider-native "model not found" details surface to the client. - `messages` (array, required): chat messages with `role` (`system` | `user` | `assistant` | `tool`) and `content`. - `stream` (boolean, optional): when `true`, respond with `text/event-stream`. - `temperature`, `top_p`, `max_tokens`, `tools`, `tool_choice`, and other standard OpenAI parameters are forwarded to the upstream provider. Responses: - `200` with a standard OpenAI chat completion JSON (non-streaming), including `usage.prompt_tokens`, `usage.completion_tokens`, `usage.total_tokens`. - `200` with SSE chunks (streaming). Vendor comment lines are stripped; a final chunk carries usage because `stream_options.include_usage` is forced on. - `401` invalid API key; `402` insufficient credit; `413` body over 1 MiB; `429` rate limited (includes `Retry-After: 10`, `X-RateLimit-Limit`, `X-RateLimit-Remaining` headers); `502` invalid upstream response; `503` upstream unavailable or gateway not configured; `504` timeout. Python (OpenAI SDK): ```python from openai import OpenAI client = OpenAI( base_url="https://sarangai.id/api/gateway/v1", api_key="sk-sarang-...", ) resp = client.chat.completions.create( model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello SarangAI!"}], ) print(resp.choices[0].message.content) ``` ### 3.2 GET /api/gateway/v1/models List available models. Public catalog data, OpenAI list format: ``` {"object": "list", "data": [ {"id": "openai/gpt-4o", ...}, ... ]} ``` - URL: `https://sarangai.id/api/gateway/v1/models` - Auth: not required. - Pricing per model is available at https://sarangai.id/models and in the `pricing` fields of the catalog entries. ### 3.3 GET /api/gateway/v1/balance Account balance and identity for the presented key. ``` curl https://sarangai.id/api/gateway/v1/balance \ -H "Authorization: Bearer sk-sarang-..." ``` Response: ```json { "balance": 12.5, "email": "dev@example.com", "name": "Dev", "tier": "DEVELOPER", "accountId": "0A1B2C3D", "currency": "CREDIT" } ``` - `balance`: remaining credits. - `accountId`: 8-character suffix shown in the dashboard and CLI (format SA-XXXXXXXX). - Auth: required (HTTP 401 otherwise). ## 4. Rate limits and billing rules - Rate limit: 10 requests / 10 seconds per API key. Invalid-key attempts are additionally throttled per IP to prevent auth flooding. - Billing: credits are deducted per request based on the served model's input and output token prices (USD-equivalent pricing with a platform markup, converted to credits). A minimum deduction applies to very small requests. - Streaming requests are billed from a tee'd copy of the provider stream, so usage settles even if the client disconnects mid-stream. - Failed upstream requests are logged but not billed. ## 5. CLI (sarangai-cli) Official terminal client. Repository: https://github.com/sarangpenyamun/sarangai-cli ```bash npm install -g sarangai-cli sarang login # device-code authorization, stores the API key locally sarang run "Hello SarangAI!" # one-shot chat with the default model ``` - `sarang login` opens https://sarangai.id/auth/cli?code=SA-XXXXXXXX and polls `GET /api/auth/cli/poll?code=SA-XXXXXXXX` until the session is approved. - `sarang run` calls `POST /api/gateway/v1/chat/completions` with the stored key. - The CLI also exposes balance info from `GET /api/gateway/v1/balance` (email, tier, accountId SA-XXXXXXXX, credit balance). ## 6. CLI device-code authentication flow (summary) 1. CLI generates a session code (format `SA-XXXXXXXX`, 8 uppercase alphanumeric characters) and prints it. 2. CLI starts polling `GET https://sarangai.id/api/auth/cli/poll?code=SA-XXXXXXXX`. 3. While unapproved, the endpoint returns `{"status": "PENDING"}` (HTTP 200). 4. The user opens the printed URL, logs into sarangai.id, and approves. 5. The server mints a dedicated API key named `SarangAI CLI Key` (`sk-sarang-` + 48 hex chars) and marks the session `APPROVED`. 6. The next poll returns `{"apiKey": "sk-sarang-..."}` exactly once; the session is then deleted, so the raw key is never served twice. 7. CLI stores the key and uses it as `Authorization: Bearer sk-sarang-...`. Full details: https://sarangai.id/auth.md ## 7. IDE / agent integrations Point any OpenAI-compatible tool at the gateway: - Cursor / Continue.dev (`~/.continue/config.json`): ```json { "models": [ { "title": "SarangAI", "provider": "openai", "model": "openai/gpt-4o", "apiKey": "sk-sarang-...", "apiBase": "https://sarangai.id/api/gateway/v1" } ] } ``` - Claude Code: ```bash export ANTHROPIC_BASE_URL="https://sarangai.id/api/gateway/v1" export ANTHROPIC_API_KEY="sk-sarang-..." ``` - Cline / Roo Code (VS Code settings JSON): ```json { "cline.apiKey": "sk-sarang-...", "cline.openAiBaseUrl": "https://sarangai.id/api/gateway/v1" } ``` ## 8. Public web tools (no auth) Free diagnostics utilities at https://sarangai.id/tools, including: DNS lookup, SSL/SSL-chain inspection, HTTP & HTTP/2 checks, response headers, PageSpeed, port scan, subdomain discovery, RDAP/whois, redirect tracing, reverse IP, CMS detection, tech stack, typosquat check, meta tags, robots/sitemap preview, broken links, link extraction, favicon, QR generator, ping, propagation, domain parking, valuation, and archive lookup. ## 9. Other public resources - Website: https://sarangai.id - Models & rates: https://sarangai.id/models - Playground: https://sarangai.id/playground - Benchmarks: https://sarangai.id/benchmarks - Blog: https://sarangai.id/blog - FAQ: https://sarangai.id/faq - Status page: https://sarangai.id/status - Terms: https://sarangai.id/terms | Privacy: https://sarangai.id/privacy - Machine-readable API catalog: https://sarangai.id/.well-known/api-catalog.json - Short LLM index: https://sarangai.id/llms.txt - This file: https://sarangai.id/llms-full.txt