Developer guide
Overview
At a glance
CLEVI models are called with a single gateway API key. To choose the ID to put in the request's model field, you need the product families and the calls each model allows; to decide which SDK to connect with, you need the base URLs of the two compatible specs.
This document covers the components, terms, product families, call paths, and console pages needed for those two decisions.
The procedure for your first call is in the Quickstart document.
Components
| Component | What it does | Details |
|---|---|---|
| Model families | CIP handles general intelligence and work automation, Cosa handles speech generation and recognition, and Cova handles information extraction from documents and images. The gateway plan table also includes Ivy, which handles embeddings and conversational inference. | Models |
| Gateway | The API entry point under platform.clevi.net. It calls models through the OpenAI-compatible path and the Anthropic-compatible path, and both paths authenticate with a single API key. | Authenticate with an API key, Make API calls |
| Console | Issue and revoke API keys, call models in the Playground, and check usage and limits. | Console pages section below |
| CleviDrive | Stores the output of the speech synthesis paths. The retention period is 3 days; after that, 404. | Make API calls |
Terms
| Term | Meaning |
|---|---|
| Gateway | The API entry point under platform.clevi.net. It provides an OpenAI-compatible path and an Anthropic-compatible path. |
| Compatible spec | A request and response format matched so that you can use the OpenAI SDK or Anthropic SDK you already have by changing only the base URL. |
| Workspace | The unit to which limits apply. Limit values differ by workspace. |
| Plan | Determines which calls each model allows. LLM_FREE, TTS_FREE, and STT_FREE are applied automatically without a request. |
| Allowed calls | The call keys a plan opens for a model, such as chat.completions, anthropic.messages, embeddings, and audio.speech. |
| Product | The billing unit for speech and voice paths. audio/speech and audio/transcriptions select it by model in the request body; voice paths specify it with the X-Product-Sku header or the product_sku query parameter. |
Product families and models
| Product family | Model ID | Primary role | Granting plan |
|---|---|---|---|
| CIP | cip-5.5-im | General-purpose work. Prioritizes response speed and throughput. 360B · 256K input · 64K output · multimodal | LLM_FREE |
| CIP | cip-5.5-mm | Long-context analysis and complex problem solving. 800B · 512K input · 64K output · multimodal | LLM_FREE |
| CIP | cip-5.5-sm | Edge servers and local work automation. Adjustable between 24B and 40B | |
| Cosa | cosa-a, cosa-b | Speech synthesis, cloning, design, and streaming | TTS_FREE |
| Cosa | cosa-asr | Speech recognition | STT_FREE |
| Ivy | ivy-4-embedding, ivy-4-embedding-mm | Embedding generation | LLM_FREE |
| Ivy | ivy-4-mm | Conversational inference | LLM_FREE |
| Cova | cova-1 | LLM-based vision OCR. Text, layout, and table extraction with brief image descriptions |
The models your account can actually call are determined by the Playground model list and the GET /models response. The plan table shows the range opened by the plans and does not include models opened through other routes. In the plan table, cip-5.5-im-chat, cip-5.5-mm-h, and cosa-tts are IDs not described in the product family document, and cip-5.5-sm and cova-1 are not in the plan table.
By connecting the product families, you can build the following flows.
- Document-based work: Cova extracts body text, tables, and image descriptions; CIP selectively checks the originals it needs and performs analysis and task processing; Cosa delivers the results as speech.
- Speech-based work: cosa-asr converts a spoken request to text, and cosa-a or cosa-b reads out the answer that CIP produced.
- On-site automation: cip-5.5-sm classifies on-site data and performs local work, and hands off material that needs further overall analysis to cip-5.5-im or cip-5.5-mm.
Call paths and authentication
Both compatible specs use the same API key, and conversational inference models can be called through either spec. Embeddings and the speech and voice paths exist only in the OpenAI-compatible spec.
| Item | Value |
|---|---|
| OpenAI-compatible base URL | https://platform.clevi.net/api/v2/aiservice/openai/v1 |
| Anthropic-compatible base URL | https://platform.clevi.net/api/v2/aiservice/anthropic (the SDK appends /v1 itself) |
| Authentication header | X-API-Key: <key> or Authorization: Bearer <key>. If both are specified, the X-API-Key value takes precedence |
| OpenAI-compatible spec paths | chat/completions, responses, completions, embeddings, models, audio, voices |
| Anthropic-compatible spec paths | v1/messages, v1/messages/count_tokens, v1/models |
| Streaming | Set stream: true in the request body to receive a Server-Sent Events response. embeddings does not stream |
| Key issuance page | The API Keys page of the console |
| Limits and usage | Differ by workspace; check the Overview, Plan and credits, and Usage pages of the console |
Errors are reported with HTTP status codes. If the upstream returns a 4xx or 5xx, that code is passed through as is, and 402 is never returned. Header rules are in the Authenticate with an API key document, per-path behavior in the Make API calls document, and what to do for each code in the Error codes document.
Console pages
| Page | What it does |
|---|---|
| Getting started guide | Walks through issuing a key to making your first call in 3 steps. |
| API Keys | Manages key issuance, revocation, and scope. |
| Playground | Shows the list of callable models and lets you call them directly from the browser. |
| Usage | Shows call volume per model and per key, and estimated credit deductions. |
| Overview | Shows the allocated usage limits applied to the workspace. |
| Plan and credits | Shows the range your plan allows and your balance. |
| Help and Support | Where to report repeated 5xx errors, together with the X-Request-Id. |
Related documents
- Quickstart: from issuing a key to your first chat/completions call and streaming
- Models: the full model list, specifications, and allowed calls
- Authenticate with an API key: authentication headers, base URLs, optional headers
- Make API calls: paths by spec, streaming rules, speech and voice rules
- Error codes: what to do for each status code, retries, error bodies