Developer guide
Models
Model summary
CLEVI models are divided into three product families: CIP, which handles general intelligence and work automation; Cosa, which handles voice generation and recognition; and Cova, which handles information extraction from documents and images. Each product can be used independently or connected into services that progress from document recognition and analysis to task execution and voice guidance.
| Product | Model ID | Key specifications | Primary role |
|---|---|---|---|
| CIP-5.5-SM | cip-5.5-sm | 24B–40B adjustable | Edge server and local work automation |
| CIP-5.5-IM | cip-5.5-im | 360B · 256K input · 64K output · multimodal | General-purpose work processing. Prioritises response speed and throughput |
| CIP-5.5-MM | cip-5.5-mm | 800B · 512K input · 64K output · multimodal | Long-context analysis and complex problem-solving |
| Cosa-A | cosa-a | Voice synthesis · cloning · design · streaming | Voice generation with detailed synthesis control |
| Cosa-B | cosa-b | Voice synthesis · cloning · design · streaming | Voice selection and customised voice generation |
| Cosa ASR | cosa-asr | Speech recognition | Converts voice input to text |
| Cova | cova-1 | LLM-based vision OCR dedicated model | Text, layout and table extraction, plus simple image descriptions |
CIP
CIP handles general intelligence and work automation. It consists of cip-5.5-im and cip-5.5-mm, accessed through the gateway, and cip-5.5-sm, deployed on edge servers and on-premises.
cip-5.5-im and cip-5.5-mm
| Category | cip-5.5-im | cip-5.5-mm |
|---|---|---|
| Scale | 360B | 800B |
| Input limit | 256K | 512K |
| Maximum output | 64K | 64K |
| Input formats | Text and visual materials (documents, screen images, tables and charts) | Text and visual materials (integrated review of code, documents and screen materials) |
| Design focus | Handles everyday knowledge work and development tasks, with a focus on response speed, cost efficiency and throughput. | Handles complex work that integrates multiple sources and conditions and connects analysis results to practical tasks. |
| Use of long inputs | Reviews multiple documents, code files and conversation histories together. | Handles extensive reference materials and work histories together, and performs agent tasks involving multiple stages of research, execution and verification. |
cip-5.5-im is suited to tasks with clear objectives and scope. Typical uses include extracting the required information from provided materials, producing outputs in a specified format, and modifying code to meet requirements. The maximum 64K output is useful for producing detailed reports or outputs consisting of multiple sections.
- Development assistance: Write functions and features, fix well-defined bugs, and implement changes in line with existing patterns.
- Test generation: Create test drafts and verification cases based on requirements and implementation.
- Document processing: Handle summarisation, classification, item extraction, format conversion and drafting.
- Multimodal analysis: Review screen images and tables and charts in documents together with relevant descriptions.
- Interactive work assistance: Answer questions based on work materials and produce the required outputs.
- Batch processing: Use for individual summaries of multiple documents, request classification and data conversion.
- Sub-agents: Handle clearly bounded subtasks such as code exploration, reviewing specific files and organising materials.
For example, it can be used to produce proposed changes, tests and change descriptions from feature requirements and related code, or to classify submitted documents and create summaries for the people responsible.
cip-5.5-mm is suited to tasks that require understanding relationships between information and maintaining consistency across the overall task. In development, it reviews the connections between requirements, design, implementation and testing; in document analysis, it synthesises the claims and evidence, differences and contradictions across multiple sources.
- Complex feature implementation: Change code, tests and documentation together while considering the contracts and behaviour of multiple modules.
- Large-scale code analysis: Review call relationships, data flows and the scope of change impacts.
- Incident root-cause analysis: Compare logs, configurations, code and execution results to organise hypotheses about the cause and procedures for verification.
- Design and decision support: Structure requirements and constraints and analyse the impact of implementation alternatives.
- Cross-document synthesis: Organise commonalities, differences and contradictions in reports and technical documents based on evidence.
- Multimodal integrated review: Interpret text and visual materials together to analyse issues and requirements.
- Multi-step agents: Use for tasks involving investigation, planning, tool execution and result review.
- Specialist document writing: Produce technical reports, design proposals and detailed review documents that reflect complex conditions and supporting evidence.
For example, it is used to review relevant codes and logs from errors occurring across multiple services, or for tasks that span requirements analysis, implementation and review of verification results for complex features.
cip-5.5-sm
cip-5.5-sm performs local inference and task automation on edge servers and in on-premises environments. As the model size can be adjusted within the 24B~40B range, you can choose a configuration that suits the computing resources of the deployment equipment and the required performance on site. As processing takes place close to where data is generated, it is used to integrate with on-site systems and utilise internal data. By configuring the required models and tools locally, it can also operate in environments with limited external connectivity.
It interprets information generated on site and connects it to processing procedures. It classifies logs and requests, extracts the required information from data, and performs repetitive tasks using connected internal tools and scripts.
- On-site event processing: Classifies equipment logs and status information, and selects events requiring attention.
- Data preprocessing: Extracts key items from work records and classifies, tags and normalises them.
- Request routing: Forwards received requests to the responsible system or processing procedure.
- Local task agent: Performs repetitive tasks by querying internal systems and running scripts.
- On-premises task support: Answers questions based on internal materials and organises task-related content.
- Follow-up analysis support: Selects materials requiring further review and summarises the key content.
For example, it is used in a workflow that classifies operational logs on a business site server, retrieves related history and then prepares an incident summary for the person in charge.
Cosa
Cosa is a suite of speech products that converts text to speech, recognises input speech as text and registers and reuses voices of your choice. It is used for service guidance, content creation, accessibility support and voice interfaces for conversational agents.
cosa-a and cosa-b handle speech generation, while cosa-asr handles speech recognition. cosa-a and cosa-b are separate models that provide their own voice lists and synthesis settings, and are selected through the same service interface. The functions shared by both models are as follows.
- Text-to-speech: synthesises the entered text using the selected voice.
- Voice cloning: registers a voice based on a reference recording for which you have permission to use it, then synthesises new sentences.
- Voice design: creates and registers a voice by describing the characteristics of the desired voice.
- Voice reuse: uses a registered voice for subsequent synthesis.
- Synthesis settings: adjusts the speaking speed, language and other options.
- Streaming delivery: synthesises text phrase by phrase as it arrives and delivers the audio sequentially.
- Playback control: supports pausing, resuming, stopping and entering additional sentences.
With progressive text input, voice playback can start before the entire response is complete. While CIP writes the response, you can configure a service in which Cosa reads out the prepared phrases first.
cosa-a
cosa-a supports text-to-speech, voice cloning, voice design and streaming output. In addition to basic functions for selecting the voice, language and speaking speed, it provides additional controls such as settings related to the number of synthesis steps, guidance strength and generation length. Use it to adjust settings and review results during voice production, or apply multiple settings to the same script and choose the result.
- Synthesises announcements and notifications for apps, kiosks and business systems.
- Creates audio versions of educational materials and user manuals.
- Produces narration for video and audio content.
- Produces recurring content using a registered voice.
- Used for streaming voice output from conversational agents.
- Used for voice production with detailed synthesis settings.
cosa-b
cosa-b supports preset voice selection, description-based voice design, reference-voice-based cloning and streaming output. Choose a voice from its own voice library or register a new voice for use in a service.
Voice cloning uses a reference recording and its corresponding text. It also supports a process that generates the reference text using speech recognition during registration, and registered voices can be reused for subsequent synthesis.
- Configures the voice persona of services and characters.
- Creates a customised voice from a description or reference recording.
- Synthesises brand announcements, character dialogue and content narration.
- Used for voice responses from conversational agents.
- Synthesises changing announcements and scripts using the same voice.
When choosing between cosa-a and cosa-b, use the desired voice, the actual synthesis results and the level of control required as your criteria.
| Item | cosa-a | cosa-b |
|---|---|---|
| Synthesis settings | Settings related to the number of synthesis steps, guidance strength and generation length, in addition to voice, language and speaking speed | Common settings such as speaking speed and language |
| Cloning input | A reference recording for which you have permission to use | A reference recording and corresponding text. Reference text can be generated through speech recognition during the registration process |
| Role | Voice generation with detailed synthesis controls | Voice selection and custom voice generation |
cosa-asr
cosa-asr converts input speech into text. It documents recorded material or converts voice requests into input that downstream business systems can process.
- Transcribe voice memos and recorded material.
- Convert voice-based questions and business requests into input.
- Generate reference text for voice cloning.
- Use it for processing that continues after transcription with summarisation, classification and search.
Passing the transcription results to CIP leads to content summarisation, request classification and business draft creation. Synthesising the processed results with cosa-a or cosa-b creates a business service with voice input and output.
Cova
cova-1 is an LLM-based model dedicated to vision OCR. It recognises text in documents and images, extracts layouts and table structures, and provides brief descriptions of visual elements.
It converts documents as viewed by people into a format that agents and business systems can use more effectively. The body text and tables are extracted as structured content, while images included in the document are provided with brief descriptions, allowing downstream agents to understand both the document's content and its visual context. By structuring a workflow that first reviews the material using the extracted text and descriptions, then examines only the original sources that require detailed verification, you can reduce the need to repeatedly input the original images and save context tokens.
| Feature | Description |
|---|---|
| Text recognition | Extracts text from documents that are scanned or photographed. |
| Layout interpretation | Separates content into blocks by taking the document layout into account. |
| Location information provision | Provides the locations of recognised blocks to link the extraction results with the original text. |
| Table structure extraction | Structures the contents of tables for search and data processing. |
| Simple image description | Briefly describes the contents of images and visual elements, providing agents with clues for interpretation. |
| Context optimisation | Reduces repeated input of the original image by conveying information centred on text, structure and image descriptions. |
| Selective detailed analysis | Helps agents determine which original images require additional review. |
| Document format conversion | Organises recognition results into HTML or Markdown. |
- Document digitisation: Converts paper documents and scanned materials into text-based resources.
- Business document intake: Extracts content from applications, supporting documents and reports.
- Tabular data processing: Uses tables included in documents for subsequent data processing.
- Pre-processing for knowledge search: Converts document images and visual materials into searchable text and structures.
- Automated document review: Delivers extraction results to CIP to verify and summarise items and compare documents.
- Agent input optimisation: Organises the body text, tables and image descriptions of long documents to deliver only the necessary information.
For example, while extracting the body text and tables from a report, it describes an embedded graph as a graph showing quarterly revenue trends. Using this description, the agent identifies the graph’s presence and purpose, and reviews the original additionally when precise figures or trends are required.
Product connections
By connecting products, you can configure the following workflows.
- Document-based tasks: Cova extracts the body text, tables and image descriptions, CIP selectively reviews the required originals to perform analysis and business processing, and Cosa delivers the results by voice.
- Voice-based tasks: Cosa ASR converts voice requests into text, and cosa-a or cosa-b reads out the answers processed by CIP.
- Field automation: cip-5.5-sm classifies field data and performs local tasks, and passes materials requiring additional comprehensive analysis to cip-5.5-im or cip-5.5-mm.
Calls allowed by the gateway
A gateway plan determines the calls permitted for each model. The table below shows the scope enabled by the LLM_FREE, TTS_FREE and STT_FREE plans, which are applied automatically without an application. It does not include models enabled through other routes.
| Model ID | Permitted calls | Basis plan |
|---|---|---|
| cip-5.5-im | anthropic.messages · chat.completions · chat.completions.batch · completions · responses | LLM_FREE |
| cip-5.5-im-chat | anthropic.messages · chat.completions · chat.completions.batch · completions · responses | LLM_FREE |
| cip-5.5-mm | anthropic.messages · chat.completions · chat.completions.batch · completions · responses | LLM_FREE |
| cip-5.5-mm-h | anthropic.messages · chat.completions · chat.completions.batch · completions · responses | LLM_FREE |
| cosa-a | audio.speech · audio.speech.clone · voices · voices.design | TTS_FREE |
| cosa-b | audio.speech · audio.speech.clone · voices · voices.design | TTS_FREE |
| cosa-tts | audio.speech · audio.speech.clone · voices · voices.design | TTS_FREE |
| cosa-asr | audio.transcriptions | STT_FREE |
| ivy-4-embedding | embeddings | LLM_FREE |
| ivy-4-embedding-mm | embeddings | LLM_FREE |
| ivy-4-mm | anthropic.messages · chat.completions · chat.completions.batch · completions · responses | LLM_FREE |
cip-5.5-sm and cova-1 are not included in this plan table. cip-5.5-im-chat, cip-5.5-mm-h, cosa-tts, ivy-4-embedding, ivy-4-embedding-mm and ivy-4-mm in the table are model IDs that are not included in the product descriptions above. The Making API Calls documentation explains which path each call key corresponds to.
