Developer guide

Models

Model Summary

CLEVI models are divided into three product families: CIP, which handles general intelligence and work automation; Cosa, which handles voice generation and recognition; and Cova, which handles information extraction from documents and images. Each product can be used independently or connected into services that progress from document recognition and analysis to task execution and voice guidance.

ProductModel IDKey specificationsPrimary role
CIP-5.5-SMcip-5.5-sm24B–40B configurableEdge servers and local work automation
CIP-5.5-IMcip-5.5-im360B · 256K input · 64K output · multimodalGeneral-purpose work processing. Prioritises response speed and throughput
CIP-5.5-MMcip-5.5-mm800B · 512K input · 64K output · multimodalLong-context analysis and complex problem-solving
Cosa-Acosa-aVoice synthesis · cloning · design · streamingVoice generation with detailed synthesis control
Cosa-Bcosa-bVoice synthesis · cloning · design · streamingVoice selection and customised voice generation
Cosa ASRcosa-asrSpeech recognitionConverting voice input into text
Covacova-1LLM-based vision OCR specialisedText, layout and table extraction, with simple image descriptions

CIP

CIP handles general intelligence and work automation. It consists of cip-5.5-im and cip-5.5-mm, which are called through a gateway, and cip-5.5-sm, which is deployed on edge servers and on-premises.

cip-5.5-im and cip-5.5-mm

Itemcip-5.5-imcip-5.5-mm
Scale360B800B
Input limit256K512K
Maximum output64K64K
Input formatsText and visual materials (documents, screen images, tables and charts)Text and visual materials (integrated review of code, documents and screen materials)
Design focusHandles everyday knowledge work and development tasks, with a focus on response speed, cost efficiency and throughput.Handles complex tasks that synthesise multiple sources and conditions and connect analysis results to practical work.
Use of long inputsReviews multiple documents, code files and conversation histories together.Handles extensive reference materials and work histories together, and performs agent tasks involving multiple stages of research, execution and verification.

cip-5.5-im is suitable for tasks with clear objectives and scope. Typical tasks include extracting the required information from the provided materials, preparing deliverables in the specified format, and modifying code according to the requirements. The maximum 64K output is useful for creating detailed reports or deliverables consisting of multiple sections.

  • Development assistance: Write functions and features, fix well-defined bugs, and implement changes in line with existing patterns.
  • Test generation: Create test drafts and validation cases based on the requirements and implementation.
  • Document processing: Handle summarisation, classification, item extraction, format conversion, and drafting.
  • Multimodal analysis: Review tables and charts in screen images and documents along with the relevant descriptions.
  • Interactive work support: Answer questions based on work materials and prepare the required deliverables.
  • Bulk processing: Use it for individual summaries of multiple documents, request classification, and data conversion.
  • Sub-agents: Assign clearly bounded sub-tasks such as code exploration, reviewing specific files, and organising materials.

For example, it can be used to receive feature requirements and related code, then prepare a proposed change, tests, and a description of the changes, or to classify submitted documents and create summaries for the relevant personnel.

cip-5.5-mm is suitable for tasks that require understanding relationships between information and maintaining consistency across the overall task. In development, it reviews the connections between requirements, design, implementation, and testing. In document analysis, it synthesises the claims and evidence, differences, and contradictions across multiple materials.

  • Complex feature implementation: Change code, tests, and documentation together while considering the contracts and behaviour of multiple modules.
  • Large-scale code analysis: Review call relationships, data flow, and the scope of change impact.
  • Incident root-cause analysis: Compare logs, configurations, code, and execution results to organise cause hypotheses and verification procedures.
  • Design and decision-making support: Structure requirements and constraints, and analyse the impact of implementation alternatives.
  • Multi-document synthesis: Summarise commonalities, differences, and contradictions in reports and technical documents based on evidence.
  • Multimodal integrated review: Interpret text and visual materials together to analyse issues and requirements.
  • Multi-step agents: Use it for tasks involving research, planning, tool execution, and result review.
  • Specialised document writing: Prepare technical reports, design proposals, and detailed review documents that reflect complex conditions and supporting evidence.

For example, it is used to review related codes and logs from errors occurring across multiple services, or for tasks that span from requirements analysis for complex features to implementation and review of validation results.

cip-5.5-sm

cip-5.5-sm performs local inference and business automation on edge servers and in on-premises environments. As the model size can be adjusted within the 24B–40B range, the configuration can be selected to suit the computing resources of the deployment equipment and the required performance at the site. Since processing takes place close to where data is generated, it is used for integration with on-site systems and internal data utilisation. By configuring the required models and tools locally, it can also operate in environments with limited external connectivity.

It interprets information generated at the site and connects it to processing procedures. It classifies logs and requests, extracts required information from data, and performs repetitive tasks using connected internal tools and scripts.

  • On-site event processing: Classifies equipment logs and status information, and identifies events that require confirmation.
  • Data preprocessing: Extracts key items from business records and performs classification, tagging and normalisation.
  • Request routing: Forwards received requests to the responsible system or processing procedure.
  • Local business agent: Performs repetitive tasks by querying internal systems and running scripts.
  • On-premises business support: Answers questions based on internal materials and organises business content.
  • Follow-up analysis support: Selects materials requiring additional review and summarises the key points.

For example, it is used in workflows that classify operational logs on a site server, look up related history, and prepare an incident summary for the person in charge.

Cosa

Cosa is a suite of voice products that converts text to speech, recognises input speech as text, and registers and reuses voices as required. It is used for service guidance, content creation, accessibility support, and voice interfaces for conversational agents.

cosa-a and cosa-b handle speech generation, while cosa-asr handles speech recognition. cosa-a and cosa-b are separate models that provide their own voice lists and synthesis settings, and they can be selected through the same service interface. The features commonly provided by both models are as follows.

  • Text-to-speech: Synthesises the entered text using the selected voice.
  • Voice cloning: Registers a voice based on a reference recording for which you have usage rights, then synthesises new sentences.
  • Voice design: Creates and registers a voice by describing the characteristics of the desired voice.
  • Voice reuse: Uses a registered voice for subsequent synthesis.
  • Synthesis settings: Adjusts the speaking speed, language and other settings.
  • Streaming delivery: Synthesises text phrase by phrase as it arrives and delivers the audio sequentially.
  • Playback control: Supports pausing, resuming, stopping and entering additional sentences.

With incremental text input, voice playback can begin before the entire response is complete. While CIP is composing the response, you can configure a service in which Cosa reads out the prepared phrases first.

cosa-a

cosa-a supports text-to-speech, voice cloning, voice design and streaming output. In addition to basic functions for selecting the voice, language and speaking speed, it provides additional controls such as settings related to the number of synthesis steps, guidance strength and generation length. It can be used to adjust settings and review results during voice production, or to apply multiple settings to the same script and select the preferred result.

  • Synthesises announcements and notifications for apps, kiosks and business systems.
  • Creates audio versions of training materials and user manuals.
  • Produces narration for video and audio content.
  • Creates recurring content using a registered voice.
  • Used for streaming voice output from conversational agents.
  • Used for voice production that involves adjusting detailed synthesis settings.

cosa-b

cosa-b supports preset voice selection, description-based voice design, reference-voice-based cloning and streaming output. You can select a voice from its built-in voice list or register a new voice for use in a service.

Voice cloning uses a reference recording and the corresponding text. It also supports a registration flow that generates the reference text through speech recognition, and registered voices can be reused for subsequent synthesis.

  • Configures the voice personas of services and characters.
  • Creates customised voices from a description or reference recording.
  • Synthesises brand announcements, character dialogue and content narration.
  • Used for voice responses from conversational agents.
  • Synthesises changing announcements and scripts using the same voice.

When choosing between cosa-a and cosa-b, use the desired voice, the actual synthesis results, and the controls you need as your criteria.

Itemcosa-acosa-b
Synthesis settingsIn addition to voice, language, and speaking speed, settings related to the number of synthesis steps, guidance strength, and generation lengthCommon settings such as speaking speed and language
Cloning inputA reference recording for which you have usage rightsA reference recording and its corresponding text. Reference text can be generated through speech recognition during the registration process
RoleVoice generation with detailed synthesis controlsVoice selection and customised voice generation

cosa-asr

cosa-asr converts input speech into text. It documents recordings and converts requests received by voice into inputs that downstream business systems can process.

  • Transcribes voice memos and recordings.
  • Converts voice-based questions and business requests into inputs.
  • Generates reference text for voice cloning.
  • Used for processing that continues with summarisation, classification, and search after transcription.

Sending the transcription results to CIP leads to content summarisation, request classification, and business draft creation. Synthesising the processed results with cosa-a or cosa-b creates a business service with voice input and output.

Cova

cova-1 is an LLM-based model dedicated to vision OCR. It recognises text in documents and images, extracts layouts and table structures, and provides brief descriptions of visual elements.

It converts documents intended for people to read into a format that agents and business systems can use more effectively. The body text and tables are extracted as structured content, while images included in the document are conveyed as brief descriptions, enabling downstream agents to understand both the document’s content and its visual context. By first reviewing the materials using the extracted text and descriptions, then examining only the original sources that require detailed verification, you can reduce the need to repeatedly submit original images and save context tokens.

FeatureDescription
Text recognitionExtracts text from documents that are scanned or photographed.
Layout interpretationSeparates content into blocks by considering the document layout.
Location information provisionProvides the locations of recognised blocks to link the extraction results with the original text.
Table structure extractionStructures table content for use in search and data processing.
Simple image descriptionBriefly describes the content of images and visual elements, providing agents with clues for interpretation.
Context optimisationReduces repeated input of the original image by conveying information centred on text, structure and image descriptions.
Selective detailed analysisHelps agents determine which original images require additional review.
Document format conversionOrganises recognition results into HTML or Markdown.
  • Document digitisation: Converts paper documents and scanned materials into text-based resources.
  • Business document intake: Extracts content from applications, supporting documents and reports.
  • Table data processing: Uses tables included in documents for subsequent data processing.
  • Pre-processing for knowledge search: Converts document images and visual materials into searchable text and structures.
  • Document review automation: Delivers extraction results to CIP to perform item verification, summarisation and comparisons across materials.
  • Agent input optimisation: Organises the body text, tables and image descriptions of lengthy documents to deliver only the necessary information.

For example, while extracting the body text and tables from a report, it describes an inserted graph as a graph showing quarterly sales trends. The agent uses this description to understand the graph's presence and purpose, and reviews the relevant original additionally when exact figures or trends are required.

Product connections

Connecting products enables you to configure the following flows.

  • Document-based operations: Cova extracts body text, tables and image descriptions, CIP selectively reviews the necessary originals to perform analysis and business operations, and Cosa delivers the results by voice.
  • Voice-based operations: Cosa ASR converts voice requests into text, and cosa-a or cosa-b reads out the answers processed by CIP.
  • On-site automation: cip-5.5-sm classifies on-site data and performs local operations, while passing materials requiring additional comprehensive analysis to cip-5.5-im or cip-5.5-mm.

Calls allowed at the gateway

Gateway plans determine the calls permitted for each model. The table below shows the scopes enabled by the LLM_FREE, TTS_FREE and STT_FREE plans, which are applied automatically without an application. Models enabled through other routes are not included.

Model IDPermitted callsBasis plan
cip-5.5-imanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cip-5.5-im-chatanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cip-5.5-mmanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cip-5.5-mm-hanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cosa-aaudio.speech · audio.speech.clone · voices · voices.designTTS_FREE
cosa-baudio.speech · audio.speech.clone · voices · voices.designTTS_FREE
cosa-ttsaudio.speech · audio.speech.clone · voices · voices.designTTS_FREE
cosa-asraudio.transcriptionsSTT_FREE
ivy-4-embeddingembeddingsLLM_FREE
ivy-4-embedding-mmembeddingsLLM_FREE
ivy-4-mmanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE

cip-5.5-sm and cova-1 are not included in this plan table. cip-5.5-im-chat, cip-5.5-mm-h, cosa-tts, ivy-4-embedding, ivy-4-embedding-mm, and ivy-4-mm in the table are model IDs that are not included in the product descriptions above. The API Calling documentation explains which endpoint each call key corresponds to.

CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5