Developer guide

Model

Model Summary

CLEVI models are divided into three product families: CIP, which handles general-purpose intelligence and business automation; Cosa, which handles voice generation and recognition; and Cova, which handles information extraction from documents and images. Each product can be used independently or connected into services that progress from document recognition and analysis to task execution and voice guidance.

ProductModel IDKey SpecificationsPrimary Role
CIP-5.5-SMcip-5.5-sm24B~40B adjustableEdge servers and local business automation
CIP-5.5-IMcip-5.5-im360B · 256K input · 64K output · multimodalGeneral-purpose business processing. Prioritizes response speed and throughput
CIP-5.5-MMcip-5.5-mm800B · 512K input · 64K output · multimodalLong-context analysis and complex problem-solving
Cosa-Acosa-aVoice synthesis · cloning · design · streamingVoice generation with detailed synthesis control
Cosa-Bcosa-bVoice synthesis · cloning · design · streamingVoice selection and customized voice generation
Cosa ASRcosa-asrSpeech recognitionConverting voice input to text
Covacova-1LLM-based vision OCR onlyText, layout, and table extraction, plus simple image descriptions

CIP

CIP handles general-purpose intelligence and business automation. It consists of cip-5.5-im and cip-5.5-mm, which are called through the gateway, and cip-5.5-sm, which is deployed on edge servers and on-premises.

cip-5.5-im and cip-5.5-mm

Categorycip-5.5-imcip-5.5-mm
Scale360B800B
Input limit256K512K
Maximum output64K64K
Input formatsText and visual materials (documents, screen images, tables and charts)Text and visual materials (integrated review of code, documents, and screen materials)
Design focusHandles everyday knowledge work and development tasks, with a focus on response speed, cost efficiency, and throughput.Handles complex tasks that integrate and analyse multiple sources and conditions, connecting analysis results to actual work.
Use of long inputsReviews multiple documents, code files, and conversation histories together.Handles extensive reference materials and work histories together, and performs agent tasks involving multiple stages of research, execution, and verification.

cip-5.5-im is suited to tasks with clear objectives and scope. Typical examples include extracting the required information from provided materials, producing outputs in a specified format, and modifying code to meet requirements. The maximum 64K output is useful for writing detailed reports or results consisting of multiple sections.

  • Development assistance: Write functions and features, fix well-defined bugs, and implement changes according to existing patterns.
  • Test generation: Draft tests and validation cases based on requirements and implementation.
  • Document processing: Handle summarization, classification, item extraction, format conversion, and drafting.
  • Multimodal analysis: Review tables and charts in screen images and documents together with relevant descriptions.
  • Interactive work support: Answer questions based on work materials and produce the required outputs.
  • High-volume processing: Use for individual summaries of multiple documents, request classification, and data transformation.
  • Subagents: Handle clearly bounded subtasks such as code exploration, reviewing specific files, and organizing materials.

For example, it can be used to receive feature requirements and related code and produce proposed changes, tests, and a change description, or to classify intake documents and create summaries for the people responsible.

cip-5.5-mm is suited to tasks that require identifying relationships between information and maintaining consistency across the overall task. In development, it reviews the connections between requirements, design, implementation, and testing; in document analysis, it synthesizes the claims and evidence, differences, and contradictions across multiple sources.

  • Complex feature implementation: Change code, tests, and documentation together while considering the contracts and behaviour of multiple modules.
  • Large-scale code analysis: Review call relationships, data flows, and the scope of change impacts.
  • Incident root-cause analysis: Compare logs, configurations, code, and execution results to organize hypotheses about the cause and verification procedures.
  • Design and decision support: Structure requirements and constraints and analyse the impact of implementation alternatives.
  • Cross-document synthesis: Organize commonalities, differences, and contradictions in reports and technical documents based on evidence.
  • Multimodal integrated review: Interpret text and visual materials together to analyse issues and requirements.
  • Multi-step agents: Use for tasks that proceed through research, planning, tool execution, and result review.
  • Specialized document writing: Write technical reports, design proposals, and detailed review documents that reflect complex conditions and supporting evidence.

For example, it is used to review related codes and logs for errors occurring across multiple services, or for tasks ranging from analysing requirements for complex functions to implementation and reviewing validation results.

cip-5.5-sm

cip-5.5-sm performs local inference and business process automation on edge servers and in on-premises environments. Since the model size can be adjusted from 24B to 40B, you can choose a configuration suited to the computing resources of the deployment equipment and the required performance at the site. Because processing takes place close to where data is generated, it is used to integrate with on-site systems and utilize internal data. By configuring the required models and tools locally, it can also operate in environments with limited external connectivity.

It interprets information generated on-site and connects it to processing procedures. It classifies logs and requests, extracts the required information from data, and performs repetitive tasks using connected internal tools and scripts.

  • On-site event processing: Classifies equipment logs and status information and selects events requiring review.
  • Data preprocessing: Extracts key items from business records and classifies, tags, and normalizes them.
  • Request routing: Forwards received requests to the responsible system or processing procedure.
  • Local business agent: Performs repetitive tasks by querying internal systems and executing scripts.
  • On-premises business support: Answers questions based on internal materials and organizes business content.
  • Follow-up analysis support: Selects materials requiring further review and summarizes the key points.

For example, it is used in a workflow that classifies operational logs on a site server, looks up related history, and prepares an incident summary for the person in charge.

Cosa

Cosa is a suite of voice products that converts text to speech, recognizes input speech as text, and registers and reuses desired voices. It is used for service guidance, content creation, accessibility support, and voice interfaces for conversational agents.

cosa-a and cosa-b handle voice generation, while cosa-asr handles speech recognition. cosa-a and cosa-b are separate models that provide their own voice lists and synthesis settings, and users select between them through the same service interface. The features provided in common by both models are as follows.

  • Text-to-speech: Synthesizes the text you enter using the selected voice.
  • Voice cloning: Registers a voice based on a reference recording you have permission to use and synthesizes new text.
  • Voice design: Creates and registers a voice by describing the characteristics of the voice you want.
  • Voice reuse: Uses a registered voice for subsequent synthesis.
  • Synthesis settings: Adjusts the speaking rate, language, and more.
  • Streaming delivery: Synthesizes text phrase by phrase as it arrives and delivers the audio sequentially.
  • Playback control: Supports pausing, resuming, stopping, and adding additional text.

With incremental text input, you can start voice playback before the entire response is complete. While CIP writes the response, you can configure a service in which Cosa reads out the phrases as they become ready.

cosa-a

cosa-a supports text-to-speech, voice cloning, voice design, and streaming output. In addition to basic features for selecting a voice, language, and speaking rate, it provides additional controls such as the number of synthesis steps, guidance strength, and settings related to generation length. Use it to adjust settings and review results during voice production, or apply multiple settings to the same script and select the preferred result.

  • Synthesizes announcements and notifications for apps, kiosks, and business systems.
  • Creates audio versions of educational materials and user manuals.
  • Produces narration for video and audio content.
  • Produces recurring content using a registered voice.
  • Used for streaming voice output from conversational agents.
  • Used for voice production that involves adjusting detailed synthesis settings.

cosa-b

cosa-b supports preset voice selection, description-based voice design, reference-voice-based cloning, and streaming output. You can choose a voice from its own voice library or register a new voice for use in your service.

Voice cloning uses a reference recording and the corresponding text. It also supports a workflow that generates the reference text through speech recognition during the registration process, and registered voices can be reused for subsequent synthesis.

  • Configures the voice personas of services and characters.
  • Creates customized voices from a description or reference recording.
  • Synthesizes brand announcements, character dialogue, and content narration.
  • Used for voice responses from conversational agents.
  • Synthesizes changing announcements and scripts with the same voice.

When choosing between cosa-a and cosa-b, use the desired voice, actual synthesis results, and required control options as your criteria.

Itemcosa-acosa-b
Synthesis settingsIn addition to voice, language, and speaking speed, settings related to the number of synthesis steps, guidance strength, and generation lengthCommon settings such as speaking speed and language
Cloning inputA reference recording for which you have usage rightsA reference recording and corresponding text. Reference text can be generated through speech recognition during registration
RoleVoice generation using detailed synthesis controlsVoice selection and customized voice generation

cosa-asr

cosa-asr converts input speech into text. It documents recordings and converts requests received by voice into input that downstream business systems can process.

  • Transcribe voice memos and recordings.
  • Convert voice-based questions and business requests into input.
  • Generate reference text for voice cloning.
  • Use it for processing that continues with summarization, classification, and search after transcription.

Delivering transcription results to CIP leads to content summarization, request classification, and business draft creation. Synthesizing the results with cosa-a or cosa-b creates a business service with voice input and output.

Cova

cova-1 is an LLM-based vision OCR model dedicated to OCR. It recognizes text in documents and images, extracts layouts and table structures, and provides brief descriptions of visual elements.

It converts documents as viewed by people into a format that agents and business systems can use more effectively. The body text and tables are extracted as structured content, while images included in the document are provided with brief descriptions, allowing downstream agents to understand both the document’s content and its visual context. By first reviewing the material using the extracted text and descriptions, then examining only the original sources that require detailed verification, you can reduce the need to repeatedly input original images and save context tokens.

FeatureDescription
Text recognitionExtracts text from documents that are scanned or photographed.
Layout interpretationSeparates content into blocks by considering the document’s layout.
Location information provisionProvides the locations of recognized blocks to link the extracted results to the original text.
Table structure extractionStructures table content for use in search and data processing.
Simple image descriptionsBriefly describes the content of images and visual elements, giving the agent clues for interpretation.
Context optimizationReduces repeated input of the original image by conveying information centred on text, structure, and image descriptions.
Selective detailed analysisHelps the agent determine which original images require additional review.
Document format conversionOrganizes recognition results into HTML or Markdown.
  • Document digitization: Converts paper documents and scanned materials into text-based resources.
  • Business document intake: Extracts content from applications, supporting documents, and reports.
  • Tabular data processing: Uses tables contained in documents for subsequent data processing.
  • Preprocessing for knowledge search: Converts document images and visual materials into searchable text and structures.
  • Document review automation: Delivers extracted results to CIP to verify items, create summaries, and compare materials.
  • Agent input optimization: Organizes the body text, tables, and image descriptions of lengthy documents to convey only the necessary information.

For example, while extracting the body text and tables from a report, it describes an inserted graph as a graph showing quarterly sales trends. The agent uses this description to identify the graph’s presence and purpose, and reviews the original additionally when exact figures or trends are required.

Product connections

Connecting products enables you to configure the following flows.

  • Document-based operations: Cova extracts the body text, tables, and image descriptions, CIP selectively reviews the necessary originals to perform analysis and business processing, and Cosa delivers the results by voice.
  • Voice-based operations: Cosa ASR converts voice requests into text, and cosa-a or cosa-b reads out the responses processed by CIP.
  • On-site automation: cip-5.5-sm classifies on-site data and performs local operations, while forwarding materials requiring additional comprehensive analysis to cip-5.5-im or cip-5.5-mm.

Calls permitted by the gateway

Gateway plans determine the calls permitted for each model. The table below shows the scope enabled by the LLM_FREE, TTS_FREE, and STT_FREE plans, which are applied automatically without an application. It does not include models enabled through other routes.

Model IDPermitted callsQualifying plan
cip-5.5-imanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cip-5.5-im-chatanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cip-5.5-mmanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cip-5.5-mm-hanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE
cosa-aaudio.speech · audio.speech.clone · voices · voices.designTTS_FREE
cosa-baudio.speech · audio.speech.clone · voices · voices.designTTS_FREE
cosa-ttsaudio.speech · audio.speech.clone · voices · voices.designTTS_FREE
cosa-asraudio.transcriptionsSTT_FREE
ivy-4-embeddingembeddingsLLM_FREE
ivy-4-embedding-mmembeddingsLLM_FREE
ivy-4-mmanthropic.messages · chat.completions · chat.completions.batch · completions · responsesLLM_FREE

cip-5.5-sm and cova-1 are not included in this plan table. cip-5.5-im-chat, cip-5.5-mm-h, cosa-tts, ivy-4-embedding, ivy-4-embedding-mm, and ivy-4-mm in the table are model IDs that are not included in the product descriptions above. The API Calling guide indicates which path each calling key corresponds to.

CLEVI

Language and region

Machine-translated languages are marked. Availability follows the published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5