COSA MODELS

Create voices and understand your users.
Meet Cosa.

Cosa is a speech AI family for converting text and speech, and registering and reusing voices.
Use it for service announcements, content creation, accessibility, and conversational agents.

Speech AI for conversational services

Combine speech synthesis, cloning, design, streaming, and recognition to build conversational services and voice interfaces.

Cosa-A

Fine control over speech synthesis

TTS Voice cloning Voice design Streaming

For voice services that need precise control
over speed, language, and synthesis methods.

Cosa-B

Voice selection and custom voice creation

Presets Description-based generation Reference voice reproduction

Create consistent brand or character voices
and reuse them across your content.

Cosa ASR

Transcription and speech recognition

Input
Live · Files
Output
Text

Turn calls, voice commands, and customer conversations into text
and connect it to CIP for processing.

Voice experiences built with Cosa

Choose speech synthesis, cloning, design, and recognition to fit your service.

Find the latest Cosa specifications in the developer docs

This page explains model capabilities and how to choose between them.
Find exact model IDs, input and output limits, supported features, and version changes in the developer docs.

Model IDs Input and output specs API examples Changelog
Cosa developer docs

Design the voice experience your service needs

CLEVI

Language and region

Machine-translated languages are translated by clevi/cip-5.5-im and marked as such. This applies to languages with a published site bundle.

136 languages

Recommended

1

East Asia

7

Southeast Asia

11

South Asia

18

Central Asia

5

Middle East and the Caucasus

10

Western and Southern Europe

16

Britain and Ireland

4

Northern Europe and the Baltics

10

Central Europe and the Balkans

14

Eastern Europe

5

East Africa and the Horn

8

West and Central Africa

9

Southern Africa

8

The Americas

5

The Pacific

5
Cosa — CLEVI