COSA MODELS

Create voices and understand your users.
Meet Cosa.

Cosa is a speech AI family for converting text and speech, and registering and reusing voices.
Use it for service announcements, content creation, accessibility, and conversational agents.

Speech AI for conversational services

Combine speech synthesis, cloning, design, streaming, and recognition to build conversational services and voice interfaces.

Cosa-A

Fine control over speech synthesis

TTS Voice cloning Voice design Streaming

For voice services that need precise control
over speed, language, and synthesis methods.

Cosa-B

Voice selection and custom voice creation

Presets Description-based generation Reference voice reproduction

Create consistent brand or character voices
and reuse them across your content.

Cosa ASR

Transcription and speech recognition

Input
Live · Files
Output
Text

Turn calls, voice commands, and customer conversations into text
and connect it to CIP for processing.

Voice experiences built with Cosa

Choose speech synthesis, cloning, design, and recognition to fit your service.

Find the latest Cosa specifications in the developer docs

This page explains model capabilities and how to choose between them.
Find exact model IDs, input and output limits, supported features, and version changes in the developer docs.

Model IDs Input and output specs API examples Changelog
Cosa developer docs

Design the voice experience your service needs

CLEVI

زبان اور علاقہ

Machine-translated languages are translated by clevi/cip-5.5-im and marked as such. This applies to languages with a published site bundle.

136 زبانیں

تجویز کردہ

1

مشرقی ایشیا

7

جنوب مشرقی ایشیا

11

جنوبی ایشیا

18

وسطی ایشیا

5

مشرق وسطیٰ اور قفقاز

10

مغربی یورپ اور جنوبی یورپ

16

سلطنت متحدہ اور آئرلینڈ

4

شمالی یورپ

10

وسطی یورپ اور بلقان

14

مشرقی یورپ

5

مشرقی افریقہ

8

مغربی افریقہ اور وسطی افریقہ

9

جنوبی افریقہ کا علاقہ

8

امیریکاز

5

اوشیانیا

5