COSA MODELS

Create voices and understand your users.
Meet Cosa.

Cosa is a speech AI family for converting text and speech, and registering and reusing voices.
Use it for service announcements, content creation, accessibility, and conversational agents.

Speech AI for conversational services

Combine speech synthesis, cloning, design, streaming, and recognition to build conversational services and voice interfaces.

Cosa-A

Fine control over speech synthesis

TTS Voice cloning Voice design Streaming

For voice services that need precise control
over speed, language, and synthesis methods.

Cosa-B

Voice selection and custom voice creation

Presets Description-based generation Reference voice reproduction

Create consistent brand or character voices
and reuse them across your content.

Cosa ASR

Transcription and speech recognition

Input
Live · Files
Output
Text

Turn calls, voice commands, and customer conversations into text
and connect it to CIP for processing.

Voice experiences built with Cosa

Choose speech synthesis, cloning, design, and recognition to fit your service.

Find the latest Cosa specifications in the developer docs

This page explains model capabilities and how to choose between them.
Find exact model IDs, input and output limits, supported features, and version changes in the developer docs.

Model IDs Input and output specs API examples Changelog
Cosa developer docs

Design the voice experience your service needs

CLEVI

Тел һәм төбәк

Machine-translated languages are translated by clevi/cip-5.5-im and marked as such. This applies to languages with a published site bundle.

136 тел

Тәкъдим ителә

1

Көнчыгыш Азия

7

Көньяк-Көнчыгыш Азия

11

Көньяк Азия

18

Үзәк Азия

5

Якын Көнчыгыш һәм Кавказ

10

Көнбатыш Европа һәм Көньяк Европа

16

Берләшкән Корольлек һәм Ирландия

4

Төньяк Европа

10

Үзәк Европа һәм Балкан

14

Көнчыгыш Европа

5

Көнчыгыш Африка

8

Көнбатыш Африка һәм Урта Африка

9

Көньяктагы Африка

8

Америка

5

Океания

5