COSA MODELS

Create voices and understand your users.
Meet Cosa.

Cosa is a speech AI family for converting text and speech, and registering and reusing voices.
Use it for service announcements, content creation, accessibility, and conversational agents.

Speech AI for conversational services

Combine speech synthesis, cloning, design, streaming, and recognition to build conversational services and voice interfaces.

Cosa-A

Fine control over speech synthesis

TTS Voice cloning Voice design Streaming

For voice services that need precise control
over speed, language, and synthesis methods.

Cosa-B

Voice selection and custom voice creation

Presets Description-based generation Reference voice reproduction

Create consistent brand or character voices
and reuse them across your content.

Cosa ASR

Transcription and speech recognition

Input
Live · Files
Output
Text

Turn calls, voice commands, and customer conversations into text
and connect it to CIP for processing.

Voice experiences built with Cosa

Choose speech synthesis, cloning, design, and recognition to fit your service.

Find the latest Cosa specifications in the developer docs

This page explains model capabilities and how to choose between them.
Find exact model IDs, input and output limits, supported features, and version changes in the developer docs.

Model IDs Input and output specs API examples Changelog
Cosa developer docs

Design the voice experience your service needs

CLEVI

שפּראַך און געגנט

שפּראַכן וואָס ווערן איבערגעזעצט דורך מאַשין ווערן איבערגעזעצט דורך clevi/cip-5.5-im און אַזוי אָנגעצייכנט. דאָס גילט פֿאַר שפּראַכן מיט אַ פֿאַרעפֿנטלעכטן וועבזײַט־פּעקל.

136 שפּראַכן

רעקאָמענדירט

1

מזרח־אַזיע

7

דרום־מזרח אַזיע

11

דרום־אַזיע

18

צענטראַל־אַזיע

5

מיטל־מזרח און קאַווקאַז

10

מערב־ און דרום־אייראָפּע

16

בריטאַניע און אירלאַנד

4

צפֿון־אייראָפּע און די באַלטן

10

מיטל־אייראָפּע און באַלקאַן

14

מזרח־אייראָפּע

5

מזרח־אַפֿריקע און דער האָרן

8

מערב־ און צענטראַל־אַפֿריקע

9

דרום־אַפֿריקע

8

אַמעריקעס

5

דער פּאַציפֿיק

5