COSA MODELS

Create voices and understand your users.
Meet Cosa.

Cosa is a speech AI family for converting text and speech, and registering and reusing voices.
Use it for service announcements, content creation, accessibility, and conversational agents.

Speech AI for conversational services

Combine speech synthesis, cloning, design, streaming, and recognition to build conversational services and voice interfaces.

Cosa-A

Fine control over speech synthesis

TTS Voice cloning Voice design Streaming

For voice services that need precise control
over speed, language, and synthesis methods.

Cosa-B

Voice selection and custom voice creation

Presets Description-based generation Reference voice reproduction

Create consistent brand or character voices
and reuse them across your content.

Cosa ASR

Transcription and speech recognition

Input
Live · Files
Output
Text

Turn calls, voice commands, and customer conversations into text
and connect it to CIP for processing.

Voice experiences built with Cosa

Choose speech synthesis, cloning, design, and recognition to fit your service.

Find the latest Cosa specifications in the developer docs

This page explains model capabilities and how to choose between them.
Find exact model IDs, input and output limits, supported features, and version changes in the developer docs.

Model IDs Input and output specs API examples Changelog
Cosa developer docs

Design the voice experience your service needs

CLEVI

Llengua i regió

Les llengües traduïdes automàticament es tradueixen mitjançant clevi/cip-5.5-im i s’hi indiquen com a tals. Això s’aplica a les llengües amb un paquet del lloc publicat.

136 llengües

Recomanat

1

Àsia oriental

7

Àsia sud-oriental

11

Àsia meridional

18

Àsia central

5

Orient Mitjà i el Caucas

10

Europa occidental i Europa meridional

16

Regne Unit i Irlanda

4

Europa septentrional

10

Europa Central i els Balcans

14

Europa oriental

5

Àfrica oriental

8

Àfrica occidental i Àfrica central

9

Àfrica meridional

8

Amèrica

5

Oceania

5