# Text to Audio (T2A)
Conjure hyper-realistic speech, music and sound effects using natural language 

## Conditional Generation (Text)

### 📄 Text Prompt
Simply write what you'd like the model to say.

### 🌍 Languages
Suno supports `language_x` to `language_x` generation out-of-the-box. 

| Language | Status |
| -------- | ------ |
| English | ✅ |
| Chinese (Mandarin) | ✅ |
| French | ✅ |
| German | ✅ |
| Hindi | ✅ |
| Italian | ✅ |
| Japanese | ✅ |
| Korean | ✅ |
| Polish | ✅ |
| Portuguese | ✅ |
| Russian | ✅ |
| Spanish | ✅ |
| Turkish | ✅ |

We're actively working on cross-lingual or `language_x` to `language_y` generation. 


### 📽 Metatags
Enrich generations with metatags. Examples include:

| Metatag | Status |
| -------- | ------ |
| [laughs] | ✅ |
| [sighs] | ✅ |
| [clears_throat] | ✅ |
| [cries] | Coming soon! |

### 🤔 Hesitations & Pauses
Use em dashes (—) to insert pauses into audio

### 🌶 Temperature
Increasing temperature increases variability of generations. Generations become more emotive, dynamic and, at times, erratic the higher you set the temperature. You're also more likely to experience emergent interjections and speaker changes. We recommend a starting default of `0.6` and tuning from there depending on your preferences.


## Conditional Generation (Text + Audio)
You can also condition on audio files of your choice as well. With just four seconds of audio, the model can emulate the style on your input file - preserving voices, emotions and acoustic environments from the input audio to produce more realistic generations. 


##Ethics Statements
Suno enables hyper-relasitic generation...
- Spoofing
- Deepfakes
