Bark
PermissiveSuno · Bark · Released 2023-04
~1B (combined sub-models)Audio
A fully generative text-to-audio model that can produce speech, music, and sound effects, but with less voice-cloning control than dedicated TTS models.
Strengths
- +MIT license, unrestricted use
- +Can generate non-speech sounds, music, and expressive tone alongside speech
- +No text-alignment/phoneme pipeline required
Limitations
- -Less reliable voice consistency/cloning than XTTS-style models
- -Slower and less predictable than dedicated TTS architectures
License
MIT
use commercially with attribution niceties
Hardware
runs on a good consumer GPU (or Apple Silicon) with quantization
Links
Stats
suno/bark— downloads·— likes
via Hugging Face