Research Note
Spirit LM Paper Code and License Record
The [Spirit LM paper](https://arxiv.org/abs/2402.05755) describes a 7 billion parameter pretrained text language model extended to speech through continued training on te
In this article
Spirit LM Paper Code and License Record
Model record
The Spirit LM paper describes a 7 billion parameter pretrained text language model extended to speech through continued training on text and speech units.
Spirit LM Base represents speech with HuBERT phonetic units. Spirit LM Expressive adds pitch and style units. Text uses subword BPE tokens. The training process interleaves aligned speech and text at word boundaries so the same autoregressive model can continue across modalities.
The paper reports training with 300 billion text tokens, 460,000 hours of speech-only data, and 110,000 hours of aligned speech and text. It also acknowledges errors in automatically produced alignments.
Evaluation boundary
The research covers speech-to-speech, speech-to-text, text-to-speech, and text-to-text generation, plus few-shot experiments such as speech recognition, speech synthesis, and intent classification.
Spirit LM did not dominate every comparison. The paper reports stronger clean-task results for some cascaded systems. Adding expressive units also creates tradeoffs in lexical and semantic modeling. Selected audio examples do not establish reliability across speakers, languages, accents, recording conditions, or accessibility needs.
Artifact and license boundary
The official repository exposes weights, inference code, evaluation scripts, and a model card. The FAIR Noncommercial Research License restricts covered materials and their outputs to noncommercial research uses and includes an acceptable-use policy.
The license is not a standard permissive software license. A public repository does not turn Spirit LM into a commercial assistant, supported service, or unrestricted voice tool.
Sources
Follow the evidence.
- about.fb.com: open source ai is the path forwardabout.fb.com
- about.fb.com: edit videos with meta aiabout.fb.com
- about.fb.com: introducing vibes ai videosabout.fb.com
- ai.meta.com: fair robotics open sourceai.meta.com
- ai.meta.com: fair news segment anything 2 1 meta spirit lm layer skip salsa linguaai.meta.com
- ai.meta.com: movie gen video sound generation blumhouseai.meta.com
- ai.meta.com: movie genai.meta.com
- ai.meta.com: movie gen a cast of media foundation modelsai.meta.com
- ai.meta.com: sparsh self supervised touch representations for vision based tactile sensingai.meta.com
- ai.meta.com: spiritlm licenseai.meta.com
- arxiv.org: 2402arxiv.org
- arxiv.org: 2410arxiv.org
- co-tracker.github.ioco-tracker.github.io
- daltonanderson.ghost.io: metas tech spree robotics video and ai releasesdaltonanderson.ghost.io
- github.com: co trackergithub.com
- github.com: sparshgithub.com
- github.com: spiritlmgithub.com
- open.spotify.com: 5OwJfB19t12yKJs4QayHy0open.spotify.com
- opensource.org: open source ai definitionopensource.org
- opensource.org: the open source initiative announces the release of the industrys first open source ai definitionopensource.org
- youtu.be: YKL shwSS Iyoutu.be