Research Note
Spirit LM Paper Code and License Record
The [Spirit LM paper](https://arxiv.org/abs/2402.05755) describes a 7 billion parameter pretrained text language model extended to speech through continued training on te
Spirit LM Paper Code and License Record
Model record
The Spirit LM paper describes a 7 billion parameter pretrained text language model extended to speech through continued training on text and speech units.
Spirit LM Base represents speech with HuBERT phonetic units. Spirit LM Expressive adds pitch and style units. Text uses subword BPE tokens. The training process interleaves aligned speech and text at word boundaries so the same autoregressive model can continue across modalities.
The paper reports training with 300 billion text tokens, 460,000 hours of speech-only data, and 110,000 hours of aligned speech and text. It also acknowledges errors in automatically produced alignments.
Evaluation boundary
The research covers speech-to-speech, speech-to-text, text-to-speech, and text-to-text generation, plus few-shot experiments such as speech recognition, speech synthesis, and intent classification.
Spirit LM did not dominate every comparison. The paper reports stronger clean-task results for some cascaded systems. Adding expressive units also creates tradeoffs in lexical and semantic modeling. Selected audio examples do not establish reliability across speakers, languages, accents, recording conditions, or accessibility needs.
Artifact and license boundary
The official repository exposes weights, inference code, evaluation scripts, and a model card. The FAIR Noncommercial Research License restricts covered materials and their outputs to noncommercial research uses and includes an acceptable-use policy.
The license is not a standard permissive software license. A public repository does not turn Spirit LM into a commercial assistant, supported service, or unrestricted voice tool.
Sources
Follow the evidence.
- youtu.be: YKL shwSS Iyoutu.be
- arxiv.org: 2402arxiv.org
- about.fb.com: open source ai is the path forwardabout.fb.com
- co-tracker.github.ioco-tracker.github.io
- ai.meta.com: sparsh self supervised touch representations for vision based tactile sensingai.meta.com
- arxiv.org: 2410arxiv.org
- github.com: co trackergithub.com
- ai.meta.com: movie gen video sound generation blumhouseai.meta.com
- ai.meta.com: movie gen a cast of media foundation modelsai.meta.com
- daltonanderson.ghost.io: metas tech spree robotics video and ai releasesdaltonanderson.ghost.io
- github.com: sparshgithub.com
- github.com: spiritlmgithub.com
- about.fb.com: edit videos with meta aiabout.fb.com
- ai.meta.com: movie genai.meta.com
- ai.meta.com: fair robotics open sourceai.meta.com
- open.spotify.com: 5OwJfB19t12yKJs4QayHy0open.spotify.com
- about.fb.com: introducing vibes ai videosabout.fb.com
- ai.meta.com: fair news segment anything 2 1 meta spirit lm layer skip salsa linguaai.meta.com
- opensource.org: the open source initiative announces the release of the industrys first open source ai definitionopensource.org
- opensource.org: open source ai definitionopensource.org
- ai.meta.com: spiritlm licenseai.meta.com