Back to the episode map

Research Note

Spirit LM Paper Code and License Record

The [Spirit LM paper](https://arxiv.org/abs/2402.05755) describes a 7 billion parameter pretrained text language model extended to speech through continued training on te

Aug 4, 20262 min readBy Dalton Anderson

Spirit LM Paper Code and License Record

Model record

The Spirit LM paper describes a 7 billion parameter pretrained text language model extended to speech through continued training on text and speech units.

Spirit LM Base represents speech with HuBERT phonetic units. Spirit LM Expressive adds pitch and style units. Text uses subword BPE tokens. The training process interleaves aligned speech and text at word boundaries so the same autoregressive model can continue across modalities.

The paper reports training with 300 billion text tokens, 460,000 hours of speech-only data, and 110,000 hours of aligned speech and text. It also acknowledges errors in automatically produced alignments.

Evaluation boundary

The research covers speech-to-speech, speech-to-text, text-to-speech, and text-to-text generation, plus few-shot experiments such as speech recognition, speech synthesis, and intent classification.

Spirit LM did not dominate every comparison. The paper reports stronger clean-task results for some cascaded systems. Adding expressive units also creates tradeoffs in lexical and semantic modeling. Selected audio examples do not establish reliability across speakers, languages, accents, recording conditions, or accessibility needs.

Artifact and license boundary

The official repository exposes weights, inference code, evaluation scripts, and a model card. The FAIR Noncommercial Research License restricts covered materials and their outputs to noncommercial research uses and includes an acceptable-use policy.

The license is not a standard permissive software license. A public repository does not turn Spirit LM into a commercial assistant, supported service, or unrestricted voice tool.

Sources

Follow the evidence.

  1. youtu.be: YKL shwSS Iyoutu.be
  2. arxiv.org: 2402arxiv.org
  3. about.fb.com: open source ai is the path forwardabout.fb.com
  4. co-tracker.github.ioco-tracker.github.io
  5. ai.meta.com: sparsh self supervised touch representations for vision based tactile sensingai.meta.com
  6. arxiv.org: 2410arxiv.org
  7. github.com: co trackergithub.com
  8. ai.meta.com: movie gen video sound generation blumhouseai.meta.com
  9. ai.meta.com: movie gen a cast of media foundation modelsai.meta.com
  10. daltonanderson.ghost.io: metas tech spree robotics video and ai releasesdaltonanderson.ghost.io
  11. github.com: sparshgithub.com
  12. github.com: spiritlmgithub.com
  13. about.fb.com: edit videos with meta aiabout.fb.com
  14. ai.meta.com: movie genai.meta.com
  15. ai.meta.com: fair robotics open sourceai.meta.com
  16. open.spotify.com: 5OwJfB19t12yKJs4QayHy0open.spotify.com
  17. about.fb.com: introducing vibes ai videosabout.fb.com
  18. ai.meta.com: fair news segment anything 2 1 meta spirit lm layer skip salsa linguaai.meta.com
  19. opensource.org: the open source initiative announces the release of the industrys first open source ai definitionopensource.org
  20. opensource.org: open source ai definitionopensource.org
  21. ai.meta.com: spiritlm licenseai.meta.com