Instructions to use facebook/seamless-m4t-v2-large with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use facebook/seamless-m4t-v2-large with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="facebook/seamless-m4t-v2-large")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("facebook/seamless-m4t-v2-large") model = AutoModel.from_pretrained("facebook/seamless-m4t-v2-large", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Clarification on the license of SeamlessM4T v2 speech encoder weights
Hi,
We are developing a Speech LLM that we intend to use commercially, and we would like to clarify the license that applies to the speech encoder weights from SeamlessM4T v2.
In our current implementation, we load the full SeamlessM4T v2 checkpoint and copy only the speech encoder weights into our own model:
full_seamless = AutoModel.from_pretrained(model_args.audio_tower_path)
model.audio_tower.load_state_dict(
full_seamless.speech_encoder.state_dict(),
strict=False
)
del full_seamless
We do not use the SeamlessM4T text decoder, T2U model, vocoder, or other model components. Only the speech_encoder parameters are used to initialize our speech encoder, after which the model is further trained as part of our own Speech LLM.
The Seamless Communication repository states that:
- W2v-BERT 2.0 speech encoder is licensed under the MIT license.
- SeamlessM4T models (v1 and v2) are licensed under CC-BY-NC 4.0.
We understand that the SeamlessM4T v2 speech encoder is based on the W2v-BERT 2.0 encoder, which was pretrained on 4.5M hours of unlabeled audio.
Could you please clarify which license applies specifically to the following weights?
full_seamless.speech_encoder.state_dict()
when full_seamless is loaded from the released SeamlessM4T v2 checkpoint.
More specifically:
- Are the
speech_encoderweights contained inside the SeamlessM4T v2 checkpoint considered part of the MIT-licensed W2v-BERT 2.0 speech encoder? - Or, because these weights are distributed as part of the SeamlessM4T v2 checkpoint, are they covered by the CC-BY-NC 4.0 license of SeamlessM4T v2?
- If we use only these speech encoder weights to initialize our own model and subsequently fine-tune/train them with our own data, would the resulting model be permitted for commercial use?
We would also appreciate confirmation on whether using the standalone facebook/w2v-bert-2.0 checkpoint instead would be the recommended approach for a commercial model.
Thank you very much for your clarification.