Video-Text-to-Text
Transformers
Safetensors
English
llava
text-generation
multimodal
Eval Results (legacy)
Instructions to use lmms-lab/LLaVA-Video-7B-Qwen2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lmms-lab/LLaVA-Video-7B-Qwen2 with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForCausalLM processor = AutoProcessor.from_pretrained("lmms-lab/LLaVA-Video-7B-Qwen2") model = AutoModelForCausalLM.from_pretrained("lmms-lab/LLaVA-Video-7B-Qwen2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Is this a newer/better model than OneVision?
#1
by ehayes-haiper - opened
Title
Yes. In terms of video. It is a video specific model
Thanks! Is inference the same as llava-OneVision? I.e. all the same tokens, dimensions etc?
Almost the same.