Instructions to use poolside-laguna-hackathon/laguna-vision with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use poolside-laguna-hackathon/laguna-vision with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="poolside-laguna-hackathon/laguna-vision")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("poolside-laguna-hackathon/laguna-vision", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Add project links to model card
Browse files
README.md
CHANGED
|
@@ -14,6 +14,8 @@ private: true
|
|
| 14 |
|
| 15 |
# Laguna Vision
|
| 16 |
|
|
|
|
|
|
|
| 17 |
Laguna Vision adds a visual input path to `poolside/Laguna-XS.2`. SigLIP encodes images, AnyRes tiling preserves screenshot/document detail, a resampler projector maps features into Laguna's embedding space, and LoRA adapters are trained with supervised visual-instruction data.
|
| 18 |
|
| 19 |
Method: **post-training multimodal adaptation via supervised fine-tuning**.
|
|
|
|
| 14 |
|
| 15 |
# Laguna Vision
|
| 16 |
|
| 17 |
+
[Open-source GitHub](https://github.com/aaronkazah/laguna-vision) 路 [Hugging Face model](https://huggingface.co/poolside-laguna-hackathon/laguna-vision)
|
| 18 |
+
|
| 19 |
Laguna Vision adds a visual input path to `poolside/Laguna-XS.2`. SigLIP encodes images, AnyRes tiling preserves screenshot/document detail, a resampler projector maps features into Laguna's embedding space, and LoRA adapters are trained with supervised visual-instruction data.
|
| 20 |
|
| 21 |
Method: **post-training multimodal adaptation via supervised fine-tuning**.
|