Instructions to use suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # pip install "transformers<5.0.0" from transformers import pipeline pipe = pipeline("translation", model="suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M # Run inference directly in the terminal: llama cli -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M # Run inference directly in the terminal: llama cli -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
Use Docker
docker model run hf.co/suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct with Ollama:
ollama run hf.co/suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct with Docker Model Runner:
docker model run hf.co/suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
- Lemonade
How to use suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct:Q4_K_M
Run and chat with the model
lemonade run user.English_to_Hinglish_fintuned_lamma_3_8b_instruct-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Llama 3 8B — English → Hinglish Translation (LoRA)
Llama 3 8B Instruct fine-tuned with QLoRA to translate standard English into Hinglish — the romanized Hindi-English code-mixed register used in everyday informal communication across India.
- Developed by: Suyash Agarwal (GitHub · LinkedIn)
- Base model:
unsloth/llama-3-8b-Instruct-bnb-4bit - Training data: News_Hinglish_English (DOI: 10.57967/hf/5120) — a curated English ↔ Hinglish parallel corpus
- License: Apache 2.0
Intended uses
- Translating English text into natural, romanized Hinglish (chat, social, news summarization for Indian audiences).
- Research on code-switched and low-resource NLP for South Asian languages.
- A starting checkpoint for further fine-tuning on other code-mixed registers.
Out of scope: formal Hindi (Devanagari) translation; Hinglish → English (this model is unidirectional); safety-critical applications without human review.
How to use
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
pip install --no-deps xformers trl peft accelerate bitsandbytes
from unsloth import FastLanguageModel
import torch
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct",
max_seq_length=2048,
dtype=None, # auto-detect; float16 for T4/V100, bfloat16 for Ampere+
load_in_4bit=True,
)
def translate(text):
prompt = """Translate the input from English to Hinglish to give the response.
### Input:
{}
### Response:
"""
inputs = tokenizer([prompt.format(text)], return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=2048, use_cache=True)
raw = tokenizer.batch_decode(outputs)[0]
return raw.split("### Response:\n")[1].split("<|eot_id|>")[0]
print(translate("This is a fine-tuned Hinglish translation model using Llama 3."))
# Yeh ek fine-tuned Hinglish translation model hai jo Llama 3 ka istemal karta hai.
GGUF quantizations are included in this repo for llama.cpp / Ollama use.
Training details
- Method: QLoRA (4-bit) using Unsloth and HuggingFace TRL — approximately 2x faster training.
- Data: English ↔ Hinglish sentence pairs from the News_Hinglish_English corpus (news domain).
- Format: instruction-tuned prompt ("Translate the input from English to Hinglish").
Training loss:
Evaluation
Qualitative side-by-side with GPT-4o on held-out news text:
English input:
Finance Minister Nirmala Sitharaman said, "There used to be a poverty index...a human development index and all of them continue, but today what is keenly watched is VIX, the volatility index of the markets." Stability of the government is important for markets to be efficient, she stated. PM Narendra Modi's third term will make markets function with stability, she added.
GPT-4o:
Finance Minister Nirmala Sitharaman ne kaha, "Pehle ek poverty index hota tha...ek human development index hota tha aur yeh sab ab bhi hain, lekin aaj jo sabse zyada dekha ja raha hai, woh hai VIX, jo markets ka volatility index hai." Unhone kaha ki sarkar ki stability markets ke efficient hone ke liye zaroori hai. PM Narendra Modi ka teesra term markets ko stability ke saath function karne mein madad karega, unhone joda.
This model:
Finance Minister Nirmala Sitharaman ne kaha, "Pehle ek poverty index hota tha... ek human development index hota tha aur sab kuch ab bhi chal raha hai, lekin aaj jo kaafi zyada dekha ja raha hai, woh VIX hai, jo markets ki volatility ka index hai." Unhone kaha ki markets ke liye sarkar ki stability zaroori hai. PM Narendra Modi ke teesre term se markets stability ke saath function karenge, unhone joda.
Quantitative benchmarks (BLEU / chrF on a held-out test split, vs GPT-4o and base Llama 3 zero-shot) are in progress and will be added here.
Limitations & bias
- Trained primarily on news-domain text; heavy slang or regional registers may be less natural.
- Outputs romanized Hinglish only — no Devanagari.
- Code-mixing ratio reflects the training corpus and may not match every audience's preference.
- Inherits biases present in Llama 3 and the source news corpus.
Citation
@misc{agarwal2024newshinglish,
author = {Agarwal, Suyash},
title = {News\_Hinglish\_English: An English--Hinglish Parallel Corpus},
year = {2024},
publisher = {Hugging Face},
doi = {10.57967/hf/5120},
url = {https://huggingface.co/datasets/suyash2739/News_Hinglish_English}
}
- Downloads last month
- 410
4-bit
Model tree for suyash2739/English_to_Hinglish_fintuned_lamma_3_8b_instruct
Base model
unsloth/llama-3-8b-Instruct-bnb-4bit
