Lafzyn
Lafzyn maps Urdu script to IPA. It is a fine-tune of Qwen/Qwen3.5-0.8B for grapheme-to-phoneme conversion.
Purpose
A pronunciation layer for Urdu text-to-speech, dictionaries, and linguistic tools. It is not a general chat model. A demo is at spaces/mahwizzzz/lafzyn, and a GGUF build is at mahwizzzz/lafzyn-gguf.
Upstream attribution
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3.5-0.8B |
| Task | Urdu grapheme-to-phoneme |
| Output | IPA |
| This repository's license | Apache-2.0 |
On 2 October 2026 the base-model repository tags included license:apache-2.0. Confirm the current Qwen license on that page before redistribution.
Data provenance
The metadata dataset id is custom, which is not a Hub dataset. This card states that training used 100,000 Urdu–IPA pairs. It does not identify the text source, the IPA author (rules, lexicon, or model), a data license, or a public dataset repository. mahwizzzz/urdu-g2p is a separate repository and is not declared as this model's training set.
Training
No trainer configuration, step count, learning rate, hardware, or seed is stored in this repository.
Evaluation
Evaluated on 500 held-out samples from the phoneme map. The metric is character-level edit distance on IPA strings (PER; lower is better). The split protocol and whether training pairs overlap these 500 strings are not further specified.
| Metric | Value |
|---|---|
| Mean PER | 16.94% |
| Perfect transcriptions (PER = 0) | 88 / 500 (17.6%) |
| Near perfect (PER < 10%) | 166 / 500 (33.2%) |
Sample outputs
Best predictions (PER = 0.000):
| Urdu | Expected IPA | Predicted IPA |
|---|---|---|
| مختار احمد | mʊxˈt̪aːr ˈæhməd |
mʊxˈt̪aːr ˈæhməd |
| قابل رحم | qɑːˈbɪl rɛˈhəm |
qɑːˈbɪl rɛˈhəm |
| رواں ہفتے | rəˈʋãː ˈhəft̪eː |
rəˈʋãː ˈhəft̪eː |
| دفاعی کمیشن | d̪ɪˈfaː.iː kəˈmɪ.ʃən |
d̪ɪˈfaː.iː kəˈmɪ.ʃən |
| نفیس المزاج | nəˈfiːs əlmɪˈzaːd͡ʒ |
nəˈfiːs əlmɪˈzaːd͡ʒ |
Highest reported errors (Arabic loanwords and rare compounds):
| Urdu | Expected IPA | Predicted IPA | PER |
|---|---|---|---|
| ائمہ سبعہ | aɪˈʔɪm.maː ˈsab.ʕa |
ˈɛːmɑː sɪˈbɑː |
0.722 |
| نفع المصرف | ˈnafʕaː alˈmasˤrif |
nəˈfɑːl ʔalˈmɪrɑːf |
0.667 |
| میٹر کیولیٹ | ˈmeːʈər ˈkjuːlɪt |
meːʈər kiːˈloːɛt̪ |
0.562 |
The previous card attributes these errors to rare Arabic-origin compounds and technical transliterations that are underrepresented in the training data. That statement is retained; the frequency counts behind it are not in this repository.
Phoneme inventory
The previous card states that the model covers this Urdu IPA set. An independent inventory check is not included here.
- Stops: b p t̪ ʈ d̪ ɖ k ɡ q ʔ
- Fricatives: f s ʃ z ʒ x ɣ ɦ h ħ ʕ
- Affricates: tʃ dʒ
- Nasals: m n ɳ ŋ n̪
- Liquids and glides: r ɾ ɽ l w j ʋ
- Vowels: ə ɪ ʊ aː iː uː eː oː ɛ ɔ æ
- Diacritics: ː (length), ̃ (nasalization), ˈ ˌ (stress)
Limitations
- Arabic-origin loanwords and rare compounds have higher error in the table above (PER > 0.5).
- The card describes the target variety as Modern Standard Urdu. Dialect and code-switching results are not reported.
- Training-data licensing is not identified, so downstream redistribution of the pairs themselves is not cleared by this model card.
License
Apache-2.0 for this repository.
Usage
from transformers import AutoTokenizer, pipeline
import torch
pipe = pipeline(
"text-generation",
model="mahwizzzz/lafzyn",
torch_dtype=torch.float16,
device_map="auto",
)
messages = [
{"role": "system", "content": "You are an expert Urdu linguist and phonetician. Convert the given Urdu word or phrase into its IPA transcription. Output only the IPA string, nothing else."},
{"role": "user", "content": "پاکستان"},
]
out = pipe(messages, max_new_tokens=64, do_sample=False)
print(out[0]["generated_text"][-1]["content"]) # → pɑːkɪsˈt̪aːn
The comment shows an illustrative IPA string from the previous card. It is not a measured output stored in this repository.
Citation
@misc{lafzyn2026,
title = {Lafzyn: Urdu Grapheme-to-Phoneme with Qwen3.5-0.8B},
author = {Mahwiz Khalil},
year = {2026},
url = {https://huggingface.co/mahwizzzz/lafzyn},
note = {Fine-tuned on 100k Urdu IPA pairs, mean PER 16.94\%}
}
- Downloads last month
- 306