Lafzyn

Lafzyn maps Urdu script to IPA. It is a fine-tune of Qwen/Qwen3.5-0.8B for grapheme-to-phoneme conversion.

Purpose

A pronunciation layer for Urdu text-to-speech, dictionaries, and linguistic tools. It is not a general chat model. A demo is at spaces/mahwizzzz/lafzyn, and a GGUF build is at mahwizzzz/lafzyn-gguf.

Upstream attribution

Property Value
Base model Qwen/Qwen3.5-0.8B
Task Urdu grapheme-to-phoneme
Output IPA
This repository's license Apache-2.0

On 2 October 2026 the base-model repository tags included license:apache-2.0. Confirm the current Qwen license on that page before redistribution.

Data provenance

The metadata dataset id is custom, which is not a Hub dataset. This card states that training used 100,000 Urdu–IPA pairs. It does not identify the text source, the IPA author (rules, lexicon, or model), a data license, or a public dataset repository. mahwizzzz/urdu-g2p is a separate repository and is not declared as this model's training set.

Training

No trainer configuration, step count, learning rate, hardware, or seed is stored in this repository.

Evaluation

Evaluated on 500 held-out samples from the phoneme map. The metric is character-level edit distance on IPA strings (PER; lower is better). The split protocol and whether training pairs overlap these 500 strings are not further specified.

Metric Value
Mean PER 16.94%
Perfect transcriptions (PER = 0) 88 / 500 (17.6%)
Near perfect (PER < 10%) 166 / 500 (33.2%)

Sample outputs

Best predictions (PER = 0.000):

Urdu Expected IPA Predicted IPA
مختار احمد mʊxˈt̪aːr ˈæhməd mʊxˈt̪aːr ˈæhməd
قابل رحم qɑːˈbɪl rɛˈhəm qɑːˈbɪl rɛˈhəm
رواں ہفتے rəˈʋãː ˈhəft̪eː rəˈʋãː ˈhəft̪eː
دفاعی کمیشن d̪ɪˈfaː.iː kəˈmɪ.ʃən d̪ɪˈfaː.iː kəˈmɪ.ʃən
نفیس المزاج nəˈfiːs əlmɪˈzaːd͡ʒ nəˈfiːs əlmɪˈzaːd͡ʒ

Highest reported errors (Arabic loanwords and rare compounds):

Urdu Expected IPA Predicted IPA PER
ائمہ سبعہ aɪˈʔɪm.maː ˈsab.ʕa ˈɛːmɑː sɪˈbɑː 0.722
نفع المصرف ˈnafʕaː alˈmasˤrif nəˈfɑːl ʔalˈmɪrɑːf 0.667
میٹر کیولیٹ ˈmeːʈər ˈkjuːlɪt meːʈər kiːˈloːɛt̪ 0.562

The previous card attributes these errors to rare Arabic-origin compounds and technical transliterations that are underrepresented in the training data. That statement is retained; the frequency counts behind it are not in this repository.

Phoneme inventory

The previous card states that the model covers this Urdu IPA set. An independent inventory check is not included here.

  • Stops: b p t̪ ʈ d̪ ɖ k ɡ q ʔ
  • Fricatives: f s ʃ z ʒ x ɣ ɦ h ħ ʕ
  • Affricates: tʃ dʒ
  • Nasals: m n ɳ ŋ n̪
  • Liquids and glides: r ɾ ɽ l w j ʋ
  • Vowels: ə ɪ ʊ aː iː uː eː oː ɛ ɔ æ
  • Diacritics: ː (length), ̃ (nasalization), ˈ ˌ (stress)

Limitations

  • Arabic-origin loanwords and rare compounds have higher error in the table above (PER > 0.5).
  • The card describes the target variety as Modern Standard Urdu. Dialect and code-switching results are not reported.
  • Training-data licensing is not identified, so downstream redistribution of the pairs themselves is not cleared by this model card.

License

Apache-2.0 for this repository.

Usage

from transformers import AutoTokenizer, pipeline
import torch

pipe = pipeline(
    "text-generation",
    model="mahwizzzz/lafzyn",
    torch_dtype=torch.float16,
    device_map="auto",
)

messages = [
    {"role": "system", "content": "You are an expert Urdu linguist and phonetician. Convert the given Urdu word or phrase into its IPA transcription. Output only the IPA string, nothing else."},
    {"role": "user", "content": "پاکستان"},
]
out = pipe(messages, max_new_tokens=64, do_sample=False)
print(out[0]["generated_text"][-1]["content"])  # → pɑːkɪsˈt̪aːn

The comment shows an illustrative IPA string from the previous card. It is not a measured output stored in this repository.

Citation

@misc{lafzyn2026,
  title   = {Lafzyn: Urdu Grapheme-to-Phoneme with Qwen3.5-0.8B},
  author  = {Mahwiz Khalil},
  year    = {2026},
  url     = {https://huggingface.co/mahwizzzz/lafzyn},
  note    = {Fine-tuned on 100k Urdu IPA pairs, mean PER 16.94\%}
}
Downloads last month
306
Safetensors
Model size
0.9B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mahwizzzz/lafzyn

Finetuned
(494)
this model
Quantizations
1 model

Space using mahwizzzz/lafzyn 1

Collection including mahwizzzz/lafzyn