Instructions to use mahwizzzz/avey-b-ur with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mahwizzzz/avey-b-ur with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="mahwizzzz/avey-b-ur", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("mahwizzzz/avey-b-ur", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Avey-B Urdu 25M
Urdu masked-language encoder trained from scratch. This is an adaptation of the Avey-B architecture to Urdu. It does not claim the benchmark results reported for Avey-B by Acharya and Hammoud.
Purpose
Use the checkpoint for Urdu classification, token classification, representation extraction, or continued pretraining. It is an encoder, not a chat model or free-text generator.
Upstream attribution
The architecture follows Avey-B (Acharya and Hammoud, ICLR 2026). Upstream code is at github.com/rimads/avey-b. Model code in this repository is included so the checkpoint can be loaded with trust_remote_code=True.
@inproceedings{2026aveyb,
title={Avey-B},
author={Acharya, Devang and Hammoud, Mohammad},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026}
}
Data provenance
| Item | Value |
|---|---|
| Corpus | HPLT 3.0 Urdu (urd_Arab), quality bins 10 and 9, Hub id HPLT/HPLT3.0 |
| Clean data | 31,297 train docs, 151 validation docs, 541,165,369 characters |
| Downstream sets named in metadata | mirfan899/imdb_urdu_reviews (sentiment, test), unimelb-nlp/wikiann config ur (NER, test) |
HPLT packages its dataset under CC0 but does not own the underlying web text. The Apache-2.0 model license does not grant rights to third-party source documents that may be reproduced by the model.
Document filters beyond the quality bins, deduplication, and the tokenizer training corpus are not stated outside MANIFEST.json. This card does not add figures from that file.
Training
| Item | Value |
|---|---|
| Sequence length | 512 |
| Token batch | 16,384 tokens/step |
| Training | 20,000 steps, BF16, fused AdamW |
| Peak learning rate | 7e-4 |
| Hardware/runtime | NVIDIA RTX 4060 8GB, approximately 20 minutes |
| Final train loss | 2.52 |
Not reported in this card: random seed, number of runs, warmup, weight decay, dropout, and gradient clipping.
Evaluation
The model-index point estimates are accuracy 0.785333 and entity F1 0.812677. The intervals below appear only in this prose table. The number of runs behind each interval is not stated. Repository ids for UrNova, HPLT Urdu BERT, and the XLM-R checkpoint are not stated.
| Model | Parameters | Sentiment accuracy | WikiANN Urdu NER F1 |
|---|---|---|---|
| Avey-B Urdu | 24.87M | 78.53% ± 0.51 | 81.27 |
| UrNova | 94.72M | 81.83% ± 0.42 | 87.79 |
| HPLT Urdu BERT | 124.36M | 83.93% ± 1.27 | 90.92 |
| XLM-R base | 278.04M | 78.80% ± 0.52 | 84.89 |
Other figures reported with this checkpoint, without a named device for the latency measurement:
- 1.71 ms encoder latency at batch 1 / length 128, and 148 MiB peak VRAM at batch 16 / length 512.
- 14.29 held-out MLM perplexity, 50.46% top-1, and 70.39% top-5 masked-token accuracy.
- 4.19 characters/token on held-out data.
Limitations
Benchmark coverage in this card is limited to masked-token prediction, movie-review sentiment, and Wikipedia-derived NER. The sentiment limitation stated previously still applies: evaluate on native, domain-specific Urdu before production use. The card describes that sentiment setup as machine-translated movie reviews.
License
Apache-2.0, as declared in the metadata above.
Usage
This model contains a custom Avey architecture, so trust_remote_code=True is required.
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
repo_id = "mahwizzzz/avey-b-ur"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(repo_id, trust_remote_code=True).eval()
text = "پاکستان کا دارالحکومت [MASK] ہے۔"
inputs = tokenizer(text, return_tensors="pt")
with torch.inference_mode():
logits = model(**inputs).logits
mask_index = (inputs.input_ids[0] == tokenizer.mask_token_id).nonzero()[0, 0]
top_ids = logits[0, mask_index].topk(5).indices
print(tokenizer.convert_ids_to_tokens(top_ids.tolist()))
Extract contextual token representations with AutoModel:
from transformers import AutoModel, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("mahwizzzz/avey-b-ur", trust_remote_code=True)
encoder = AutoModel.from_pretrained("mahwizzzz/avey-b-ur", trust_remote_code=True)
hidden = encoder(**tokenizer("یہ ایک اردو جملہ ہے۔", return_tensors="pt")).last_hidden_state
print(hidden.shape) # [batch, tokens, 384]
- Downloads last month
- 117
Dataset used to train mahwizzzz/avey-b-ur
Collection including mahwizzzz/avey-b-ur
Paper for mahwizzzz/avey-b-ur
Evaluation results
- Accuracy on IMDb Urdu Reviewstest set self-reported0.785
- Entity F1 on WikiANN Urdutest set self-reported0.813

