Qwen3.5-2B-Base-Pashto-Plus

Extended-vocabulary build of Qwen/Qwen3.5-2B-Base with 53 newly added tokens covering Pashto, Sindhi, and Balochi script forms that the stock Qwen3.5 tokenizer does not represent as single tokens.

⚠️ This is a base model. It is not instruction-tuned and will not chat. Its purpose is to serve as a starting point for fine-tuning on Pashto / Sindhi / Balochi text with an already-correct character vocabulary.

Why this exists

Qwen3.5's vocabulary (248,044 BPE base + 33 specials = 248,077 total) covers Arabic, Persian, and Urdu well, but splits many Pashto, Sindhi, and Balochi letters into multiple sub-tokens. That means a single Pashto character like ښ costs 2–3 tokens and fragments the model's attention during training and inference.

This repository fixes that by adding those characters as first-class single tokens.

Added tokens (53)

Group Tokens
Pashto-specific څ ځ ښ ږ ټ ډ ڼ ۍ
Sindhi ٻ ڀ ٿ ٽ ٺ ڄ ڃ ڇ ڌ ڏ ڊ ڍ ڙ ڦ ڻ ڱ ڳ
Balochi ۈ ۉ ۊ ۋ ۏ
Arabic-extended ۃ ۂ
Eastern Arabic-Indic digits ۰ ۱ ۲ ۳ ۴ ۵ ۶ ۷ ۸ ۹
Arabic decimal separator ؉
Bidi / formatting controls U+200D (ZWJ) · U+200F (RLM) · U+202A–U+202E
Arabic combining marks ٓ ٔ ٕ

All 53 tokens were freshly initialized with the base model's initializer_range = 0.02 and are distinct from each other (max pairwise cosine ≈ 0.075).

Vocabulary layout

Component Count Range
BPE base vocab 247,? (unchanged) 0 – 248,043
Original specials (33) 33 248,044 – 248,076
New Pashto/Sindhi/Balochi tokens 53 248,077 – 248,129
Total len(tokenizer) 248,130

Model embedding table: (248130, 2048), input/output tied.

File inventory

File Purpose
config.json Qwen3_5ForConditionalGeneration architecture config
generation_config.json default generation parameters
model.safetensors 3.76 GB bf16 weights (with resized + fresh-init rows)
tokenizer.json full tokenizer, trim_offsets: false
tokenizer_config.json chat template + special tokens
vocab.json BPE vocabulary (248,130 entries)
merges.txt BPE merges (247,587 rules, unchanged)
special_tokens_map.json bos / eos / unk / pad definitions
chat_template.jinja Qwen3.5 chat template

The tokenizer is deliberately GGUF-safe: ByteLevel.trim_offsets is set to false to match upstream Qwen3.5, and vocab.json / merges.txt are included so convert_hf_to_gguf.py works out of the box.

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

repo = "nassimjp/Qwen3.5-2B-Base-Pashto-Plus"
tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)

# Verify a Pashto letter now encodes as one token
ids = tok.encode("ښ", add_special_tokens=False)
print(ids)          # [248079]  ← single ID

# Base model: this continues text, it does not "answer"
inputs = tok("د افغانستان پلازمېنه", return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=40)
print(tok.decode(out[0], skip_special_tokens=True))

GGUF conversion

python llama.cpp/convert_hf_to_gguf.py \
    ./Qwen3.5-2B-Base-Pashto-Plus \
    --outfile Qwen3.5-2B-Base-Pashto-Plus.BF16.gguf \
    --outtype bf16

ℹ️ llama.cpp versions older than the one that added Qwen3.5 support will fail with BPE pre-tokenizer was not recognized. Update base.py via convert_hf_to_gguf_update.py with the tokenizer's chkhsh if needed.

Intended use

  • ✅ Continued pre-training on Pashto / Sindhi / Balochi corpora
  • ✅ LoRA / QLoRA fine-tuning with a Pashto-correct vocabulary
  • ✅ Research on tokenizer extension for low-resource scripts
  • ✅ Building a chat model on top via SFT

Out of scope

  • ❌ Direct chat / instruction following (this is a base model)
  • ❌ Text generation without a task prompt
  • ❌ Vision tasks — the vision tower is included but was not retrained

Technical notes

  • Base architecture: Qwen3_5ForConditionalGeneration (hybrid Gated DeltaNet + Gated Attention, 24 layers, hidden 2048)
  • Embeddings: tie_word_embeddings: true — verified tied after resize
  • New rows: independently sampled from N(0, 0.02²), no duplicates (max pairwise cosine 0.0747, all norms ~0.88–0.94)
  • No weights from the base model were altered — only 53 rows were appended

Provenance

  • Base: Qwen/Qwen3.5-2B-Base
  • Method: deterministic token-surgery script that audits the 175-atom Pashto/ Urdu/Persian/Balochi/Sindhi inventory, adds the 53 missing single characters, resizes embeddings, initializes new rows independently, and verifies every added token encodes to exactly one ID.
  • Transformers: 5.18.0.dev0
  • Forensic gate: ✅ passed

Citation

If you use this model, cite the original Qwen3.5 release:

@misc{qwen3.5,
  title  = {{Qwen3.5}: Towards Native Multimodal Agents},
  author = {{Qwen Team}},
  month  = {February},
  year   = {2026},
  url    = {https://qwen.ai/blog?id=qwen3.5}
}

License

Apache 2.0, inherited from the base model.


### How to use this

**Option A — via the HF web editor (the URL you linked)**

Open `https://huggingface.co/nassimjp/Qwen3.5-2B-Base-Pashto-Plus/new/main?filename=README.md`, paste the whole block above, commit.

**Option B — locally, then upload**

Save the block to `./Qwen3.5-2B-Base-Pashto-Plus/README.md` on disk, then either:

```bash
huggingface-cli upload nassimjp/Qwen3.5-2B-Base-Pashto-Plus \
  ./Qwen3.5-2B-Base-Pashto-Plus \
  --repo-type model \
  --exclude "*.pre_fix.bak" ".ipynb_checkpoints/*" "__pycache__/*"
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nassimjp/Qwen3.5-2B-Base-Pashto-Plus

Quantized
(29)
this model