Model: jaimalleshk/Qwen3-0.6B-EN-MonoPrune Source: Original Platform
license, base_model, base_model_relation, language, library_name, pipeline_tag, tags
| license | base_model | base_model_relation | language | library_name | pipeline_tag | tags | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen3-0.6B | quantized |
|
gguf | text-generation |
|
Qwen3-0.6B-EN-MonoPrune (English-specialized, training-free)
⚠️ ENGLISH-ONLY. Non-English input and output degrade to byte-level artifacts BY DESIGN. That is the specialization, not a defect. If you serve any non-English traffic, do not use this model.
Qwen3-0.6B with 31.6% of its vocabulary (48,016 non-Latin-script tokens) and the corresponding embedding and tied LM-head rows structurally removed. No weights were trained, fine-tuned, or numerically altered — every retained parameter is byte-identical to the original.
Measured results
All figures come from a single environment: Intel i7-10750H (6C/12T laptop), Windows 11, 6 threads, llama.cpp b-221f0f635 built with the Vulkan backend registered but run with n_gpu_layers=0, so all layers compute on the CPU. Timings are medians over 20 strictly interleaved base/pruned pairs, reproduced in a second session at a 28% different host speed.
| Metric | Original | This model | Δ |
|---|---|---|---|
| Vocabulary | 151,936 | 103,920 | −31.6% |
| Parameters | 596.0 M | 546.9 M | −8.2% |
| File size (BF16 GGUF) | 1,198 MB | 1,098 MB | −8.4% |
| Peak RAM | 1,323 MB | 1,212 MB | −8.4% |
| Decode throughput | 22.67 t/s | 24.27 t/s | +7.3% (95% CI +5.5% to +8.6%) |
| Prefill throughput | 920.3 t/s | 914.3 t/s | −0.7% (unchanged, as predicted) |
| English perplexity | — | — | −0.03% |
| English generations | — | 52/52 byte-identical | under greedy decoding |
Read the speed number carefully. Twenty paired repetitions on one machine, none negative, with the effect reproduced at a second host speed 28% apart — so it is real. It is also modest. We report the median; the mean is +7.9%, inflated by one repetition whose base run stalled. Expect roughly +5% to +9% on similar CPU hardware, and treat the −8.4% memory saving as the more dependable benefit, since it is structural rather than measured.
Prefill is unchanged and should be. The output head runs once per prefill but once per generated token, so only decode can benefit. A prefill speedup would indicate a measurement error.
Why English output is byte-identical
English text tokenizes to the exact same token sequence after pruning — guaranteed, not sampled, because the BPE merge list is closed under the retained vocabulary (every retained merged token keeps both of its parts). Retained weights are unchanged, so the transformer computation is bit-equal; only the final softmax is over fewer entries. Under greedy decoding the argmax is therefore identical.
This guarantee does not extend to sampling. Temperature and top-p decoding over a different support was never tested and is not bounded by this result.
Retention policy (en_v2)
Kept: ASCII English, Latin-extended (résumé/café-class loanwords), digits, punctuation, Unicode symbols and math, Greek letters (π, α…), all 256 byte-fallback tokens, all control and chat tokens — plus 178 partial-UTF-8 fragments re-admitted by merge closure.
Removed: CJK, Kana, Hangul, Cyrillic, Arabic, Hebrew, Thai, Devanagari, other scripts, and unreferenced partial-UTF-8 fragments.
Greek is worth a note. Under policy v1 it was removed, and the model responded by writing \pi in LaTeX instead of π — it did not crash, it adapted, producing confident and subtly different output. Retaining Greek costs 244 tokens (0.16% of the vocabulary) and fixes it. Retention should be decided by semantic class, not by corpus frequency.
Full per-token classification in vocab_classes.csv; old→new ID map in the idmap JSON.
Known behaviours (from the 20-case battery)
| Behaviour | Detail |
|---|---|
| English text, code, math, numbers, diacritics | Identical to base |
| Non-English prompts | Degraded byte-artifact output, no crash |
| Mixed-language prompts | English stays fluent; conditioning on foreign words weakens — output can drift semantically while still reading well |
| Emoji | Mostly fine via byte-fallback; occasional variant-selector drift |
The mixed-language case is the one to watch in production: the failure is fluent, not obvious.
Composability
Pruning commutes byte-exactly with Q8_0 quantization — prune-then-quantize and quantize-then-prune produce byte-identical embeddings, so apply them in either order. Verified for Q8_0 only; Q4_K uses 256-weight super-blocks and needs a separate check.
Limitations
One model, one family, one size, one hardware environment, greedy decoding only.
How much you save depends on your model, and not simply on its size. The saving is (removable share of the vocabulary) × (embedding share), where embedding share ≈ V / (12·L·d) — vocabulary over depth × width. Across the four models we pruned: Gemma-3-4B −13.1%, Qwen3-0.6B −8.4%, Phi-4-mini −4.6%, Qwen3-4B −3.1%. A large multilingual vocabulary on a modest-width model pays best; parameter count alone predicts the wrong ordering. Batched GPU serving will see less decode benefit than batch-1 CPU, since the output-head cost amortises across the batch.
Provenance
Produced by the MonoPrune pipeline (training-free vocabulary pruning with BPE merge closure and byte-exact row surgery). Code, measurements and the full write-up: github.com/jaimalleshk/AI-Models-Mono-Language-Pruning. An arXiv preprint is in preparation; this card will be updated with the identifier when it is announced.
The repository includes a claims ledger recording every claim we withdrew, including three decode-speed figures on larger models withdrawn for lack of surviving logs. Read it before citing anything here.
Created 2026-08-02, revised 2026-08-04. Author: Jai Mallesh K R, Independent Researcher. Implementation and measurement were carried out with AI agents under the author's direction.
Attribution
Derivative of Qwen3-0.6B by the Qwen team (Alibaba Cloud), Apache License 2.0. Modification: structural removal of vocabulary entries and the corresponding embedding rows only. No weight values were altered.