--- license: mit base_model: - WeiboAI/VibeThinker-3B library_name: gguf pipeline_tag: text-generation tags: - gguf - llama.cpp - quantized - unsloth - lora - trl - sft --- # VibeThinker-3B-SFT-Fable5-Glint-GGUF GGUF quantizations of a LoRA fine-tune of [`WeiboAI/VibeThinker-3B`](https://huggingface.co/WeiboAI/VibeThinker-3B), supervised fine-tuned on `ermiaazarkhalili/Fable-5-Glint-Clean` (private). Quantized from [`ermiaazarkhalili/VibeThinker-3B-SFT-Fable5-Glint`](https://huggingface.co/ermiaazarkhalili/VibeThinker-3B-SFT-Fable5-Glint). See that repository for the full-precision weights. | | | | --- | --- | | **Base model** | [`WeiboAI/VibeThinker-3B`](https://huggingface.co/WeiboAI/VibeThinker-3B) | | **Training data** | `ermiaazarkhalili/Fable-5-Glint-Clean` (private) | | **Method** | LoRA supervised fine-tuning via [Unsloth](https://github.com/unslothai/unsloth) + [TRL](https://github.com/huggingface/trl) | | **License** | `mit` (inherited from the base model) | ## Available quantizations | File | Size | | --- | --- | | `vibethinker-3b-sft-fable5-glint.q4_k_m.gguf` | 1.93 GB | | `vibethinker-3b-sft-fable5-glint.q5_k_m.gguf` | 2.22 GB | | `vibethinker-3b-sft-fable5-glint.q8_0.gguf` | 3.29 GB | ## Usage ### llama.cpp ```bash huggingface-cli download ermiaazarkhalili/VibeThinker-3B-SFT-Fable5-Glint-GGUF vibethinker-3b-sft-fable5-glint.q4_k_m.gguf --local-dir . llama-cli -m vibethinker-3b-sft-fable5-glint.q4_k_m.gguf -p "Explain gradient checkpointing in two sentences." -n 256 ``` ### Ollama ```bash echo 'FROM ./vibethinker-3b-sft-fable5-glint.q4_k_m.gguf' > Modelfile ollama create vibethinker-3b-sft-fable5-glint-gguf -f Modelfile ollama run vibethinker-3b-sft-fable5-glint-gguf ``` ## Training configuration | Setting | Value | | --- | --- | | LoRA rank (r) | 16 | | LoRA alpha | 16 | | Learning rate | 0.0002 | | Epochs | 3 | | Effective batch size | 8 (2 x 4 grad accum) | | Max sequence length | 4096 | | Base precision | 4-bit (QLoRA) | | Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | ## Observed training loss Measured from our SLURM logs for this configuration. These are training-loss observations only — no downstream benchmark evaluation has been run on this model, so they should not be read as a quality claim. | SLURM job | Steps | First loss | Final loss | | --- | --- | --- | --- | | `45987994` | 1,554 | 3.3731 | 1.1936 | | `46021015` | 1,554 | 3.3731 | 1.1904 | ## Limitations - No benchmark evaluation has been run on this checkpoint. The only reported numbers are training-loss observations. - Inherits the biases, knowledge cutoff and failure modes of the base model. - Fine-tuned on a single instruction-following dataset; behaviour outside that distribution is untested. - LoRA adapters were merged into the base weights, so the merged model cannot be detached from this fine-tune. ## Reproducing Trained by `notebooks/fable_distillation_vibethinker-3b_fable-glint_unsloth.ipynb`, executed non-interactively with papermill on a SLURM H100 partition (Unsloth + TRL, LoRA). --- *Card generated from the training run's own configuration and logs by* *`scripts/generate_hub_model_card.py`.*