--- base_model: - microsoft/FastContext-1.0-4B-SFT library_name: transformers pipeline_tag: text-generation tags: - unsloth - lora - trl - sft --- # FastContext-4B-SFT_base-SFT-Claude-Opus-Reasoning-Unsloth A LoRA fine-tune of [`microsoft/FastContext-1.0-4B-SFT`](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT), supervised fine-tuned on `ermiaazarkhalili/claude-reasoning-distillation` (private) (config `sft`). | | | | --- | --- | | **Base model** | [`microsoft/FastContext-1.0-4B-SFT`](https://huggingface.co/microsoft/FastContext-1.0-4B-SFT) | | **Architecture** | `Qwen3ForCausalLM` | | **Parameters** | 4.0B | | **Training data** | `ermiaazarkhalili/claude-reasoning-distillation` (private) (config `sft`) | | **Method** | LoRA supervised fine-tuning via [Unsloth](https://github.com/unslothai/unsloth) + [TRL](https://github.com/huggingface/trl) | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "ermiaazarkhalili/FastContext-4B-SFT_base-SFT-Claude-Opus-Reasoning-Unsloth" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, dtype='auto', device_map='auto') messages = [{"role": "user", "content": "Explain gradient checkpointing in two sentences."}] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors='pt' ).to(model.device) outputs = model.generate(inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ## Training configuration | Setting | Value | | --- | --- | | LoRA rank (r) | 16 | | LoRA alpha | 16 | | Learning rate | 0.0002 | | Epochs | 1 | | Effective batch size | 8 (2 x 4 grad accum) | | Max sequence length | 2048 | | Base precision | 4-bit (QLoRA) | | Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | ## Observed training loss Measured from our SLURM logs for this configuration. These are training-loss observations only — no downstream benchmark evaluation has been run on this model, so they should not be read as a quality claim. | SLURM job | Steps | First loss | Final loss | | --- | --- | --- | --- | | `45169147` | 1,310 | 1.7844 | 0.7491 | ## Limitations - No benchmark evaluation has been run on this checkpoint. The only reported numbers are training-loss observations. - Inherits the biases, knowledge cutoff and failure modes of the base model. - Fine-tuned on a single instruction-following dataset; behaviour outside that distribution is untested. - LoRA adapters were merged into the base weights, so the merged model cannot be detached from this fine-tune. ## Reproducing Trained by `notebooks/sft_distillation_fastcontext-4b-sft_unsloth.ipynb`, executed non-interactively with papermill on a SLURM H100 partition (Unsloth + TRL, LoRA). --- *Card generated from the training run's own configuration and logs by* *`scripts/generate_hub_model_card.py`.*