fromtransformersimportAutoModelForCausalLM,AutoTokenizermodel_id="ermiaazarkhalili/Qwen3-4B-SFT-Fable5-Glint"tokenizer=AutoTokenizer.from_pretrained(model_id)model=AutoModelForCausalLM.from_pretrained(model_id,dtype='auto',device_map='auto')messages=[{"role":"user","content":"Explain gradient checkpointing in two sentences."}]inputs=tokenizer.apply_chat_template(messages,add_generation_prompt=True,return_tensors='pt').to(model.device)outputs=model.generate(inputs,max_new_tokens=256)print(tokenizer.decode(outputs[0],skip_special_tokens=True))
Next-token accuracy on a deterministic held-out split of ermiaazarkhalili/Fable-5-Glint-Clean (n = 199 samples), scored against the base model unsloth/Qwen3-4B. Assistant tokens only; both models are scored identically.
Metric
Base
This model
Δ
Top-1 accuracy
0.5917
0.6910
+0.0993
Top-5 accuracy
0.8352
0.9125
+0.0773
A delta measures how far fine-tuning moved this model from its own starting point; it is not a ranking against other models, which start from different baselines.
Observed training loss
Measured from our SLURM logs for this configuration. These are training-loss
observations only — see the held-out evaluation above for measured accuracy.
SLURM job
Steps
First loss
Final loss
unlabelled
1,554
1.4223
0.8867
Limitations
No benchmark evaluation has been run on this checkpoint. The only reported
numbers are training-loss observations.
Inherits the biases, knowledge cutoff and failure modes of the base model.
Fine-tuned on a single instruction-following dataset; behaviour outside that
distribution is untested.
LoRA adapters were merged into the base weights, so the merged model cannot
be detached from this fine-tune.
Reproducing
Trained by notebooks/fable_distillation_qwen3-4b_fable-glint_unsloth.ipynb, executed non-interactively with
papermill on a SLURM H100 partition (Unsloth + TRL, LoRA).
Card generated from the training run's own configuration and logs byscripts/generate_hub_model_card.py.