初始化项目,由ModelHub XC社区提供模型
Model: Mwanzau/Mazgu_Llama-1B-V2-Knowledge-16bit Source: Original Platform
This commit is contained in:
85
README.md
Normal file
85
README.md
Normal file
@@ -0,0 +1,85 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: unsloth/Llama-3.2-1B-Instruct
|
||||
tags:
|
||||
- safetensors
|
||||
- llama-3
|
||||
- merged
|
||||
- tumbuka
|
||||
- bantu
|
||||
- malawi
|
||||
language:
|
||||
- tum
|
||||
- en
|
||||
pipeline_tag: text-generation
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# Mazgu_Llama-1B-V2-Knowledge-16bit (Standalone Model)
|
||||
|
||||
**Mazgu_Llama-1B-V2-Knowledge-16bit** is a fully merged, standalone 16-bit `model.safetensors` release of Llama-3-1B, fine-tuned specifically for the Tumbuka language.
|
||||
|
||||
Because the LoRA weights are merged directly into the base architecture, this repository **does not require loading external adapter weights at runtime**. It serves as an ideal baseline for:
|
||||
* Direct inference via Hugging Face `transformers` or `vLLM`.
|
||||
* Continual pre-training or further SFT runs.
|
||||
* Full-model export and custom quantization pipelines.
|
||||
|
||||
---
|
||||
|
||||
## 💻 Quickstart Guide
|
||||
|
||||
### Option 1: Standard Transformers
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "Mwanzau/Mazgu_Llama-1B-V2-Knowledge-16bit"
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_id,
|
||||
torch_dtype=torch.float16,
|
||||
device_map="auto"
|
||||
)
|
||||
|
||||
prompt = "### Instruction:\nLongosolani vya Yesu mu Mateyo 24.\n\n### Response:\n"
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
|
||||
|
||||
outputs = model.generate(
|
||||
**inputs,
|
||||
max_new_tokens=150,
|
||||
temperature=0.6,
|
||||
repetition_penalty=1.2
|
||||
)
|
||||
|
||||
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
||||
|
||||
---
|
||||
|
||||
### Option 2: Optimized 4-Bit Loading with Unsloth (Colab / Kaggle)
|
||||
|
||||
from unsloth import FastLanguageModel
|
||||
|
||||
model, tokenizer = FastLanguageModel.from_pretrained(
|
||||
model_name = "Mwanzau/Mazgu_Llama-1B-V2-Knowledge-16bit",
|
||||
max_seq_length = 2048,
|
||||
load_in_4bit = True, # Saves VRAM on T4 GPUs
|
||||
)
|
||||
|
||||
FastLanguageModel.for_inference(model)
|
||||
|
||||
prompt = "### Instruction:\nLongosolani vyakurya ivyo vili bwino ku munda.\n\n### Response:\n"
|
||||
inputs = tokenizer([prompt], return_tensors="pt").to("cuda")
|
||||
|
||||
outputs = model.generate(**inputs, max_new_tokens=150, temperature=0.6, repetition_penalty=1.2)
|
||||
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
||||
|
||||
---
|
||||
|
||||
### 📊 Summary
|
||||
Format: Unsharded Native safetensors
|
||||
Precision: 16-bit Float (bfloat16/float16)
|
||||
Vocabulary Focus: Tumbuka instruction-following and factual knowledge retrieval.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user