ae0db33715009b6687770f2aa9572b411cfd2429
Model: wuyonghui0810/text-generation Source: Original Platform
license, language, tasks, frameworks, base_model, base_model_relation, widgets, datasets, tags
| license | language | tasks | frameworks | base_model | base_model_relation | widgets | datasets | tags | |||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Apache License 2.0 |
|
|
PyTorch |
|
finetune |
|
|
|
Qwen2.5-0.5B-Instruct 微调模型
模型概述
这是一个基于 Qwen2.5-0.5B-Instruct 模型的微调版本,使用 LoRA (低秩适应) 训练方法在特定领域数据上进行了定制化训练,涵盖数学、英语、科学、化学、物理等1000条自定义数据进行训练,以增强模型在特定任务上的性能。
模型详情
- 基础模型: Qwen/Qwen2.5-0.5B-Instruct
- 模型类型: qwen2_5
- 训练方法: LoRA 微调
- 模板: qwen2_5
- 模型大小: 0.5B 参数
- 检查点: checkpoint-186-merged
- 训练日期: 2025-08-05
训练配置
- 最大序列长度: 1024 个 token
- 学习率: 1e-4
- 梯度累积步数: 16
- 训练轮数: 3
- 最终训练损失: 0.0802
- 最终评估损失: 0.00013508
调用方法
快速开始
这里提供了一个使用 wuyonghui0810/text-generation模型的代码片段,展示了如何加载分词器和模型以及如何生成内容。
from modelscope import AutoModelForCausalLM, AutoTokenizer
model_name = "wuyonghui0810/text-generation"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
prompt = "光的速度是多少?"
messages = [
{"role": "system", "content": "You are Vkey, created by yonghui, You are a helpful assistant."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=512
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
训练前
光的传播速度是多少:
- 在标准条件下,光的传播速度约为 299792458 米/秒。这个速度是恒定不变的,无论光源的位置、距离多远,都保持不变。这与物理定律和实验数据一致,证明了光速是宇宙中的基本物理常数。
训练后
- <<< 光的传播速度是多少
- 用户问的是关于物理的问题:'光的传播速度是多少'。我需要根据物理学原理给出准确的答案。对于这个问题,正确的回答是:光在真空中的传播速度约为299,792,458米/秒。。
- 光在真空中的传播速度约为299,792,458米/秒。
- <<< 光的速度
- 用户问的是关于物理的问题:'光的速度'。我需要根据物理学原理给出准确的答案。对于这个问题,正确的回答是:光在真空中的速度约为299,792,458米/秒。。
- 光在真空中的速度约为299,792,458米/秒。
微调方法
使用 Swift 进行训练认知微调
CUDA_VISIBLE_DEVICES=0 \
swift sft \
--torch_dtype 'float16' \
--model 'Qwen/Qwen2.5-0.5B-Instruct' \
--model_type 'qwen2_5' \
--template 'qwen2_5' \
--system 'You are Vkey, created by yonghui. You are a helpful assistant.' \
--dataset '/mnt/workspace/myipynb/transformers/datasets/training_data.jsonl' \
--max_length '1024' \
--init_weights 'True' \
--learning_rate '1e-4' \
--gradient_accumulation_steps '16' \
--eval_steps '500' \
--truncation_strategy 'delete' \
--report_to 'tensorboard' \
--add_version False \
--output_dir /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718 \
--logging_dir /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/runs \
--ignore_args_error True \
--device_map 'cpu' \
> /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/runs/run.log 2>&1 &
使用 swift 框架进行推理
CUDA_VISIBLE_DEVICES=0 \
swift infer \
--model /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/checkpoint-186-merged \
--merge_lora true \
--infer_backend vllm \
--temperature 0 \
--max_new_tokens 2048
使用 swift 框架进行导出
--adapters 参数:用于导出适配器权重(如 LoRA),需要配合基础模型使用 --model 参数:用于导出完整模型,可以直接使用 由于 checkpoint-186-merged 是一个完整的合并模型(包含基础模型和适配器权重),所以应该使用 --model 参数。
方案一(使用的这个):
CUDA_VISIBLE_DEVICES=0 swift export \
--model /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/checkpoint-186-merged \
--push_to_hub true \
--hub_model_id 'wuyonghui0810/text-generation' \
--hub_token 'ms-c9f5013e-8343-4d26-a53b-4d5a75f8a973' \
--use_hf false
方案二:
CUDA_VISIBLE_DEVICES=0 swift export \
--adapters /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/checkpoint-186 \
--push_to_hub true \
--hub_model_id 'wuyonghui0810/text-generation' \
--hub_token 'ms-c9f5013e-8343-4d26-a53b-4d5a75f8a973' \
--use_hf false
关键推理参数
--model: 用于导出完整模型,可以直接使用(或者--adapters:训练好的适配器/LoRA 权重路径)--merge_lora true: 将 LoRA 权重与基础模型合并--infer_backend vllm: 使用 vLLM 进行推理加速--temperature 0: 确定性输出(贪婪解码)--max_new_tokens 2048: 最大生成 token 数量
性能表现
- 评估样本每秒处理数: 18.64
- 训练样本每秒处理数: 5.045
- 内存使用: 模型权重约 0.9276 GB
- 推理速度: 使用 vLLM 和 CUDA 图形优化
系统要求
- GPU: 支持 CUDA 的 GPU(推荐)
- 显存: 至少 4-8GB 以获得最佳性能
- 软件环境:
- Python 3.11+
- Swift 框架
- vLLM 库
- CUDA 工具包
模型能力
此微调模型设计用于:
- 回答特定领域问题
- 准确遵循指令
- 提供一致且确定性的响应(temperature=0时)
- 处理长上下文输入(最多8192个token)
局限性
- 当 temperature=0 时响应是确定性的
- 性能取决于微调数据的质量
- 在训练领域外可能泛化能力有限
如需技术支持或有关此模型的问题,请参考模型目录中包含的训练日志和配置文件。
Description
Languages
Jinja
100%