ModelHub XC ae0db33715 初始化项目,由ModelHub XC社区提供模型
Model: wuyonghui0810/text-generation
Source: Original Platform
2026-09-29 16:19:14 +08:00

license, language, tasks, frameworks, base_model, base_model_relation, widgets, datasets, tags
license language tasks frameworks base_model base_model_relation widgets datasets tags
Apache License 2.0
zh
text-generation
PyTorch
Qwen/Qwen2.5-0.5B-Instruct
finetune
version task inputs output examples
1 text-generation
type displayType validator
text TextArea
max_words
128
displayType displayValueMapping
Text text
name title inputs
1 示例1
name data
text 光的速度是多少?
wuyonghui0810/General-Knowledge
文本生成
科技探索
永辉

Qwen2.5-0.5B-Instruct 微调模型

模型概述

这是一个基于 Qwen2.5-0.5B-Instruct 模型的微调版本,使用 LoRA (低秩适应) 训练方法在特定领域数据上进行了定制化训练,涵盖数学、英语、科学、化学、物理等1000条自定义数据进行训练,以增强模型在特定任务上的性能。

模型详情

  • 基础模型: Qwen/Qwen2.5-0.5B-Instruct
  • 模型类型: qwen2_5
  • 训练方法: LoRA 微调
  • 模板: qwen2_5
  • 模型大小: 0.5B 参数
  • 检查点: checkpoint-186-merged
  • 训练日期: 2025-08-05

训练配置

  • 最大序列长度: 1024 个 token
  • 学习率: 1e-4
  • 梯度累积步数: 16
  • 训练轮数: 3
  • 最终训练损失: 0.0802
  • 最终评估损失: 0.00013508

调用方法

快速开始

这里提供了一个使用 wuyonghui0810/text-generation模型的代码片段,展示了如何加载分词器和模型以及如何生成内容。

from modelscope import AutoModelForCausalLM, AutoTokenizer
model_name = "wuyonghui0810/text-generation"
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
prompt = "光的速度是多少?"
messages = [
    {"role": "system", "content": "You are Vkey, created by yonghui, You are a helpful assistant."},
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=512
)
generated_ids = [
    output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)

训练前

光的传播速度是多少:

  • 在标准条件下,光的传播速度约为 299792458 米/秒。这个速度是恒定不变的,无论光源的位置、距离多远,都保持不变。这与物理定律和实验数据一致,证明了光速是宇宙中的基本物理常数。

训练后

  • <<< 光的传播速度是多少
  • 用户问的是关于物理的问题:'光的传播速度是多少'。我需要根据物理学原理给出准确的答案。对于这个问题,正确的回答是:光在真空中的传播速度约为299,792,458米/秒。。
  • 光在真空中的传播速度约为299,792,458米/秒。
  • <<< 光的速度
  • 用户问的是关于物理的问题:'光的速度'。我需要根据物理学原理给出准确的答案。对于这个问题,正确的回答是:光在真空中的速度约为299,792,458米/秒。。
  • 光在真空中的速度约为299,792,458米/秒。

微调方法

使用 Swift 进行训练认知微调

CUDA_VISIBLE_DEVICES=0 \
swift sft \
    --torch_dtype 'float16' \
    --model 'Qwen/Qwen2.5-0.5B-Instruct' \
    --model_type 'qwen2_5' \
    --template 'qwen2_5' \
    --system 'You are Vkey, created by yonghui. You are a helpful assistant.' \
    --dataset '/mnt/workspace/myipynb/transformers/datasets/training_data.jsonl' \
    --max_length '1024' \
    --init_weights 'True' \
    --learning_rate '1e-4' \
    --gradient_accumulation_steps '16' \
    --eval_steps '500' \
    --truncation_strategy 'delete' \
    --report_to 'tensorboard' \
    --add_version False \
    --output_dir /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718 \
    --logging_dir /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/runs \
    --ignore_args_error True \
    --device_map 'cpu' \
    > /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/runs/run.log 2>&1 &

使用 swift 框架进行推理

CUDA_VISIBLE_DEVICES=0 \
swift infer \
    --model /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/checkpoint-186-merged \
    --merge_lora true \
    --infer_backend vllm \
    --temperature 0 \
    --max_new_tokens 2048

使用 swift 框架进行导出

--adapters 参数:用于导出适配器权重(如 LoRA),需要配合基础模型使用 --model 参数:用于导出完整模型,可以直接使用 由于 checkpoint-186-merged 是一个完整的合并模型(包含基础模型和适配器权重),所以应该使用 --model 参数。

方案一(使用的这个):

CUDA_VISIBLE_DEVICES=0 swift export \
--model /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/checkpoint-186-merged \
--push_to_hub true \
--hub_model_id 'wuyonghui0810/text-generation' \
--hub_token 'ms-c9f5013e-8343-4d26-a53b-4d5a75f8a973' \
--use_hf false

方案二:

CUDA_VISIBLE_DEVICES=0 swift export \
--adapters /mnt/workspace/myipynb/transformers/ms-swift/output/Qwen2.5-0.5B-Instruct/v0-20250805-170718/checkpoint-186 \
--push_to_hub true \
--hub_model_id 'wuyonghui0810/text-generation' \
--hub_token 'ms-c9f5013e-8343-4d26-a53b-4d5a75f8a973' \
--use_hf false

关键推理参数

  • --model: 用于导出完整模型,可以直接使用(或者--adapters:训练好的适配器/LoRA 权重路径)
  • --merge_lora true: 将 LoRA 权重与基础模型合并
  • --infer_backend vllm: 使用 vLLM 进行推理加速
  • --temperature 0: 确定性输出(贪婪解码)
  • --max_new_tokens 2048: 最大生成 token 数量

性能表现

  • 评估样本每秒处理数: 18.64
  • 训练样本每秒处理数: 5.045
  • 内存使用: 模型权重约 0.9276 GB
  • 推理速度: 使用 vLLM 和 CUDA 图形优化

系统要求

  • GPU: 支持 CUDA 的 GPU(推荐)
  • 显存: 至少 4-8GB 以获得最佳性能
  • 软件环境:
    • Python 3.11+
    • Swift 框架
    • vLLM 库
    • CUDA 工具包

模型能力

此微调模型设计用于:

  • 回答特定领域问题
  • 准确遵循指令
  • 提供一致且确定性的响应(temperature=0时)
  • 处理长上下文输入(最多8192个token)

局限性

  • 当 temperature=0 时响应是确定性的
  • 性能取决于微调数据的质量
  • 在训练领域外可能泛化能力有限

如需技术支持或有关此模型的问题,请参考模型目录中包含的训练日志和配置文件。

Description
Model synced from source: wuyonghui0810/text-generation
Readme 4.4 MiB
Languages
Jinja 100%