初始化项目,由ModelHub XC社区提供模型

Model: heli-stand/Qwen2.5-7B-hh-rlhf-sft
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-24 07:39:16 +08:00
commit 1df86c61f0
15 changed files with 152117 additions and 0 deletions

31
README.md Normal file
View File

@@ -0,0 +1,31 @@
---
library_name: transformers
tags:
- qwen2.5
- hh-rlhf
---
# Model Card for SFT Model Fine-tuned on hh-rlhf Dataset from Qwen2.5-7B
This model is supervised fine-tuned from Qwen2.5-7B based on the anthropic helpfulness-harmlessness rlhf datasets.
### Model Sources
- **Repository:** https://huggingface.co/Qwen/Qwen2.5-7B
## Training Procedure
### Training Data
- **Repository:** https://huggingface.co/datasets/heli-stand/hh-rlhf
### Training Hyperparameters
- **Optimizer:** Adamw
- **Learning Rate:** 1e-5
- **Batch Size:** 128
- **Epochs:** 1
- **Max_Length:** 1024
## Evaluation Results
- **Best Evaluation Loss:** 1.678