初始化项目,由ModelHub XC社区提供模型
Model: heli-stand/Qwen2.5-7B-hh-rlhf-sft Source: Original Platform
This commit is contained in:
31
README.md
Normal file
31
README.md
Normal file
@@ -0,0 +1,31 @@
|
||||
---
|
||||
library_name: transformers
|
||||
tags:
|
||||
- qwen2.5
|
||||
- hh-rlhf
|
||||
---
|
||||
# Model Card for SFT Model Fine-tuned on hh-rlhf Dataset from Qwen2.5-7B
|
||||
|
||||
This model is supervised fine-tuned from Qwen2.5-7B based on the anthropic helpfulness-harmlessness rlhf datasets.
|
||||
|
||||
### Model Sources
|
||||
|
||||
- **Repository:** https://huggingface.co/Qwen/Qwen2.5-7B
|
||||
|
||||
## Training Procedure
|
||||
|
||||
### Training Data
|
||||
|
||||
- **Repository:** https://huggingface.co/datasets/heli-stand/hh-rlhf
|
||||
|
||||
### Training Hyperparameters
|
||||
|
||||
- **Optimizer:** Adamw
|
||||
- **Learning Rate:** 1e-5
|
||||
- **Batch Size:** 128
|
||||
- **Epochs:** 1
|
||||
- **Max_Length:** 1024
|
||||
|
||||
## Evaluation Results
|
||||
|
||||
- **Best Evaluation Loss:** 1.678
|
||||
Reference in New Issue
Block a user