Model: heli-stand/Qwen2.5-7B-hh-rlhf-sft Source: Original Platform
library_name, tags
| library_name | tags | ||
|---|---|---|---|
| transformers |
|
Model Card for SFT Model Fine-tuned on hh-rlhf Dataset from Qwen2.5-7B
This model is supervised fine-tuned from Qwen2.5-7B based on the anthropic helpfulness-harmlessness rlhf datasets.
Model Sources
- Repository: https://huggingface.co/Qwen/Qwen2.5-7B
Training Procedure
Training Data
Training Hyperparameters
- Optimizer: Adamw
- Learning Rate: 1e-5
- Batch Size: 128
- Epochs: 1
- Max_Length: 1024
Evaluation Results
- Best Evaluation Loss: 1.678
Description