52 lines
1.6 KiB
Markdown
52 lines
1.6 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
tags:
|
|||
|
|
- unsloth
|
|||
|
|
- trl
|
|||
|
|
- sft
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
base_model:
|
|||
|
|
- Qwen/Qwen3-0.6B
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
library_name: transformers
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# Qwen3-0.6B-PsychSupport-Expert
|
|||
|
|
|
|||
|
|
This project performs full fine-tuning on the **Qwen3-0.6B** language model to enhance its psychological support reasoning and empathetic response capabilities. The model was optimized using the bfloat16 (bf16) data type.
|
|||
|
|
|
|||
|
|
## Training Procedure
|
|||
|
|
|
|||
|
|
1. **Dataset Preparation**
|
|||
|
|
|
|||
|
|
* Dataset: Containing paired patient emotional context descriptions and step-by-step empathetic support responses.
|
|||
|
|
|
|||
|
|
2. **Model Loading and Configuration**
|
|||
|
|
|
|||
|
|
* Base model: **Qwen3-0.6B**, loaded with the `unsloth` library in bf16 precision.
|
|||
|
|
* Full fine-tuning (`full_finetuning=True`) applied to all layers to adapt the model for psychological support tasks.
|
|||
|
|
|
|||
|
|
3. **Supervised Fine-Tuning (SFT)**
|
|||
|
|
|
|||
|
|
* Utilized the Hugging Face TRL library with the Supervised Fine-Tuning approach.
|
|||
|
|
* The model was trained to generate both intermediate empathetic reasoning steps and final supportive messages.
|
|||
|
|
* Training hyperparameters:
|
|||
|
|
|
|||
|
|
* Epochs: 2
|
|||
|
|
* Learning rate: 2e-5
|
|||
|
|
* Batch size: 8
|
|||
|
|
|
|||
|
|
## Purpose and Outcome
|
|||
|
|
|
|||
|
|
* Enhanced the model’s ability to provide empathetic, context-aware psychological support to users.
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|
* Performance was measured on a held-out validation set with the following metric:
|
|||
|
|
|
|||
|
|
* **Support Coherence:** Rated 74.32% similarity to expert-generated responses.
|
|||
|
|
|
|||
|
|
## License
|
|||
|
|
|
|||
|
|
This project is licensed under the Apache License 2.0. See the [LICENSE](./LICENSE) file for details.
|