初始化项目,由ModelHub XC社区提供模型
Model: iamthewalrus67/kulyk-uk-en-grpo Source: Original Platform
This commit is contained in:
62
README.md
Normal file
62
README.md
Normal file
@@ -0,0 +1,62 @@
|
||||
---
|
||||
language:
|
||||
- uk
|
||||
- en
|
||||
license_name: lfm1.0
|
||||
license_link: https://www.liquid.ai/lfm-license
|
||||
tags:
|
||||
- translation
|
||||
- grpo
|
||||
- ukrainian
|
||||
- lfm2
|
||||
base_model: Yehor/kulyk-uk-en
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||
# kulyk-uk-en-grpo
|
||||
|
||||
Ukrainian-to-English translation model based on [Yehor/kulyk-uk-en](https://huggingface.co/Yehor/kulyk-uk-en) (LFM2-350M), improved with GRPO using calibrated guardrail rewards on WikiMatrix data.
|
||||
|
||||
## Results
|
||||
|
||||
### FLoRes+ devtest (sentence-by-sentence, greedy, repetition_penalty=1.05)
|
||||
|
||||
| Model | BLEU | chrF | CometKiwi |
|
||||
|-------|:----:|:----:|:---------:|
|
||||
| kulyk-uk-en (baseline) | 36.25 | 62.54 | 0.7401 |
|
||||
| **kulyk-uk-en-grpo** | **37.03** | **63.25** | **0.7436** |
|
||||
|
||||
### WMT24 uk-en (news domain, out-of-distribution)
|
||||
|
||||
| Model | BLEU | chrF | CometKiwi |
|
||||
|-------|:----:|:----:|:---------:|
|
||||
| kulyk-uk-en (baseline) | 27.90 | 54.49 | 0.6571 |
|
||||
| **kulyk-uk-en-grpo** | **28.25** | **54.54** | **0.6598** |
|
||||
|
||||
## Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained("iamthewalrus67/kulyk-uk-en-grpo", trust_remote_code=True)
|
||||
tokenizer = AutoTokenizer.from_pretrained("iamthewalrus67/kulyk-uk-en-grpo", trust_remote_code=True)
|
||||
|
||||
prompt = "Translate the text to English:\nПогода сьогодні чудова."
|
||||
input_ids = tokenizer.apply_chat_template(
|
||||
[{"role": "user", "content": prompt}],
|
||||
add_generation_prompt=True, return_tensors="pt", tokenize=True
|
||||
).to(model.device)
|
||||
|
||||
output = model.generate(input_ids, max_new_tokens=256, do_sample=False, repetition_penalty=1.05)
|
||||
print(tokenizer.decode(output[0][input_ids.shape[1]:], skip_special_tokens=True))
|
||||
```
|
||||
|
||||
## Training Details
|
||||
|
||||
- **Method**: GRPO with calibrated guardrail rewards
|
||||
- **Data**: WikiMatrix uk-en (132K pairs)
|
||||
- **Rewards**: chrF (0.30) + BLEU (0.25) + CometKiwi (0.20) + 5x calibrated guardrails (0.05 each)
|
||||
- **Training**: Full fine-tune (no LoRA), single GPU, 300 steps
|
||||
- **Best checkpoint**: step 100
|
||||
|
||||
See [reward-driven-translation](https://github.com/iamthewalrus67/reward-driven-translation) for full reproduction code.
|
||||
Reference in New Issue
Block a user