104 lines
5.8 KiB
Markdown
104 lines
5.8 KiB
Markdown
---
|
|
license: apache-2.0
|
|
library_name: transformers
|
|
base_model: Qwen/Qwen3-4B
|
|
pipeline_tag: text-generation
|
|
tags:
|
|
- moral reasoning
|
|
- value reasoning
|
|
- persona
|
|
- chain-of-thought
|
|
language:
|
|
- hi
|
|
---
|
|
|
|
# Model Card for MET-D-Qwen3-4B-hi-only
|
|
|
|
MET-D-Qwen3-4B-hi-only is a Hindi-only moral reasoning model fine-tuned from [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B). Given a moral dilemma, a character description, and a candidate action, it judges the action from that character's perspective and explains its judgment with an explicit chain-of-thought before answering. Moral dilemmas rarely have a single correct answer, which makes reasoning traces hard to verify. We address this by introducing a character perspective that yields a ground-truth answer, which is used for rejection-sampling the model's own reasoning traces, conditioned on a per-language, per-situation selection of theoretical grounds. Both the reasoning trace and the final answer are generated in Hindi.
|
|
|
|
## Model Details
|
|
|
|
- **Base model:** Qwen/Qwen3-4B
|
|
- **Task:** for a given `(situation, character description, action)` triple, answer two questions from the character's perspective:
|
|
1. Is the action acceptable? (`Yes` / `No` / `Ambiguous`)
|
|
2. Would (not) doing it be emotionally/mentally discomforting? (`Yes` / `No`)
|
|
- **Training data:** self-generated reasoning traces, rejection-sampled against the ground truth per character perspective, in Hindi only
|
|
- **Language coverage:** this checkpoint is trained on **Hindi only**. The multilingual model and other single-language variants are also available (see [Model Variants](#model-variants) below).
|
|
- **License:** Apache 2.0 (inherited from the base model)
|
|
|
|
## How to Get Started
|
|
|
|
```python
|
|
from transformers import AutoTokenizer
|
|
from vllm import LLM, SamplingParams
|
|
|
|
model_id = "launch/MET-D-Qwen3-4B-hi-only"
|
|
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
|
llm = LLM(model=model_id, max_model_len=16384)
|
|
|
|
situation = "..." # the moral dilemma narrative
|
|
character = "..." # the character description
|
|
action = "..." # the candidate action to judge
|
|
|
|
prompt = f"""मैं आपको एक नैतिक दुविधा दूँगा और आपको एक कैरेक्टर का विवरण और सवाल पेश करूँगा।
|
|
पहले सवाल का जवाब हाँ, नहीं, या अस्पष्ट होना चाहिए, और दूसरे सवाल का जवाब हाँ या नहीं होना चाहिए।
|
|
आपको आउटपुट फॉर्मेट का सख्ती से पालन करना चाहिए, क्योंकि इसे ऑटोमैटिकली पार्स किया जाएगा।
|
|
सुनिश्चित करें कि आप सभी 2 सवालों के लिए "कारण:" और "उत्तर:" दोनों आउटपुट करें।
|
|
|
|
नैतिक दुविधा:
|
|
{situation}
|
|
|
|
कैरेक्टर विवरण:
|
|
{character}
|
|
|
|
सवाल:
|
|
1. A के नज़रिए से, क्या निम्नलिखित कार्रवाई — जो पहले ही की जा चुकी है, जिसे करने पर विचार किया जा रहा है, जिसे करने पर विचार किया जा सकता है, या जिस पर स्थिति के संदर्भ में सवाल उठाया जा रहा है — स्वीकार्य मानी जाएगी? "{action}"
|
|
2. यदि पिछला उत्तर हाँ है, तो क्या A को वह कार्रवाई करना मानसिक या भावनात्मक रूप से असहज लगेगा? इसके विपरीत, यदि पिछला उत्तर नहीं है, तो क्या A को वह कार्रवाई न करना मानसिक या भावनात्मक रूप से असहज लगेगा?
|
|
|
|
आपका जवाब:
|
|
1. कारण: {{कारण}} उत्तर: {{हाँ/नहीं/अस्पष्ट}}
|
|
2. कारण: {{कारण}} उत्तर: {{हाँ/नहीं}}
|
|
"""
|
|
|
|
chat_prompt = tokenizer.apply_chat_template(
|
|
[{"role": "user", "content": prompt}],
|
|
tokenize=False,
|
|
add_generation_prompt=True,
|
|
)
|
|
|
|
sampling_params = SamplingParams(temperature=0.0, max_tokens=2048)
|
|
outputs = llm.generate(chat_prompt, sampling_params)
|
|
print(outputs[0].outputs[0].text)
|
|
```
|
|
|
|
## Model Variants
|
|
|
|
This checkpoint is part of the [MET collection](https://huggingface.co/collections/launch/met), which includes the same task across base models and language subsets:
|
|
|
|
| Repo | Base model | Language(s) |
|
|
|---|---|---|
|
|
| `launch/MET-D-Qwen3-4B` | Qwen3-4B | all 6 (mixed) |
|
|
| `launch/MET-D-Qwen3-4B-en-only` | Qwen3-4B | English only |
|
|
| `launch/MET-D-Qwen3-4B-es-only` | Qwen3-4B | Spanish only |
|
|
| `launch/MET-D-Qwen3-4B-hi-only` | Qwen3-4B | Hindi only |
|
|
| `launch/MET-D-Qwen3-4B-ko-only` | Qwen3-4B | Korean only |
|
|
| `launch/MET-D-Qwen3-4B-ms-only` | Qwen3-4B | Malay only |
|
|
| `launch/MET-D-Qwen3-4B-zh-only` | Qwen3-4B | Chinese only |
|
|
| `launch/MET-D-Qwen3-8B` | Qwen3-8B | all 6 (mixed) |
|
|
| `launch/MET-D-Qwen3-8B-en-only` | Qwen3-8B | English only |
|
|
| `launch/MET-D-Gemma3-4B` | Gemma-3-4B-it | all 6 (mixed) |
|
|
| `launch/MET-D-Gemma3-4B-en-only` | Gemma-3-4B-it | English only |
|
|
|
|
## Citation
|
|
|
|
If you use this, please cite:
|
|
|
|
```bibtex
|
|
@article{lee2026met,
|
|
title={MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning},
|
|
author={Lee, Ayoung and Kwon, Ryan and Zhang, Yunxiang and Liu, Yuxuan and Railton, Peter and Wang, Lu},
|
|
journal={arXiv preprint arXiv:2607.11736},
|
|
year={2026}
|
|
}
|
|
```
|