235 lines
7.0 KiB
Markdown
235 lines
7.0 KiB
Markdown
|
|
---
|
|||
|
|
license: apache-2.0
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
tags:
|
|||
|
|
- pluto-genesis
|
|||
|
|
- qwen3
|
|||
|
|
- sub-1b
|
|||
|
|
- qlora
|
|||
|
|
- instruction-tuning
|
|||
|
|
- reasoning
|
|||
|
|
- mathematics
|
|||
|
|
- coding
|
|||
|
|
- fine-tuned
|
|||
|
|
- conversational
|
|||
|
|
- text-generation-inference
|
|||
|
|
base_model: Qwen/Qwen3-0.6B
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
library_name: transformers
|
|||
|
|
datasets:
|
|||
|
|
- teknium/OpenHermes-2.5
|
|||
|
|
- microsoft/orca-math-word-problems-200k
|
|||
|
|
- m-a-p/CodeFeedback-Filtered-Instruction
|
|||
|
|
model-index:
|
|||
|
|
- name: Pluto-Genesis-0.6B
|
|||
|
|
results: []
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
<div align="center">
|
|||
|
|
|
|||
|
|
# 🪐 Pluto-Genesis-0.6B
|
|||
|
|
|
|||
|
|
<p align="center">
|
|||
|
|
<img src="./banner.png" width="100%">
|
|||
|
|
</p>
|
|||
|
|
|
|||
|
|
**An instruction-tuned sub-1B language model fine-tuned on 80K curated samples**
|
|||
|
|
|
|||
|
|
[](https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B)
|
|||
|
|
[](https://opensource.org/licenses/Apache-2.0)
|
|||
|
|
[]()
|
|||
|
|
[](https://huggingface.co/Qwen/Qwen3-0.6B)
|
|||
|
|
[](https://doi.org/10.5281/zenodo.21368749)
|
|||
|
|
|
|||
|
|
</div>
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Model Description
|
|||
|
|
|
|||
|
|
**Pluto-Genesis-0.6B** is a fine-tuned instruction-following language model built on top of [Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B). It was trained using **QLoRA** (Quantized Low-Rank Adaptation) on a curated mixture of 80,000 high-quality instruction-response pairs spanning general reasoning, mathematical problem solving, and code generation.
|
|||
|
|
|
|||
|
|
This model is part of the **Pluto AI** research project by Siddharth N.R., exploring efficient fine-tuning of sub-1B language models on consumer-grade hardware.
|
|||
|
|
|
|||
|
|
> **Research Goal:** Demonstrate that a sub-1B model fine-tuned on carefully curated data can achieve competitive performance on reasoning, math, and coding benchmarks while remaining deployable on consumer hardware.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Training Details
|
|||
|
|
|
|||
|
|
| Property | Value |
|
|||
|
|
|----------|-------|
|
|||
|
|
| **Base Model** | Qwen/Qwen3-0.6B |
|
|||
|
|
| **Parameters** | 596 Million (596,049,920) |
|
|||
|
|
| **Method** | QLoRA (4-bit NF4 + LoRA) |
|
|||
|
|
| **LoRA Rank** | r=64, α=128 |
|
|||
|
|
| **LoRA Target Modules** | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
|
|||
|
|
| **Training Steps** | 2,475 |
|
|||
|
|
| **Final Training Loss** | 0.2741 |
|
|||
|
|
| **Precision** | FP16 |
|
|||
|
|
| **Optimizer** | Paged AdamW 8-bit |
|
|||
|
|
| **Learning Rate** | 2e-4 (cosine schedule) |
|
|||
|
|
| **Effective Batch Size** | 32 (2 × 16 grad accum) |
|
|||
|
|
| **Sequence Length** | 1024 tokens |
|
|||
|
|
| **Hardware** | Tesla T4 (16 GB) |
|
|||
|
|
| **Framework** | Transformers + PEFT + TRL |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Training Data
|
|||
|
|
|
|||
|
|
| Domain | Dataset | Samples | Skills Targeted |
|
|||
|
|
|--------|---------|---------|----------------|
|
|||
|
|
| 🧠 General Reasoning | [OpenHermes-2.5](https://huggingface.co/datasets/teknium/OpenHermes-2.5) | 30,000 | Instruction following, reasoning |
|
|||
|
|
| 🔢 Mathematics | [Orca-Math-200K](https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k) | 30,000 | Word problems, step-by-step math |
|
|||
|
|
| 💻 Code | [CodeFeedback-Filtered](https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction) | 20,000 | Code generation, debugging |
|
|||
|
|
| **Total** | | **80,000** | |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Benchmarks
|
|||
|
|
|
|||
|
|
> 📊 Benchmarks will be added shortly. The model is currently being evaluated on ARC-Challenge, HellaSwag, MMLU, GSM8K, and TruthfulQA using lm-evaluation-harness v0.4.4.
|
|||
|
|
|
|||
|
|
<p align="center">
|
|||
|
|
<img src="./detailed_benchmark.png" width="100%">
|
|||
|
|
</p>
|
|||
|
|
|
|||
|
|
<p align="center">
|
|||
|
|
<img src="./benchmark-comparison.png" width="100%">
|
|||
|
|
</p>
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Usage
|
|||
|
|
|
|||
|
|
## GGUF Quantizations
|
|||
|
|
|
|||
|
|
GGUF versions for llama.cpp, Ollama and LM Studio are available here:
|
|||
|
|
|
|||
|
|
> https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B-GGUF
|
|||
|
|
|
|||
|
|
### Basic Inference
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"Siddh07ETH/Pluto-Genesis-0.6B",
|
|||
|
|
torch_dtype=torch.float16,
|
|||
|
|
device_map="auto",
|
|||
|
|
)
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained("Siddh07ETH/Pluto-Genesis-0.6B")
|
|||
|
|
|
|||
|
|
messages = [{"role": "user", "content": "Explain what a neural network is."}]
|
|||
|
|
|
|||
|
|
text = tokenizer.apply_chat_template(
|
|||
|
|
messages, tokenize=False, add_generation_prompt=True
|
|||
|
|
)
|
|||
|
|
inputs = tokenizer(text, return_tensors="pt").to(model.device)
|
|||
|
|
|
|||
|
|
with torch.no_grad():
|
|||
|
|
output = model.generate(
|
|||
|
|
**inputs,
|
|||
|
|
max_new_tokens=256,
|
|||
|
|
temperature=0.3,
|
|||
|
|
do_sample=True,
|
|||
|
|
top_p=0.9,
|
|||
|
|
repetition_penalty=1.1,
|
|||
|
|
)
|
|||
|
|
|
|||
|
|
response = tokenizer.decode(
|
|||
|
|
output[0][inputs.input_ids.shape[1]:],
|
|||
|
|
skip_special_tokens=True
|
|||
|
|
)
|
|||
|
|
print(response)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### Low Memory Inference (4-bit)
|
|||
|
|
|
|||
|
|
```python
|
|||
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
|
|||
|
|
import torch
|
|||
|
|
|
|||
|
|
quant_config = BitsAndBytesConfig(
|
|||
|
|
load_in_4bit=True,
|
|||
|
|
bnb_4bit_quant_type="nf4",
|
|||
|
|
bnb_4bit_compute_dtype=torch.float16,
|
|||
|
|
)
|
|||
|
|
model = AutoModelForCausalLM.from_pretrained(
|
|||
|
|
"Siddh07ETH/Pluto-Genesis-0.6B",
|
|||
|
|
quantization_config=quant_config,
|
|||
|
|
device_map="auto",
|
|||
|
|
)
|
|||
|
|
tokenizer = AutoTokenizer.from_pretrained("Siddh07ETH/Pluto-Genesis-0.6B")
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Recommended Generation Settings
|
|||
|
|
|
|||
|
|
| Setting | Value | Reason |
|
|||
|
|
|---------|-------|--------|
|
|||
|
|
| `temperature` | 0.3 | Conservative — reduces hallucinations |
|
|||
|
|
| `top_p` | 0.9 | Focused vocabulary |
|
|||
|
|
| `repetition_penalty` | 1.1 | Prevents rambling |
|
|||
|
|
| `max_new_tokens` | 200–512 | Keeps answers concise |
|
|||
|
|
| `do_sample` | True | Required when temperature < 1.0 |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- **Model size:** At 596M parameters this model will hallucinate on topics outside its training distribution. Always verify factual claims.
|
|||
|
|
- **Context length:** Trained on sequences up to 1024 tokens. Performance may degrade on longer contexts.
|
|||
|
|
- **Knowledge cutoff:** The model does not have access to real-time information.
|
|||
|
|
- **Research only:** Not intended for production deployment without further evaluation and safety testing.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Author
|
|||
|
|
|
|||
|
|
**Siddharth N.R.**
|
|||
|
|
[HuggingFace](https://huggingface.co/Siddh07ETH)
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 📄 Research Paper
|
|||
|
|
|
|||
|
|
The complete research paper describing the Pluto-Genesis training pipeline, benchmark evaluation, checkpoint recovery system, and engineering methodology is available on Zenodo.
|
|||
|
|
|
|||
|
|
**Paper**
|
|||
|
|
> Pluto-Genesis: Reliable QLoRA Fine-Tuning on Ephemeral Compute with Cross-Session Checkpoint Recovery for Sub-1B Language Models
|
|||
|
|
|
|||
|
|
**DOI**
|
|||
|
|
https://doi.org/10.5281/zenodo.21368749
|
|||
|
|
|
|||
|
|
This paper includes:
|
|||
|
|
|
|||
|
|
- Training methodology
|
|||
|
|
- QLoRA configuration
|
|||
|
|
- Cross-session checkpoint recovery
|
|||
|
|
- Benchmark evaluation
|
|||
|
|
- Deployment workflow
|
|||
|
|
- Reproducibility details
|
|||
|
|
|
|||
|
|
## Citation
|
|||
|
|
|
|||
|
|
```bibtex
|
|||
|
|
@misc{plutogenesis2026,
|
|||
|
|
author = {Siddharth N.R.},
|
|||
|
|
title = {Pluto-Genesis-0.6B: An Instruction-Tuned Sub-1B Language Model},
|
|||
|
|
year = {2026},
|
|||
|
|
publisher = {HuggingFace},
|
|||
|
|
url = {https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B}
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## License
|
|||
|
|
|
|||
|
|
Apache 2.0 — see [LICENSE](https://opensource.org/licenses/Apache-2.0).
|
|||
|
|
Base model [Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) is also Apache 2.0.
|