Files
Pluto-Genesis-0.6B/README.md
ModelHub XC f5e7c3d58d 初始化项目,由ModelHub XC社区提供模型
Model: Pluto-AI-Labs/Pluto-Genesis-0.6B
Source: Original Platform
2026-08-30 01:44:18 +08:00

235 lines
7.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: apache-2.0
language:
- en
tags:
- pluto-genesis
- qwen3
- sub-1b
- qlora
- instruction-tuning
- reasoning
- mathematics
- coding
- fine-tuned
- conversational
- text-generation-inference
base_model: Qwen/Qwen3-0.6B
pipeline_tag: text-generation
library_name: transformers
datasets:
- teknium/OpenHermes-2.5
- microsoft/orca-math-word-problems-200k
- m-a-p/CodeFeedback-Filtered-Instruction
model-index:
- name: Pluto-Genesis-0.6B
results: []
---
<div align="center">
# 🪐 Pluto-Genesis-0.6B
<p align="center">
<img src="./banner.png" width="100%">
</p>
**An instruction-tuned sub-1B language model fine-tuned on 80K curated samples**
[![HuggingFace](https://img.shields.io/badge/🤗-HuggingFace-yellow)](https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue)](https://opensource.org/licenses/Apache-2.0)
[![Model Size](https://img.shields.io/badge/Parameters-596M-green)]()
[![Base Model](https://img.shields.io/badge/Base-Qwen3--0.6B-orange)](https://huggingface.co/Qwen/Qwen3-0.6B)
[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.21368749.svg)](https://doi.org/10.5281/zenodo.21368749)
</div>
---
## Model Description
**Pluto-Genesis-0.6B** is a fine-tuned instruction-following language model built on top of [Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B). It was trained using **QLoRA** (Quantized Low-Rank Adaptation) on a curated mixture of 80,000 high-quality instruction-response pairs spanning general reasoning, mathematical problem solving, and code generation.
This model is part of the **Pluto AI** research project by Siddharth N.R., exploring efficient fine-tuning of sub-1B language models on consumer-grade hardware.
> **Research Goal:** Demonstrate that a sub-1B model fine-tuned on carefully curated data can achieve competitive performance on reasoning, math, and coding benchmarks while remaining deployable on consumer hardware.
---
## Training Details
| Property | Value |
|----------|-------|
| **Base Model** | Qwen/Qwen3-0.6B |
| **Parameters** | 596 Million (596,049,920) |
| **Method** | QLoRA (4-bit NF4 + LoRA) |
| **LoRA Rank** | r=64, α=128 |
| **LoRA Target Modules** | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| **Training Steps** | 2,475 |
| **Final Training Loss** | 0.2741 |
| **Precision** | FP16 |
| **Optimizer** | Paged AdamW 8-bit |
| **Learning Rate** | 2e-4 (cosine schedule) |
| **Effective Batch Size** | 32 (2 × 16 grad accum) |
| **Sequence Length** | 1024 tokens |
| **Hardware** | Tesla T4 (16 GB) |
| **Framework** | Transformers + PEFT + TRL |
---
## Training Data
| Domain | Dataset | Samples | Skills Targeted |
|--------|---------|---------|----------------|
| 🧠 General Reasoning | [OpenHermes-2.5](https://huggingface.co/datasets/teknium/OpenHermes-2.5) | 30,000 | Instruction following, reasoning |
| 🔢 Mathematics | [Orca-Math-200K](https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k) | 30,000 | Word problems, step-by-step math |
| 💻 Code | [CodeFeedback-Filtered](https://huggingface.co/datasets/m-a-p/CodeFeedback-Filtered-Instruction) | 20,000 | Code generation, debugging |
| **Total** | | **80,000** | |
---
## Benchmarks
> 📊 Benchmarks will be added shortly. The model is currently being evaluated on ARC-Challenge, HellaSwag, MMLU, GSM8K, and TruthfulQA using lm-evaluation-harness v0.4.4.
<p align="center">
<img src="./detailed_benchmark.png" width="100%">
</p>
<p align="center">
<img src="./benchmark-comparison.png" width="100%">
</p>
---
## Usage
## GGUF Quantizations
GGUF versions for llama.cpp, Ollama and LM Studio are available here:
> https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B-GGUF
### Basic Inference
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"Siddh07ETH/Pluto-Genesis-0.6B",
torch_dtype=torch.float16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Siddh07ETH/Pluto-Genesis-0.6B")
messages = [{"role": "user", "content": "Explain what a neural network is."}]
text = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=256,
temperature=0.3,
do_sample=True,
top_p=0.9,
repetition_penalty=1.1,
)
response = tokenizer.decode(
output[0][inputs.input_ids.shape[1]:],
skip_special_tokens=True
)
print(response)
```
### Low Memory Inference (4-bit)
```python
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch
quant_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16,
)
model = AutoModelForCausalLM.from_pretrained(
"Siddh07ETH/Pluto-Genesis-0.6B",
quantization_config=quant_config,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Siddh07ETH/Pluto-Genesis-0.6B")
```
---
## Recommended Generation Settings
| Setting | Value | Reason |
|---------|-------|--------|
| `temperature` | 0.3 | Conservative — reduces hallucinations |
| `top_p` | 0.9 | Focused vocabulary |
| `repetition_penalty` | 1.1 | Prevents rambling |
| `max_new_tokens` | 200–512 | Keeps answers concise |
| `do_sample` | True | Required when temperature < 1.0 |
---
## Limitations
- **Model size:** At 596M parameters this model will hallucinate on topics outside its training distribution. Always verify factual claims.
- **Context length:** Trained on sequences up to 1024 tokens. Performance may degrade on longer contexts.
- **Knowledge cutoff:** The model does not have access to real-time information.
- **Research only:** Not intended for production deployment without further evaluation and safety testing.
---
## Author
**Siddharth N.R.**
[HuggingFace](https://huggingface.co/Siddh07ETH)
---
## 📄 Research Paper
The complete research paper describing the Pluto-Genesis training pipeline, benchmark evaluation, checkpoint recovery system, and engineering methodology is available on Zenodo.
**Paper**
> Pluto-Genesis: Reliable QLoRA Fine-Tuning on Ephemeral Compute with Cross-Session Checkpoint Recovery for Sub-1B Language Models
**DOI**
https://doi.org/10.5281/zenodo.21368749
This paper includes:
- Training methodology
- QLoRA configuration
- Cross-session checkpoint recovery
- Benchmark evaluation
- Deployment workflow
- Reproducibility details
## Citation
```bibtex
@misc{plutogenesis2026,
author = {Siddharth N.R.},
title = {Pluto-Genesis-0.6B: An Instruction-Tuned Sub-1B Language Model},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/Siddh07ETH/Pluto-Genesis-0.6B}
}
```
---
## License
Apache 2.0 — see [LICENSE](https://opensource.org/licenses/Apache-2.0).
Base model [Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) is also Apache 2.0.