初始化项目,由ModelHub XC社区提供模型
Model: yezdata/SmolLM2-1.7B-Instruct-DocstringGenerator Source: Original Platform
This commit is contained in:
36
.gitattributes
vendored
Normal file
36
.gitattributes
vendored
Normal file
@@ -0,0 +1,36 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
smollm2_1_7b_instruct_merged-q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
221
README.md
Normal file
221
README.md
Normal file
@@ -0,0 +1,221 @@
|
||||
---
|
||||
language:
|
||||
- en
|
||||
license: apache-2.0
|
||||
library_name: transformers
|
||||
tags:
|
||||
- code
|
||||
- python
|
||||
- docstring
|
||||
- documentation
|
||||
- code-generation
|
||||
- lora
|
||||
- qlora
|
||||
- smollm2
|
||||
- instruct
|
||||
- causal-lm
|
||||
base_model: HuggingFaceTB/SmolLM2-1.7B-Instruct
|
||||
pipeline_tag: text-generation
|
||||
model-index:
|
||||
- name: SmolLM2-1.7B-Instruct-DocstringGenerator
|
||||
results: []
|
||||
datasets:
|
||||
- codeparrot/codeparrot-clean
|
||||
---
|
||||
|
||||
# SmolLM2-1.7B-Instruct · DocstringGenerator
|
||||
|
||||
> A fine-tuned **[SmolLM2-1.7B-Instruct](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct)** specialised in writing **concise, high-level Python docstrings** for functions, methods and classes.
|
||||
> This model is the backbone of the **[PyDoctor](https://github.com/yezdata/pydoctor)** CLI — a fully local, LLM-powered tool that automatically writes and manages docstrings in your Python codebase.
|
||||
|
||||
[](https://github.com/yezdata/pydoctor)
|
||||
[](LICENSE)
|
||||
[](https://www.python.org/)
|
||||
[](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct)
|
||||
|
||||
---
|
||||
|
||||
## Intended Use
|
||||
|
||||
The model generates **summary-style docstrings** — single-paragraph, plain-English descriptions of a Python code block's purpose and architectural role. It does **not** produce `Args:`, `Returns:`, or `Raises:` sections by design.
|
||||
|
||||
**Suitable for:**
|
||||
- Automated docstring generation in CI/CD pipelines
|
||||
- Interactive IDE plugins
|
||||
- Local, privacy-preserving documentation workflows via llama.cpp / GGUF
|
||||
|
||||
**Not suitable for:**
|
||||
- General-purpose code generation
|
||||
- Generating full NumPy/Google-style docstrings with parameter tables (explicitly omitted)
|
||||
- Non-Python languages
|
||||
|
||||
---
|
||||
|
||||
## Quick Start
|
||||
### With llama.cpp (GGUF · recommended for local use)
|
||||
|
||||
```bash
|
||||
# Download the Q8_0 GGUF
|
||||
huggingface-cli download \
|
||||
yezdata/SmolLM2-1.7B-Instruct-DocstringGenerator \
|
||||
smollm2_1_7b_instruct_merged-q8_0.gguf \
|
||||
--local-dir ./models
|
||||
|
||||
# Run inference
|
||||
llama-cli \
|
||||
-m ./models/smollm2_1_7b_instruct_merged-q8_0.gguf \
|
||||
--chat-template chatml \
|
||||
-p "..."
|
||||
```
|
||||
|
||||
> **Tip:** The [PyDoctor CLI](https://github.com/yezdata/pydoctor) handles prompt construction, parsing, and atomic file rewrites out of the box.
|
||||
|
||||
---
|
||||
|
||||
## Prompt Format (ChatML)
|
||||
|
||||
The model uses the **ChatML** template native to SmolLM2-Instruct:
|
||||
|
||||
```
|
||||
<|im_start|>system
|
||||
{SYSTEM_PROMPT}<|im_end|>
|
||||
<|im_start|>user
|
||||
CONTEXT
|
||||
{context_code}
|
||||
|
||||
TARGET CODE
|
||||
{target_code}<|im_end|>
|
||||
<|im_start|>assistant
|
||||
```
|
||||
|
||||
The model then generates only the raw docstring text, terminated by `<|im_end|>`.
|
||||
|
||||
**Context definition:**
|
||||
- **function** target -> context = "Independent code block"
|
||||
- **method** target → context = `__init__` signature of its enclosing class
|
||||
- **class** target → context = signatures of its methods
|
||||
|
||||
---
|
||||
|
||||
## Training Pipeline
|
||||
|
||||
### Stage 1 — Code Extraction
|
||||
|
||||
Raw Python source files were streamed from **[codeparrot/codeparrot-clean](https://huggingface.co/datasets/codeparrot/codeparrot-clean)** (~200 k samples). Each file passed a quality filter that rejected:
|
||||
|
||||
| Filter | Threshold |
|
||||
|---|---|
|
||||
| Too few lines | < 3 non-empty lines |
|
||||
| Minified code | avg line length > 150 chars |
|
||||
| Low alphabetic ratio | < 15 % (binary / machine-generated) |
|
||||
| Repetitive boilerplate | unique line ratio < 10 % |
|
||||
| Oversized files | > 50 000 characters |
|
||||
|
||||
Surviving files were parsed with **[LibCST](https://libcst.readthedocs.io/)** producing `(target, context)` pairs.
|
||||
|
||||
### Stage 2 — Synthetic Docstring Generation
|
||||
|
||||
`(target, context)` pairs were labelled in parallel using **DeepSeek V4 Flash** (via OpenRouter):
|
||||
|
||||
The teacher-model system prompt enforced:
|
||||
1. Describe semantic purpose and architectural role, not implementation details
|
||||
2. Use context to disambiguate class membership
|
||||
|
||||
### Stage 3 — Instruct Data Preparation & Tokenisation
|
||||
|
||||
Synthetic batches were assembled into ChatML prompt/completion pairs:
|
||||
|
||||
```python
|
||||
prompt = (
|
||||
f"<|im_start|>system\n{SYSTEM_PROMPT}<|im_end|>\n"
|
||||
f"<|im_start|>user\nCONTEXT\n{context}\n\nTARGET CODE\n{target}<|im_end|>\n"
|
||||
f"<|im_start|>assistant\n"
|
||||
)
|
||||
completion = f"{docstring}<|im_end|>"
|
||||
```
|
||||
|
||||
Labels were constructed so that **only completion tokens** are trained on — prompt tokens are masked from cross-entropy loss.
|
||||
|
||||
### Stage 4 — QLoRA Fine-tuning
|
||||
|
||||
Fine-tuning was performed on Kaggle kernels (`instruct_finetune.py`):
|
||||
|
||||
| Hyperparameter | Value |
|
||||
|---|---|
|
||||
| Quantisation | 4-bit NF4, double quant, fp16 compute |
|
||||
| LoRA rank `r` | 32 |
|
||||
| LoRA alpha `α` | 64 |
|
||||
| LoRA dropout | 0.2 |
|
||||
| LoRA bias | none |
|
||||
| Target modules | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` |
|
||||
| Optimizer | AdamW 8-bit (bitsandbytes) |
|
||||
| Learning rate | 2e-4 |
|
||||
| LR schedule | Cosine with 5 % warmup |
|
||||
| Weight decay | 0.01 |
|
||||
| Batch size | 8 per device |
|
||||
| Gradient accumulation | 8 steps → effective batch 64 |
|
||||
| Epochs | 1 |
|
||||
| Max sequence length | 1 024 tokens (95th-pct filter) |
|
||||
| Validation split | 1 % held-out, evaluated each epoch |
|
||||
| Seed | 1337 |
|
||||
|
||||
Loss = next-token cross-entropy, **prompt tokens ignored** via label mask.
|
||||
|
||||
### Stage 5 — LoRA Merge & GGUF Export
|
||||
|
||||
After training, LoRA adapters were merged back into the base model weights and converted to **Q8_0 GGUF** using `llama.cpp`:
|
||||
|
||||
```
|
||||
LoRA adapter (epoch 1, safetensors)
|
||||
│
|
||||
▼ merge_and_unload()
|
||||
│
|
||||
merged fp16 safetensors
|
||||
│
|
||||
▼ llama.cpp convert_hf_to_gguf.py --outtype q8_0
|
||||
▼
|
||||
smollm2_1_7b_instruct_merged-q8_0.gguf
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Files
|
||||
|
||||
| File | Description |
|
||||
|---|---|
|
||||
| `smollm2_1_7b_instruct_merged-q8_0.gguf` | Q8_0 GGUF for llama.cpp — recommended for local use |
|
||||
| `safetensors/model.safetensors` | Merged fp16 weights |
|
||||
| `safetensors/config.json` | HuggingFace model configuration |
|
||||
| `safetensors/tokenizer.json` / `safetensors/tokenizer_config.json` | SmolLM2-1.7B-Instruct tokenizer |
|
||||
|
||||
---
|
||||
|
||||
## Limitations & Bias
|
||||
|
||||
- **Summary-only style:** the model is trained to output a single-paragraph summary. It will not produce `Args:` / `Returns:` sections.
|
||||
- **Python only:** trained exclusively on Python source code from codeparrot-clean.
|
||||
- **Context dependency:** quality improves when the correct context string is provided. Passing an empty context for class methods may reduce coherence.
|
||||
- **Teacher model bias:** docstring style reflects DeepSeek V4 Flash's preferences filtered through the strict prompt rules. Unusual code idioms may yield generic descriptions.
|
||||
- **Not a general assistant:** the model is heavily specialised and will likely perform poorly on tasks other than docstring generation.
|
||||
|
||||
---
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{pydoctor2026,
|
||||
author = {yezdata},
|
||||
title = {PyDoctor: Local LLM-powered Python Docstring Generator},
|
||||
year = {2026},
|
||||
howpublished = {\url{https://github.com/yezdata/pydoctor}},
|
||||
note = {Fine-tuned SmolLM2-1.7B-Instruct model available at
|
||||
\url{https://huggingface.co/yezdata/SmolLM2-1.7B-Instruct-DocstringGenerator}}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## License
|
||||
|
||||
This model is released under the **Apache 2.0** license, matching the base `SmolLM2-1.7B-Instruct` model.
|
||||
Training data originates from `codeparrot/codeparrot-clean` (MIT)
|
||||
6
safetensors/chat_template.jinja
Normal file
6
safetensors/chat_template.jinja
Normal file
@@ -0,0 +1,6 @@
|
||||
{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
|
||||
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
|
||||
' }}{% endif %}{{'<|im_start|>' + message['role'] + '
|
||||
' + message['content'] + '<|im_end|>' + '
|
||||
'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
|
||||
' }}{% endif %}
|
||||
43
safetensors/config.json
Normal file
43
safetensors/config.json
Normal file
@@ -0,0 +1,43 @@
|
||||
{
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 1,
|
||||
"dtype": "float16",
|
||||
"eos_token_id": 2,
|
||||
"head_dim": 64,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 2048,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 8192,
|
||||
"max_position_embeddings": 8192,
|
||||
"mlp_bias": false,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 32,
|
||||
"num_hidden_layers": 24,
|
||||
"num_key_value_heads": 32,
|
||||
"pad_token_id": 2,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 130000,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"tie_word_embeddings": true,
|
||||
"transformers.js_config": {
|
||||
"dtype": "q4",
|
||||
"kv_cache_dtype": {
|
||||
"fp16": "float16",
|
||||
"q4f16": "float16"
|
||||
},
|
||||
"use_external_data_format": {
|
||||
"model.onnx": true,
|
||||
"model_fp16.onnx": true
|
||||
}
|
||||
},
|
||||
"transformers_version": "5.13.0",
|
||||
"use_cache": true,
|
||||
"vocab_size": 49152
|
||||
}
|
||||
7
safetensors/generation_config.json
Normal file
7
safetensors/generation_config.json
Normal file
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": 2,
|
||||
"pad_token_id": 2,
|
||||
"transformers_version": "5.13.0"
|
||||
}
|
||||
3
safetensors/model.safetensors
Normal file
3
safetensors/model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:aa3f38dd222a93f5675531f5b33cc6e83f76e706bba64455d8890b5bddc2e25e
|
||||
size 3422777736
|
||||
244965
safetensors/tokenizer.json
Normal file
244965
safetensors/tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
19
safetensors/tokenizer_config.json
Normal file
19
safetensors/tokenizer_config.json
Normal file
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": "<|im_start|>",
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"extra_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>"
|
||||
],
|
||||
"is_local": false,
|
||||
"local_files_only": false,
|
||||
"model_max_length": 8192,
|
||||
"pad_token": "<|im_end|>",
|
||||
"tokenizer_class": "GPT2Tokenizer",
|
||||
"unk_token": "<|endoftext|>",
|
||||
"vocab_size": 49152
|
||||
}
|
||||
3
smollm2_1_7b_instruct_merged-q8_0.gguf
Normal file
3
smollm2_1_7b_instruct_merged-q8_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:bf2ef7dbd77adb7ca8691172cd2e3813a4742bde1eb19cd4d5585e8118dcd0f2
|
||||
size 1820414496
|
||||
Reference in New Issue
Block a user