初始化项目,由ModelHub XC社区提供模型

Model: rodrigoramosrs/qwen3-4b-dotnet-specialist
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-28 01:02:17 +08:00
commit bacb895ff4
29 changed files with 303633 additions and 0 deletions

204
README.md Normal file
View File

@@ -0,0 +1,204 @@
---
license: cc-by-4.0
datasets:
- rodrigoramosrs/dotnet
language:
- en
- pt
base_model:
- Qwen/Qwen3-4B-Instruct-2507
tags:
- dotnet
- c#
- tech
- net
- develop
- documentation
- rodrigoramosrs
---
# 🧠 qwen3-4b-dotnet-specialist
### Fine-tuned model for technical reasoning and .NET documentation understanding
[![Model Type](https://img.shields.io/badge/Model-Qwen3--4B-green)]() [![Framework](https://img.shields.io/badge/Framework-.NET-blue)]() [![Dataset](https://img.shields.io/badge/Dataset-dotnet--QnA--Curated-orange)](https://huggingface.co/datasets/rodrigoramosrs/dotnet) [![License](https://img.shields.io/badge/License-CC%20BY--SA%204.0-lightgrey)]()
---
## 📘 Overview
`qwen3-4b-dotnet-specialist` is a **fine-tuned variant of Qwen 3 (4B parameters)**, specialized in understanding and generating **accurate, structured, and deeply technical content** related to the **.NET ecosystem**, including C#, ASP.NET Core, EF Core, CLI tools, documentation standards, and advanced runtime concepts.
This model was trained with a **highly curated dataset of 70,000 question-answer pairs**, derived from the official Microsoft documentation repository ([dotnet/docs](https://github.com/dotnet/docs)).
---
## ☀️ A Sustainable Experiment in AI Engineering
This project was built under a guiding principle:
> “Good science is not made of answers, but of the **right questions**.”
Every question in the dataset was algorithmically generated to test **specific technical reasoning paths**, and each answer was produced through a **Retrieval-Augmented Generation (RAG)** process — retrieving context from the entire documentation dataset rather than from the paragraph that originated the question.
That design choice produced **richer, contextually consistent, and cross-referenced answers** — making this model particularly strong in **documentation synthesis, reasoning across APIs, and multi-version comparison tasks**.
The entire curation, training, and evaluation process was powered using **solar energy**, highlighting that *research-grade AI can be done locally, sustainably, and accessibly*.
---
## 🧩 Dataset
📦 Dataset used: [**rodrigoramosrs/dotnet**](https://huggingface.co/datasets/rodrigoramosrs/dotnet)
- **Source:** Extracted and processed from [`github.com/dotnet/docs`](https://github.com/dotnet/docs)
- **Original size:** ~300 MB of unstructured text
- **Post-curation size:** ~60 MB
- **Format:** JSONL with `instruction`, `input`, and `output` keys
- **Samples:** ~70,000 Q&A pairs
- **Language:** English
- **Domain:** .NET / C# / Microsoft Docs structure
Each entry follows this structure:
```json
{
"instruction": "Explain how to organize tutorials in the .NET documentation portal.",
"input": "",
"output": "Detailed, step-by-step answer using the DocFX structure and YAML front-matter conventions."
}
````
---
## ⚙️ Training
* **Base model:** Qwen3-4B (Instruct variant)
* **Training method:** LoRA fine-tuning
* **Context length:** -
* **Batch size:** 8
* **Precision:** bfloat16
* **Epochs:** 6.0
* **Optimizer:** AdamW (8-bit)
* **Scheduler:** Cosine decay with warmup
* **Learning Rate:** 2e-4
* **Warmup Ratio:** 0.7
* **Gradient Accumulation Steps:** 6
* **Infrastructure:** Local GPU (RTX 5080)
* **Power source:** Off-grid solar system
### 🧮 Data Pipeline
1. **Extraction** – Crawled markdown files from [`dotnet/docs`](https://github.com/dotnet/docs)
2. **Cleaning** – Removed metadata, HTML, and outdated versions
3. **Segmentation** – Split long sections into atomic topics
4. **Question Generation** – Built synthetic instructions using a tuned model focused on documentation comprehension
5. **Answer Generation (RAG)** – Retrieved context from the *entire dataset* before generating final answers
6. **Ranking & Filtering** – Applied cross-encoder ranking and manual curation to ensure quality
7. **Finalization** – Consolidated into clean, versioned JSONL format
### 📊 Training Configuration
```python
Config:
trainer = SFTTrainer(
model=model,
train_dataset=train_ds,
tokenizer=tokenizer,
formatting_func=formatting_func,
args=SFTConfig(
per_device_train_batch_size=8,
gradient_accumulation_steps=6,
num_train_epochs=6.0,
learning_rate=2e-4,
lr_scheduler_type="cosine",
warmup_ratio=0.7,
logging_steps=10,
save_strategy="steps",
save_steps=200,
eval_steps=200,
output_dir=output_dir,
push_to_hub=False, # 🚫 impede upload automático
hub_model_id=None, # 🚫 não referencia repositório remoto
bf16=True,
gradient_checkpointing=True,
optim="adamw_8bit",
weight_decay=0.001,
max_grad_norm=1.0,
dataloader_pin_memory=False,
dataloader_num_workers=4,
report_to=None,
ddp_find_unused_parameters=True, # True ajuda em multi-GPU,
),
eval_dataset=eval_ds, # ✅ Inclui o dataset de avaliação (opcional, mas recomendado)
)
```
### 📈 Training Results
- **Final Loss:** 0.742800
- **Training Steps:** 3930
- **Evaluation Metrics:**
- **Perplexity:** Good
- **Factual Accuracy (manual):** ~High
- **Response Consistency:** High
- **Formatting Accuracy:** High
---
## 🧠 Intended Use
The model excels at:
* Explaining .NET concepts, frameworks, and internal mechanics
* Answering developer documentation questions
* Summarizing and rewriting technical guides
* Generating structured technical explanations
* Acting as a documentation assistant for software engineers
---
## 🚫 Limitations
* Limited to **.NET and related ecosystems** — not designed for general-purpose conversation.
* May occasionally produce overly detailed explanations when prompted ambiguously.
* Not a replacement for Microsoft’s official documentation — rather a **complementary reasoning model**.
---
## 💬 Example Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "rodrigoramosrs/qwen3-4b-dotnet-specialist"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16).eval()
prompt = """Explain how to publish an ASP.NET Core app using the .NET CLI."""
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=600, temperature=0.3, top_p=0.9)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
---
## 🔖 License & Attribution
* **Model License:** CC BY-SA 4.0
* **Dataset License:** CC BY-SA 4.0
* **Data Source:** Official Microsoft .NET documentation ([dotnet/docs](https://github.com/dotnet/docs))
* **Creator:** [Rodrigo Ramos (@rodrigoramosrs)](https://huggingface.co/rodrigoramosrs)
---
## 🌍 Closing Note
This project is a proof that **precision and sustainability** can coexist in AI research.
It demonstrates that with the **right questions**, good data, and discipline, one person — powered by sunlight — can build a specialized model that truly understands a complex technical domain.
> *Built locally. Trained on clean data. Powered by the sun.* ☀️
---
**Model:** [`rodrigoramosrs/qwen3-4b-dotnet-specialist`](https://huggingface.co/rodrigoramosrs/qwen3-4b-dotnet-specialist)
**Dataset:** [`rodrigoramosrs/dotnet`](https://huggingface.co/datasets/rodrigoramosrs/dotnet)