Files
ModelHub XC bacb895ff4 初始化项目,由ModelHub XC社区提供模型
Model: rodrigoramosrs/qwen3-4b-dotnet-specialist
Source: Original Platform
2026-08-28 01:02:17 +08:00

204 lines
7.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
license: cc-by-4.0
datasets:
- rodrigoramosrs/dotnet
language:
- en
- pt
base_model:
- Qwen/Qwen3-4B-Instruct-2507
tags:
- dotnet
- c#
- tech
- net
- develop
- documentation
- rodrigoramosrs
---
# 🧠 qwen3-4b-dotnet-specialist
### Fine-tuned model for technical reasoning and .NET documentation understanding
[![Model Type](https://img.shields.io/badge/Model-Qwen3--4B-green)]() [![Framework](https://img.shields.io/badge/Framework-.NET-blue)]() [![Dataset](https://img.shields.io/badge/Dataset-dotnet--QnA--Curated-orange)](https://huggingface.co/datasets/rodrigoramosrs/dotnet) [![License](https://img.shields.io/badge/License-CC%20BY--SA%204.0-lightgrey)]()
---
## 📘 Overview
`qwen3-4b-dotnet-specialist` is a **fine-tuned variant of Qwen 3 (4B parameters)**, specialized in understanding and generating **accurate, structured, and deeply technical content** related to the **.NET ecosystem**, including C#, ASP.NET Core, EF Core, CLI tools, documentation standards, and advanced runtime concepts.
This model was trained with a **highly curated dataset of 70,000 question-answer pairs**, derived from the official Microsoft documentation repository ([dotnet/docs](https://github.com/dotnet/docs)).
---
## ☀️ A Sustainable Experiment in AI Engineering
This project was built under a guiding principle:
> “Good science is not made of answers, but of the **right questions**.”
Every question in the dataset was algorithmically generated to test **specific technical reasoning paths**, and each answer was produced through a **Retrieval-Augmented Generation (RAG)** process — retrieving context from the entire documentation dataset rather than from the paragraph that originated the question.
That design choice produced **richer, contextually consistent, and cross-referenced answers** — making this model particularly strong in **documentation synthesis, reasoning across APIs, and multi-version comparison tasks**.
The entire curation, training, and evaluation process was powered using **solar energy**, highlighting that *research-grade AI can be done locally, sustainably, and accessibly*.
---
## 🧩 Dataset
📦 Dataset used: [**rodrigoramosrs/dotnet**](https://huggingface.co/datasets/rodrigoramosrs/dotnet)
- **Source:** Extracted and processed from [`github.com/dotnet/docs`](https://github.com/dotnet/docs)
- **Original size:** ~300 MB of unstructured text
- **Post-curation size:** ~60 MB
- **Format:** JSONL with `instruction`, `input`, and `output` keys
- **Samples:** ~70,000 Q&A pairs
- **Language:** English
- **Domain:** .NET / C# / Microsoft Docs structure
Each entry follows this structure:
```json
{
"instruction": "Explain how to organize tutorials in the .NET documentation portal.",
"input": "",
"output": "Detailed, step-by-step answer using the DocFX structure and YAML front-matter conventions."
}
````
---
## ⚙️ Training
* **Base model:** Qwen3-4B (Instruct variant)
* **Training method:** LoRA fine-tuning
* **Context length:** -
* **Batch size:** 8
* **Precision:** bfloat16
* **Epochs:** 6.0
* **Optimizer:** AdamW (8-bit)
* **Scheduler:** Cosine decay with warmup
* **Learning Rate:** 2e-4
* **Warmup Ratio:** 0.7
* **Gradient Accumulation Steps:** 6
* **Infrastructure:** Local GPU (RTX 5080)
* **Power source:** Off-grid solar system
### 🧮 Data Pipeline
1. **Extraction** – Crawled markdown files from [`dotnet/docs`](https://github.com/dotnet/docs)
2. **Cleaning** – Removed metadata, HTML, and outdated versions
3. **Segmentation** – Split long sections into atomic topics
4. **Question Generation** – Built synthetic instructions using a tuned model focused on documentation comprehension
5. **Answer Generation (RAG)** – Retrieved context from the *entire dataset* before generating final answers
6. **Ranking & Filtering** – Applied cross-encoder ranking and manual curation to ensure quality
7. **Finalization** – Consolidated into clean, versioned JSONL format
### 📊 Training Configuration
```python
Config:
trainer = SFTTrainer(
model=model,
train_dataset=train_ds,
tokenizer=tokenizer,
formatting_func=formatting_func,
args=SFTConfig(
per_device_train_batch_size=8,
gradient_accumulation_steps=6,
num_train_epochs=6.0,
learning_rate=2e-4,
lr_scheduler_type="cosine",
warmup_ratio=0.7,
logging_steps=10,
save_strategy="steps",
save_steps=200,
eval_steps=200,
output_dir=output_dir,
push_to_hub=False, # 🚫 impede upload automático
hub_model_id=None, # 🚫 não referencia repositório remoto
bf16=True,
gradient_checkpointing=True,
optim="adamw_8bit",
weight_decay=0.001,
max_grad_norm=1.0,
dataloader_pin_memory=False,
dataloader_num_workers=4,
report_to=None,
ddp_find_unused_parameters=True, # True ajuda em multi-GPU,
),
eval_dataset=eval_ds, # ✅ Inclui o dataset de avaliação (opcional, mas recomendado)
)
```
### 📈 Training Results
- **Final Loss:** 0.742800
- **Training Steps:** 3930
- **Evaluation Metrics:**
- **Perplexity:** Good
- **Factual Accuracy (manual):** ~High
- **Response Consistency:** High
- **Formatting Accuracy:** High
---
## 🧠 Intended Use
The model excels at:
* Explaining .NET concepts, frameworks, and internal mechanics
* Answering developer documentation questions
* Summarizing and rewriting technical guides
* Generating structured technical explanations
* Acting as a documentation assistant for software engineers
---
## 🚫 Limitations
* Limited to **.NET and related ecosystems** — not designed for general-purpose conversation.
* May occasionally produce overly detailed explanations when prompted ambiguously.
* Not a replacement for Microsoft’s official documentation — rather a **complementary reasoning model**.
---
## 💬 Example Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_name = "rodrigoramosrs/qwen3-4b-dotnet-specialist"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16).eval()
prompt = """Explain how to publish an ASP.NET Core app using the .NET CLI."""
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=600, temperature=0.3, top_p=0.9)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
---
## 🔖 License & Attribution
* **Model License:** CC BY-SA 4.0
* **Dataset License:** CC BY-SA 4.0
* **Data Source:** Official Microsoft .NET documentation ([dotnet/docs](https://github.com/dotnet/docs))
* **Creator:** [Rodrigo Ramos (@rodrigoramosrs)](https://huggingface.co/rodrigoramosrs)
---
## 🌍 Closing Note
This project is a proof that **precision and sustainability** can coexist in AI research.
It demonstrates that with the **right questions**, good data, and discipline, one person — powered by sunlight — can build a specialized model that truly understands a complex technical domain.
> *Built locally. Trained on clean data. Powered by the sun.* ☀️
---
**Model:** [`rodrigoramosrs/qwen3-4b-dotnet-specialist`](https://huggingface.co/rodrigoramosrs/qwen3-4b-dotnet-specialist)
**Dataset:** [`rodrigoramosrs/dotnet`](https://huggingface.co/datasets/rodrigoramosrs/dotnet)