--- license: cc-by-4.0 datasets: - rodrigoramosrs/dotnet language: - en - pt base_model: - Qwen/Qwen3-4B-Instruct-2507 tags: - dotnet - c# - tech - net - develop - documentation - rodrigoramosrs --- # 🧠 qwen3-4b-dotnet-specialist ### Fine-tuned model for technical reasoning and .NET documentation understanding [![Model Type](https://img.shields.io/badge/Model-Qwen3--4B-green)]() [![Framework](https://img.shields.io/badge/Framework-.NET-blue)]() [![Dataset](https://img.shields.io/badge/Dataset-dotnet--QnA--Curated-orange)](https://huggingface.co/datasets/rodrigoramosrs/dotnet) [![License](https://img.shields.io/badge/License-CC%20BY--SA%204.0-lightgrey)]() --- ## 📘 Overview `qwen3-4b-dotnet-specialist` is a **fine-tuned variant of Qwen 3 (4B parameters)**, specialized in understanding and generating **accurate, structured, and deeply technical content** related to the **.NET ecosystem**, including C#, ASP.NET Core, EF Core, CLI tools, documentation standards, and advanced runtime concepts. This model was trained with a **highly curated dataset of 70,000 question-answer pairs**, derived from the official Microsoft documentation repository ([dotnet/docs](https://github.com/dotnet/docs)). --- ## ☀️ A Sustainable Experiment in AI Engineering This project was built under a guiding principle: > “Good science is not made of answers, but of the **right questions**.” Every question in the dataset was algorithmically generated to test **specific technical reasoning paths**, and each answer was produced through a **Retrieval-Augmented Generation (RAG)** process — retrieving context from the entire documentation dataset rather than from the paragraph that originated the question. That design choice produced **richer, contextually consistent, and cross-referenced answers** — making this model particularly strong in **documentation synthesis, reasoning across APIs, and multi-version comparison tasks**. The entire curation, training, and evaluation process was powered using **solar energy**, highlighting that *research-grade AI can be done locally, sustainably, and accessibly*. --- ## 🧩 Dataset 📦 Dataset used: [**rodrigoramosrs/dotnet**](https://huggingface.co/datasets/rodrigoramosrs/dotnet) - **Source:** Extracted and processed from [`github.com/dotnet/docs`](https://github.com/dotnet/docs) - **Original size:** ~300 MB of unstructured text - **Post-curation size:** ~60 MB - **Format:** JSONL with `instruction`, `input`, and `output` keys - **Samples:** ~70,000 Q&A pairs - **Language:** English - **Domain:** .NET / C# / Microsoft Docs structure Each entry follows this structure: ```json { "instruction": "Explain how to organize tutorials in the .NET documentation portal.", "input": "", "output": "Detailed, step-by-step answer using the DocFX structure and YAML front-matter conventions." } ```` --- ## ⚙️ Training * **Base model:** Qwen3-4B (Instruct variant) * **Training method:** LoRA fine-tuning * **Context length:** - * **Batch size:** 8 * **Precision:** bfloat16 * **Epochs:** 6.0 * **Optimizer:** AdamW (8-bit) * **Scheduler:** Cosine decay with warmup * **Learning Rate:** 2e-4 * **Warmup Ratio:** 0.7 * **Gradient Accumulation Steps:** 6 * **Infrastructure:** Local GPU (RTX 5080) * **Power source:** Off-grid solar system ### 🧮 Data Pipeline 1. **Extraction** – Crawled markdown files from [`dotnet/docs`](https://github.com/dotnet/docs) 2. **Cleaning** – Removed metadata, HTML, and outdated versions 3. **Segmentation** – Split long sections into atomic topics 4. **Question Generation** – Built synthetic instructions using a tuned model focused on documentation comprehension 5. **Answer Generation (RAG)** – Retrieved context from the *entire dataset* before generating final answers 6. **Ranking & Filtering** – Applied cross-encoder ranking and manual curation to ensure quality 7. **Finalization** – Consolidated into clean, versioned JSONL format ### 📊 Training Configuration ```python Config: trainer = SFTTrainer( model=model, train_dataset=train_ds, tokenizer=tokenizer, formatting_func=formatting_func, args=SFTConfig( per_device_train_batch_size=8, gradient_accumulation_steps=6, num_train_epochs=6.0, learning_rate=2e-4, lr_scheduler_type="cosine", warmup_ratio=0.7, logging_steps=10, save_strategy="steps", save_steps=200, eval_steps=200, output_dir=output_dir, push_to_hub=False, # 🚫 impede upload automático hub_model_id=None, # 🚫 não referencia repositório remoto bf16=True, gradient_checkpointing=True, optim="adamw_8bit", weight_decay=0.001, max_grad_norm=1.0, dataloader_pin_memory=False, dataloader_num_workers=4, report_to=None, ddp_find_unused_parameters=True, # True ajuda em multi-GPU, ), eval_dataset=eval_ds, # ✅ Inclui o dataset de avaliação (opcional, mas recomendado) ) ``` ### 📈 Training Results - **Final Loss:** 0.742800 - **Training Steps:** 3930 - **Evaluation Metrics:** - **Perplexity:** Good - **Factual Accuracy (manual):** ~High - **Response Consistency:** High - **Formatting Accuracy:** High --- ## 🧠 Intended Use The model excels at: * Explaining .NET concepts, frameworks, and internal mechanics * Answering developer documentation questions * Summarizing and rewriting technical guides * Generating structured technical explanations * Acting as a documentation assistant for software engineers --- ## 🚫 Limitations * Limited to **.NET and related ecosystems** — not designed for general-purpose conversation. * May occasionally produce overly detailed explanations when prompted ambiguously. * Not a replacement for Microsoft’s official documentation — rather a **complementary reasoning model**. --- ## 💬 Example Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_name = "rodrigoramosrs/qwen3-4b-dotnet-specialist" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.bfloat16).eval() prompt = """Explain how to publish an ASP.NET Core app using the .NET CLI.""" inputs = tokenizer(prompt, return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_new_tokens=600, temperature=0.3, top_p=0.9) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` --- ## 🔖 License & Attribution * **Model License:** CC BY-SA 4.0 * **Dataset License:** CC BY-SA 4.0 * **Data Source:** Official Microsoft .NET documentation ([dotnet/docs](https://github.com/dotnet/docs)) * **Creator:** [Rodrigo Ramos (@rodrigoramosrs)](https://huggingface.co/rodrigoramosrs) --- ## 🌍 Closing Note This project is a proof that **precision and sustainability** can coexist in AI research. It demonstrates that with the **right questions**, good data, and discipline, one person — powered by sunlight — can build a specialized model that truly understands a complex technical domain. > *Built locally. Trained on clean data. Powered by the sun.* ☀️ --- **Model:** [`rodrigoramosrs/qwen3-4b-dotnet-specialist`](https://huggingface.co/rodrigoramosrs/qwen3-4b-dotnet-specialist) **Dataset:** [`rodrigoramosrs/dotnet`](https://huggingface.co/datasets/rodrigoramosrs/dotnet)