# Data Summary for qwen3-4b-dotnet-specialist ## 1. General information ### 1.0.1 Version of the Summary: 1.0 ### 1.0.2 Last update: 24-Nov-2025 ## 1.1 Model Developer Identification ### 1.1.1 Model Developer name and contact details: Rodrigo Ramos (@rodrigoramosrs) Contact: [Hugging Face Profile](https://huggingface.co/rodrigoramosrs) ## 1.2 Model Identification ### 1.2.1 Versioned model name(s): qwen3-4b-dotnet-specialist ### 1.2.2 Model release date: Nov-2025 ## 1.3 Overall training data size and characteristics ### 1.3.1 Size of dataset and characteristics #### 1.3.1.A Text training data size: ~70,000 question-answer pairs #### 1.3.1.B Text training data content: Training data is derived from the official Microsoft documentation repository (dotnet/docs) and includes: 1. Extracted and processed markdown files from GitHub repository dotnet/docs 2. Structured technical content covering .NET, C#, ASP.NET Core, EF Core, CLI tools, documentation standards, and advanced runtime concepts 3. High-quality question-answer pairs generated algorithmically to test specific technical reasoning paths 4. Answers produced through Retrieval-Augmented Generation (RAG) process using context from the entire documentation dataset rather than just the originating paragraph #### 1.3.1.C Image training data size: Not applicable. Images are not part of the training #### 1.3.1.D Image training data content: Not applicable #### 1.3.1.E Audio training data size: Not applicable. Audio data is not part of the training data #### 1.3.1.F Audio training data content: Not applicable #### 1.3.1.G Video training data size: Not applicable. Video data is not part of the training data #### 1.3.1.H Video training data content: Not applicable #### 1.3.1.I Other training data size: Not applicable #### 1.3.1.J Other training data content: Not applicable ### 1.3.2 Latest date of data acquisition/collection for model training: Not specified in the provided content ### 1.3.3 Is data collection ongoing to update the model with new data collection after deployment? No ### 1.3.4 Date the training dataset was first used to train the model: Not specified in the provided content ### 1.3.5 Rationale or purpose of data selection: Datasets were selected to maximize high-quality technical reasoning and problem-solving capabilities within the .NET ecosystem. The mixture emphasizes carefully curated, algorithmically generated synthetic question-answer pairs derived from official documentation to improve technical understanding, documentation synthesis, and reasoning across APIs while maintaining factual accuracy. ## 2. List of data sources ### 2.1 Publicly available datasets #### 2.1.1 Have you used publicly available datasets to train the model? Yes Source: Official Microsoft .NET documentation repository (dotnet/docs) ### 2.2 Private non-publicly available datasets obtained from third parties #### 2.2.1 Datasets commercially licensed by rights holders or their representatives Not applicable - dataset is derived from public GitHub repository #### 2.2.2 Private datasets obtained from other third-parties Not applicable - dataset is derived from public GitHub repository ### 2.3 Personal Information #### 2.3.1 Was personal data used to train the model? No personal data was used for training this model. ### 2.4 Synthetic data #### 2.4.1 Was any synthetic AI-generated data used to train the model? Yes - algorithmically generated question-answer pairs based on technical documentation content ## 3. Data processing aspects ### 3.1 Respect of reservation of rights from text and data mining exception or limitation #### 3.1.1 Does this dataset include any data protected by copyright, trademark, or patent? The dataset is derived from the public dotnet/docs GitHub repository which has appropriate licensing for reuse. ### 3.2 Other information #### 3.2.1 Does the dataset include information about consumer groups without revealing individual consumer identities? No personal or consumer identity information is included in the dataset. #### 3.2.2 Was the dataset cleaned or modified before model training? Yes - the dataset was cleaned and processed through: 1. Extraction from markdown files 2. Removal of metadata, HTML, and outdated versions 3. Segmentation into atomic topics 4. Algorithmic question generation 5. RAG-based answer generation using full documentation context 6. Cross-encoder ranking and manual curation for quality assurance 7. Consolidation into clean, versioned JSONL format