Files
qwen3-4b-dotnet-specialist/data_summary_card.md
ModelHub XC bacb895ff4 初始化项目,由ModelHub XC社区提供模型
Model: rodrigoramosrs/qwen3-4b-dotnet-specialist
Source: Original Platform
2026-08-28 01:02:17 +08:00

4.4 KiB

Data Summary for qwen3-4b-dotnet-specialist

1. General information

1.0.1 Version of the Summary: 1.0

1.0.2 Last update: 24-Nov-2025

1.1 Model Developer Identification

1.1.1 Model Developer name and contact details:

Rodrigo Ramos (@rodrigoramosrs)
Contact: Hugging Face Profile

1.2 Model Identification

1.2.1 Versioned model name(s):

qwen3-4b-dotnet-specialist

1.2.2 Model release date:

Nov-2025

1.3 Overall training data size and characteristics

1.3.1 Size of dataset and characteristics

1.3.1.A Text training data size:

~70,000 question-answer pairs

1.3.1.B Text training data content:

Training data is derived from the official Microsoft documentation repository (dotnet/docs) and includes:

  1. Extracted and processed markdown files from GitHub repository dotnet/docs
  2. Structured technical content covering .NET, C#, ASP.NET Core, EF Core, CLI tools, documentation standards, and advanced runtime concepts
  3. High-quality question-answer pairs generated algorithmically to test specific technical reasoning paths
  4. Answers produced through Retrieval-Augmented Generation (RAG) process using context from the entire documentation dataset rather than just the originating paragraph

1.3.1.C Image training data size:

Not applicable. Images are not part of the training

1.3.1.D Image training data content:

Not applicable

1.3.1.E Audio training data size:

Not applicable. Audio data is not part of the training data

1.3.1.F Audio training data content:

Not applicable

1.3.1.G Video training data size:

Not applicable. Video data is not part of the training data

1.3.1.H Video training data content:

Not applicable

1.3.1.I Other training data size:

Not applicable

1.3.1.J Other training data content:

Not applicable

1.3.2 Latest date of data acquisition/collection for model training:

Not specified in the provided content

1.3.3 Is data collection ongoing to update the model with new data collection after deployment?

No

1.3.4 Date the training dataset was first used to train the model:

Not specified in the provided content

1.3.5 Rationale or purpose of data selection:

Datasets were selected to maximize high-quality technical reasoning and problem-solving capabilities within the .NET ecosystem. The mixture emphasizes carefully curated, algorithmically generated synthetic question-answer pairs derived from official documentation to improve technical understanding, documentation synthesis, and reasoning across APIs while maintaining factual accuracy.

2. List of data sources

2.1 Publicly available datasets

2.1.1 Have you used publicly available datasets to train the model?

Yes

Source: Official Microsoft .NET documentation repository (dotnet/docs)

2.2 Private non-publicly available datasets obtained from third parties

2.2.1 Datasets commercially licensed by rights holders or their representatives

Not applicable - dataset is derived from public GitHub repository

2.2.2 Private datasets obtained from other third-parties

Not applicable - dataset is derived from public GitHub repository

2.3 Personal Information

2.3.1 Was personal data used to train the model?

No personal data was used for training this model.

2.4 Synthetic data

2.4.1 Was any synthetic AI-generated data used to train the model?

Yes - algorithmically generated question-answer pairs based on technical documentation content

3. Data processing aspects

3.1 Respect of reservation of rights from text and data mining exception or limitation

The dataset is derived from the public dotnet/docs GitHub repository which has appropriate licensing for reuse.

3.2 Other information

3.2.1 Does the dataset include information about consumer groups without revealing individual consumer identities?

No personal or consumer identity information is included in the dataset.

3.2.2 Was the dataset cleaned or modified before model training?

Yes - the dataset was cleaned and processed through:

  1. Extraction from markdown files
  2. Removal of metadata, HTML, and outdated versions
  3. Segmentation into atomic topics
  4. Algorithmic question generation
  5. RAG-based answer generation using full documentation context
  6. Cross-encoder ranking and manual curation for quality assurance
  7. Consolidation into clean, versioned JSONL format