初始化项目,由ModelHub XC社区提供模型
Model: reaperdoesntknow/TopologicalQwen Source: Original Platform
This commit is contained in:
36
.gitattributes
vendored
Normal file
36
.gitattributes
vendored
Normal file
@@ -0,0 +1,36 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
251
README.md
Normal file
251
README.md
Normal file
@@ -0,0 +1,251 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
tags:
|
||||
- qwen3
|
||||
- sft
|
||||
- trl
|
||||
- topological-knowledge-distillation
|
||||
- disc
|
||||
- convergent-intelligence
|
||||
- convergentintel
|
||||
- edge
|
||||
- distillation
|
||||
- knowledge-distillation
|
||||
base_model:
|
||||
- reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT
|
||||
---
|
||||
|
||||
# TopologicalQwen
|
||||
|
||||
**Topology-Aware Knowledge Distillation from Qwen3-30B-A3B → 1.7B**
|
||||
|
||||
*Convergent Intelligence LLC: Research Division*
|
||||
|
||||
---
|
||||
|
||||
## What This Is
|
||||
|
||||
TopologicalQwen is a 1.7B parameter model distilled from Qwen3-30B-A3B using **Topological Knowledge Distillation (TKD)** — a methodology that treats the teacher's output distribution over a concatenated token stream as a bounded variation (BV) function and decomposes knowledge transfer into three channels via the Mesh Fundamental Identity:
|
||||
|
||||
1. **Smooth distillation (AC component)** — Standard KL divergence over regions where the teacher's distribution varies continuously. This is what every other KD method does and stops at.
|
||||
2. **Jump corrections (D^j f)** — Explicit correction terms at conceptual boundaries where the teacher's distribution exhibits discontinuities. These are the points where topic, register, or reasoning mode shifts — standard KD smears across them, losing structural information.
|
||||
3. **Drift corrections (D^c f)** — The Cantor/singular-continuous component capturing gradual distributional drift that neither the smooth nor jump terms account for. This is the residual structure that emerges in generation quality.
|
||||
|
||||
Standard knowledge distillation only handles term (1). TKD captures all three.
|
||||
|
||||
## Architecture
|
||||
|
||||
| Parameter | Value |
|
||||
|-----------|-------|
|
||||
| Architecture | Qwen3ForCausalLM |
|
||||
| Parameters | ~2.03B (1.7B effective) |
|
||||
| Hidden Size | 2048 |
|
||||
| Layers | 28 |
|
||||
| Attention Heads | 16 (Q) / 8 (KV) — GQA |
|
||||
| Intermediate | 6144 |
|
||||
| Context Length | 40,960 tokens |
|
||||
| Vocabulary | 151,936 |
|
||||
| Precision | FP32 training, BF16/FP16 inference |
|
||||
|
||||
## Training
|
||||
|
||||
**Student:** [Disctil-Qwen3-1.7B](https://huggingface.co/reaperdoesntknow/Disctil-Qwen3-1.7B) (DISC-refined uncensored Qwen3)
|
||||
**Teacher:** Qwen3-30B-A3B-Thinking-2507
|
||||
|
||||
**Datasets (physics CoT, ~1,599 samples):**
|
||||
- CoT Differential Equations (636 examples)
|
||||
- CoT Theoretical Mechanics (307 examples)
|
||||
- CoT Electromagnetism (580 examples)
|
||||
- CoT General Relativity (76 examples)
|
||||
|
||||
**DualMind format** — each training sample is restructured into `<explore>` (derivation), `<examine>` (verification/self-critique), and `<response>` (clean answer) blocks. The model learns a cognitive loop: generate reasoning, then critique it, then synthesize.
|
||||
|
||||
### TKD Pipeline (4 phases)
|
||||
|
||||
**Phase 1 — Teacher logit caching:** Single forward pass through the 30B teacher with top-64 logit compression to disk. One pass, no repeated teacher inference.
|
||||
|
||||
**Phase 2 — DISC topology pass:** Vectorized discrepancy operator maps the knowledge manifold. Jump detection at 3σ threshold with 1.25× amplification. Gap energy density computed over 64-token windows.
|
||||
|
||||
**Phase 3 — Topology-guided adaptive windowing:** 512-token windows cut at low-discrepancy positions (overlap 32–128) rather than fixed stride. The topology tells you where to cut without losing information across boundaries.
|
||||
|
||||
**Phase 4 — Curriculum-ordered continuous KD:** 4-phase curriculum (easiest 30% first). Proof-weighted loss: 2.25× → 1.1× decaying weights on reasoning tokens. KD alpha ramps from 0 → 0.45 (starting at 15% of training, reaching target at 45%). KL divergence at T=2.0. Effective batch size 32 (2 × 16 grad accumulation). Cosine LR: 5e-6 → 5e-7.
|
||||
|
||||
### Hyperparameters
|
||||
|
||||
| Parameter | Value |
|
||||
|-----------|-------|
|
||||
| Effective batch size | 32 (2 × 16 accum) |
|
||||
| Learning rate | 5e-6 → 5e-7 (cosine) |
|
||||
| Warmup steps | 30 |
|
||||
| Weight decay | 1e-3 |
|
||||
| Gradient clip | 1.0 |
|
||||
| Temperature | 2.0 |
|
||||
| KD target α | 0.45 |
|
||||
| Proof weight | 2.25 → 1.1 |
|
||||
| Jump threshold | 3σ |
|
||||
| Jump amplifier | 1.25× |
|
||||
| Precision | BF16 (autocast) |
|
||||
|
||||
Full methodology: [Structure Over Scale (DOI: 10.57967/hf/8165)](https://doi.org/10.57967/hf/8165)
|
||||
|
||||
## Usage
|
||||
|
||||
The model responds in DualMind format: `<explore>` → `<examine>` → `<response>`.
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
"reaperdoesntknow/TopologicalQwen",
|
||||
torch_dtype="auto",
|
||||
device_map="auto"
|
||||
)
|
||||
tokenizer = AutoTokenizer.from_pretrained("reaperdoesntknow/TopologicalQwen")
|
||||
|
||||
# Prompt with DualMind format — start the explore block
|
||||
prompt = (
|
||||
"##USER:\n"
|
||||
"Prove that every convergent sequence is a Cauchy sequence.\n\n"
|
||||
"<explore>\n"
|
||||
)
|
||||
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
||||
output = model.generate(
|
||||
**inputs, max_new_tokens=2048, do_sample=True,
|
||||
top_p=0.9, temperature=0.6, repetition_penalty=1.15
|
||||
)
|
||||
result = tokenizer.decode(output[0], skip_special_tokens=True)
|
||||
print(result)
|
||||
|
||||
# Verify mode transitions
|
||||
assert "<explore>" in result and "</explore>" in result # derivation
|
||||
assert "<examine>" in result and "</examine>" in result # self-critique
|
||||
assert "<response>" in result and "</response>" in result # clean answer
|
||||
```
|
||||
|
||||
### What the Output Looks Like
|
||||
|
||||
```
|
||||
<explore>
|
||||
[Unconstrained derivation — the model works through the proof freely]
|
||||
</explore>
|
||||
|
||||
<examine>
|
||||
[Adversarial self-response — the model critiques its own derivation]
|
||||
</examine>
|
||||
|
||||
<response>
|
||||
[Clean final answer synthesized from the internal dialogue]
|
||||
</response>
|
||||
```
|
||||
|
||||
This is the multi-model collision array collapsed into a single architecture. The dialectical structure that produces novel insights from architectural diversity is recreated through role-conditioned generation on shared weights.
|
||||
|
||||
## Distillation Chain
|
||||
|
||||
```
|
||||
Qwen3-1.7B (base)
|
||||
→ DiStil-Qwen3-1.7B-uncensored (uncensored SFT)
|
||||
→ Disctil-Qwen3-1.7B (DISC refinement)
|
||||
→ TopologicalQwen (TKD from 30B-Thinking teacher + DualMind format) ← you are here
|
||||
```
|
||||
|
||||
## What Makes This Different
|
||||
|
||||
The broader Convergent Intelligence portfolio ([49 models, 22,500+ downloads](https://huggingface.co/reaperdoesntknow)) was trained on CPU at FP32 for a total compute cost of $24. That proves the methodology — structure beats scale.
|
||||
|
||||
**This model is the exception.** TopologicalQwen was trained on Colab H100 at BF16 precision with a 30B-parameter teacher. Same TKD methodology, premium compute. This is the DistilQwen collection's answer to "what happens when you give this pipeline real hardware?"
|
||||
|
||||
The result: a 1.7B model that exhibits dual-mental-modality reasoning (explore → examine → respond) with structural quality that standard distillation at any precision doesn't produce. The methodology is the constant. The hardware is the variable. Both produce results that shouldn't exist at this parameter count.
|
||||
|
||||
Every knowledge distillation method in the literature treats the teacher's output as a smooth function and minimizes KL divergence globally. This works for the easy parts — regions where the teacher's distribution varies slowly. But language has structure: topic shifts, reasoning mode transitions, register changes. At these boundaries, the teacher's distribution jumps. Standard KD averages across these jumps, teaching the student a blurred version of the teacher's actual knowledge.
|
||||
|
||||
TKD uses the DISC (Discrepancy Calculus) framework to detect these structural features before training, then allocates capacity and loss weight accordingly. The result is a student that preserves the teacher's structural understanding, not just its surface statistics.
|
||||
|
||||
The empirical evidence: this model at 1.7B consistently produces responses with structural reasoning quality that standard distillation at the same parameter count does not achieve.
|
||||
|
||||
## Mathematical Foundations: Discrepancy Calculus (DISC)
|
||||
|
||||
TKD is grounded in Discrepancy Calculus — a measure-theoretic framework that treats singularities as primary structure rather than pathology. The full theory is developed in *"On the Formal Analysis of Discrepancy Calculus"* (Colca, 2026; Convergent Intelligence LLC: Research Division).
|
||||
|
||||
**The Core Operator.** The discrepancy operator quantifies local mismatch between integration and differentiation:
|
||||
|
||||
$$Df(x) = \lim_{\varepsilon \downarrow 0} \frac{1}{\varepsilon} \int_x^{x+\varepsilon} \frac{|f(t) - f(x)|}{|t - x|}\, dt$$
|
||||
|
||||
For smooth $f$: $Df(x) = |f'(x)|$ (classical recovery). For rough $f$: $D$ localizes irregularity to null sets while preserving integral structure.
|
||||
|
||||
**The Mesh Fundamental Identity.** Every function of bounded variation decomposes as:
|
||||
|
||||
$$f(b) - f(a) = \underbrace{\int_a^b f'(x)\,dx}_{\text{smooth (AC)}} + \underbrace{\sum_{x \in J_f} \Delta f(x)}_{\text{jumps}} + \underbrace{D^c f(I)}_{\text{Cantor drift}}$$
|
||||
|
||||
This is the theoretical backbone of TKD. Standard knowledge distillation captures only the first term. TKD preserves all three.
|
||||
|
||||
**TKD Application.** The teacher's output distribution $p_T(x)$ over a concatenated token stream is treated as a BV function. The DISC topology pass computes:
|
||||
|
||||
1. **Discrepancy energy** $E_{\text{disc}}[p_T] = \frac{1}{2}\int w(x)(Dp_T(x))^2 dx$ — identifies regions of high structural information density
|
||||
2. **Jump set** $J_{p_T} = \{x : Dp_T(x) > 3\sigma\}$ — locates conceptual boundaries (topic shifts, reasoning transitions)
|
||||
3. **Gap energy density** over 64-token windows — measures Cantor-type drift invisible to both smooth and jump analysis
|
||||
|
||||
Windows are cut at low-discrepancy positions rather than fixed stride. Loss weight is amplified at jump positions (1.25×). The topology tells you where the knowledge has architecture.
|
||||
|
||||
**Why This Matters (Meta-Discrepancy Theorem).** Theorem 11.15 of the DISC monograph proves: when the gap measure $\mu_{\text{gap}} > 0$ and discrepancy energy $E_{\text{disc}} > 0$, the classical FTC/MVT/chain-rule package is *impossible* on positive measure. Standard KD — which implicitly assumes smooth teacher distributions — provably cannot capture the structural information that TKD preserves. This is not a heuristic argument. It is a mathematical impossibility result.
|
||||
|
||||
## Related Models
|
||||
|
||||
| Model | Description | Downloads |
|
||||
|-------|-------------|-----------|
|
||||
| [Qwen3-1.7B-Thinking-Distil](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil) | TKD with Thinking teacher | 1,188 |
|
||||
| [Qwen3-1.7B-Coder-Distilled-SFT](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT) | TKD with Coder teacher | 966 |
|
||||
| [DiStil-Qwen3-1.7B-uncensored](https://huggingface.co/reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored) | Uncensored base for DISC chain | 1,030 |
|
||||
| [DualMind](https://huggingface.co/reaperdoesntknow/DualMind) | Dual cognition on shared weights | 260 |
|
||||
| [Dualmind-Qwen-1.7B-Thinking](https://huggingface.co/reaperdoesntknow/Dualmind-Qwen-1.7B-Thinking) | Opus 4.6 reasoning traces → 1.7B | New |
|
||||
|
||||
**[DistilQwen Collection](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c)** — Full proof-weighted distillation series (9 models)
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@misc{colca2026topologicalqwen,
|
||||
title={TopologicalQwen: Topology-Aware Knowledge Distillation via Bounded Variation Decomposition},
|
||||
author={Colca, Roy S.},
|
||||
year={2026},
|
||||
publisher={HuggingFace},
|
||||
url={https://huggingface.co/reaperdoesntknow/TopologicalQwen},
|
||||
note={Convergent Intelligence LLC: Research Division}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
|
||||
<!-- CIX-CROSSLINK-START -->
|
||||
|
||||
---
|
||||
|
||||
## From the Convergent Intelligence Portfolio
|
||||
|
||||
**[DistilQwen Collection](https://huggingface.co/collections/reaperdoesntknow/distilqwen-69bf40ec669117e3f069ef1c)** — Our only BF16 series. Proof-weighted distillation from Qwen3-30B-A3B → 1.7B and 0.6B on H100. Three teacher variants (Instruct, Thinking, Coder), nine models. The rest of the portfolio proves structure beats scale on CPU. This collection shows what happens when you give the methodology real hardware.
|
||||
|
||||
Top model: [Qwen3-1.7B-Thinking-Distil](https://huggingface.co/reaperdoesntknow/Qwen3-1.7B-Thinking-Distil) — 1,188 downloads
|
||||
|
||||
**[DualMind Collection](https://huggingface.co/collections/reaperdoesntknow/dualmind-67e6e07f4de0f45b0dca0dc4)** — Dual cognition architecture. Single model, two internal voices, three cognitive phases. Five models including [Dualmind-Qwen-1.7B-Thinking](https://huggingface.co/reaperdoesntknow/Dualmind-Qwen-1.7B-Thinking) (Opus 4.6 reasoning variant).
|
||||
|
||||
Full methodology: [Structure Over Scale (DOI: 10.57967/hf/8165)](https://doi.org/10.57967/hf/8165) | [From Three Teachers to Dual Cognition (DOI: 10.57967/hf/8184)](https://doi.org/10.57967/hf/8184)
|
||||
|
||||
*Convergent Intelligence LLC: Research Division*
|
||||
|
||||
<!-- CIX-CROSSLINK-END -->
|
||||
|
||||
*Convergent Intelligence LLC: Research Division*
|
||||
*"Where classical analysis fails to see, we begin."*
|
||||
|
||||
---
|
||||
<sub>Part of the [reaperdoesntknow research portfolio](https://huggingface.co/reaperdoesntknow) — 49 models, 22,598 total downloads | Last refreshed: 2026-03-30 12:02 UTC</sub>
|
||||
<!-- cix-keeper-ts:2026-07-10T13:17:15Z -->
|
||||
<!-- card-refresh: 2026-03-30 -->
|
||||
|
||||
---
|
||||
|
||||
*Last updated: 2026-03-31 by Convergent Intelligence LLC: Research Division*
|
||||
89
chat_template.jinja
Normal file
89
chat_template.jinja
Normal file
@@ -0,0 +1,89 @@
|
||||
{%- if tools %}
|
||||
{{- '<|im_start|>system\n' }}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- messages[0].content + '\n\n' }}
|
||||
{%- endif %}
|
||||
{{- "# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
|
||||
{%- for tool in tools %}
|
||||
{{- "\n" }}
|
||||
{{- tool | tojson }}
|
||||
{%- endfor %}
|
||||
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
|
||||
{%- else %}
|
||||
{%- if messages[0].role == 'system' %}
|
||||
{{- '<|im_start|>system\n' + messages[0].content + '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
||||
{%- for message in messages[::-1] %}
|
||||
{%- set index = (messages|length - 1) - loop.index0 %}
|
||||
{%- if ns.multi_step_tool and message.role == "user" and message.content is string and not(message.content.startswith('<tool_response>') and message.content.endswith('</tool_response>')) %}
|
||||
{%- set ns.multi_step_tool = false %}
|
||||
{%- set ns.last_query_index = index %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- for message in messages %}
|
||||
{%- if message.content is string %}
|
||||
{%- set content = message.content %}
|
||||
{%- else %}
|
||||
{%- set content = '' %}
|
||||
{%- endif %}
|
||||
{%- if (message.role == "user") or (message.role == "system" and not loop.first) %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
||||
{%- elif message.role == "assistant" %}
|
||||
{%- set reasoning_content = '' %}
|
||||
{%- if message.reasoning_content is string %}
|
||||
{%- set reasoning_content = message.reasoning_content %}
|
||||
{%- else %}
|
||||
{%- if '</think>' in content %}
|
||||
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
||||
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- if loop.index0 > ns.last_query_index %}
|
||||
{%- if loop.last or (not loop.last and reasoning_content) %}
|
||||
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content.strip('\n') + '\n</think>\n\n' + content.lstrip('\n') }}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- else %}
|
||||
{{- '<|im_start|>' + message.role + '\n' + content }}
|
||||
{%- endif %}
|
||||
{%- if message.tool_calls %}
|
||||
{%- for tool_call in message.tool_calls %}
|
||||
{%- if (loop.first and content) or (not loop.first) %}
|
||||
{{- '\n' }}
|
||||
{%- endif %}
|
||||
{%- if tool_call.function %}
|
||||
{%- set tool_call = tool_call.function %}
|
||||
{%- endif %}
|
||||
{{- '<tool_call>\n{"name": "' }}
|
||||
{{- tool_call.name }}
|
||||
{{- '", "arguments": ' }}
|
||||
{%- if tool_call.arguments is string %}
|
||||
{{- tool_call.arguments }}
|
||||
{%- else %}
|
||||
{{- tool_call.arguments | tojson }}
|
||||
{%- endif %}
|
||||
{{- '}\n</tool_call>' }}
|
||||
{%- endfor %}
|
||||
{%- endif %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- elif message.role == "tool" %}
|
||||
{%- if loop.first or (messages[loop.index0 - 1].role != "tool") %}
|
||||
{{- '<|im_start|>user' }}
|
||||
{%- endif %}
|
||||
{{- '\n<tool_response>\n' }}
|
||||
{{- content }}
|
||||
{{- '\n</tool_response>' }}
|
||||
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
|
||||
{{- '<|im_end|>\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
{%- endfor %}
|
||||
{%- if add_generation_prompt %}
|
||||
{{- '<|im_start|>assistant\n' }}
|
||||
{%- if enable_thinking is defined and enable_thinking is false %}
|
||||
{{- '<think>\n\n</think>\n\n' }}
|
||||
{%- endif %}
|
||||
{%- endif %}
|
||||
63
config.json
Normal file
63
config.json
Normal file
@@ -0,0 +1,63 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Qwen3ForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": null,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 151645,
|
||||
"head_dim": 128,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 2048,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 6144,
|
||||
"layer_types": [
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention"
|
||||
],
|
||||
"max_position_embeddings": 40960,
|
||||
"max_window_layers": 28,
|
||||
"model_type": "qwen3",
|
||||
"num_attention_heads": 16,
|
||||
"num_hidden_layers": 28,
|
||||
"num_key_value_heads": 8,
|
||||
"pad_token_id": 151643,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 1000000,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"sliding_window": null,
|
||||
"tie_word_embeddings": false,
|
||||
"transformers_version": "5.0.0",
|
||||
"use_cache": false,
|
||||
"use_sliding_window": false,
|
||||
"vocab_size": 151936
|
||||
}
|
||||
3
events.out.tfevents.1774658438.f65acf799204.8471.0
Normal file
3
events.out.tfevents.1774658438.f65acf799204.8471.0
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:bab2ee681d40fe790f91a278a14814f1ceed24b2f3393efc9eae0ea02c3ef042
|
||||
size 202371
|
||||
12
generation_config.json
Normal file
12
generation_config.json
Normal file
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"do_sample": true,
|
||||
"eos_token_id": [
|
||||
151645,
|
||||
151643
|
||||
],
|
||||
"pad_token_id": 151643,
|
||||
"temperature": 0.6,
|
||||
"top_k": 20,
|
||||
"top_p": 0.95,
|
||||
"transformers_version": "5.0.0"
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:6ce2945722a7d4cb6cbeb6a16ae9755a0612a01a8b098a61da68d5ca6bd1739b
|
||||
size 4063515640
|
||||
3
tokenizer.json
Normal file
3
tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:be75606093db2094d7cd20f3c2f385c212750648bd6ea4fb2bf507a6a4c55506
|
||||
size 11422650
|
||||
29
tokenizer_config.json
Normal file
29
tokenizer_config.json
Normal file
@@ -0,0 +1,29 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": null,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"extra_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>",
|
||||
"<|object_ref_start|>",
|
||||
"<|object_ref_end|>",
|
||||
"<|box_start|>",
|
||||
"<|box_end|>",
|
||||
"<|quad_start|>",
|
||||
"<|quad_end|>",
|
||||
"<|vision_start|>",
|
||||
"<|vision_end|>",
|
||||
"<|vision_pad|>",
|
||||
"<|image_pad|>",
|
||||
"<|video_pad|>"
|
||||
],
|
||||
"is_local": true,
|
||||
"model_max_length": 131072,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"split_special_tokens": false,
|
||||
"tokenizer_class": "Qwen2Tokenizer",
|
||||
"unk_token": null
|
||||
}
|
||||
5198
trainer_state (3).json
Normal file
5198
trainer_state (3).json
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user