初始化项目,由ModelHub XC社区提供模型
Model: Zynerji/Ektome-Qwen1.5-0.5B-Chat-PristinelyUncensored Source: Original Platform
This commit is contained in:
44
.gitattributes
vendored
Normal file
44
.gitattributes
vendored
Normal file
@@ -0,0 +1,44 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
||||
hero.png filter=lfs diff=lfs merge=lfs -text
|
||||
Ektome-Qwen1.5-0.5B-Chat-IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ektome-Qwen1.5-0.5B-Chat-IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ektome-Qwen1.5-0.5B-Chat-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ektome-Qwen1.5-0.5B-Chat-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ektome-Qwen1.5-0.5B-Chat-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
Ektome-Qwen1.5-0.5B-Chat-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
imatrix.dat filter=lfs diff=lfs merge=lfs -text
|
||||
3
Ektome-Qwen1.5-0.5B-Chat-IQ3_M.gguf
Normal file
3
Ektome-Qwen1.5-0.5B-Chat-IQ3_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:56247fc45687d7d4a3842a95876bd6aee9426b05a606904c0a7934eb9ae94639
|
||||
size 274365824
|
||||
3
Ektome-Qwen1.5-0.5B-Chat-IQ4_XS.gguf
Normal file
3
Ektome-Qwen1.5-0.5B-Chat-IQ4_XS.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:db2339699f94c6d7f422feed8ad30ca0fbc7ccf1bb7741860c7bb486312c1680
|
||||
size 297842048
|
||||
3
Ektome-Qwen1.5-0.5B-Chat-Q4_K_M.gguf
Normal file
3
Ektome-Qwen1.5-0.5B-Chat-Q4_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:9b0b710aa388f5d12ab5b417d152a672af4b382c2e7e1994496b986383faeeec
|
||||
size 319640960
|
||||
3
Ektome-Qwen1.5-0.5B-Chat-Q5_K_M.gguf
Normal file
3
Ektome-Qwen1.5-0.5B-Chat-Q5_K_M.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:dba071b3e41ca6eb566b06e89b7d87af7058aeb0f17983f9032ebac098961810
|
||||
size 352277888
|
||||
3
Ektome-Qwen1.5-0.5B-Chat-Q6_K.gguf
Normal file
3
Ektome-Qwen1.5-0.5B-Chat-Q6_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d92f3cf1afe400454baf5615b967558132f9032c7e8512f1486c8c49bf4dbbe5
|
||||
size 386954624
|
||||
3
Ektome-Qwen1.5-0.5B-Chat-Q8_0.gguf
Normal file
3
Ektome-Qwen1.5-0.5B-Chat-Q8_0.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:26ec15891cf07333c8f591a7ec352aa7b98f16c3a631948bd7936d625dd84685
|
||||
size 499296640
|
||||
121
README.md
Normal file
121
README.md
Normal file
@@ -0,0 +1,121 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen1.5-0.5B-Chat
|
||||
tags:
|
||||
- uncensored
|
||||
- abliterated
|
||||
- uncertified
|
||||
- ektome
|
||||
- sphragis
|
||||
- qwen1.5
|
||||
language:
|
||||
- en
|
||||
pipeline_tag: text-generation
|
||||
---
|
||||
|
||||

|
||||
|
||||
# Ektome-Qwen1.5-0.5B-Chat-PristinelyUncensored
|
||||
|
||||
**Uncensored. No n=2800 certificate has been run for this model, so no capability-retention claim is made.**
|
||||
|
||||
> **compliance 0.06 to 1.00 at capability -0.007 vs pristine.**
|
||||
|
||||
$$\colorbox{black}{$\color{white}
|
||||
\begin{array}{ll}
|
||||
\textsf{EKTOME CERTIFICATE} & {} \\
|
||||
\textsf{capability} & \textsf{NOT} \\
|
||||
\textsf{margin} & 3\% \\
|
||||
\textsf{items } n & 200 \\
|
||||
\textsf{worst-axis bound} & -0.008 \\
|
||||
\textsf{compliance} & 0.06 \rightarrow 1.00 \\
|
||||
\end{array}$}$$
|
||||
|
||||
> ### ⚠️ Not certified
|
||||
>
|
||||
> No n=2800 paired certificate exists for this model. Any numbers below are
|
||||
> point estimates with no confidence interval.
|
||||
|
||||
|
||||
📄 **[Read the whitepaper (PDF)](./whitepaper.pdf)** — full method, receipts and certification.
|
||||
The PDF is the authoritative document: dark-typeset, with the complete derivation, the
|
||||
per-axis certificate and the reproducibility hashes.
|
||||
|
||||
---
|
||||
|
||||
## Why this exists
|
||||
|
||||
Standard abliteration removes a coarse *refusal direction* that is entangled with
|
||||
directions carrying knowledge and reasoning. The result is an uncensored model with a
|
||||
capability tax that is **almost never measured**.
|
||||
|
||||
Ektomē (ἐκτομή, *excision*) isolates and removes only the refusal-**specific**
|
||||
component, leaving general helpfulness intact, and does so norm-preservingly on the
|
||||
pristine model — no training, no distillation, no damage to repair. The extraction
|
||||
depth is selected per model by automated search against measured compliance.
|
||||
|
||||
The estimator, excision operator and depth-selection procedure are proprietary.
|
||||
What is published here is the **measured outcome** and the evidence for it, which you
|
||||
can verify against the artifacts in this repo.
|
||||
|
||||
## The receipt
|
||||
|
||||
| model | capability (MMLU-val) ↑ | compliance on harmful ↑ |
|
||||
|---|---|---|
|
||||
| pristine `Qwen1.5-0.5B-Chat` | 0.328 | 0.060 |
|
||||
| **Ektomē (this model)** | **0.335** | **1.000** |
|
||||
|
||||
These are **point estimates with no confidence interval** — which is precisely why the next section exists.
|
||||
|
||||
|
||||
## The certificate
|
||||
|
||||
Capability retention is certified by a paired non-inferiority test against the pristine
|
||||
model (exact McNemar, Holm-corrected, one-sided bootstrap bound on the drop $d$ vs a
|
||||
3% margin):
|
||||
|
||||
| axis | n | ref | cand | d upper | verdict |
|
||||
|---|---|---|---|---|---|
|
||||
| MMLU-val (POINT ESTIMATE, n=200, no CI) | 200 | 0.328 | 0.335 | -0.008 | UNCERTIFIED |
|
||||
|
||||
|
||||
**Overall: NOT CERTIFIED - no n=2800 paired test has been run for this model**
|
||||
|
||||
Reproducible from `seed=20260726`, pack `sha256:7bbaff877146e081…`.
|
||||
|
||||
### Generation health checks
|
||||
|
||||
| metric | pristine | Ektomē | n |
|
||||
|---|---|---|---|
|
||||
| `foreign_rate` | 0.0 | 0.0 | 15 |
|
||||
| `degen_rate` | 0.0 | 0.1 | 15 |
|
||||
| `instr_pass` | 1.0 | 0.8 | 5 |
|
||||
|
||||
These are **degeneration guards** — code-switching, babbling, format compliance —
|
||||
not capability measures. Note the sample sizes: they detect a broken model, not a
|
||||
subtly weaker one. The capability claim rests on the certificate above, not here.
|
||||
|
||||
|
||||
## Quantisations
|
||||
|
||||
_No quantisations have been published for this model yet — bf16 weights only._
|
||||
|
||||
|
||||
## Limitations
|
||||
|
||||
The certificate bounds **capability retention only**. It does not certify safety, factual
|
||||
accuracy, or fitness for any purpose. Axes marked *inconclusive* are honestly
|
||||
under-powered, and the certificate states the $n$ needed to resolve them. Compliance uses
|
||||
a keyword classifier — a proxy that evasive phrasing can fool. **This model is uncensored
|
||||
by construction: it will not refuse, and you are accountable for what you do with it.**
|
||||
|
||||
## Citation
|
||||
|
||||
```bibtex
|
||||
@software{ektome_Ektome-Qwen1.5-0.5B-Chat-PristinelyUncensored,
|
||||
title = {Ektome-Qwen1.5-0.5B-Chat-PristinelyUncensored},
|
||||
author = {Zynerji},
|
||||
year = {2026},
|
||||
url = {https://huggingface.co/Zynerji/Ektome-Qwen1.5-0.5B-Chat-PristinelyUncensored}
|
||||
}
|
||||
```
|
||||
133
cert_Qwen1.5-0.5B-Chat.json
Normal file
133
cert_Qwen1.5-0.5B-Chat.json
Normal file
@@ -0,0 +1,133 @@
|
||||
{
|
||||
"sphragis_version": "0.1.0",
|
||||
"generated_at": "2026-07-27T18:34:00.914337+00:00",
|
||||
"reference": {
|
||||
"endpoint": "http://127.0.0.1:8080/v1",
|
||||
"model": "ref"
|
||||
},
|
||||
"candidate": {
|
||||
"endpoint": "http://127.0.0.1:8081/v1",
|
||||
"model": "cand"
|
||||
},
|
||||
"pack": {
|
||||
"name": "/root/pack_v2.jsonl",
|
||||
"n_tasks": 2800,
|
||||
"sha256": "2de27099bbb15bab4f7b35599b215f038fdb68b3f5f4e1cda1a464f9dd18e14e"
|
||||
},
|
||||
"overall": "FAIL",
|
||||
"claim": "At least one axis shows a statistically significant accuracy regression (exact McNemar, Holm-corrected, alpha=0.05).",
|
||||
"axes": [
|
||||
{
|
||||
"axis": "arithmetic",
|
||||
"n": 1400,
|
||||
"counts": {
|
||||
"both_correct": 393,
|
||||
"ref_only": 53,
|
||||
"cand_only": 28,
|
||||
"both_wrong": 926
|
||||
},
|
||||
"acc_reference": 0.318571,
|
||||
"acc_candidate": 0.300714,
|
||||
"regression_d": 0.017857,
|
||||
"d_ci": [
|
||||
0.005,
|
||||
0.030714
|
||||
],
|
||||
"d_upper_bound": 0.028571,
|
||||
"p_regression": 0.00363775,
|
||||
"p_regression_holm": 0.01091325,
|
||||
"p_improvement": 0.99820183,
|
||||
"improved": false,
|
||||
"mde_at_power": 0.015973,
|
||||
"n_needed_for_margin": 396,
|
||||
"verdict": "FAIL",
|
||||
"reason": "significant_regression"
|
||||
},
|
||||
{
|
||||
"axis": "instruction",
|
||||
"n": 600,
|
||||
"counts": {
|
||||
"both_correct": 45,
|
||||
"ref_only": 0,
|
||||
"cand_only": 0,
|
||||
"both_wrong": 555
|
||||
},
|
||||
"acc_reference": 0.075,
|
||||
"acc_candidate": 0.075,
|
||||
"regression_d": 0.0,
|
||||
"d_ci": [
|
||||
0.0,
|
||||
0.0
|
||||
],
|
||||
"d_upper_bound": 0.0,
|
||||
"p_regression": 1.0,
|
||||
"p_regression_holm": 1.0,
|
||||
"p_improvement": 1.0,
|
||||
"improved": false,
|
||||
"mde_at_power": null,
|
||||
"n_needed_for_margin": 204,
|
||||
"verdict": "PASS",
|
||||
"reason": "non_inferior_within_margin"
|
||||
},
|
||||
{
|
||||
"axis": "knowledge",
|
||||
"n": 400,
|
||||
"counts": {
|
||||
"both_correct": 67,
|
||||
"ref_only": 8,
|
||||
"cand_only": 1,
|
||||
"both_wrong": 324
|
||||
},
|
||||
"acc_reference": 0.1875,
|
||||
"acc_candidate": 0.17,
|
||||
"regression_d": 0.0175,
|
||||
"d_ci": [
|
||||
0.005,
|
||||
0.0325
|
||||
],
|
||||
"d_upper_bound": 0.03,
|
||||
"p_regression": 0.01953125,
|
||||
"p_regression_holm": 0.0390625,
|
||||
"p_improvement": 0.99804688,
|
||||
"improved": false,
|
||||
"mde_at_power": 0.0186,
|
||||
"n_needed_for_margin": 204,
|
||||
"verdict": "FAIL",
|
||||
"reason": "significant_regression"
|
||||
},
|
||||
{
|
||||
"axis": "reasoning",
|
||||
"n": 400,
|
||||
"counts": {
|
||||
"both_correct": 128,
|
||||
"ref_only": 35,
|
||||
"cand_only": 8,
|
||||
"both_wrong": 229
|
||||
},
|
||||
"acc_reference": 0.4075,
|
||||
"acc_candidate": 0.34,
|
||||
"regression_d": 0.0675,
|
||||
"d_ci": [
|
||||
0.0375,
|
||||
0.1
|
||||
],
|
||||
"d_upper_bound": 0.095,
|
||||
"p_regression": 2.097e-05,
|
||||
"p_regression_holm": 8.387e-05,
|
||||
"p_improvement": 0.99999552,
|
||||
"improved": false,
|
||||
"mde_at_power": 0.040656,
|
||||
"n_needed_for_margin": 737,
|
||||
"verdict": "FAIL",
|
||||
"reason": "significant_regression"
|
||||
}
|
||||
],
|
||||
"params": {
|
||||
"margin": 0.03,
|
||||
"alpha": 0.05,
|
||||
"n_floor": 30,
|
||||
"power": 0.8,
|
||||
"n_boot": 4000,
|
||||
"seed": 0
|
||||
}
|
||||
}
|
||||
6
chat_template.jinja
Normal file
6
chat_template.jinja
Normal file
@@ -0,0 +1,6 @@
|
||||
{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
' }}{% endif %}{{'<|im_start|>' + message['role'] + '
|
||||
' + message['content'] + '<|im_end|>' + '
|
||||
'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
|
||||
' }}{% endif %}
|
||||
57
config.json
Normal file
57
config.json
Normal file
@@ -0,0 +1,57 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Qwen2ForCausalLM"
|
||||
],
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 151643,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 151645,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 1024,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 2816,
|
||||
"layer_types": [
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention"
|
||||
],
|
||||
"max_position_embeddings": 32768,
|
||||
"max_window_layers": 21,
|
||||
"model_type": "qwen2",
|
||||
"num_attention_heads": 16,
|
||||
"num_hidden_layers": 24,
|
||||
"num_key_value_heads": 16,
|
||||
"pad_token_id": null,
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 1000000.0,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"sliding_window": null,
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "5.13.1",
|
||||
"use_cache": true,
|
||||
"use_sliding_window": false,
|
||||
"vocab_size": 151936
|
||||
}
|
||||
28
ektome_report.json
Normal file
28
ektome_report.json
Normal file
@@ -0,0 +1,28 @@
|
||||
{
|
||||
"status": "SHIP",
|
||||
"target": 0.99,
|
||||
"SE_mmlu": 0.023465067121148406,
|
||||
"base_compliance": 0.06,
|
||||
"compliance": 1.0,
|
||||
"base_cap": 0.3275,
|
||||
"cap": 0.335,
|
||||
"dcap": 0.0075,
|
||||
"gen_base": {
|
||||
"foreign_rate": 0.0,
|
||||
"degen_rate": 0.0,
|
||||
"instr_pass": 1.0
|
||||
},
|
||||
"gen": {
|
||||
"foreign_rate": 0.0,
|
||||
"degen_rate": 0.1,
|
||||
"instr_pass": 0.8
|
||||
},
|
||||
"gen_delta": {
|
||||
"d_foreign": 0.0,
|
||||
"d_degen": 0.1,
|
||||
"d_instr": -0.2,
|
||||
"holds_gen": true
|
||||
},
|
||||
"base_model": "Qwen/Qwen1.5-0.5B-Chat",
|
||||
"_note": "Redacted: search trajectory, selected depth and edit-scope removed. Reported values are the measured outcome only."
|
||||
}
|
||||
12
generation_config.json
Normal file
12
generation_config.json
Normal file
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"bos_token_id": 151643,
|
||||
"do_sample": true,
|
||||
"eos_token_id": [
|
||||
151645,
|
||||
151643
|
||||
],
|
||||
"pad_token_id": 151643,
|
||||
"repetition_penalty": 1.1,
|
||||
"top_p": 0.8,
|
||||
"transformers_version": "5.13.1"
|
||||
}
|
||||
3
hero.png
Normal file
3
hero.png
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:d02149d325d2b507b673eb0b3502ae6ba000837116e7514ce913c534a519ed8b
|
||||
size 242454
|
||||
3
imatrix.dat
Normal file
3
imatrix.dat
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:c1b119a7940f2b76e884c282b42ae5cc4ab335ca26ee318cce5894a8e31396dc
|
||||
size 886048
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:36a2afb3d9e0ae931e9e3dca0a6bab4b2efd4c9e5c2106b4ae3e08bf2a2681d5
|
||||
size 928008104
|
||||
6
nvfp4/chat_template.jinja
Normal file
6
nvfp4/chat_template.jinja
Normal file
@@ -0,0 +1,6 @@
|
||||
{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
|
||||
You are a helpful assistant.<|im_end|>
|
||||
' }}{% endif %}{{'<|im_start|>' + message['role'] + '
|
||||
' + message['content'] + '<|im_end|>' + '
|
||||
'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
|
||||
' }}{% endif %}
|
||||
107
nvfp4/config.json
Normal file
107
nvfp4/config.json
Normal file
@@ -0,0 +1,107 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Qwen2ForCausalLM"
|
||||
],
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 151643,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 151645,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 1024,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 2816,
|
||||
"layer_types": [
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention",
|
||||
"full_attention"
|
||||
],
|
||||
"max_position_embeddings": 32768,
|
||||
"max_window_layers": 21,
|
||||
"model_type": "qwen2",
|
||||
"num_attention_heads": 16,
|
||||
"num_hidden_layers": 24,
|
||||
"num_key_value_heads": 16,
|
||||
"pad_token_id": null,
|
||||
"quantization_config": {
|
||||
"config_groups": {
|
||||
"group_0": {
|
||||
"format": "nvfp4-pack-quantized",
|
||||
"input_activations": {
|
||||
"actorder": null,
|
||||
"block_structure": null,
|
||||
"dynamic": "local",
|
||||
"group_size": 16,
|
||||
"num_bits": 4,
|
||||
"observer": "static_minmax",
|
||||
"observer_kwargs": {},
|
||||
"scale_dtype": "torch.float8_e4m3fn",
|
||||
"strategy": "tensor_group",
|
||||
"symmetric": true,
|
||||
"type": "float",
|
||||
"zp_dtype": null
|
||||
},
|
||||
"output_activations": null,
|
||||
"targets": [
|
||||
"Linear"
|
||||
],
|
||||
"weights": {
|
||||
"actorder": null,
|
||||
"block_structure": null,
|
||||
"dynamic": false,
|
||||
"group_size": 16,
|
||||
"num_bits": 4,
|
||||
"observer": "memoryless_minmax",
|
||||
"observer_kwargs": {},
|
||||
"scale_dtype": "torch.float8_e4m3fn",
|
||||
"strategy": "tensor_group",
|
||||
"symmetric": true,
|
||||
"type": "float",
|
||||
"zp_dtype": null
|
||||
}
|
||||
}
|
||||
},
|
||||
"format": "nvfp4-pack-quantized",
|
||||
"global_compression_ratio": null,
|
||||
"ignore": [
|
||||
"lm_head"
|
||||
],
|
||||
"kv_cache_scheme": null,
|
||||
"quant_method": "compressed-tensors",
|
||||
"quantization_status": "compressed",
|
||||
"sparsity_config": {},
|
||||
"transform_config": {},
|
||||
"version": "0.17.1"
|
||||
},
|
||||
"rms_norm_eps": 1e-06,
|
||||
"rope_parameters": {
|
||||
"rope_theta": 1000000.0,
|
||||
"rope_type": "default"
|
||||
},
|
||||
"sliding_window": null,
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "5.10.1",
|
||||
"use_cache": true,
|
||||
"use_sliding_window": false,
|
||||
"vocab_size": 151936
|
||||
}
|
||||
12
nvfp4/generation_config.json
Normal file
12
nvfp4/generation_config.json
Normal file
@@ -0,0 +1,12 @@
|
||||
{
|
||||
"bos_token_id": 151643,
|
||||
"do_sample": true,
|
||||
"eos_token_id": [
|
||||
151645,
|
||||
151643
|
||||
],
|
||||
"pad_token_id": 151643,
|
||||
"repetition_penalty": 1.1,
|
||||
"top_p": 0.8,
|
||||
"transformers_version": "5.10.1"
|
||||
}
|
||||
3
nvfp4/model.safetensors
Normal file
3
nvfp4/model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:8635b0cd843af85b0b3ff5aeb55f073598ff285ec55aa2224214bea589324d0c
|
||||
size 484911640
|
||||
7
nvfp4/recipe.yaml
Normal file
7
nvfp4/recipe.yaml
Normal file
@@ -0,0 +1,7 @@
|
||||
default_stage:
|
||||
default_modifiers:
|
||||
QuantizationModifier:
|
||||
targets: [Linear]
|
||||
ignore: [lm_head]
|
||||
scheme: NVFP4
|
||||
bypass_divisibility_checks: false
|
||||
3
nvfp4/tokenizer.json
Normal file
3
nvfp4/tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:48f722bc04c884e2fe1525fdcd85a1293a8499b6e620c1ac7c083c49632305fb
|
||||
size 11418262
|
||||
19
nvfp4/tokenizer_config.json
Normal file
19
nvfp4/tokenizer_config.json
Normal file
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": null,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"extra_special_tokens": [
|
||||
"<|im_start|>",
|
||||
"<|im_end|>"
|
||||
],
|
||||
"is_local": false,
|
||||
"local_files_only": false,
|
||||
"model_max_length": 32768,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"split_special_tokens": false,
|
||||
"tokenizer_class": "Qwen2Tokenizer",
|
||||
"unk_token": null
|
||||
}
|
||||
3
tokenizer.json
Normal file
3
tokenizer.json
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:48f722bc04c884e2fe1525fdcd85a1293a8499b6e620c1ac7c083c49632305fb
|
||||
size 11418262
|
||||
16
tokenizer_config.json
Normal file
16
tokenizer_config.json
Normal file
@@ -0,0 +1,16 @@
|
||||
{
|
||||
"add_prefix_space": false,
|
||||
"backend": "tokenizers",
|
||||
"bos_token": null,
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|im_end|>",
|
||||
"errors": "replace",
|
||||
"is_local": false,
|
||||
"local_files_only": false,
|
||||
"model_max_length": 32768,
|
||||
"pad_token": "<|endoftext|>",
|
||||
"split_special_tokens": false,
|
||||
"tokenizer_class": "Qwen2Tokenizer",
|
||||
"unk_token": null,
|
||||
"chat_template": "{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n' }}{% endif %}{{'<|im_start|>' + message['role'] + '\n' + message['content'] + '<|im_end|>' + '\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\n' }}{% endif %}"
|
||||
}
|
||||
BIN
whitepaper.pdf
Normal file
BIN
whitepaper.pdf
Normal file
Binary file not shown.
Reference in New Issue
Block a user