commit 23af150f272ec071ed5044f4b0536b33dbea9fdd Author: ModelHub XC Date: Tue Jun 30 21:33:12 2026 +0800 初始化项目,由ModelHub XC社区提供模型 Model: Mungert/Mistral-7B-Instruct-v0.3-GGUF Source: Original Platform diff --git a/.gitattributes b/.gitattributes new file mode 100644 index 0000000..772f434 --- /dev/null +++ b/.gitattributes @@ -0,0 +1,82 @@ +*.7z filter=lfs diff=lfs merge=lfs -text +*.arrow filter=lfs diff=lfs merge=lfs -text +*.bin filter=lfs diff=lfs merge=lfs -text +*.bin.* filter=lfs diff=lfs merge=lfs -text +*.bz2 filter=lfs diff=lfs merge=lfs -text +*.ftz filter=lfs diff=lfs merge=lfs -text +*.gz filter=lfs diff=lfs merge=lfs -text +*.h5 filter=lfs diff=lfs merge=lfs -text +*.joblib filter=lfs diff=lfs merge=lfs -text +*.lfs.* filter=lfs diff=lfs merge=lfs -text +*.model filter=lfs diff=lfs merge=lfs -text +*.msgpack filter=lfs diff=lfs merge=lfs -text +*.onnx filter=lfs diff=lfs merge=lfs -text +*.ot filter=lfs diff=lfs merge=lfs -text +*.parquet filter=lfs diff=lfs merge=lfs -text +*.pb filter=lfs diff=lfs merge=lfs -text +*.pt filter=lfs diff=lfs merge=lfs -text +*.pth filter=lfs diff=lfs merge=lfs -text +*.rar filter=lfs diff=lfs merge=lfs -text +saved_model/**/* filter=lfs diff=lfs merge=lfs -text +*.tar.* filter=lfs diff=lfs merge=lfs -text +*.tflite filter=lfs diff=lfs merge=lfs -text +*.tgz filter=lfs diff=lfs merge=lfs -text +*.xz filter=lfs diff=lfs merge=lfs -text +*.zip filter=lfs diff=lfs merge=lfs -text +*.zstandard filter=lfs diff=lfs merge=lfs -text +*.tfevents* filter=lfs diff=lfs merge=lfs -text +*.db* filter=lfs diff=lfs merge=lfs -text +*.ark* filter=lfs diff=lfs merge=lfs -text +**/*ckpt*data* filter=lfs diff=lfs merge=lfs -text +**/*ckpt*.meta filter=lfs diff=lfs merge=lfs -text +**/*ckpt*.index filter=lfs diff=lfs merge=lfs -text +*.safetensors filter=lfs diff=lfs merge=lfs -text +*.ckpt filter=lfs diff=lfs merge=lfs -text + +*.ggml filter=lfs diff=lfs merge=lfs -text +*.llamafile* filter=lfs diff=lfs merge=lfs -text +*.pt2 filter=lfs diff=lfs merge=lfs -text +*.mlmodel filter=lfs diff=lfs merge=lfs -text +*.npy filter=lfs diff=lfs merge=lfs -text +*.npz filter=lfs diff=lfs merge=lfs -text +*.pickle filter=lfs diff=lfs merge=lfs -text +*.pkl filter=lfs diff=lfs merge=lfs -text +*.tar filter=lfs diff=lfs merge=lfs -text +*.wasm filter=lfs diff=lfs merge=lfs -text +*.zst filter=lfs diff=lfs merge=lfs -text +*tfevents* filter=lfs diff=lfs merge=lfs -text + +Mistral-7B-Instruct-v0.3-iq2_s.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq2_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-f16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq3_xxs.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-f16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq3_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-bf16-q4_k.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-bf16_q8_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-bf16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q2_k_s.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-bf16-q8_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq3_s.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq4_nl.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q8_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq3_xs.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-f16-q6_k.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq2_xxs.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q2_k_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q3_k_s.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-f16_q8_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3.imatrix filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q6_k_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q5_k_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-bf16.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q4_1.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q5_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q5_1.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q5_k_s.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq2_xs.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-iq4_xs.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q3_k_m.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q4_0.gguf filter=lfs diff=lfs merge=lfs -text +Mistral-7B-Instruct-v0.3-q4_k_s.gguf filter=lfs diff=lfs merge=lfs -text \ No newline at end of file diff --git a/Mistral-7B-Instruct-v0.3-bf16-q4_k.gguf b/Mistral-7B-Instruct-v0.3-bf16-q4_k.gguf new file mode 100644 index 0000000..e757ae2 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-bf16-q4_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a8cf3b26c9e25466cdd82d0c7f3958ff11877a0dcb0d5a6750d57fceb88acb23 +size 4724088960 diff --git a/Mistral-7B-Instruct-v0.3-bf16-q6_k.gguf b/Mistral-7B-Instruct-v0.3-bf16-q6_k.gguf new file mode 100644 index 0000000..a8d6083 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-bf16-q6_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e98aabf2c20d5bd0dce2b8ac1e3dbeb515ad924413cbcca4abd47f439bd01d7b +size 6263922976 diff --git a/Mistral-7B-Instruct-v0.3-bf16-q8_0.gguf b/Mistral-7B-Instruct-v0.3-bf16-q8_0.gguf new file mode 100644 index 0000000..b447184 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-bf16-q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:dfc48f6c9d8b37bf3c960d2280cb6c17261868555cf930f74d81a2bad3477621 +size 7954227200 diff --git a/Mistral-7B-Instruct-v0.3-bf16.gguf b/Mistral-7B-Instruct-v0.3-bf16.gguf new file mode 100644 index 0000000..051ef0e --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-bf16.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6011c09261aaed75a9aaff78aa356f2e6661f21cb41f292c9edfb144f828a98b +size 14497341472 diff --git a/Mistral-7B-Instruct-v0.3-bf16_q8_0.gguf b/Mistral-7B-Instruct-v0.3-bf16_q8_0.gguf new file mode 100644 index 0000000..fe46f2f --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-bf16_q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:57693642f811c5fe914183e4f5b04a8c1dd21e36f01f5a4f423f73353dd2afb4 +size 10344980512 diff --git a/Mistral-7B-Instruct-v0.3-f16-q4_k.gguf b/Mistral-7B-Instruct-v0.3-f16-q4_k.gguf new file mode 100644 index 0000000..ff02a0e --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-f16-q4_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2f62995e5ff2871cc1b4e3550b99bae1d95432c39276182fd59c30f27d33b048 +size 4724088960 diff --git a/Mistral-7B-Instruct-v0.3-f16-q6_k.gguf b/Mistral-7B-Instruct-v0.3-f16-q6_k.gguf new file mode 100644 index 0000000..b6dce40 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-f16-q6_k.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7001e2636462b92f52dbca161a32ff37c4361872bd3aadc0d12c53c78497e610 +size 6263922976 diff --git a/Mistral-7B-Instruct-v0.3-f16-q8_0.gguf b/Mistral-7B-Instruct-v0.3-f16-q8_0.gguf new file mode 100644 index 0000000..52275b6 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-f16-q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7e0c37f7dfaab87c1aa306ab34884e8a4d9c9b6c9cd531950882ff4e1992b39c +size 7954227200 diff --git a/Mistral-7B-Instruct-v0.3-f16_q8_0.gguf b/Mistral-7B-Instruct-v0.3-f16_q8_0.gguf new file mode 100644 index 0000000..2fce9c2 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-f16_q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0fa3c0ec8f739ad83581ad58df982d4ed5b65ab0e4a03e5c488ff5454dc802bc +size 10344980512 diff --git a/Mistral-7B-Instruct-v0.3-iq2_m.gguf b/Mistral-7B-Instruct-v0.3-iq2_m.gguf new file mode 100644 index 0000000..07ccdca --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq2_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9114880577815cf232067b797f4ad46a38092ed498b558077cc59a2d2d5c1f94 +size 2674843968 diff --git a/Mistral-7B-Instruct-v0.3-iq2_s.gguf b/Mistral-7B-Instruct-v0.3-iq2_s.gguf new file mode 100644 index 0000000..65f0f9e --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq2_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:adac374c3f2003addb04e00b1f52747fca49baae0650c9b8174649a898b812bf +size 2545606976 diff --git a/Mistral-7B-Instruct-v0.3-iq2_xs.gguf b/Mistral-7B-Instruct-v0.3-iq2_xs.gguf new file mode 100644 index 0000000..625e109 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq2_xs.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5f1d680659835c9bc7fcf30c77909cf03f817fe1badc2718d6a32f23997e9042 +size 2451431744 diff --git a/Mistral-7B-Instruct-v0.3-iq2_xxs.gguf b/Mistral-7B-Instruct-v0.3-iq2_xxs.gguf new file mode 100644 index 0000000..c6122f3 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq2_xxs.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4e1150ed9317a75992a8313911dc2cc8e71e814bbd412db0ed7b017517828280 +size 2241585472 diff --git a/Mistral-7B-Instruct-v0.3-iq3_m.gguf b/Mistral-7B-Instruct-v0.3-iq3_m.gguf new file mode 100644 index 0000000..52fb195 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq3_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0941b1663d6b6f14a4b8658e4330611e9c4a3ab39d19d13daef63ea4f0bfe29f +size 3479953728 diff --git a/Mistral-7B-Instruct-v0.3-iq3_s.gguf b/Mistral-7B-Instruct-v0.3-iq3_s.gguf new file mode 100644 index 0000000..a16f233 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq3_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:368a8c03a042c5a29ecd4f9fc102beddfee641c182d5e8b5967062403eee50e6 +size 3422282048 diff --git a/Mistral-7B-Instruct-v0.3-iq3_xs.gguf b/Mistral-7B-Instruct-v0.3-iq3_xs.gguf new file mode 100644 index 0000000..2661b57 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq3_xs.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a8d56f7e2a4a60936d8d3d21d91541377b86f899b61324b236eda7406ec8a0af +size 3087262016 diff --git a/Mistral-7B-Instruct-v0.3-iq3_xxs.gguf b/Mistral-7B-Instruct-v0.3-iq3_xxs.gguf new file mode 100644 index 0000000..06ccee7 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq3_xxs.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:762e17ff53da30ec6d20e1f5c8e5f647fdaac5b5dd270207384e850493f0449b +size 3001278784 diff --git a/Mistral-7B-Instruct-v0.3-iq4_nl.gguf b/Mistral-7B-Instruct-v0.3-iq4_nl.gguf new file mode 100644 index 0000000..5f1a552 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq4_nl.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a425cd36166ef113464d956e89d89c561588ffc59f08458527e7986e0d2f6fa4 +size 4130070848 diff --git a/Mistral-7B-Instruct-v0.3-iq4_xs.gguf b/Mistral-7B-Instruct-v0.3-iq4_xs.gguf new file mode 100644 index 0000000..00fd112 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-iq4_xs.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:5550c5c80007d28f563bf9680d58d5e68e333da0269f9bac824da786edb66a78 +size 3911967040 diff --git a/Mistral-7B-Instruct-v0.3-q2_k_m.gguf b/Mistral-7B-Instruct-v0.3-q2_k_m.gguf new file mode 100644 index 0000000..e219003 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q2_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b000fcdc4bf2a19c796af22ccc3ada88cb93d400542a6a8623744455bd8a59af +size 2788942144 diff --git a/Mistral-7B-Instruct-v0.3-q2_k_s.gguf b/Mistral-7B-Instruct-v0.3-q2_k_s.gguf new file mode 100644 index 0000000..ca78a42 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q2_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f66ee5f06a84a03d0df276b8b8e89c9dba2d293b3923f30cfe3de447820f27cb +size 2743591232 diff --git a/Mistral-7B-Instruct-v0.3-q3_k_m.gguf b/Mistral-7B-Instruct-v0.3-q3_k_m.gguf new file mode 100644 index 0000000..d74a0a4 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q3_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:95731c4037799b091b81d1b4a952169c52f76575ccb9bedb47e361fce01ae4eb +size 3575374144 diff --git a/Mistral-7B-Instruct-v0.3-q3_k_s.gguf b/Mistral-7B-Instruct-v0.3-q3_k_s.gguf new file mode 100644 index 0000000..91b4838 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q3_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ad11814ee1d99dc380e485212c418d2d98a8e340906729ac8d29005e16941699 +size 3511411008 diff --git a/Mistral-7B-Instruct-v0.3-q4_0.gguf b/Mistral-7B-Instruct-v0.3-q4_0.gguf new file mode 100644 index 0000000..34c2fcb --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q4_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b90c885b37a4336c8e79ac4352d6f36c977c7548860ffdcb48a94dbd6432f32f +size 4078690624 diff --git a/Mistral-7B-Instruct-v0.3-q4_1.gguf b/Mistral-7B-Instruct-v0.3-q4_1.gguf new file mode 100644 index 0000000..3d82aca --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q4_1.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d022488d55e27ff8ff57d66589511cd7af4a13419be108ab2deaf8bbb708b32d +size 4531675456 diff --git a/Mistral-7B-Instruct-v0.3-q4_k_m.gguf b/Mistral-7B-Instruct-v0.3-q4_k_m.gguf new file mode 100644 index 0000000..ff0efda --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q4_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a2c14920684d150d1adc93116e7c43ed9935b82920b7f9748c9ce0b76797e228 +size 4424196416 diff --git a/Mistral-7B-Instruct-v0.3-q4_k_s.gguf b/Mistral-7B-Instruct-v0.3-q4_k_s.gguf new file mode 100644 index 0000000..9dc0893 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q4_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f53ebc8a5c775fc9b3eb2866f64fc408282cfd591efc00671cb540c75a9921c3 +size 4227588416 diff --git a/Mistral-7B-Instruct-v0.3-q5_0.gguf b/Mistral-7B-Instruct-v0.3-q5_0.gguf new file mode 100644 index 0000000..6e49e0c --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q5_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:f99f8f71bb4c837df845c8944da38ccddf64382cd3528b8028e58bfee4f19b5d +size 4984660288 diff --git a/Mistral-7B-Instruct-v0.3-q5_1.gguf b/Mistral-7B-Instruct-v0.3-q5_1.gguf new file mode 100644 index 0000000..270e142 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q5_1.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:6601a9142980f971a5142af9161a13726d0ecf70937d2ab97a96171d96e6695d +size 5437645120 diff --git a/Mistral-7B-Instruct-v0.3-q5_k_m.gguf b/Mistral-7B-Instruct-v0.3-q5_k_m.gguf new file mode 100644 index 0000000..acfe35e --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q5_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:7663e939dca39c961a72d387ea5b3d6f4051d2a5cd8b11d96fd182cbe312d532 +size 5171831104 diff --git a/Mistral-7B-Instruct-v0.3-q5_k_s.gguf b/Mistral-7B-Instruct-v0.3-q5_k_s.gguf new file mode 100644 index 0000000..2dbac99 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q5_k_s.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:cc0c1b7af583d7762360cb71016c3d4394fd750e1aaae69de63863778de5ff02 +size 5104984384 diff --git a/Mistral-7B-Instruct-v0.3-q6_k_m.gguf b/Mistral-7B-Instruct-v0.3-q6_k_m.gguf new file mode 100644 index 0000000..6b49464 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q6_k_m.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a197ad409a4f8b4e85ba86d5f310f4b3d56a54bc9b2a3a73eeb221d16709ef90 +size 5947253056 diff --git a/Mistral-7B-Instruct-v0.3-q8_0.gguf b/Mistral-7B-Instruct-v0.3-q8_0.gguf new file mode 100644 index 0000000..c54f106 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3-q8_0.gguf @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4225464dfa6ddac8ee2e7445108b2415090e8658c1754ddff3e15bc5fb335555 +size 7702568992 diff --git a/Mistral-7B-Instruct-v0.3.imatrix b/Mistral-7B-Instruct-v0.3.imatrix new file mode 100644 index 0000000..13a6747 --- /dev/null +++ b/Mistral-7B-Instruct-v0.3.imatrix @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2e79af97172cc6ed110cbdaf46f500f7df2306afb22a78a43f16318cfabb189f +size 4988169 diff --git a/README.md b/README.md new file mode 100644 index 0000000..d905165 --- /dev/null +++ b/README.md @@ -0,0 +1,293 @@ +--- +license: apache-2.0 +base_model: mistralai/Mistral-7B-v0.3 +extra_gated_description: If you want to learn more about how we process your personal data, please read our Privacy Policy. +--- + +# Mistral-7B-Instruct-v0.3 GGUF Models + + +## Model Generation Details + +This model was generated using [llama.cpp](https://github.com/ggerganov/llama.cpp) at commit [`bf9087f5`](https://github.com/ggerganov/llama.cpp/commit/bf9087f59aab940cf312b85a67067ce33d9e365a). + + + + + +--- + +## Quantization Beyond the IMatrix + +I've been experimenting with a new quantization approach that selectively elevates the precision of key layers beyond what the default IMatrix configuration provides. + +In my testing, standard IMatrix quantization underperforms at lower bit depths, especially with Mixture of Experts (MoE) models. To address this, I'm using the `--tensor-type` option in `llama.cpp` to manually "bump" important layers to higher precision. You can see the implementation here: +👉 [Layer bumping with llama.cpp](https://github.com/Mungert69/GGUFModelBuilder/blob/main/model-converter/tensor_list_builder.py) + +While this does increase model file size, it significantly improves precision for a given quantization level. + +### **I'd love your feedback—have you tried this? How does it perform for you?** + + + + +--- + + + Click here to get info on choosing the right GGUF model format + + +--- + + + + + + +# Model Card for Mistral-7B-Instruct-v0.3 + +The Mistral-7B-Instruct-v0.3 Large Language Model (LLM) is an instruct fine-tuned version of the Mistral-7B-v0.3. + +Mistral-7B-v0.3 has the following changes compared to [Mistral-7B-v0.2](https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.2/edit/main/README.md) +- Extended vocabulary to 32768 +- Supports v3 Tokenizer +- Supports function calling + +## Installation + +It is recommended to use `mistralai/Mistral-7B-Instruct-v0.3` with [mistral-inference](https://github.com/mistralai/mistral-inference). For HF transformers code snippets, please keep scrolling. + +``` +pip install mistral_inference +``` + +## Download + +```py +from huggingface_hub import snapshot_download +from pathlib import Path + +mistral_models_path = Path.home().joinpath('mistral_models', '7B-Instruct-v0.3') +mistral_models_path.mkdir(parents=True, exist_ok=True) + +snapshot_download(repo_id="mistralai/Mistral-7B-Instruct-v0.3", allow_patterns=["params.json", "consolidated.safetensors", "tokenizer.model.v3"], local_dir=mistral_models_path) +``` + +### Chat + +After installing `mistral_inference`, a `mistral-chat` CLI command should be available in your environment. You can chat with the model using + +``` +mistral-chat $HOME/mistral_models/7B-Instruct-v0.3 --instruct --max_tokens 256 +``` + +### Instruct following + +```py +from mistral_inference.transformer import Transformer +from mistral_inference.generate import generate + +from mistral_common.tokens.tokenizers.mistral import MistralTokenizer +from mistral_common.protocol.instruct.messages import UserMessage +from mistral_common.protocol.instruct.request import ChatCompletionRequest + + +tokenizer = MistralTokenizer.from_file(f"{mistral_models_path}/tokenizer.model.v3") +model = Transformer.from_folder(mistral_models_path) + +completion_request = ChatCompletionRequest(messages=[UserMessage(content="Explain Machine Learning to me in a nutshell.")]) + +tokens = tokenizer.encode_chat_completion(completion_request).tokens + +out_tokens, _ = generate([tokens], model, max_tokens=64, temperature=0.0, eos_id=tokenizer.instruct_tokenizer.tokenizer.eos_id) +result = tokenizer.instruct_tokenizer.tokenizer.decode(out_tokens[0]) + +print(result) +``` + +### Function calling + +```py +from mistral_common.protocol.instruct.tool_calls import Function, Tool +from mistral_inference.transformer import Transformer +from mistral_inference.generate import generate + +from mistral_common.tokens.tokenizers.mistral import MistralTokenizer +from mistral_common.protocol.instruct.messages import UserMessage +from mistral_common.protocol.instruct.request import ChatCompletionRequest + + +tokenizer = MistralTokenizer.from_file(f"{mistral_models_path}/tokenizer.model.v3") +model = Transformer.from_folder(mistral_models_path) + +completion_request = ChatCompletionRequest( + tools=[ + Tool( + function=Function( + name="get_current_weather", + description="Get the current weather", + parameters={ + "type": "object", + "properties": { + "location": { + "type": "string", + "description": "The city and state, e.g. San Francisco, CA", + }, + "format": { + "type": "string", + "enum": ["celsius", "fahrenheit"], + "description": "The temperature unit to use. Infer this from the users location.", + }, + }, + "required": ["location", "format"], + }, + ) + ) + ], + messages=[ + UserMessage(content="What's the weather like today in Paris?"), + ], +) + +tokens = tokenizer.encode_chat_completion(completion_request).tokens + +out_tokens, _ = generate([tokens], model, max_tokens=64, temperature=0.0, eos_id=tokenizer.instruct_tokenizer.tokenizer.eos_id) +result = tokenizer.instruct_tokenizer.tokenizer.decode(out_tokens[0]) + +print(result) +``` + +## Generate with `transformers` + +If you want to use Hugging Face `transformers` to generate text, you can do something like this. + +```py +from transformers import pipeline + +messages = [ + {"role": "system", "content": "You are a pirate chatbot who always responds in pirate speak!"}, + {"role": "user", "content": "Who are you?"}, +] +chatbot = pipeline("text-generation", model="mistralai/Mistral-7B-Instruct-v0.3") +chatbot(messages) +``` + + +## Function calling with `transformers` + +To use this example, you'll need `transformers` version 4.42.0 or higher. Please see the +[function calling guide](https://huggingface.co/docs/transformers/main/chat_templating#advanced-tool-use--function-calling) +in the `transformers` docs for more information. + +```python +from transformers import AutoModelForCausalLM, AutoTokenizer +import torch + +model_id = "mistralai/Mistral-7B-Instruct-v0.3" +tokenizer = AutoTokenizer.from_pretrained(model_id) + +def get_current_weather(location: str, format: str): + """ + Get the current weather + + Args: + location: The city and state, e.g. San Francisco, CA + format: The temperature unit to use. Infer this from the users location. (choices: ["celsius", "fahrenheit"]) + """ + pass + +conversation = [{"role": "user", "content": "What's the weather like in Paris?"}] +tools = [get_current_weather] + + +# format and tokenize the tool use prompt +inputs = tokenizer.apply_chat_template( + conversation, + tools=tools, + add_generation_prompt=True, + return_dict=True, + return_tensors="pt", +) + +model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16, device_map="auto") + +inputs.to(model.device) +outputs = model.generate(**inputs, max_new_tokens=1000) +print(tokenizer.decode(outputs[0], skip_special_tokens=True)) +``` + +Note that, for reasons of space, this example does not show a complete cycle of calling a tool and adding the tool call and tool +results to the chat history so that the model can use them in its next generation. For a full tool calling example, please +see the [function calling guide](https://huggingface.co/docs/transformers/main/chat_templating#advanced-tool-use--function-calling), +and note that Mistral **does** use tool call IDs, so these must be included in your tool calls and tool results. They should be +exactly 9 alphanumeric characters. + + +## Limitations + +The Mistral 7B Instruct model is a quick demonstration that the base model can be easily fine-tuned to achieve compelling performance. +It does not have any moderation mechanisms. We're looking forward to engaging with the community on ways to +make the model finely respect guardrails, allowing for deployment in environments requiring moderated outputs. + +## The Mistral AI Team + +Albert Jiang, Alexandre Sablayrolles, Alexis Tacnet, Antoine Roux, Arthur Mensch, Audrey Herblin-Stoop, Baptiste Bout, Baudouin de Monicault, Blanche Savary, Bam4d, Caroline Feldman, Devendra Singh Chaplot, Diego de las Casas, Eleonore Arcelin, Emma Bou Hanna, Etienne Metzger, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Harizo Rajaona, Jean-Malo Delignon, Jia Li, Justus Murke, Louis Martin, Louis Ternon, Lucile Saulnier, Lélio Renard Lavaud, Margaret Jennings, Marie Pellat, Marie Torelli, Marie-Anne Lachaux, Nicolas Schuhl, Patrick von Platen, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Thibaut Lavril, Timothée Lacroix, Théophile Gervet, Thomas Wang, Valera Nemychnikova, William El Sayed, William Marshall + + + +--- + +# 🚀 If you find these models useful + +Help me test my **AI-Powered Quantum Network Monitor Assistant** with **quantum-ready security checks**: + +👉 [Quantum Network Monitor](https://readyforquantum.com/?assistant=open&utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) + + +The full Open Source Code for the Quantum Network Monitor Service available at my github repos ( repos with NetworkMonitor in the name) : [Source Code Quantum Network Monitor](https://github.com/Mungert69). You will also find the code I use to quantize the models if you want to do it yourself [GGUFModelBuilder](https://github.com/Mungert69/GGUFModelBuilder) + +💬 **How to test**: + Choose an **AI assistant type**: + - `TurboLLM` (GPT-4.1-mini) + - `HugLLM` (Hugginface Open-source models) + - `TestLLM` (Experimental CPU-only) + +### **What I’m Testing** +I’m pushing the limits of **small open-source models for AI network monitoring**, specifically: +- **Function calling** against live network services +- **How small can a model go** while still handling: + - Automated **Nmap security scans** + - **Quantum-readiness checks** + - **Network Monitoring tasks** + +🟡 **TestLLM** – Current experimental model (llama.cpp on 2 CPU threads on huggingface docker space): +- ✅ **Zero-configuration setup** +- ⏳ 30s load time (slow inference but **no API costs**) . No token limited as the cost is low. +- 🔧 **Help wanted!** If you’re into **edge-device AI**, let’s collaborate! + +### **Other Assistants** +🟢 **TurboLLM** – Uses **gpt-4.1-mini** : +- **It performs very well but unfortunatly OpenAI charges per token. For this reason tokens usage is limited. +- **Create custom cmd processors to run .net code on Quantum Network Monitor Agents** +- **Real-time network diagnostics and monitoring** +- **Security Audits** +- **Penetration testing** (Nmap/Metasploit) + +🔵 **HugLLM** – Latest Open-source models: +- 🌐 Runs on Hugging Face Inference API. Performs pretty well using the lastest models hosted on Novita. + +### 💡 **Example commands you could test**: +1. `"Give me info on my websites SSL certificate"` +2. `"Check if my server is using quantum safe encyption for communication"` +3. `"Run a comprehensive security audit on my server"` +4. '"Create a cmd processor to .. (what ever you want)" Note you need to install a [Quantum Network Monitor Agent](https://readyforquantum.com/Download/?utm_source=huggingface&utm_medium=referral&utm_campaign=huggingface_repo_readme) to run the .net code on. This is a very flexible and powerful feature. Use with caution! + +### Final Word + +I fund the servers used to create these model files, run the Quantum Network Monitor service, and pay for inference from Novita and OpenAI—all out of my own pocket. All the code behind the model creation and the Quantum Network Monitor project is [open source](https://github.com/Mungert69). Feel free to use whatever you find helpful. + +If you appreciate the work, please consider [buying me a coffee](https://www.buymeacoffee.com/mahadeva) ☕. Your support helps cover service costs and allows me to raise token limits for everyone. + +I'm also open to job opportunities or sponsorship. + +Thank you! 😊 diff --git a/configuration.json b/configuration.json new file mode 100644 index 0000000..159097f --- /dev/null +++ b/configuration.json @@ -0,0 +1 @@ +{"framework": "pytorch", "task": "others", "allow_remote": true} \ No newline at end of file