初始化项目,由ModelHub XC社区提供模型
Model: eaddario/gemma-3-12b-it-pruned-GGUF Source: Original Platform
This commit is contained in:
13
scores/gemma-3-12b-it-F16.arc
Normal file
13
scores/gemma-3-12b-it-F16.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 69.2000 +/- 1.6869
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 11768.83 ms
|
||||
llama_perf_context_print: prompt eval time = 74484.60 ms / 35749 tokens ( 2.08 ms per token, 479.95 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 75938.16 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-F16.hsw
Normal file
12
scores/gemma-3-12b-it-F16.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 81.20000000% [78.2474%, 83.8346%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 1473.70 ms
|
||||
llama_perf_context_print: prompt eval time = 281267.73 ms / 125814 tokens ( 2.24 ms per token, 447.31 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 287221.72 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
1813
scores/gemma-3-12b-it-F16.md
Normal file
1813
scores/gemma-3-12b-it-F16.md
Normal file
File diff suppressed because it is too large
Load Diff
13
scores/gemma-3-12b-it-F16.mmlu
Normal file
13
scores/gemma-3-12b-it-F16.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 45.3333 +/- 1.8190
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 1708.63 ms
|
||||
llama_perf_context_print: prompt eval time = 168929.92 ms / 67555 tokens ( 2.50 ms per token, 399.90 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 171048.28 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-F16.tqa
Normal file
13
scores/gemma-3-12b-it-F16.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 41.4667 +/- 1.8002
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 1540.31 ms
|
||||
llama_perf_context_print: prompt eval time = 245844.98 ms / 50067 tokens ( 4.91 ms per token, 203.65 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 248430.74 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-F16.wng
Normal file
11
scores/gemma-3-12b-it-F16.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 74.8000 +/- 1.5864
|
||||
|
||||
llama_perf_context_print: load time = 1465.03 ms
|
||||
llama_perf_context_print: prompt eval time = 58933.69 ms / 21857 tokens ( 2.70 ms per token, 370.87 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 59720.41 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
1746
scores/gemma-3-12b-it-IQ3_M.md
Normal file
1746
scores/gemma-3-12b-it-IQ3_M.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-IQ3_S.md
Normal file
1746
scores/gemma-3-12b-it-IQ3_S.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-IQ4_NL.md
Normal file
1746
scores/gemma-3-12b-it-IQ4_NL.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q3_K_L.md
Normal file
1746
scores/gemma-3-12b-it-Q3_K_L.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q3_K_M.md
Normal file
1746
scores/gemma-3-12b-it-Q3_K_M.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q3_K_S.md
Normal file
1746
scores/gemma-3-12b-it-Q3_K_S.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q4_K_M.md
Normal file
1746
scores/gemma-3-12b-it-Q4_K_M.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q4_K_S.md
Normal file
1746
scores/gemma-3-12b-it-Q4_K_S.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q5_K_M.md
Normal file
1746
scores/gemma-3-12b-it-Q5_K_M.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q5_K_S.md
Normal file
1746
scores/gemma-3-12b-it-Q5_K_S.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q6_K.md
Normal file
1746
scores/gemma-3-12b-it-Q6_K.md
Normal file
File diff suppressed because it is too large
Load Diff
1746
scores/gemma-3-12b-it-Q8_0.md
Normal file
1746
scores/gemma-3-12b-it-Q8_0.md
Normal file
File diff suppressed because it is too large
Load Diff
13
scores/gemma-3-12b-it-iq3_m.arc
Normal file
13
scores/gemma-3-12b-it-iq3_m.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 66.9333 +/- 1.7190
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2521.51 ms
|
||||
llama_perf_context_print: prompt eval time = 87825.27 ms / 35749 tokens ( 2.46 ms per token, 407.05 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 89287.36 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-iq3_m.hsw
Normal file
12
scores/gemma-3-12b-it-iq3_m.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 74.66666667% [71.4337%, 77.6482%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 520.21 ms
|
||||
llama_perf_context_print: prompt eval time = 302150.78 ms / 125814 tokens ( 2.40 ms per token, 416.39 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 308129.94 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-iq3_m.mmlu
Normal file
13
scores/gemma-3-12b-it-iq3_m.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 42.0000 +/- 1.8034
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 518.01 ms
|
||||
llama_perf_context_print: prompt eval time = 157589.37 ms / 67555 tokens ( 2.33 ms per token, 428.68 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 159704.02 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-iq3_m.ppx
Normal file
37
scores/gemma-3-12b-it-iq3_m.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 35.478449 ± 0.411747
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 80.18%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.366958 ± 0.007075
|
||||
Mean PPL(Q)/PPL(base) : 3.923398 ± 0.027759
|
||||
Mean PPL(Q)-PPL(base) : 26.435664 ± 0.356985
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.229937 ± 0.004277
|
||||
Maximum KLD: 25.876987
|
||||
99.9% KLD: 13.510958
|
||||
99.0% KLD: 8.435304
|
||||
99.0% KLD: 8.435304
|
||||
Median KLD: 0.728065
|
||||
10.0% KLD: 0.020280
|
||||
5.0% KLD: 0.004677
|
||||
1.0% KLD: 0.000320
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -6.562 ± 0.077 %
|
||||
Maximum Δp: 99.860%
|
||||
99.9% Δp: 90.203%
|
||||
99.0% Δp: 66.591%
|
||||
95.0% Δp: 36.644%
|
||||
90.0% Δp: 19.461%
|
||||
75.0% Δp: 1.507%
|
||||
Median Δp: -0.332%
|
||||
25.0% Δp: -9.977%
|
||||
10.0% Δp: -45.830%
|
||||
5.0% Δp: -75.458%
|
||||
1.0% Δp: -98.595%
|
||||
0.1% Δp: -99.993%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 30.166 ± 0.089 %
|
||||
Same top p: 65.259 ± 0.124 %
|
||||
13
scores/gemma-3-12b-it-iq3_m.tqa
Normal file
13
scores/gemma-3-12b-it-iq3_m.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 39.4667 +/- 1.7860
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 525.08 ms
|
||||
llama_perf_context_print: prompt eval time = 127081.64 ms / 50067 tokens ( 2.54 ms per token, 393.98 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 129691.56 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-iq3_m.wng
Normal file
11
scores/gemma-3-12b-it-iq3_m.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 67.4667 +/- 1.7119
|
||||
|
||||
llama_perf_context_print: load time = 522.07 ms
|
||||
llama_perf_context_print: prompt eval time = 54601.23 ms / 21857 tokens ( 2.50 ms per token, 400.30 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 55373.40 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-iq3_s.arc
Normal file
13
scores/gemma-3-12b-it-iq3_s.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 65.0667 +/- 1.7420
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2418.65 ms
|
||||
llama_perf_context_print: prompt eval time = 88136.62 ms / 35749 tokens ( 2.47 ms per token, 405.61 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 89561.06 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-iq3_s.hsw
Normal file
12
scores/gemma-3-12b-it-iq3_s.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 74.80000000% [71.5718%, 77.7755%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 510.06 ms
|
||||
llama_perf_context_print: prompt eval time = 312098.17 ms / 125814 tokens ( 2.48 ms per token, 403.12 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 317986.17 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-iq3_s.mmlu
Normal file
13
scores/gemma-3-12b-it-iq3_s.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 42.8000 +/- 1.8079
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 499.63 ms
|
||||
llama_perf_context_print: prompt eval time = 160758.88 ms / 67555 tokens ( 2.38 ms per token, 420.23 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 162859.07 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-iq3_s.ppx
Normal file
37
scores/gemma-3-12b-it-iq3_s.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 32.557804 ± 0.369351
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 80.53%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.281050 ± 0.006837
|
||||
Mean PPL(Q)/PPL(base) : 3.600418 ± 0.024617
|
||||
Mean PPL(Q)-PPL(base) : 23.515018 ± 0.314640
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.196823 ± 0.004146
|
||||
Maximum KLD: 24.545359
|
||||
99.9% KLD: 12.464450
|
||||
99.0% KLD: 8.054839
|
||||
99.0% KLD: 8.054839
|
||||
Median KLD: 0.702317
|
||||
10.0% KLD: 0.019899
|
||||
5.0% KLD: 0.004923
|
||||
1.0% KLD: 0.000363
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -6.314 ± 0.076 %
|
||||
Maximum Δp: 99.585%
|
||||
99.9% Δp: 90.128%
|
||||
99.0% Δp: 66.718%
|
||||
95.0% Δp: 36.821%
|
||||
90.0% Δp: 19.531%
|
||||
75.0% Δp: 1.614%
|
||||
Median Δp: -0.338%
|
||||
25.0% Δp: -9.890%
|
||||
10.0% Δp: -44.374%
|
||||
5.0% Δp: -73.795%
|
||||
1.0% Δp: -98.369%
|
||||
0.1% Δp: -99.991%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 29.810 ± 0.088 %
|
||||
Same top p: 65.550 ± 0.124 %
|
||||
13
scores/gemma-3-12b-it-iq3_s.tqa
Normal file
13
scores/gemma-3-12b-it-iq3_s.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.4000 +/- 1.7930
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 509.01 ms
|
||||
llama_perf_context_print: prompt eval time = 125818.94 ms / 50067 tokens ( 2.51 ms per token, 397.93 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 128433.23 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-iq3_s.wng
Normal file
11
scores/gemma-3-12b-it-iq3_s.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 66.1333 +/- 1.7292
|
||||
|
||||
llama_perf_context_print: load time = 501.85 ms
|
||||
llama_perf_context_print: prompt eval time = 53507.05 ms / 21857 tokens ( 2.45 ms per token, 408.49 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 54292.47 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-iq4_nl.arc
Normal file
13
scores/gemma-3-12b-it-iq4_nl.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 68.2667 +/- 1.7007
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2844.87 ms
|
||||
llama_perf_context_print: prompt eval time = 91988.56 ms / 35749 tokens ( 2.57 ms per token, 388.62 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 93449.81 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-iq4_nl.hsw
Normal file
12
scores/gemma-3-12b-it-iq4_nl.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 75.46666667% [72.2626%, 78.4112%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 559.13 ms
|
||||
llama_perf_context_print: prompt eval time = 328969.37 ms / 125814 tokens ( 2.61 ms per token, 382.45 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 334877.24 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-iq4_nl.mmlu
Normal file
13
scores/gemma-3-12b-it-iq4_nl.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 42.6667 +/- 1.8072
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 571.85 ms
|
||||
llama_perf_context_print: prompt eval time = 171936.29 ms / 67555 tokens ( 2.55 ms per token, 392.91 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 174058.79 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-iq4_nl.ppx
Normal file
37
scores/gemma-3-12b-it-iq4_nl.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 36.879518 ± 0.436763
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 80.79%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.405689 ± 0.007174
|
||||
Mean PPL(Q)/PPL(base) : 4.078336 ± 0.029260
|
||||
Mean PPL(Q)-PPL(base) : 27.836733 ± 0.381333
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.190922 ± 0.004215
|
||||
Maximum KLD: 25.421314
|
||||
99.9% KLD: 13.739526
|
||||
99.0% KLD: 8.520005
|
||||
99.0% KLD: 8.520005
|
||||
Median KLD: 0.710480
|
||||
10.0% KLD: 0.017773
|
||||
5.0% KLD: 0.003969
|
||||
1.0% KLD: 0.000258
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -6.063 ± 0.075 %
|
||||
Maximum Δp: 99.690%
|
||||
99.9% Δp: 88.494%
|
||||
99.0% Δp: 65.535%
|
||||
95.0% Δp: 36.662%
|
||||
90.0% Δp: 19.570%
|
||||
75.0% Δp: 1.644%
|
||||
Median Δp: -0.273%
|
||||
25.0% Δp: -9.153%
|
||||
10.0% Δp: -43.591%
|
||||
5.0% Δp: -73.479%
|
||||
1.0% Δp: -98.555%
|
||||
0.1% Δp: -99.995%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 29.578 ± 0.088 %
|
||||
Same top p: 66.655 ± 0.123 %
|
||||
13
scores/gemma-3-12b-it-iq4_nl.tqa
Normal file
13
scores/gemma-3-12b-it-iq4_nl.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.1333 +/- 1.7910
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 557.58 ms
|
||||
llama_perf_context_print: prompt eval time = 133950.32 ms / 50067 tokens ( 2.68 ms per token, 373.77 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 136572.17 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-iq4_nl.wng
Normal file
11
scores/gemma-3-12b-it-iq4_nl.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 68.5333 +/- 1.6968
|
||||
|
||||
llama_perf_context_print: load time = 558.43 ms
|
||||
llama_perf_context_print: prompt eval time = 56861.97 ms / 21857 tokens ( 2.60 ms per token, 384.39 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 57651.50 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q3_k_l.arc
Normal file
13
scores/gemma-3-12b-it-q3_k_l.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 67.8667 +/- 1.7063
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2488.69 ms
|
||||
llama_perf_context_print: prompt eval time = 93489.45 ms / 35749 tokens ( 2.62 ms per token, 382.39 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 94938.12 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q3_k_l.hsw
Normal file
12
scores/gemma-3-12b-it-q3_k_l.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 75.73333333% [72.5391%, 78.6653%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 525.75 ms
|
||||
llama_perf_context_print: prompt eval time = 338885.65 ms / 125814 tokens ( 2.69 ms per token, 371.26 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 344811.83 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q3_k_l.mmlu
Normal file
13
scores/gemma-3-12b-it-q3_k_l.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 42.5333 +/- 1.8065
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 523.91 ms
|
||||
llama_perf_context_print: prompt eval time = 174640.97 ms / 67555 tokens ( 2.59 ms per token, 386.82 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 176748.74 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q3_k_l.ppx
Normal file
37
scores/gemma-3-12b-it-q3_k_l.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 33.537202 ± 0.371564
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 79.10%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.310688 ± 0.006832
|
||||
Mean PPL(Q)/PPL(base) : 3.708725 ± 0.025338
|
||||
Mean PPL(Q)-PPL(base) : 24.494416 ± 0.318026
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.263077 ± 0.004253
|
||||
Maximum KLD: 23.683029
|
||||
99.9% KLD: 12.829600
|
||||
99.0% KLD: 8.069357
|
||||
99.0% KLD: 8.069357
|
||||
Median KLD: 0.755642
|
||||
10.0% KLD: 0.018898
|
||||
5.0% KLD: 0.004577
|
||||
1.0% KLD: 0.000375
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -8.200 ± 0.078 %
|
||||
Maximum Δp: 99.800%
|
||||
99.9% Δp: 90.394%
|
||||
99.0% Δp: 66.393%
|
||||
95.0% Δp: 34.806%
|
||||
90.0% Δp: 16.907%
|
||||
75.0% Δp: 0.806%
|
||||
Median Δp: -0.550%
|
||||
25.0% Δp: -12.684%
|
||||
10.0% Δp: -50.926%
|
||||
5.0% Δp: -77.962%
|
||||
1.0% Δp: -98.305%
|
||||
0.1% Δp: -99.990%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 31.084 ± 0.088 %
|
||||
Same top p: 62.253 ± 0.126 %
|
||||
13
scores/gemma-3-12b-it-q3_k_l.tqa
Normal file
13
scores/gemma-3-12b-it-q3_k_l.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.2667 +/- 1.7920
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 525.48 ms
|
||||
llama_perf_context_print: prompt eval time = 136845.27 ms / 50067 tokens ( 2.73 ms per token, 365.87 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 139475.22 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q3_k_l.wng
Normal file
11
scores/gemma-3-12b-it-q3_k_l.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 67.0667 +/- 1.7172
|
||||
|
||||
llama_perf_context_print: load time = 528.70 ms
|
||||
llama_perf_context_print: prompt eval time = 58406.45 ms / 21857 tokens ( 2.67 ms per token, 374.22 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 59199.95 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q3_k_m.arc
Normal file
13
scores/gemma-3-12b-it-q3_k_m.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 68.6667 +/- 1.6949
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2459.83 ms
|
||||
llama_perf_context_print: prompt eval time = 93086.65 ms / 35749 tokens ( 2.60 ms per token, 384.04 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 94546.19 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q3_k_m.hsw
Normal file
12
scores/gemma-3-12b-it-q3_k_m.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 75.20000000% [71.9861%, 78.1570%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 506.32 ms
|
||||
llama_perf_context_print: prompt eval time = 333957.19 ms / 125814 tokens ( 2.65 ms per token, 376.74 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 339911.25 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q3_k_m.mmlu
Normal file
13
scores/gemma-3-12b-it-q3_k_m.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 44.4000 +/- 1.8155
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 510.44 ms
|
||||
llama_perf_context_print: prompt eval time = 172922.20 ms / 67555 tokens ( 2.56 ms per token, 390.67 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 175045.07 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q3_k_m.ppx
Normal file
37
scores/gemma-3-12b-it-q3_k_m.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 34.029948 ± 0.376082
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 78.62%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.325274 ± 0.006875
|
||||
Mean PPL(Q)/PPL(base) : 3.763215 ± 0.025870
|
||||
Mean PPL(Q)-PPL(base) : 24.987163 ± 0.322907
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.300974 ± 0.004321
|
||||
Maximum KLD: 23.795326
|
||||
99.9% KLD: 12.854443
|
||||
99.0% KLD: 8.215378
|
||||
99.0% KLD: 8.215378
|
||||
Median KLD: 0.782525
|
||||
10.0% KLD: 0.019878
|
||||
5.0% KLD: 0.004841
|
||||
1.0% KLD: 0.000391
|
||||
Minimum KLD: -0.000001
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -8.606 ± 0.079 %
|
||||
Maximum Δp: 99.896%
|
||||
99.9% Δp: 91.072%
|
||||
99.0% Δp: 67.090%
|
||||
95.0% Δp: 34.759%
|
||||
90.0% Δp: 16.501%
|
||||
75.0% Δp: 0.726%
|
||||
Median Δp: -0.601%
|
||||
25.0% Δp: -13.523%
|
||||
10.0% Δp: -52.748%
|
||||
5.0% Δp: -79.067%
|
||||
1.0% Δp: -98.555%
|
||||
0.1% Δp: -99.990%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 31.537 ± 0.088 %
|
||||
Same top p: 61.420 ± 0.127 %
|
||||
13
scores/gemma-3-12b-it-q3_k_m.tqa
Normal file
13
scores/gemma-3-12b-it-q3_k_m.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.0000 +/- 1.7900
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 504.87 ms
|
||||
llama_perf_context_print: prompt eval time = 134889.64 ms / 50067 tokens ( 2.69 ms per token, 371.17 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 137522.13 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q3_k_m.wng
Normal file
11
scores/gemma-3-12b-it-q3_k_m.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 65.8667 +/- 1.7325
|
||||
|
||||
llama_perf_context_print: load time = 501.08 ms
|
||||
llama_perf_context_print: prompt eval time = 56606.33 ms / 21857 tokens ( 2.59 ms per token, 386.12 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 57401.53 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q3_k_s.arc
Normal file
13
scores/gemma-3-12b-it-q3_k_s.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 67.6000 +/- 1.7100
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2303.10 ms
|
||||
llama_perf_context_print: prompt eval time = 91534.38 ms / 35749 tokens ( 2.56 ms per token, 390.55 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 92983.34 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q3_k_s.hsw
Normal file
12
scores/gemma-3-12b-it-q3_k_s.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 75.20000000% [71.9861%, 78.1570%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 501.43 ms
|
||||
llama_perf_context_print: prompt eval time = 325090.51 ms / 125814 tokens ( 2.58 ms per token, 387.01 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 331071.94 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q3_k_s.mmlu
Normal file
13
scores/gemma-3-12b-it-q3_k_s.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 43.7333 +/- 1.8126
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 508.75 ms
|
||||
llama_perf_context_print: prompt eval time = 171219.64 ms / 67555 tokens ( 2.53 ms per token, 394.55 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 173347.99 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q3_k_s.ppx
Normal file
37
scores/gemma-3-12b-it-q3_k_s.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 36.103787 ± 0.402250
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 77.13%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.384430 ± 0.007124
|
||||
Mean PPL(Q)/PPL(base) : 3.992551 ± 0.028445
|
||||
Mean PPL(Q)-PPL(base) : 27.061001 ± 0.350072
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.395578 ± 0.004559
|
||||
Maximum KLD: 24.088387
|
||||
99.9% KLD: 13.768757
|
||||
99.0% KLD: 8.603606
|
||||
99.0% KLD: 8.603606
|
||||
Median KLD: 0.858334
|
||||
10.0% KLD: 0.021285
|
||||
5.0% KLD: 0.005174
|
||||
1.0% KLD: 0.000393
|
||||
Minimum KLD: -0.000002
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -9.212 ± 0.080 %
|
||||
Maximum Δp: 99.844%
|
||||
99.9% Δp: 91.248%
|
||||
99.0% Δp: 67.876%
|
||||
95.0% Δp: 34.787%
|
||||
90.0% Δp: 16.139%
|
||||
75.0% Δp: 0.593%
|
||||
Median Δp: -0.696%
|
||||
25.0% Δp: -14.997%
|
||||
10.0% Δp: -54.930%
|
||||
5.0% Δp: -79.831%
|
||||
1.0% Δp: -98.568%
|
||||
0.1% Δp: -99.992%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 32.089 ± 0.088 %
|
||||
Same top p: 59.830 ± 0.128 %
|
||||
13
scores/gemma-3-12b-it-q3_k_s.tqa
Normal file
13
scores/gemma-3-12b-it-q3_k_s.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 39.8667 +/- 1.7890
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 501.19 ms
|
||||
llama_perf_context_print: prompt eval time = 132421.97 ms / 50067 tokens ( 2.64 ms per token, 378.09 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 135070.83 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q3_k_s.wng
Normal file
11
scores/gemma-3-12b-it-q3_k_s.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 66.5333 +/- 1.7242
|
||||
|
||||
llama_perf_context_print: load time = 501.82 ms
|
||||
llama_perf_context_print: prompt eval time = 56390.15 ms / 21857 tokens ( 2.58 ms per token, 387.60 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 57175.02 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q4_k_m.arc
Normal file
13
scores/gemma-3-12b-it-q4_k_m.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 69.8667 +/- 1.6766
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2958.46 ms
|
||||
llama_perf_context_print: prompt eval time = 95508.61 ms / 35749 tokens ( 2.67 ms per token, 374.30 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 97132.91 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q4_k_m.hsw
Normal file
12
scores/gemma-3-12b-it-q4_k_m.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 77.46666667% [74.3409%, 80.3125%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 567.97 ms
|
||||
llama_perf_context_print: prompt eval time = 339481.51 ms / 125814 tokens ( 2.70 ms per token, 370.61 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 345440.37 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q4_k_m.mmlu
Normal file
13
scores/gemma-3-12b-it-q4_k_m.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 43.6000 +/- 1.8119
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 576.27 ms
|
||||
llama_perf_context_print: prompt eval time = 176741.59 ms / 67555 tokens ( 2.62 ms per token, 382.22 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 178859.75 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q4_k_m.ppx
Normal file
37
scores/gemma-3-12b-it-q4_k_m.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 31.219358 ± 0.342580
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 80.83%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.239071 ± 0.006532
|
||||
Mean PPL(Q)/PPL(base) : 3.452405 ± 0.022552
|
||||
Mean PPL(Q)-PPL(base) : 22.176573 ± 0.287880
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.106608 ± 0.003932
|
||||
Maximum KLD: 20.797123
|
||||
99.9% KLD: 12.310450
|
||||
99.0% KLD: 7.690277
|
||||
99.0% KLD: 7.690277
|
||||
Median KLD: 0.642402
|
||||
10.0% KLD: 0.017366
|
||||
5.0% KLD: 0.004376
|
||||
1.0% KLD: 0.000366
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -7.494 ± 0.075 %
|
||||
Maximum Δp: 99.726%
|
||||
99.9% Δp: 87.733%
|
||||
99.0% Δp: 63.430%
|
||||
95.0% Δp: 33.085%
|
||||
90.0% Δp: 16.533%
|
||||
75.0% Δp: 0.946%
|
||||
Median Δp: -0.461%
|
||||
25.0% Δp: -11.162%
|
||||
10.0% Δp: -47.083%
|
||||
5.0% Δp: -74.484%
|
||||
1.0% Δp: -98.197%
|
||||
0.1% Δp: -99.992%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 29.739 ± 0.088 %
|
||||
Same top p: 65.218 ± 0.124 %
|
||||
13
scores/gemma-3-12b-it-q4_k_m.tqa
Normal file
13
scores/gemma-3-12b-it-q4_k_m.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.8000 +/- 1.7958
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 569.78 ms
|
||||
llama_perf_context_print: prompt eval time = 138110.48 ms / 50067 tokens ( 2.76 ms per token, 362.51 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 140746.53 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q4_k_m.wng
Normal file
11
scores/gemma-3-12b-it-q4_k_m.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 68.4000 +/- 1.6988
|
||||
|
||||
llama_perf_context_print: load time = 573.82 ms
|
||||
llama_perf_context_print: prompt eval time = 59035.21 ms / 21857 tokens ( 2.70 ms per token, 370.24 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 59825.11 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q4_k_s.arc
Normal file
13
scores/gemma-3-12b-it-q4_k_s.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 69.8667 +/- 1.6766
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 2939.37 ms
|
||||
llama_perf_context_print: prompt eval time = 94796.83 ms / 35749 tokens ( 2.65 ms per token, 377.11 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 96249.68 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q4_k_s.hsw
Normal file
12
scores/gemma-3-12b-it-q4_k_s.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 77.33333333% [74.2021%, 80.1860%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 567.43 ms
|
||||
llama_perf_context_print: prompt eval time = 336113.08 ms / 125814 tokens ( 2.67 ms per token, 374.32 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 342078.71 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q4_k_s.mmlu
Normal file
13
scores/gemma-3-12b-it-q4_k_s.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 43.4667 +/- 1.8113
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 557.41 ms
|
||||
llama_perf_context_print: prompt eval time = 175467.57 ms / 67555 tokens ( 2.60 ms per token, 385.00 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 177580.76 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q4_k_s.ppx
Normal file
37
scores/gemma-3-12b-it-q4_k_s.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 31.245217 ± 0.343080
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 80.86%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.239899 ± 0.006533
|
||||
Mean PPL(Q)/PPL(base) : 3.455265 ± 0.022573
|
||||
Mean PPL(Q)-PPL(base) : 22.202432 ± 0.288349
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.105266 ± 0.003926
|
||||
Maximum KLD: 20.987593
|
||||
99.9% KLD: 12.310256
|
||||
99.0% KLD: 7.698290
|
||||
99.0% KLD: 7.698290
|
||||
Median KLD: 0.643550
|
||||
10.0% KLD: 0.017519
|
||||
5.0% KLD: 0.004389
|
||||
1.0% KLD: 0.000363
|
||||
Minimum KLD: -0.000002
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -7.460 ± 0.075 %
|
||||
Maximum Δp: 99.266%
|
||||
99.9% Δp: 87.347%
|
||||
99.0% Δp: 63.610%
|
||||
95.0% Δp: 33.077%
|
||||
90.0% Δp: 16.572%
|
||||
75.0% Δp: 0.938%
|
||||
Median Δp: -0.465%
|
||||
25.0% Δp: -11.152%
|
||||
10.0% Δp: -46.828%
|
||||
5.0% Δp: -74.300%
|
||||
1.0% Δp: -98.187%
|
||||
0.1% Δp: -99.991%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 29.698 ± 0.088 %
|
||||
Same top p: 65.245 ± 0.124 %
|
||||
13
scores/gemma-3-12b-it-q4_k_s.tqa
Normal file
13
scores/gemma-3-12b-it-q4_k_s.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.8000 +/- 1.7958
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 580.69 ms
|
||||
llama_perf_context_print: prompt eval time = 137720.71 ms / 50067 tokens ( 2.75 ms per token, 363.54 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 140323.32 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q4_k_s.wng
Normal file
11
scores/gemma-3-12b-it-q4_k_s.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 69.6000 +/- 1.6807
|
||||
|
||||
llama_perf_context_print: load time = 571.52 ms
|
||||
llama_perf_context_print: prompt eval time = 58871.14 ms / 21857 tokens ( 2.69 ms per token, 371.27 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 59658.60 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q5_k_m.arc
Normal file
13
scores/gemma-3-12b-it-q5_k_m.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 68.8000 +/- 1.6929
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 3504.83 ms
|
||||
llama_perf_context_print: prompt eval time = 95720.53 ms / 35749 tokens ( 2.68 ms per token, 373.47 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 97175.56 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q5_k_m.hsw
Normal file
12
scores/gemma-3-12b-it-q5_k_m.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 77.86666667% [74.7577%, 80.6916%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 627.41 ms
|
||||
llama_perf_context_print: prompt eval time = 342154.54 ms / 125814 tokens ( 2.72 ms per token, 367.71 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 348106.77 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q5_k_m.mmlu
Normal file
13
scores/gemma-3-12b-it-q5_k_m.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 45.3333 +/- 1.8190
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 644.49 ms
|
||||
llama_perf_context_print: prompt eval time = 177127.61 ms / 67555 tokens ( 2.62 ms per token, 381.39 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 179250.91 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q5_k_m.ppx
Normal file
37
scores/gemma-3-12b-it-q5_k_m.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 30.668770 ± 0.335889
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 81.60%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.221278 ± 0.006415
|
||||
Mean PPL(Q)/PPL(base) : 3.391518 ± 0.021755
|
||||
Mean PPL(Q)-PPL(base) : 21.625984 ± 0.280606
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.039101 ± 0.003848
|
||||
Maximum KLD: 22.598866
|
||||
99.9% KLD: 12.543374
|
||||
99.0% KLD: 7.697062
|
||||
99.0% KLD: 7.697062
|
||||
Median KLD: 0.596171
|
||||
10.0% KLD: 0.017018
|
||||
5.0% KLD: 0.004340
|
||||
1.0% KLD: 0.000363
|
||||
Minimum KLD: -0.000000
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -7.245 ± 0.074 %
|
||||
Maximum Δp: 99.814%
|
||||
99.9% Δp: 85.186%
|
||||
99.0% Δp: 61.478%
|
||||
95.0% Δp: 32.505%
|
||||
90.0% Δp: 16.393%
|
||||
75.0% Δp: 0.971%
|
||||
Median Δp: -0.442%
|
||||
25.0% Δp: -10.578%
|
||||
10.0% Δp: -45.165%
|
||||
5.0% Δp: -72.951%
|
||||
1.0% Δp: -98.402%
|
||||
0.1% Δp: -99.994%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 29.177 ± 0.088 %
|
||||
Same top p: 66.780 ± 0.123 %
|
||||
13
scores/gemma-3-12b-it-q5_k_m.tqa
Normal file
13
scores/gemma-3-12b-it-q5_k_m.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 41.2000 +/- 1.7984
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 653.09 ms
|
||||
llama_perf_context_print: prompt eval time = 137695.68 ms / 50067 tokens ( 2.75 ms per token, 363.61 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 140323.22 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q5_k_m.wng
Normal file
11
scores/gemma-3-12b-it-q5_k_m.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 68.8000 +/- 1.6929
|
||||
|
||||
llama_perf_context_print: load time = 647.06 ms
|
||||
llama_perf_context_print: prompt eval time = 58456.53 ms / 21857 tokens ( 2.67 ms per token, 373.90 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 59248.75 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q5_k_s.arc
Normal file
13
scores/gemma-3-12b-it-q5_k_s.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 68.9333 +/- 1.6909
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 3403.68 ms
|
||||
llama_perf_context_print: prompt eval time = 94455.25 ms / 35749 tokens ( 2.64 ms per token, 378.48 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 95912.08 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q5_k_s.hsw
Normal file
12
scores/gemma-3-12b-it-q5_k_s.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 77.73333333% [74.6188%, 80.5653%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 627.98 ms
|
||||
llama_perf_context_print: prompt eval time = 334461.40 ms / 125814 tokens ( 2.66 ms per token, 376.17 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 340403.29 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q5_k_s.mmlu
Normal file
13
scores/gemma-3-12b-it-q5_k_s.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 45.0667 +/- 1.8180
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 666.76 ms
|
||||
llama_perf_context_print: prompt eval time = 174509.11 ms / 67555 tokens ( 2.58 ms per token, 387.11 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 176638.86 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q5_k_s.ppx
Normal file
37
scores/gemma-3-12b-it-q5_k_s.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 30.644285 ± 0.335716
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 81.62%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.220479 ± 0.006413
|
||||
Mean PPL(Q)/PPL(base) : 3.388810 ± 0.021733
|
||||
Mean PPL(Q)-PPL(base) : 21.601499 ± 0.280413
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.038541 ± 0.003836
|
||||
Maximum KLD: 22.424044
|
||||
99.9% KLD: 12.381549
|
||||
99.0% KLD: 7.687484
|
||||
99.0% KLD: 7.687484
|
||||
Median KLD: 0.597444
|
||||
10.0% KLD: 0.016928
|
||||
5.0% KLD: 0.004340
|
||||
1.0% KLD: 0.000359
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -7.203 ± 0.074 %
|
||||
Maximum Δp: 99.718%
|
||||
99.9% Δp: 84.842%
|
||||
99.0% Δp: 61.523%
|
||||
95.0% Δp: 32.607%
|
||||
90.0% Δp: 16.381%
|
||||
75.0% Δp: 0.982%
|
||||
Median Δp: -0.437%
|
||||
25.0% Δp: -10.550%
|
||||
10.0% Δp: -45.107%
|
||||
5.0% Δp: -72.871%
|
||||
1.0% Δp: -98.353%
|
||||
0.1% Δp: -99.993%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 29.119 ± 0.088 %
|
||||
Same top p: 66.765 ± 0.123 %
|
||||
13
scores/gemma-3-12b-it-q5_k_s.tqa
Normal file
13
scores/gemma-3-12b-it-q5_k_s.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 41.3333 +/- 1.7993
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 633.84 ms
|
||||
llama_perf_context_print: prompt eval time = 135548.17 ms / 50067 tokens ( 2.71 ms per token, 369.37 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 138178.93 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q5_k_s.wng
Normal file
11
scores/gemma-3-12b-it-q5_k_s.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 68.6667 +/- 1.6949
|
||||
|
||||
llama_perf_context_print: load time = 639.95 ms
|
||||
llama_perf_context_print: prompt eval time = 57785.71 ms / 21857 tokens ( 2.64 ms per token, 378.24 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 58579.25 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q6_k.arc
Normal file
13
scores/gemma-3-12b-it-q6_k.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 68.9333 +/- 1.6909
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 4062.76 ms
|
||||
llama_perf_context_print: prompt eval time = 95561.27 ms / 35749 tokens ( 2.67 ms per token, 374.10 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 97004.29 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q6_k.hsw
Normal file
12
scores/gemma-3-12b-it-q6_k.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 78.66666667% [75.5926%, 81.4486%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 712.37 ms
|
||||
llama_perf_context_print: prompt eval time = 341004.43 ms / 125814 tokens ( 2.71 ms per token, 368.95 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 346935.77 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q6_k.mmlu
Normal file
13
scores/gemma-3-12b-it-q6_k.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 45.4667 +/- 1.8194
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 755.78 ms
|
||||
llama_perf_context_print: prompt eval time = 181685.88 ms / 67555 tokens ( 2.69 ms per token, 371.82 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 183820.78 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q6_k.ppx
Normal file
37
scores/gemma-3-12b-it-q6_k.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 30.107026 ± 0.329175
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 81.96%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.202791 ± 0.006352
|
||||
Mean PPL(Q)/PPL(base) : 3.329397 ± 0.021147
|
||||
Mean PPL(Q)-PPL(base) : 21.064240 ± 0.273651
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.009233 ± 0.003785
|
||||
Maximum KLD: 22.312464
|
||||
99.9% KLD: 12.292943
|
||||
99.0% KLD: 7.613831
|
||||
99.0% KLD: 7.613831
|
||||
Median KLD: 0.574097
|
||||
10.0% KLD: 0.016605
|
||||
5.0% KLD: 0.004251
|
||||
1.0% KLD: 0.000354
|
||||
Minimum KLD: -0.000000
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -7.022 ± 0.073 %
|
||||
Maximum Δp: 99.488%
|
||||
99.9% Δp: 84.278%
|
||||
99.0% Δp: 60.487%
|
||||
95.0% Δp: 32.046%
|
||||
90.0% Δp: 16.402%
|
||||
75.0% Δp: 1.023%
|
||||
Median Δp: -0.418%
|
||||
25.0% Δp: -10.160%
|
||||
10.0% Δp: -43.946%
|
||||
5.0% Δp: -72.034%
|
||||
1.0% Δp: -98.340%
|
||||
0.1% Δp: -99.994%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 28.800 ± 0.088 %
|
||||
Same top p: 67.476 ± 0.122 %
|
||||
13
scores/gemma-3-12b-it-q6_k.tqa
Normal file
13
scores/gemma-3-12b-it-q6_k.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 40.8000 +/- 1.7958
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 719.91 ms
|
||||
llama_perf_context_print: prompt eval time = 141181.13 ms / 50067 tokens ( 2.82 ms per token, 354.63 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 143821.77 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q6_k.wng
Normal file
11
scores/gemma-3-12b-it-q6_k.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 68.6667 +/- 1.6949
|
||||
|
||||
llama_perf_context_print: load time = 720.44 ms
|
||||
llama_perf_context_print: prompt eval time = 59839.70 ms / 21857 tokens ( 2.74 ms per token, 365.26 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 60641.60 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q8_0.arc
Normal file
13
scores/gemma-3-12b-it-q8_0.arc
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 68.8000 +/- 1.6929
|
||||
Random chance: 25.0083 +/- 1.5824
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 5336.21 ms
|
||||
llama_perf_context_print: prompt eval time = 90355.25 ms / 35749 tokens ( 2.53 ms per token, 395.65 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 91808.86 ms / 35750 tokens
|
||||
ggml_metal_free: deallocating
|
||||
12
scores/gemma-3-12b-it-q8_0.hsw
Normal file
12
scores/gemma-3-12b-it-q8_0.hsw
Normal file
@@ -0,0 +1,12 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
|
||||
|
||||
750 78.40000000% [75.3141%, 81.1964%]
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 857.12 ms
|
||||
llama_perf_context_print: prompt eval time = 326655.40 ms / 125814 tokens ( 2.60 ms per token, 385.16 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 332582.43 ms / 125815 tokens
|
||||
ggml_metal_free: deallocating
|
||||
13
scores/gemma-3-12b-it-q8_0.mmlu
Normal file
13
scores/gemma-3-12b-it-q8_0.mmlu
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 45.4667 +/- 1.8194
|
||||
Random chance: 25.0000 +/- 1.5822
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 847.03 ms
|
||||
llama_perf_context_print: prompt eval time = 167726.56 ms / 67555 tokens ( 2.48 ms per token, 402.77 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 169830.02 ms / 67556 tokens
|
||||
ggml_metal_free: deallocating
|
||||
37
scores/gemma-3-12b-it-q8_0.ppx
Normal file
37
scores/gemma-3-12b-it-q8_0.ppx
Normal file
@@ -0,0 +1,37 @@
|
||||
====== Perplexity statistics ======
|
||||
Mean PPL(Q) : 30.014168 ± 0.327772
|
||||
Mean PPL(base) : 9.042786 ± 0.071503
|
||||
Cor(ln(PPL(Q)), ln(PPL(base))): 82.04%
|
||||
Mean ln(PPL(Q)/PPL(base)) : 1.199702 ± 0.006332
|
||||
Mean PPL(Q)/PPL(base) : 3.319129 ± 0.021018
|
||||
Mean PPL(Q)-PPL(base) : 20.971382 ± 0.272199
|
||||
|
||||
====== KL divergence statistics ======
|
||||
Mean KLD: 1.004325 ± 0.003776
|
||||
Maximum KLD: 22.233902
|
||||
99.9% KLD: 12.361748
|
||||
99.0% KLD: 7.606510
|
||||
99.0% KLD: 7.606510
|
||||
Median KLD: 0.570352
|
||||
10.0% KLD: 0.016693
|
||||
5.0% KLD: 0.004282
|
||||
1.0% KLD: 0.000351
|
||||
Minimum KLD: -0.000003
|
||||
|
||||
====== Token probability statistics ======
|
||||
Mean Δp: -7.018 ± 0.073 %
|
||||
Maximum Δp: 99.437%
|
||||
99.9% Δp: 83.997%
|
||||
99.0% Δp: 60.615%
|
||||
95.0% Δp: 32.020%
|
||||
90.0% Δp: 16.340%
|
||||
75.0% Δp: 1.021%
|
||||
Median Δp: -0.417%
|
||||
25.0% Δp: -10.112%
|
||||
10.0% Δp: -43.816%
|
||||
5.0% Δp: -71.991%
|
||||
1.0% Δp: -98.342%
|
||||
0.1% Δp: -99.994%
|
||||
Minimum Δp: -100.000%
|
||||
RMS Δp : 28.772 ± 0.088 %
|
||||
Same top p: 67.589 ± 0.122 %
|
||||
13
scores/gemma-3-12b-it-q8_0.tqa
Normal file
13
scores/gemma-3-12b-it-q8_0.tqa
Normal file
@@ -0,0 +1,13 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final result: 41.7333 +/- 1.8018
|
||||
Random chance: 19.8992 +/- 1.4588
|
||||
|
||||
|
||||
llama_perf_context_print: load time = 861.28 ms
|
||||
llama_perf_context_print: prompt eval time = 133009.67 ms / 50067 tokens ( 2.66 ms per token, 376.42 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 135633.15 ms / 50068 tokens
|
||||
ggml_metal_free: deallocating
|
||||
11
scores/gemma-3-12b-it-q8_0.wng
Normal file
11
scores/gemma-3-12b-it-q8_0.wng
Normal file
@@ -0,0 +1,11 @@
|
||||
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
|
||||
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
|
||||
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
|
||||
|
||||
Final Winogrande score(750 tasks): 68.8000 +/- 1.6929
|
||||
|
||||
llama_perf_context_print: load time = 853.87 ms
|
||||
llama_perf_context_print: prompt eval time = 55909.06 ms / 21857 tokens ( 2.56 ms per token, 390.94 tokens per second)
|
||||
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
|
||||
llama_perf_context_print: total time = 56690.38 ms / 21858 tokens
|
||||
ggml_metal_free: deallocating
|
||||
Reference in New Issue
Block a user