初始化项目,由ModelHub XC社区提供模型

Model: eaddario/gemma-3-12b-it-pruned-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-21 18:50:10 +08:00
commit 66ebbdcd29
114 changed files with 24940 additions and 0 deletions

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
Final result: 69.2000 +/- 1.6869
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 11768.83 ms
llama_perf_context_print: prompt eval time = 74484.60 ms / 35749 tokens ( 2.08 ms per token, 479.95 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 75938.16 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
750 81.20000000% [78.2474%, 83.8346%]
llama_perf_context_print: load time = 1473.70 ms
llama_perf_context_print: prompt eval time = 281267.73 ms / 125814 tokens ( 2.24 ms per token, 447.31 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 287221.72 ms / 125815 tokens
ggml_metal_free: deallocating

1813
scores/gemma-3-12b-it-F16.md Normal file

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
Final result: 45.3333 +/- 1.8190
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 1708.63 ms
llama_perf_context_print: prompt eval time = 168929.92 ms / 67555 tokens ( 2.50 ms per token, 399.90 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 171048.28 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
Final result: 41.4667 +/- 1.8002
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 1540.31 ms
llama_perf_context_print: prompt eval time = 245844.98 ms / 50067 tokens ( 4.91 ms per token, 203.65 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 248430.74 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 40 key-value pairs and 626 tensors from ./gemma-3-12b-it-F16.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 74.8000 +/- 1.5864
llama_perf_context_print: load time = 1465.03 ms
llama_perf_context_print: prompt eval time = 58933.69 ms / 21857 tokens ( 2.70 ms per token, 370.87 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 59720.41 ms / 21858 tokens
ggml_metal_free: deallocating

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
Final result: 66.9333 +/- 1.7190
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2521.51 ms
llama_perf_context_print: prompt eval time = 87825.27 ms / 35749 tokens ( 2.46 ms per token, 407.05 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 89287.36 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
750 74.66666667% [71.4337%, 77.6482%]
llama_perf_context_print: load time = 520.21 ms
llama_perf_context_print: prompt eval time = 302150.78 ms / 125814 tokens ( 2.40 ms per token, 416.39 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 308129.94 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
Final result: 42.0000 +/- 1.8034
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 518.01 ms
llama_perf_context_print: prompt eval time = 157589.37 ms / 67555 tokens ( 2.33 ms per token, 428.68 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 159704.02 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 35.478449 ± 0.411747
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 80.18%
Mean ln(PPL(Q)/PPL(base)) : 1.366958 ± 0.007075
Mean PPL(Q)/PPL(base) : 3.923398 ± 0.027759
Mean PPL(Q)-PPL(base) : 26.435664 ± 0.356985
====== KL divergence statistics ======
Mean KLD: 1.229937 ± 0.004277
Maximum KLD: 25.876987
99.9% KLD: 13.510958
99.0% KLD: 8.435304
99.0% KLD: 8.435304
Median KLD: 0.728065
10.0% KLD: 0.020280
5.0% KLD: 0.004677
1.0% KLD: 0.000320
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -6.562 ± 0.077 %
Maximum Δp: 99.860%
99.9% Δp: 90.203%
99.0% Δp: 66.591%
95.0% Δp: 36.644%
90.0% Δp: 19.461%
75.0% Δp: 1.507%
Median Δp: -0.332%
25.0% Δp: -9.977%
10.0% Δp: -45.830%
5.0% Δp: -75.458%
1.0% Δp: -98.595%
0.1% Δp: -99.993%
Minimum Δp: -100.000%
RMS Δp : 30.166 ± 0.089 %
Same top p: 65.259 ± 0.124 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
Final result: 39.4667 +/- 1.7860
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 525.08 ms
llama_perf_context_print: prompt eval time = 127081.64 ms / 50067 tokens ( 2.54 ms per token, 393.98 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 129691.56 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_M.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 67.4667 +/- 1.7119
llama_perf_context_print: load time = 522.07 ms
llama_perf_context_print: prompt eval time = 54601.23 ms / 21857 tokens ( 2.50 ms per token, 400.30 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 55373.40 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
Final result: 65.0667 +/- 1.7420
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2418.65 ms
llama_perf_context_print: prompt eval time = 88136.62 ms / 35749 tokens ( 2.47 ms per token, 405.61 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 89561.06 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
750 74.80000000% [71.5718%, 77.7755%]
llama_perf_context_print: load time = 510.06 ms
llama_perf_context_print: prompt eval time = 312098.17 ms / 125814 tokens ( 2.48 ms per token, 403.12 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 317986.17 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
Final result: 42.8000 +/- 1.8079
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 499.63 ms
llama_perf_context_print: prompt eval time = 160758.88 ms / 67555 tokens ( 2.38 ms per token, 420.23 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 162859.07 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 32.557804 ± 0.369351
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 80.53%
Mean ln(PPL(Q)/PPL(base)) : 1.281050 ± 0.006837
Mean PPL(Q)/PPL(base) : 3.600418 ± 0.024617
Mean PPL(Q)-PPL(base) : 23.515018 ± 0.314640
====== KL divergence statistics ======
Mean KLD: 1.196823 ± 0.004146
Maximum KLD: 24.545359
99.9% KLD: 12.464450
99.0% KLD: 8.054839
99.0% KLD: 8.054839
Median KLD: 0.702317
10.0% KLD: 0.019899
5.0% KLD: 0.004923
1.0% KLD: 0.000363
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -6.314 ± 0.076 %
Maximum Δp: 99.585%
99.9% Δp: 90.128%
99.0% Δp: 66.718%
95.0% Δp: 36.821%
90.0% Δp: 19.531%
75.0% Δp: 1.614%
Median Δp: -0.338%
25.0% Δp: -9.890%
10.0% Δp: -44.374%
5.0% Δp: -73.795%
1.0% Δp: -98.369%
0.1% Δp: -99.991%
Minimum Δp: -100.000%
RMS Δp : 29.810 ± 0.088 %
Same top p: 65.550 ± 0.124 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
Final result: 40.4000 +/- 1.7930
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 509.01 ms
llama_perf_context_print: prompt eval time = 125818.94 ms / 50067 tokens ( 2.51 ms per token, 397.93 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 128433.23 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ3_S.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 66.1333 +/- 1.7292
llama_perf_context_print: load time = 501.85 ms
llama_perf_context_print: prompt eval time = 53507.05 ms / 21857 tokens ( 2.45 ms per token, 408.49 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 54292.47 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
Final result: 68.2667 +/- 1.7007
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2844.87 ms
llama_perf_context_print: prompt eval time = 91988.56 ms / 35749 tokens ( 2.57 ms per token, 388.62 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 93449.81 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
750 75.46666667% [72.2626%, 78.4112%]
llama_perf_context_print: load time = 559.13 ms
llama_perf_context_print: prompt eval time = 328969.37 ms / 125814 tokens ( 2.61 ms per token, 382.45 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 334877.24 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
Final result: 42.6667 +/- 1.8072
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 571.85 ms
llama_perf_context_print: prompt eval time = 171936.29 ms / 67555 tokens ( 2.55 ms per token, 392.91 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 174058.79 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 36.879518 ± 0.436763
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 80.79%
Mean ln(PPL(Q)/PPL(base)) : 1.405689 ± 0.007174
Mean PPL(Q)/PPL(base) : 4.078336 ± 0.029260
Mean PPL(Q)-PPL(base) : 27.836733 ± 0.381333
====== KL divergence statistics ======
Mean KLD: 1.190922 ± 0.004215
Maximum KLD: 25.421314
99.9% KLD: 13.739526
99.0% KLD: 8.520005
99.0% KLD: 8.520005
Median KLD: 0.710480
10.0% KLD: 0.017773
5.0% KLD: 0.003969
1.0% KLD: 0.000258
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -6.063 ± 0.075 %
Maximum Δp: 99.690%
99.9% Δp: 88.494%
99.0% Δp: 65.535%
95.0% Δp: 36.662%
90.0% Δp: 19.570%
75.0% Δp: 1.644%
Median Δp: -0.273%
25.0% Δp: -9.153%
10.0% Δp: -43.591%
5.0% Δp: -73.479%
1.0% Δp: -98.555%
0.1% Δp: -99.995%
Minimum Δp: -100.000%
RMS Δp : 29.578 ± 0.088 %
Same top p: 66.655 ± 0.123 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
Final result: 40.1333 +/- 1.7910
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 557.58 ms
llama_perf_context_print: prompt eval time = 133950.32 ms / 50067 tokens ( 2.68 ms per token, 373.77 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 136572.17 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-IQ4_NL.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 68.5333 +/- 1.6968
llama_perf_context_print: load time = 558.43 ms
llama_perf_context_print: prompt eval time = 56861.97 ms / 21857 tokens ( 2.60 ms per token, 384.39 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 57651.50 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
Final result: 67.8667 +/- 1.7063
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2488.69 ms
llama_perf_context_print: prompt eval time = 93489.45 ms / 35749 tokens ( 2.62 ms per token, 382.39 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 94938.12 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
750 75.73333333% [72.5391%, 78.6653%]
llama_perf_context_print: load time = 525.75 ms
llama_perf_context_print: prompt eval time = 338885.65 ms / 125814 tokens ( 2.69 ms per token, 371.26 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 344811.83 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
Final result: 42.5333 +/- 1.8065
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 523.91 ms
llama_perf_context_print: prompt eval time = 174640.97 ms / 67555 tokens ( 2.59 ms per token, 386.82 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 176748.74 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 33.537202 ± 0.371564
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 79.10%
Mean ln(PPL(Q)/PPL(base)) : 1.310688 ± 0.006832
Mean PPL(Q)/PPL(base) : 3.708725 ± 0.025338
Mean PPL(Q)-PPL(base) : 24.494416 ± 0.318026
====== KL divergence statistics ======
Mean KLD: 1.263077 ± 0.004253
Maximum KLD: 23.683029
99.9% KLD: 12.829600
99.0% KLD: 8.069357
99.0% KLD: 8.069357
Median KLD: 0.755642
10.0% KLD: 0.018898
5.0% KLD: 0.004577
1.0% KLD: 0.000375
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -8.200 ± 0.078 %
Maximum Δp: 99.800%
99.9% Δp: 90.394%
99.0% Δp: 66.393%
95.0% Δp: 34.806%
90.0% Δp: 16.907%
75.0% Δp: 0.806%
Median Δp: -0.550%
25.0% Δp: -12.684%
10.0% Δp: -50.926%
5.0% Δp: -77.962%
1.0% Δp: -98.305%
0.1% Δp: -99.990%
Minimum Δp: -100.000%
RMS Δp : 31.084 ± 0.088 %
Same top p: 62.253 ± 0.126 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
Final result: 40.2667 +/- 1.7920
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 525.48 ms
llama_perf_context_print: prompt eval time = 136845.27 ms / 50067 tokens ( 2.73 ms per token, 365.87 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 139475.22 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_L.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 67.0667 +/- 1.7172
llama_perf_context_print: load time = 528.70 ms
llama_perf_context_print: prompt eval time = 58406.45 ms / 21857 tokens ( 2.67 ms per token, 374.22 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 59199.95 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
Final result: 68.6667 +/- 1.6949
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2459.83 ms
llama_perf_context_print: prompt eval time = 93086.65 ms / 35749 tokens ( 2.60 ms per token, 384.04 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 94546.19 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
750 75.20000000% [71.9861%, 78.1570%]
llama_perf_context_print: load time = 506.32 ms
llama_perf_context_print: prompt eval time = 333957.19 ms / 125814 tokens ( 2.65 ms per token, 376.74 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 339911.25 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
Final result: 44.4000 +/- 1.8155
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 510.44 ms
llama_perf_context_print: prompt eval time = 172922.20 ms / 67555 tokens ( 2.56 ms per token, 390.67 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 175045.07 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 34.029948 ± 0.376082
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 78.62%
Mean ln(PPL(Q)/PPL(base)) : 1.325274 ± 0.006875
Mean PPL(Q)/PPL(base) : 3.763215 ± 0.025870
Mean PPL(Q)-PPL(base) : 24.987163 ± 0.322907
====== KL divergence statistics ======
Mean KLD: 1.300974 ± 0.004321
Maximum KLD: 23.795326
99.9% KLD: 12.854443
99.0% KLD: 8.215378
99.0% KLD: 8.215378
Median KLD: 0.782525
10.0% KLD: 0.019878
5.0% KLD: 0.004841
1.0% KLD: 0.000391
Minimum KLD: -0.000001
====== Token probability statistics ======
Mean Δp: -8.606 ± 0.079 %
Maximum Δp: 99.896%
99.9% Δp: 91.072%
99.0% Δp: 67.090%
95.0% Δp: 34.759%
90.0% Δp: 16.501%
75.0% Δp: 0.726%
Median Δp: -0.601%
25.0% Δp: -13.523%
10.0% Δp: -52.748%
5.0% Δp: -79.067%
1.0% Δp: -98.555%
0.1% Δp: -99.990%
Minimum Δp: -100.000%
RMS Δp : 31.537 ± 0.088 %
Same top p: 61.420 ± 0.127 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
Final result: 40.0000 +/- 1.7900
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 504.87 ms
llama_perf_context_print: prompt eval time = 134889.64 ms / 50067 tokens ( 2.69 ms per token, 371.17 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 137522.13 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_M.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 65.8667 +/- 1.7325
llama_perf_context_print: load time = 501.08 ms
llama_perf_context_print: prompt eval time = 56606.33 ms / 21857 tokens ( 2.59 ms per token, 386.12 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 57401.53 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
Final result: 67.6000 +/- 1.7100
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2303.10 ms
llama_perf_context_print: prompt eval time = 91534.38 ms / 35749 tokens ( 2.56 ms per token, 390.55 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 92983.34 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
750 75.20000000% [71.9861%, 78.1570%]
llama_perf_context_print: load time = 501.43 ms
llama_perf_context_print: prompt eval time = 325090.51 ms / 125814 tokens ( 2.58 ms per token, 387.01 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 331071.94 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
Final result: 43.7333 +/- 1.8126
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 508.75 ms
llama_perf_context_print: prompt eval time = 171219.64 ms / 67555 tokens ( 2.53 ms per token, 394.55 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 173347.99 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 36.103787 ± 0.402250
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 77.13%
Mean ln(PPL(Q)/PPL(base)) : 1.384430 ± 0.007124
Mean PPL(Q)/PPL(base) : 3.992551 ± 0.028445
Mean PPL(Q)-PPL(base) : 27.061001 ± 0.350072
====== KL divergence statistics ======
Mean KLD: 1.395578 ± 0.004559
Maximum KLD: 24.088387
99.9% KLD: 13.768757
99.0% KLD: 8.603606
99.0% KLD: 8.603606
Median KLD: 0.858334
10.0% KLD: 0.021285
5.0% KLD: 0.005174
1.0% KLD: 0.000393
Minimum KLD: -0.000002
====== Token probability statistics ======
Mean Δp: -9.212 ± 0.080 %
Maximum Δp: 99.844%
99.9% Δp: 91.248%
99.0% Δp: 67.876%
95.0% Δp: 34.787%
90.0% Δp: 16.139%
75.0% Δp: 0.593%
Median Δp: -0.696%
25.0% Δp: -14.997%
10.0% Δp: -54.930%
5.0% Δp: -79.831%
1.0% Δp: -98.568%
0.1% Δp: -99.992%
Minimum Δp: -100.000%
RMS Δp : 32.089 ± 0.088 %
Same top p: 59.830 ± 0.128 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
Final result: 39.8667 +/- 1.7890
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 501.19 ms
llama_perf_context_print: prompt eval time = 132421.97 ms / 50067 tokens ( 2.64 ms per token, 378.09 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 135070.83 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q3_K_S.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 66.5333 +/- 1.7242
llama_perf_context_print: load time = 501.82 ms
llama_perf_context_print: prompt eval time = 56390.15 ms / 21857 tokens ( 2.58 ms per token, 387.60 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 57175.02 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
Final result: 69.8667 +/- 1.6766
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2958.46 ms
llama_perf_context_print: prompt eval time = 95508.61 ms / 35749 tokens ( 2.67 ms per token, 374.30 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 97132.91 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
750 77.46666667% [74.3409%, 80.3125%]
llama_perf_context_print: load time = 567.97 ms
llama_perf_context_print: prompt eval time = 339481.51 ms / 125814 tokens ( 2.70 ms per token, 370.61 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 345440.37 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
Final result: 43.6000 +/- 1.8119
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 576.27 ms
llama_perf_context_print: prompt eval time = 176741.59 ms / 67555 tokens ( 2.62 ms per token, 382.22 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 178859.75 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 31.219358 ± 0.342580
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 80.83%
Mean ln(PPL(Q)/PPL(base)) : 1.239071 ± 0.006532
Mean PPL(Q)/PPL(base) : 3.452405 ± 0.022552
Mean PPL(Q)-PPL(base) : 22.176573 ± 0.287880
====== KL divergence statistics ======
Mean KLD: 1.106608 ± 0.003932
Maximum KLD: 20.797123
99.9% KLD: 12.310450
99.0% KLD: 7.690277
99.0% KLD: 7.690277
Median KLD: 0.642402
10.0% KLD: 0.017366
5.0% KLD: 0.004376
1.0% KLD: 0.000366
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -7.494 ± 0.075 %
Maximum Δp: 99.726%
99.9% Δp: 87.733%
99.0% Δp: 63.430%
95.0% Δp: 33.085%
90.0% Δp: 16.533%
75.0% Δp: 0.946%
Median Δp: -0.461%
25.0% Δp: -11.162%
10.0% Δp: -47.083%
5.0% Δp: -74.484%
1.0% Δp: -98.197%
0.1% Δp: -99.992%
Minimum Δp: -100.000%
RMS Δp : 29.739 ± 0.088 %
Same top p: 65.218 ± 0.124 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
Final result: 40.8000 +/- 1.7958
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 569.78 ms
llama_perf_context_print: prompt eval time = 138110.48 ms / 50067 tokens ( 2.76 ms per token, 362.51 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 140746.53 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_M.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 68.4000 +/- 1.6988
llama_perf_context_print: load time = 573.82 ms
llama_perf_context_print: prompt eval time = 59035.21 ms / 21857 tokens ( 2.70 ms per token, 370.24 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 59825.11 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
Final result: 69.8667 +/- 1.6766
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 2939.37 ms
llama_perf_context_print: prompt eval time = 94796.83 ms / 35749 tokens ( 2.65 ms per token, 377.11 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 96249.68 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
750 77.33333333% [74.2021%, 80.1860%]
llama_perf_context_print: load time = 567.43 ms
llama_perf_context_print: prompt eval time = 336113.08 ms / 125814 tokens ( 2.67 ms per token, 374.32 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 342078.71 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
Final result: 43.4667 +/- 1.8113
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 557.41 ms
llama_perf_context_print: prompt eval time = 175467.57 ms / 67555 tokens ( 2.60 ms per token, 385.00 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 177580.76 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 31.245217 ± 0.343080
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 80.86%
Mean ln(PPL(Q)/PPL(base)) : 1.239899 ± 0.006533
Mean PPL(Q)/PPL(base) : 3.455265 ± 0.022573
Mean PPL(Q)-PPL(base) : 22.202432 ± 0.288349
====== KL divergence statistics ======
Mean KLD: 1.105266 ± 0.003926
Maximum KLD: 20.987593
99.9% KLD: 12.310256
99.0% KLD: 7.698290
99.0% KLD: 7.698290
Median KLD: 0.643550
10.0% KLD: 0.017519
5.0% KLD: 0.004389
1.0% KLD: 0.000363
Minimum KLD: -0.000002
====== Token probability statistics ======
Mean Δp: -7.460 ± 0.075 %
Maximum Δp: 99.266%
99.9% Δp: 87.347%
99.0% Δp: 63.610%
95.0% Δp: 33.077%
90.0% Δp: 16.572%
75.0% Δp: 0.938%
Median Δp: -0.465%
25.0% Δp: -11.152%
10.0% Δp: -46.828%
5.0% Δp: -74.300%
1.0% Δp: -98.187%
0.1% Δp: -99.991%
Minimum Δp: -100.000%
RMS Δp : 29.698 ± 0.088 %
Same top p: 65.245 ± 0.124 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
Final result: 40.8000 +/- 1.7958
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 580.69 ms
llama_perf_context_print: prompt eval time = 137720.71 ms / 50067 tokens ( 2.75 ms per token, 363.54 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 140323.32 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q4_K_S.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 69.6000 +/- 1.6807
llama_perf_context_print: load time = 571.52 ms
llama_perf_context_print: prompt eval time = 58871.14 ms / 21857 tokens ( 2.69 ms per token, 371.27 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 59658.60 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
Final result: 68.8000 +/- 1.6929
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 3504.83 ms
llama_perf_context_print: prompt eval time = 95720.53 ms / 35749 tokens ( 2.68 ms per token, 373.47 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 97175.56 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
750 77.86666667% [74.7577%, 80.6916%]
llama_perf_context_print: load time = 627.41 ms
llama_perf_context_print: prompt eval time = 342154.54 ms / 125814 tokens ( 2.72 ms per token, 367.71 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 348106.77 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
Final result: 45.3333 +/- 1.8190
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 644.49 ms
llama_perf_context_print: prompt eval time = 177127.61 ms / 67555 tokens ( 2.62 ms per token, 381.39 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 179250.91 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 30.668770 ± 0.335889
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 81.60%
Mean ln(PPL(Q)/PPL(base)) : 1.221278 ± 0.006415
Mean PPL(Q)/PPL(base) : 3.391518 ± 0.021755
Mean PPL(Q)-PPL(base) : 21.625984 ± 0.280606
====== KL divergence statistics ======
Mean KLD: 1.039101 ± 0.003848
Maximum KLD: 22.598866
99.9% KLD: 12.543374
99.0% KLD: 7.697062
99.0% KLD: 7.697062
Median KLD: 0.596171
10.0% KLD: 0.017018
5.0% KLD: 0.004340
1.0% KLD: 0.000363
Minimum KLD: -0.000000
====== Token probability statistics ======
Mean Δp: -7.245 ± 0.074 %
Maximum Δp: 99.814%
99.9% Δp: 85.186%
99.0% Δp: 61.478%
95.0% Δp: 32.505%
90.0% Δp: 16.393%
75.0% Δp: 0.971%
Median Δp: -0.442%
25.0% Δp: -10.578%
10.0% Δp: -45.165%
5.0% Δp: -72.951%
1.0% Δp: -98.402%
0.1% Δp: -99.994%
Minimum Δp: -100.000%
RMS Δp : 29.177 ± 0.088 %
Same top p: 66.780 ± 0.123 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
Final result: 41.2000 +/- 1.7984
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 653.09 ms
llama_perf_context_print: prompt eval time = 137695.68 ms / 50067 tokens ( 2.75 ms per token, 363.61 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 140323.22 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_M.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 68.8000 +/- 1.6929
llama_perf_context_print: load time = 647.06 ms
llama_perf_context_print: prompt eval time = 58456.53 ms / 21857 tokens ( 2.67 ms per token, 373.90 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 59248.75 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
Final result: 68.9333 +/- 1.6909
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 3403.68 ms
llama_perf_context_print: prompt eval time = 94455.25 ms / 35749 tokens ( 2.64 ms per token, 378.48 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 95912.08 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
750 77.73333333% [74.6188%, 80.5653%]
llama_perf_context_print: load time = 627.98 ms
llama_perf_context_print: prompt eval time = 334461.40 ms / 125814 tokens ( 2.66 ms per token, 376.17 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 340403.29 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
Final result: 45.0667 +/- 1.8180
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 666.76 ms
llama_perf_context_print: prompt eval time = 174509.11 ms / 67555 tokens ( 2.58 ms per token, 387.11 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 176638.86 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 30.644285 ± 0.335716
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 81.62%
Mean ln(PPL(Q)/PPL(base)) : 1.220479 ± 0.006413
Mean PPL(Q)/PPL(base) : 3.388810 ± 0.021733
Mean PPL(Q)-PPL(base) : 21.601499 ± 0.280413
====== KL divergence statistics ======
Mean KLD: 1.038541 ± 0.003836
Maximum KLD: 22.424044
99.9% KLD: 12.381549
99.0% KLD: 7.687484
99.0% KLD: 7.687484
Median KLD: 0.597444
10.0% KLD: 0.016928
5.0% KLD: 0.004340
1.0% KLD: 0.000359
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -7.203 ± 0.074 %
Maximum Δp: 99.718%
99.9% Δp: 84.842%
99.0% Δp: 61.523%
95.0% Δp: 32.607%
90.0% Δp: 16.381%
75.0% Δp: 0.982%
Median Δp: -0.437%
25.0% Δp: -10.550%
10.0% Δp: -45.107%
5.0% Δp: -72.871%
1.0% Δp: -98.353%
0.1% Δp: -99.993%
Minimum Δp: -100.000%
RMS Δp : 29.119 ± 0.088 %
Same top p: 66.765 ± 0.123 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
Final result: 41.3333 +/- 1.7993
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 633.84 ms
llama_perf_context_print: prompt eval time = 135548.17 ms / 50067 tokens ( 2.71 ms per token, 369.37 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 138178.93 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q5_K_S.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 68.6667 +/- 1.6949
llama_perf_context_print: load time = 639.95 ms
llama_perf_context_print: prompt eval time = 57785.71 ms / 21857 tokens ( 2.64 ms per token, 378.24 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 58579.25 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
Final result: 68.9333 +/- 1.6909
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 4062.76 ms
llama_perf_context_print: prompt eval time = 95561.27 ms / 35749 tokens ( 2.67 ms per token, 374.10 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 97004.29 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
750 78.66666667% [75.5926%, 81.4486%]
llama_perf_context_print: load time = 712.37 ms
llama_perf_context_print: prompt eval time = 341004.43 ms / 125814 tokens ( 2.71 ms per token, 368.95 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 346935.77 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
Final result: 45.4667 +/- 1.8194
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 755.78 ms
llama_perf_context_print: prompt eval time = 181685.88 ms / 67555 tokens ( 2.69 ms per token, 371.82 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 183820.78 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 30.107026 ± 0.329175
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 81.96%
Mean ln(PPL(Q)/PPL(base)) : 1.202791 ± 0.006352
Mean PPL(Q)/PPL(base) : 3.329397 ± 0.021147
Mean PPL(Q)-PPL(base) : 21.064240 ± 0.273651
====== KL divergence statistics ======
Mean KLD: 1.009233 ± 0.003785
Maximum KLD: 22.312464
99.9% KLD: 12.292943
99.0% KLD: 7.613831
99.0% KLD: 7.613831
Median KLD: 0.574097
10.0% KLD: 0.016605
5.0% KLD: 0.004251
1.0% KLD: 0.000354
Minimum KLD: -0.000000
====== Token probability statistics ======
Mean Δp: -7.022 ± 0.073 %
Maximum Δp: 99.488%
99.9% Δp: 84.278%
99.0% Δp: 60.487%
95.0% Δp: 32.046%
90.0% Δp: 16.402%
75.0% Δp: 1.023%
Median Δp: -0.418%
25.0% Δp: -10.160%
10.0% Δp: -43.946%
5.0% Δp: -72.034%
1.0% Δp: -98.340%
0.1% Δp: -99.994%
Minimum Δp: -100.000%
RMS Δp : 28.800 ± 0.088 %
Same top p: 67.476 ± 0.122 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
Final result: 40.8000 +/- 1.7958
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 719.91 ms
llama_perf_context_print: prompt eval time = 141181.13 ms / 50067 tokens ( 2.82 ms per token, 354.63 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 143821.77 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q6_K.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 68.6667 +/- 1.6949
llama_perf_context_print: load time = 720.44 ms
llama_perf_context_print: prompt eval time = 59839.70 ms / 21857 tokens ( 2.74 ms per token, 365.26 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 60641.60 ms / 21858 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
Final result: 68.8000 +/- 1.6929
Random chance: 25.0083 +/- 1.5824
llama_perf_context_print: load time = 5336.21 ms
llama_perf_context_print: prompt eval time = 90355.25 ms / 35749 tokens ( 2.53 ms per token, 395.65 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 91808.86 ms / 35750 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,12 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
750 78.40000000% [75.3141%, 81.1964%]
llama_perf_context_print: load time = 857.12 ms
llama_perf_context_print: prompt eval time = 326655.40 ms / 125814 tokens ( 2.60 ms per token, 385.16 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 332582.43 ms / 125815 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
Final result: 45.4667 +/- 1.8194
Random chance: 25.0000 +/- 1.5822
llama_perf_context_print: load time = 847.03 ms
llama_perf_context_print: prompt eval time = 167726.56 ms / 67555 tokens ( 2.48 ms per token, 402.77 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 169830.02 ms / 67556 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,37 @@
====== Perplexity statistics ======
Mean PPL(Q) : 30.014168 ± 0.327772
Mean PPL(base) : 9.042786 ± 0.071503
Cor(ln(PPL(Q)), ln(PPL(base))): 82.04%
Mean ln(PPL(Q)/PPL(base)) : 1.199702 ± 0.006332
Mean PPL(Q)/PPL(base) : 3.319129 ± 0.021018
Mean PPL(Q)-PPL(base) : 20.971382 ± 0.272199
====== KL divergence statistics ======
Mean KLD: 1.004325 ± 0.003776
Maximum KLD: 22.233902
99.9% KLD: 12.361748
99.0% KLD: 7.606510
99.0% KLD: 7.606510
Median KLD: 0.570352
10.0% KLD: 0.016693
5.0% KLD: 0.004282
1.0% KLD: 0.000351
Minimum KLD: -0.000003
====== Token probability statistics ======
Mean Δp: -7.018 ± 0.073 %
Maximum Δp: 99.437%
99.9% Δp: 83.997%
99.0% Δp: 60.615%
95.0% Δp: 32.020%
90.0% Δp: 16.340%
75.0% Δp: 1.021%
Median Δp: -0.417%
25.0% Δp: -10.112%
10.0% Δp: -43.816%
5.0% Δp: -71.991%
1.0% Δp: -98.342%
0.1% Δp: -99.994%
Minimum Δp: -100.000%
RMS Δp : 28.772 ± 0.088 %
Same top p: 67.589 ± 0.122 %

View File

@@ -0,0 +1,13 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
Final result: 41.7333 +/- 1.8018
Random chance: 19.8992 +/- 1.4588
llama_perf_context_print: load time = 861.28 ms
llama_perf_context_print: prompt eval time = 133009.67 ms / 50067 tokens ( 2.66 ms per token, 376.42 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 135633.15 ms / 50068 tokens
ggml_metal_free: deallocating

View File

@@ -0,0 +1,11 @@
build: 5540 (291f2b69) with Apple clang version 17.0.0 (clang-1700.0.13.3) for arm64-apple-darwin24.4.0
llama_model_load_from_file_impl: using device Metal (Apple M4 Max) - 49151 MiB free
llama_model_loader: loaded meta data with 45 key-value pairs and 600 tensors from ./gemma-3-12b-it-Q8_0.gguf (version GGUF V3 (latest))
Final Winogrande score(750 tasks): 68.8000 +/- 1.6929
llama_perf_context_print: load time = 853.87 ms
llama_perf_context_print: prompt eval time = 55909.06 ms / 21857 tokens ( 2.56 ms per token, 390.94 tokens per second)
llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second)
llama_perf_context_print: total time = 56690.38 ms / 21858 tokens
ggml_metal_free: deallocating