初始化项目,由ModelHub XC社区提供模型

Model: appvoid/palmer-006
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-26 21:34:42 +08:00
commit e313f32a5d
23 changed files with 191677 additions and 0 deletions

View File

@@ -0,0 +1,157 @@
{
"model": "appvoid/palmer-006",
"tokenizer": "appvoid/palmer-006",
"params": 91131072,
"dataset": {
"path": "/home/appvoid/arithmark-3.jsonl",
"sha256": "bf8ab1a5193d52cdf0e05ff0b3ca226bdfcf416cb6e75562dcbe72e7e4559435",
"items": 1000
},
"runtime": {
"device": "cuda",
"dtype": "torch.bfloat16",
"batch_size": 32,
"max_context": 1024,
"attention_implementation": "auto",
"compiled": false,
"torch_version": "2.9.1+rocm7.2.0.git7e1940d4"
},
"primary_metric": "acc_norm",
"results": {
"arithmark-3": {
"acc": 53.2,
"acc_norm": 52.7,
"raw_correct": 532,
"norm_correct": 527,
"total": 1000,
"categories": {
"elementary_school_math_continuation::addition::grades_1_2::easy": {
"acc": 32.03125,
"acc_norm": 31.25,
"raw_correct": 41,
"norm_correct": 40,
"total": 128
},
"elementary_school_math_continuation::comparison::grades_2_3::medium": {
"acc": 40.909090909090914,
"acc_norm": 43.18181818181818,
"raw_correct": 18,
"norm_correct": 19,
"total": 44
},
"elementary_school_math_continuation::comparison_difference::grades_2_3::medium": {
"acc": 47.91666666666667,
"acc_norm": 45.83333333333333,
"raw_correct": 23,
"norm_correct": 22,
"total": 48
},
"elementary_school_math_continuation::data::grades_2_3::easy": {
"acc": 32.55813953488372,
"acc_norm": 30.23255813953488,
"raw_correct": 14,
"norm_correct": 13,
"total": 43
},
"elementary_school_math_continuation::division::grades_3_4::medium": {
"acc": 87.03703703703704,
"acc_norm": 87.03703703703704,
"raw_correct": 47,
"norm_correct": 47,
"total": 54
},
"elementary_school_math_continuation::fractions_counting::grades_3_4::medium": {
"acc": 28.000000000000004,
"acc_norm": 26.0,
"raw_correct": 14,
"norm_correct": 13,
"total": 50
},
"elementary_school_math_continuation::geometry_area::grades_4_5::medium": {
"acc": 94.23076923076923,
"acc_norm": 94.23076923076923,
"raw_correct": 49,
"norm_correct": 49,
"total": 52
},
"elementary_school_math_continuation::geometry_perimeter::grades_4_5::medium": {
"acc": 64.44444444444444,
"acc_norm": 64.44444444444444,
"raw_correct": 29,
"norm_correct": 29,
"total": 45
},
"elementary_school_math_continuation::measurement::grades_2_3::easy": {
"acc": 68.42105263157895,
"acc_norm": 67.10526315789474,
"raw_correct": 52,
"norm_correct": 51,
"total": 76
},
"elementary_school_math_continuation::money::grades_3_4::medium": {
"acc": 37.5,
"acc_norm": 37.5,
"raw_correct": 24,
"norm_correct": 24,
"total": 64
},
"elementary_school_math_continuation::multiplication::grades_3_4::medium": {
"acc": 81.08108108108108,
"acc_norm": 81.08108108108108,
"raw_correct": 60,
"norm_correct": 60,
"total": 74
},
"elementary_school_math_continuation::patterns::grades_3_4::medium": {
"acc": 64.15094339622641,
"acc_norm": 64.15094339622641,
"raw_correct": 34,
"norm_correct": 34,
"total": 53
},
"elementary_school_math_continuation::subtraction::grades_1_2::easy": {
"acc": 37.60683760683761,
"acc_norm": 37.60683760683761,
"raw_correct": 44,
"norm_correct": 44,
"total": 117
},
"elementary_school_math_continuation::time::grades_2_3::easy": {
"acc": 100.0,
"acc_norm": 100.0,
"raw_correct": 55,
"norm_correct": 55,
"total": 55
},
"elementary_school_math_continuation::two_step_add_subtract::grades_2_3::medium": {
"acc": 26.08695652173913,
"acc_norm": 26.08695652173913,
"raw_correct": 12,
"norm_correct": 12,
"total": 46
},
"elementary_school_math_continuation::two_step_addition::grades_2_3::medium": {
"acc": 42.10526315789473,
"acc_norm": 42.10526315789473,
"raw_correct": 8,
"norm_correct": 8,
"total": 19
},
"elementary_school_math_continuation::two_step_subtraction::grades_2_3::medium": {
"acc": 25.0,
"acc_norm": 21.875,
"raw_correct": 8,
"norm_correct": 7,
"total": 32
}
},
"timing": {
"tokenization_seconds": 0.19381939000004422,
"evaluation_seconds": 17.806661402999907,
"examples_per_second": 56.15875864475802
},
"primary_metric": "acc_norm",
"primary_acc": 52.7
}
}
}

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,94 @@
================================================================================
RUN 1 — ARC-Easy, ARC-Challenge, Winogrande, PIQA
================================================================================
2026-07-26:19:07:26 INFO [_cli.run:388] Selected Tasks: ['arc_easy', 'arc_challenge', 'winogrande', 'piqa']
2026-07-26:19:07:28 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
2026-07-26:19:07:28 INFO [evaluator:239] Initializing hf model, with arguments: {'pretrained': 'appvoid/palmer-006'}
2026-07-26:19:07:49 INFO [models.huggingface:286] Using device 'cuda:0'
config.json: 1.62kB [00:00, 3.38MB/s]
tokenizer_config.json: 99.6kB [00:00, 54.3MB/s]
tokenizer.json: 2.35MB [00:00, 11.4MB/s]
special_tokens_map.json: 7.42kB [00:00, 19.4MB/s]
2026-07-26:19:07:52 INFO [models.huggingface:579] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
model.safetensors: 100%|█████████████████████| 182M/182M [00:08<00:00, 22.1MB/s]
[transformers] The fast path is not available because one of `(selective_state_update, causal_conv1d_fn, causal_conv1d_update)` is None. Falling back to the naive implementation. To install follow https://github.com/state-spaces/mamba/#installation and https://github.com/Dao-AILab/causal-conv1d
Loading weights: 100%|██████████████████████| 386/386 [00:00<00:00, 2314.98it/s]
generation_config.json: 100%|███████████████████| 133/133 [00:00<00:00, 571kB/s]
README.md: 9.00kB [00:00, 161kB/s]
ARC-Easy/train-00000-of-00001.parquet: 100%|██| 331k/331k [00:00<00:00, 803kB/s]
ARC-Easy/test-00000-of-00001.parquet: 100%|██| 346k/346k [00:00<00:00, 1.63MB/s]
ARC-Easy/validation-00000-of-00001.parqu(…): 100%|█| 86.1k/86.1k [00:00<00:00, 4
Generating train split: 100%|█████| 2251/2251 [00:00<00:00, 24351.27 examples/s]
Generating test split: 100%|█████| 2376/2376 [00:00<00:00, 393371.21 examples/s]
Generating validation split: 100%|█| 570/570 [00:00<00:00, 154851.56 examples/s]
ARC-Challenge/train-00000-of-00001.parqu(…): 100%|█| 190k/190k [00:00<00:00, 896
ARC-Challenge/test-00000-of-00001.parque(…): 100%|█| 204k/204k [00:00<00:00, 959
ARC-Challenge/validation-00000-of-00001.(…): 100%|█| 55.7k/55.7k [00:00<00:00, 2
Generating train split: 100%|████| 1119/1119 [00:00<00:00, 177512.34 examples/s]
Generating test split: 100%|█████| 1172/1172 [00:00<00:00, 224575.10 examples/s]
Generating validation split: 100%|█| 299/299 [00:00<00:00, 101431.32 examples/s]
README.md: 11.2kB [00:00, 25.5MB/s]
winogrande_xl/train-00000-of-00001.parqu(…): 100%|█| 2.06M/2.06M [00:00<00:00, 4
winogrande_xl/test-00000-of-00001.parque(…): 100%|█| 118k/118k [00:00<00:00, 284
winogrande_xl/validation-00000-of-00001.(…): 100%|█| 85.9k/85.9k [00:00<00:00, 2
Generating train split: 100%|█| 40398/40398 [00:00<00:00, 1299756.78 examples/s]
Generating test split: 100%|█████| 1767/1767 [00:00<00:00, 430541.14 examples/s]
Generating validation split: 100%|█| 1267/1267 [00:00<00:00, 403108.79 examples/
piqa_train.parquet: 100%|██████████████████| 2.64M/2.64M [00:00<00:00, 3.25MB/s]
piqa_validation.parquet: 100%|████████████████| 300k/300k [00:00<00:00, 726kB/s]
piqa_test.parquet: 100%|█████████████████████| 496k/496k [00:00<00:00, 1.20MB/s]
Generating train split: 100%|██| 16113/16113 [00:00<00:00, 686092.14 examples/s]
Generating validation split: 100%|█| 1838/1838 [00:00<00:00, 398137.21 examples/
Generating test split: 100%|█████| 3084/3084 [00:00<00:00, 528054.93 examples/s]
2026-07-26:19:08:16 INFO [evaluator_utils:446] Selected tasks:
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: arc_challenge (arc/arc_challenge.yaml)
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: arc_easy (arc/arc_easy.yaml)
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: piqa (piqa/piqa.yaml)
2026-07-26:19:08:16 INFO [evaluator_utils:480] Task: winogrande (winogrande/default.yaml)
2026-07-26:19:08:16 INFO [api.task:312] Building contexts for arc_easy on rank 0...
100%|██████████████████████████████████████| 2376/2376 [00:04<00:00, 507.62it/s]
2026-07-26:19:08:21 INFO [api.task:312] Building contexts for arc_challenge on rank 0...
100%|██████████████████████████████████████| 1172/1172 [00:02<00:00, 505.21it/s]
2026-07-26:19:08:24 INFO [api.task:312] Building contexts for winogrande on rank 0...
100%|████████████████████████████████████| 1267/1267 [00:00<00:00, 73584.28it/s]
2026-07-26:19:08:24 INFO [api.task:312] Building contexts for piqa on rank 0...
100%|██████████████████████████████████████| 1838/1838 [00:02<00:00, 691.67it/s]
2026-07-26:19:08:26 INFO [evaluator:585] Running loglikelihood requests
Tokenizing inputs: 0%| | 0/20398 [00:00<?, ?it/s][transformers] Ignoring clean_up_tokenization_spaces=True for BPE tokenizer TokenizersBackend. The clean_up_tokenization post-processing step is designed for WordPiece tokenizers and is destructive for BPE (it strips spaces before punctuation). Set clean_up_tokenization_spaces=False to suppress this warning, or set clean_up_tokenization_spaces_for_bpe_even_though_it_will_corrupt_output=True to force cleanup anyway.
Tokenizing inputs: 100%|████████████████| 20398/20398 [00:06<00:00, 3197.87it/s]
Running loglikelihood requests: 100%|█████| 20398/20398 [04:21<00:00, 78.09it/s]
2026-07-26:19:13:03 INFO [loggers.evaluation_tracker:316] Output path not provided, skipping saving results aggregated
hf ({'pretrained': 'appvoid/palmer-006'}), gen_kwargs: ({}), limit: None, num_fewshot: None, batch_size: 16
| Tasks |Version|Filter|n-shot| Metric | |Value | |Stderr|
|-------------|------:|------|-----:|--------|---|-----:|---|-----:|
|arc_challenge| 1|none | 0|acc |↑ |0.2619|± |0.0128|
| | |none | 0|acc_norm|↑ |0.2927|± |0.0133|
|arc_easy | 1|none | 0|acc |↑ |0.5097|± |0.0103|
| | |none | 0|acc_norm|↑ |0.4482|± |0.0102|
|piqa | 1|none | 0|acc |↑ |0.6376|± |0.0112|
| | |none | 0|acc_norm|↑ |0.6360|± |0.0112|
|winogrande | 1|none | 0|acc |↑ |0.5036|± |0.0141|
================================================================================
RUN 2 — HellaSwag
================================================================================
2026-07-26:20:16:54 INFO [_cli.run:388] Selected Tasks: ['hellaswag']
2026-07-26:20:16:56 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
2026-07-26:20:16:56 INFO [evaluator:239] Initializing hf model, with arguments: {'pretrained': 'appvoid/palmer-006'}
2026-07-26:20:17:01 INFO [models.huggingface:286] Using device 'cuda:0'
2026-07-26:20:17:02 INFO [models.huggingface:579] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
The fast path is not available because one of `(selective_state_update, causal_conv1d_fn, causal_conv1d_update)` is None. Falling back to the naive implementation. To install follow https://github.com/state-spaces/mamba/#installation and https://github.com/Dao-AILab/causal-conv1d
Loading weights: 100%|█| 386/386 [00:00<00:00, 1408.36it/s, Materializing param=
2026-07-26:20:17:11 INFO [evaluator_utils:446] Selected tasks:
2026-07-26:20:17:11 INFO [evaluator_utils:480] Task: hellaswag (hellaswag/hellaswag.yaml)
2026-07-26:20:17:11 INFO [api.task:312] Building contexts for hellaswag on rank 0...
100%|███████████████████████████████████| 10042/10042 [00:07<00:00, 1317.53it/s]
2026-07-26:20:17:19 INFO [evaluator:585] Running loglikelihood requests
Tokenizing inputs: 100%|████████████████| 40168/40168 [00:23<00:00, 1726.55it/s]
Running loglikelihood requests: 100%|█████| 40168/40168 [17:38<00:00, 37.96it/s]
2026-07-26:20:35:31 INFO [loggers.evaluation_tracker:316] Output path not provided, skipping saving results aggregated
hf ({'pretrained': 'appvoid/palmer-006'}), gen_kwargs: ({}), limit: None, num_fewshot: None, batch_size: 64
| Tasks |Version|Filter|n-shot| Metric | |Value | |Stderr|
|---------|------:|------|-----:|--------|---|-----:|---|-----:|
|hellaswag| 1|none | 0|acc |↑ |0.3299|± |0.0047|
| | |none | 0|acc_norm|↑ |0.3841|± |0.0049|