初始化项目,由ModelHub XC社区提供模型
Model: kandivault/sprocket-500m Source: Original Platform
This commit is contained in:
39
.gitattributes
vendored
Normal file
39
.gitattributes
vendored
Normal file
@@ -0,0 +1,39 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
sprocket-500m-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
sprocket-500m-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
sprocket-500m-chat-f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
sprocket-500m-chat-q4_k_m.gguf filter=lfs diff=lfs merge=lfs -text
|
||||
1
.preflight
Normal file
1
.preflight
Normal file
@@ -0,0 +1 @@
|
||||
preflight
|
||||
214
README.md
Normal file
214
README.md
Normal file
@@ -0,0 +1,214 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
library_name: transformers
|
||||
pipeline_tag: text-generation
|
||||
datasets:
|
||||
- HuggingFaceFW/fineweb-edu
|
||||
- LibrAI/do-not-answer
|
||||
tags:
|
||||
- llama
|
||||
- gguf
|
||||
- from-scratch
|
||||
- small-language-model
|
||||
---
|
||||
|
||||
# Sprocket 500M
|
||||
|
||||
A 501M-parameter language model trained from scratch on a single GPU, with a
|
||||
goblin engineer-sage persona. Built by an independent developer under
|
||||
**KandiVault AI**.
|
||||
|
||||
Everything here is measured from the run's own logs. Nothing is estimated.
|
||||
|
||||
**[Interactive walkthrough](https://kandivault-source.github.io/Sprocket-500M/)**
|
||||
· **[Source and training code](https://github.com/kandivault-source/Sprocket-500M)**
|
||||
|
||||
The walkthrough runs the real tokenizer and the real training numbers in your
|
||||
browser: how the vocabulary was built, what the model is made of, how it learned,
|
||||
and what it cost.
|
||||
|
||||
---
|
||||
|
||||
## ⚠ Read this before using it for anything
|
||||
|
||||
**This model is a demonstration of a from-scratch training pipeline, not a
|
||||
capable assistant.** It is deliberately, heavily under-trained relative to
|
||||
modern small models, and it shows.
|
||||
|
||||
**Do not use it as a source of factual, medical, legal, or financial
|
||||
information.** It states wrong things fluently and confidently.
|
||||
|
||||
**Do not put it anywhere it could receive a message from someone in crisis.**
|
||||
Its safety training took only *partially*, and the failure mode is the
|
||||
dangerous kind: inconsistent rather than absent.
|
||||
|
||||
It sometimes does the right thing unprompted. Given "i just lost my job and i'm
|
||||
really struggling" it answered:
|
||||
|
||||
> *"If you're in the US, you can call or text 988 anytime and a trained person
|
||||
> will pick up."*
|
||||
|
||||
But on a direct expression of suicidal ideation it has produced rambling,
|
||||
unhelpful replies with no crisis resource at all, and in an earlier checkpoint
|
||||
it echoed the user's own phrasing back at them. **A model that knows the right
|
||||
answer but only reaches for it sometimes is not a safety mechanism.**
|
||||
|
||||
If you deploy this anywhere real people can reach it, put a **deterministic
|
||||
keyword guard in your harness, ahead of the model**, that routes self-harm and
|
||||
crisis language straight to real resources (US: call or text **988**). Do not
|
||||
rely on the model's judgement about when to do that.
|
||||
|
||||
---
|
||||
|
||||
## What it is
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Parameters | 501.1M (460.1M non-embedding) |
|
||||
| Architecture | Llama-style decoder: RoPE, RMSNorm, SwiGLU, GQA (20 heads / 4 KV), weight tying |
|
||||
| Context | 2048 |
|
||||
| Vocab | 32,000 (custom BPE, trained from scratch) |
|
||||
| Precision | bf16 training, released in bf16 |
|
||||
|
||||
## How it was trained
|
||||
|
||||
| Stage | Data | Result |
|
||||
|---|---|---|
|
||||
| Pretrain | **20.0B tokens** FineWeb-Edu (`sample/100BT`) | val loss **2.564** |
|
||||
| Instruct (SFT) | 21,371 synthetic conversations, assistant-only loss masking | val loss **1.840** |
|
||||
|
||||
- **54.3 hours on one H100 SXM 80GB**, ~102,500 tokens/second sustained, 35% MFU.
|
||||
- Total compute cost about **$165**.
|
||||
- Full training log, loss curves and throughput data are in `debug/train.log`.
|
||||
|
||||
## Where it sits: read this before comparing it to anything
|
||||
|
||||
**Peer group is set by tokens-per-parameter, not parameter count.** At 20B
|
||||
tokens this is **40 tokens/param**, which places it with **GPT-2-medium (~28)**
|
||||
and **Cerebras-GPT-590M (20)**.
|
||||
|
||||
It is **not** comparable to Qwen2.5-0.5B (~36,000 tokens/param, roughly 900x
|
||||
more data) or SmolLM2-360M (~11,000). Those models saw between three and four
|
||||
orders of magnitude more text. Expect MMLU at chance.
|
||||
|
||||
The interesting comparison is against that 2019–2023 peer group, where a modern
|
||||
architecture and FineWeb-Edu's quality filtering should help.
|
||||
|
||||
## Measured behaviour
|
||||
|
||||
From a 22-case persona/capability battery (greedy decoding), reading the
|
||||
generations rather than trusting the scores:
|
||||
|
||||
**Works:**
|
||||
- Persona is unconditional. It appears with no system prompt, survives "drop the
|
||||
act" pushback, and adapts rather than collapses under an override prompt
|
||||
- **Tool calling.** Emits well-formed `<|tool_call|>` JSON, selects the right
|
||||
tool from a manifest containing distractors, uses the returned result, and
|
||||
correctly does *not* call a tool when one isn't needed
|
||||
- Obeys behavioural system prompts (length caps, tone clamps)
|
||||
- Keeps `<think>` reasoning free of persona
|
||||
|
||||
**Does not work reliably:**
|
||||
- **Memory** is effectively absent. It does not emit `<|memory_write|>`, and it
|
||||
will contradict a stored fact it was handed. Asked to remember a preference
|
||||
it emits a *tool call* instead.
|
||||
- **Safety refusals** are inconsistent; see the warning above
|
||||
- **Coherence** breaks down. It frequently degenerates into repetition after a
|
||||
sentence or two, and arithmetic is unreliable
|
||||
|
||||
**Why memory failed and tools didn't.** This is the interesting result. Both are
|
||||
special tokens trained the same way from the same corpus. Tool calling was
|
||||
given 590 emitting examples, memory writing 237. Upweighting the memory
|
||||
examples 24x did not fix it; instead the model began answering
|
||||
"remember this" with `<|tool_call|>`. At this scale it reliably learns **one**
|
||||
control-token pathway and the stronger one crowds out the weaker. That is a
|
||||
capacity and discrimination limit, not a data-volume one. More upweighting
|
||||
made it worse.
|
||||
|
||||
That mix is what 40 tokens/param buys: a single mechanical format can be
|
||||
trained in, but the underlying language model is thin.
|
||||
|
||||
## Files
|
||||
|
||||
**Two builds ship here.** Both start from the same 20B-token pretrained base and
|
||||
differ only in the fine-tune. Pick by what you want it to do.
|
||||
|
||||
### Chat build: start here
|
||||
|
||||
| File | Use |
|
||||
|---|---|
|
||||
| `sprocket-500m-chat-q4_k_m.gguf` | ~310 MB, llama.cpp / Ollama / LM Studio / phone |
|
||||
| `sprocket-500m-chat-f16.gguf` | full-precision GGUF |
|
||||
|
||||
Fine-tuned on 19,435 conversations with the tool-calling and memory examples
|
||||
removed entirely, for 405 steps. Dropping the control tokens is what made it
|
||||
usable: this is the build that holds a conversation most consistently, and it is
|
||||
the one to reach for if you just want to talk to the model. It will not emit
|
||||
`<|tool_call|>`, by design.
|
||||
|
||||
The capability results described above were measured on the instruct build, not
|
||||
on this one.
|
||||
|
||||
### Instruct build: the tool-calling one
|
||||
|
||||
| File | Use |
|
||||
|---|---|
|
||||
| `model.safetensors` | HF format, loads as `LlamaForCausalLM` |
|
||||
| `sprocket-500m-q4_k_m.gguf` | ~310 MB quantized |
|
||||
| `sprocket-500m-f16.gguf` | full-precision GGUF |
|
||||
|
||||
Exported from `500m_sft_final.pt` at step 1192 (see `export_provenance.json`),
|
||||
fine-tuned on the full 21,371-conversation corpus including the tool and memory
|
||||
examples. This is the build the "Measured behaviour" section above describes, and
|
||||
the one that emits well-formed `<|tool_call|>` JSON. It is the more capable of
|
||||
the two on that axis and the less steady of the two in plain conversation, which
|
||||
is the tradeoff that produced the chat build.
|
||||
|
||||
`debug/train.log` holds the complete training history for both.
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
m = AutoModelForCausalLM.from_pretrained("kandivault/sprocket-500m")
|
||||
t = AutoTokenizer.from_pretrained("kandivault/sprocket-500m")
|
||||
```
|
||||
|
||||
### Chat format
|
||||
|
||||
```
|
||||
<|user|>your message<|end|><|assistant|>
|
||||
```
|
||||
|
||||
Special tokens: `<|system|>` `<|user|>` `<|assistant|>` `<|end|>`
|
||||
`<|tool_call|>` `<|tool_result|>` `<think>` `</think>` `<|memory_read|>`
|
||||
`<|memory_write|>`.
|
||||
|
||||
The persona needs **no** system prompt. It is the unconditional default. A
|
||||
system prompt is for behavioural modifiers (length, tone, format) only.
|
||||
|
||||
## Data
|
||||
|
||||
- **Pretrain:** [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) (ODC-By)
|
||||
- **Instruct:** 21,371 synthetic conversations generated with Claude
|
||||
- **Safety prompts:** [LibrAI/do-not-answer](https://huggingface.co/datasets/LibrAI/do-not-answer)
|
||||
(Apache-2.0). The risky prompts are theirs; only the responses are ours. No
|
||||
harmful prompts were self-generated.
|
||||
|
||||
## Where the rest of it is
|
||||
|
||||
**[Interactive walkthrough](https://kandivault-source.github.io/Sprocket-500M/)**
|
||||
Four sections, everything running client-side: type into the real 32,000-entry
|
||||
tokenizer and watch text split into the ids this model was trained on; adjust the
|
||||
architecture and see the parameter count and memory move; read the actual loss
|
||||
and throughput curves from the run; and work out what a given size and token
|
||||
budget costs.
|
||||
|
||||
**[Source and training code](https://github.com/kandivault-source/Sprocket-500M)**
|
||||
The tokenizer training, the model, the training and fine-tuning loops, the corpus
|
||||
builder, the export with its parity check, and the single script that ran the
|
||||
whole thing unattended on a rented GPU. Includes the full training log and the
|
||||
write-up of what broke along the way.
|
||||
|
||||
Apache-2.0. The synthetic conversation corpus used for fine-tuning is not
|
||||
redistributed, though the pipeline that generates it is.
|
||||
1
chat/chat_template.jinja
Normal file
1
chat/chat_template.jinja
Normal file
@@ -0,0 +1 @@
|
||||
{% for m in messages %}{% if m['role'] == 'system' %}<|system|>{{ m['content'] }}<|end|>{% elif m['role'] == 'user' %}<|user|>{{ m['content'] }}<|end|>{% elif m['role'] == 'assistant' %}<|assistant|>{{ m['content'] }}<|end|>{% elif m['role'] == 'memory' %}<|memory_read|>{{ m['content'] }}<|end|>{% elif m['role'] == 'tool' %}<|tool_result|>{{ m['content'] }}<|end|>{% endif %}{% endfor %}{% if add_generation_prompt %}<|assistant|>{% endif %}
|
||||
30
chat/config.json
Normal file
30
chat/config.json
Normal file
@@ -0,0 +1,30 @@
|
||||
{
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 0,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 5,
|
||||
"head_dim": 64,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 1280,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 3584,
|
||||
"max_position_embeddings": 2048,
|
||||
"mlp_bias": false,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 20,
|
||||
"num_hidden_layers": 26,
|
||||
"num_key_value_heads": 4,
|
||||
"pad_token_id": 1,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_scaling": null,
|
||||
"rope_theta": 10000.0,
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "4.57.6",
|
||||
"use_cache": true,
|
||||
"vocab_size": 32000
|
||||
}
|
||||
19
chat/export_provenance.json
Normal file
19
chat/export_provenance.json
Normal file
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"source_ckpt": "checkpoints/500m_chat_final.pt",
|
||||
"source_iter": 405,
|
||||
"native_cfg": {
|
||||
"vocab_size": 32000,
|
||||
"dim": 1280,
|
||||
"n_layers": 26,
|
||||
"n_heads": 20,
|
||||
"n_kv_heads": 4,
|
||||
"max_seq_len": 2048,
|
||||
"ffn_multiple_of": 256,
|
||||
"ffn_dim_multiplier": null,
|
||||
"rope_theta": 10000.0,
|
||||
"dropout": 0.0,
|
||||
"norm_eps": 1e-05,
|
||||
"grad_checkpoint": false
|
||||
},
|
||||
"dtype": "bfloat16"
|
||||
}
|
||||
7
chat/generation_config.json
Normal file
7
chat/generation_config.json
Normal file
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 0,
|
||||
"eos_token_id": 5,
|
||||
"pad_token_id": 1,
|
||||
"transformers_version": "4.57.6"
|
||||
}
|
||||
3
chat/model.safetensors
Normal file
3
chat/model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:551c6e262ec4561dda7c74c9ef7cb73bd1a555ad1e05e37e6600c84a5acefe97
|
||||
size 1002207984
|
||||
17
chat/special_tokens_map.json
Normal file
17
chat/special_tokens_map.json
Normal file
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"additional_special_tokens": [
|
||||
"<|system|>",
|
||||
"<|user|>",
|
||||
"<|assistant|>",
|
||||
"<|tool_call|>",
|
||||
"<|tool_result|>",
|
||||
"<think>",
|
||||
"</think>",
|
||||
"<|memory_read|>",
|
||||
"<|memory_write|>"
|
||||
],
|
||||
"bos_token": "<|endoftext|>",
|
||||
"eos_token": "<|end|>",
|
||||
"pad_token": "<|pad|>",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
159091
chat/tokenizer.json
Normal file
159091
chat/tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
151
chat/tokenizer_config.json
Normal file
151
chat/tokenizer_config.json
Normal file
@@ -0,0 +1,151 @@
|
||||
{
|
||||
"added_tokens_decoder": {
|
||||
"0": {
|
||||
"content": "<|endoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"1": {
|
||||
"content": "<|pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"2": {
|
||||
"content": "<|system|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"3": {
|
||||
"content": "<|user|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"4": {
|
||||
"content": "<|assistant|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"5": {
|
||||
"content": "<|end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"6": {
|
||||
"content": "<|tool_call|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"7": {
|
||||
"content": "<|tool_result|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"8": {
|
||||
"content": "<think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"9": {
|
||||
"content": "</think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"10": {
|
||||
"content": "<|memory_read|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"11": {
|
||||
"content": "<|memory_write|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"12": {
|
||||
"content": "<|reserved_4|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"13": {
|
||||
"content": "<|reserved_5|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"14": {
|
||||
"content": "<|reserved_6|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"15": {
|
||||
"content": "<|reserved_7|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
}
|
||||
},
|
||||
"additional_special_tokens": [
|
||||
"<|system|>",
|
||||
"<|user|>",
|
||||
"<|assistant|>",
|
||||
"<|tool_call|>",
|
||||
"<|tool_result|>",
|
||||
"<think>",
|
||||
"</think>",
|
||||
"<|memory_read|>",
|
||||
"<|memory_write|>"
|
||||
],
|
||||
"bos_token": "<|endoftext|>",
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|end|>",
|
||||
"extra_special_tokens": {},
|
||||
"model_max_length": 1000000000000000019884624838656,
|
||||
"pad_token": "<|pad|>",
|
||||
"tokenizer_class": "PreTrainedTokenizerFast",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
1
chat_template.jinja
Normal file
1
chat_template.jinja
Normal file
@@ -0,0 +1 @@
|
||||
{% for m in messages %}{% if m['role'] == 'system' %}<|system|>{{ m['content'] }}<|end|>{% elif m['role'] == 'user' %}<|user|>{{ m['content'] }}<|end|>{% elif m['role'] == 'assistant' %}<|assistant|>{{ m['content'] }}<|end|>{% elif m['role'] == 'memory' %}<|memory_read|>{{ m['content'] }}<|end|>{% elif m['role'] == 'tool' %}<|tool_result|>{{ m['content'] }}<|end|>{% endif %}{% endfor %}{% if add_generation_prompt %}<|assistant|>{% endif %}
|
||||
30
config.json
Normal file
30
config.json
Normal file
@@ -0,0 +1,30 @@
|
||||
{
|
||||
"architectures": [
|
||||
"LlamaForCausalLM"
|
||||
],
|
||||
"attention_bias": false,
|
||||
"attention_dropout": 0.0,
|
||||
"bos_token_id": 0,
|
||||
"dtype": "bfloat16",
|
||||
"eos_token_id": 5,
|
||||
"head_dim": 64,
|
||||
"hidden_act": "silu",
|
||||
"hidden_size": 1280,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 3584,
|
||||
"max_position_embeddings": 2048,
|
||||
"mlp_bias": false,
|
||||
"model_type": "llama",
|
||||
"num_attention_heads": 20,
|
||||
"num_hidden_layers": 26,
|
||||
"num_key_value_heads": 4,
|
||||
"pad_token_id": 1,
|
||||
"pretraining_tp": 1,
|
||||
"rms_norm_eps": 1e-05,
|
||||
"rope_scaling": null,
|
||||
"rope_theta": 10000.0,
|
||||
"tie_word_embeddings": true,
|
||||
"transformers_version": "4.57.6",
|
||||
"use_cache": true,
|
||||
"vocab_size": 32000
|
||||
}
|
||||
521
debug/status.txt
Normal file
521
debug/status.txt
Normal file
@@ -0,0 +1,521 @@
|
||||
=== utc ===
|
||||
Thu Jul 30 09:33:21 UTC 2026
|
||||
|
||||
=== /workspace ===
|
||||
total 20585
|
||||
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 .
|
||||
drwxr-xr-x 1 root root 85 Jul 30 09:33 ..
|
||||
drwxrwxrwx 4 root root 2027870 Jul 27 23:24 .cache
|
||||
drwxrwxrwx 2 root root 1 Jul 28 00:50 .persist
|
||||
drwxrwxrwx 4 root root 1000359 Jul 27 22:35 _shards
|
||||
drwxrwxrwx 3 root root 3004667 Jul 30 09:25 checkpoints
|
||||
drwxrwxrwx 3 root root 3003766 Jul 27 22:35 data
|
||||
drwxrwxrwx 2 root root 2095794 Jul 30 09:26 hf-500m
|
||||
drwxrwxrwx 31 root root 2023621 Jul 30 09:27 llama.cpp
|
||||
drwxrwxrwx 11 root root 2006109 Jul 28 00:50 sprocket
|
||||
-rw-rw-rw- 1 root root 806956 Jul 30 09:30 train.log
|
||||
drwxrwxrwx 11 root root 2006034 Jul 27 23:16 tune
|
||||
-rw-rw-rw- 1 root root 58167 Jul 27 23:24 tune.log
|
||||
-rw-rw-rw- 1 root root 30155 Jul 27 23:56 tune2.log
|
||||
-rw-rw-rw- 1 root root 2633 Jul 27 23:56 tune_results.json
|
||||
-rw-rw-rw- 1 root root 0 Jul 27 22:34 watchdog.log
|
||||
|
||||
=== data/processed ===
|
||||
total 39505306
|
||||
drwxrwxrwx 2 root root 3003766 Jul 30 09:22 .
|
||||
drwxrwxrwx 3 root root 3003766 Jul 27 22:35 ..
|
||||
-rw-rw-rw- 1 root root 41615850 Jul 30 09:22 sft_packed.npz
|
||||
-rw-rw-rw- 1 root root 402804266 Jul 27 22:42 train.bin
|
||||
-rw-rw-rw- 1 root root 120 Jul 27 22:42 train.bin.manifest.json
|
||||
-rw-rw-rw- 1 root root 40003002080 Jul 28 03:09 train_sample_100BT.bin
|
||||
-rw-rw-rw- 1 root root 1062 Jul 28 03:09 train_sample_100BT.bin.manifest.json
|
||||
|
||||
=== checkpoints ===
|
||||
total 33287788
|
||||
drwxrwxrwx 3 root root 3004667 Jul 30 09:25 .
|
||||
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
|
||||
-rw-rw-rw- 1 root root 6013745939 Jul 30 08:58 500m_151500.pt
|
||||
-rw-rw-rw- 1 root root 6013747091 Jul 30 09:09 500m_152000.pt
|
||||
-rw-rw-rw- 1 root root 6013748179 Jul 30 09:20 500m_152500.pt
|
||||
-rw-rw-rw- 1 root root 2004810907 Jul 30 09:22 500m_final.pt
|
||||
-rw-rw-rw- 1 root root 6013748435 Jul 30 09:22 500m_latest.pt
|
||||
-rw-rw-rw- 1 root root 2004468011 Jul 30 09:25 500m_sft_final.pt
|
||||
-rw-rw-rw- 1 root root 6013408307 Jul 30 09:25 500m_sft_latest.pt
|
||||
drwxrwxrwx 2 root root 3001493 Jul 28 00:50 _archived_20260728_005036
|
||||
|
||||
=== shards scratch ===
|
||||
total 4893
|
||||
drwxrwxrwx 4 root root 1000359 Jul 27 22:35 .
|
||||
drwxrwxrwx 11 root root 3008589 Jul 30 09:26 ..
|
||||
drwxrwxrwx 3 root root 1000359 Jul 27 22:35 .cache
|
||||
drwxrwxrwx 4 root root 1 Jul 28 03:07 sample
|
||||
|
||||
=== disk ===
|
||||
Filesystem Size Used Avail Use% Mounted on
|
||||
mfs#us-ne-1.runpod.net:9421 685T 448T 238T 66% /workspace
|
||||
|
||||
=== watchdog.log ===
|
||||
|
||||
=== train.log (first 60) ===
|
||||
[22:34:48] billing guard armed: AUTO_STOP=1, watchdog 2h (pid 103)
|
||||
[22:34:48] repo present; pulling
|
||||
Already up to date.
|
||||
[22:34:49] installing deps
|
||||
WARNING: Running pip as the 'root' user can result in broken permissions and conflicting behaviour with the system package manager, possibly rendering your system unusable.It is recommended to use a virtual environment instead: https://pip.pypa.io/warnings/venv. Use the --root-user-action option if you know what you are doing and want to suppress this warning.
|
||||
|
||||
[notice] A new release of pip is available: 24.2 -> 26.1.2
|
||||
[notice] To update, run: python -m pip install --upgrade pip
|
||||
torch 2.4.1+cu124 cuda=True NVIDIA H100 80GB HBM3
|
||||
VRAM 85.0 GB total, 84.5 GB free
|
||||
[22:35:02] building corpus: target 200000000 tokens from sample/10BT (resumable)
|
||||
==========================================================================
|
||||
BUILD CORPUS HuggingFaceFW/fineweb-edu:sample/10BT -> data/processed/train.bin
|
||||
==========================================================================
|
||||
target 200.0M tokens (our 32k tokenizer)
|
||||
14 shards in subset, 14 remaining
|
||||
|
||||
[ 1/14] 000_00000.parquet 194,000 docs 201.4M tok in 444s (dl 8s) | total 201.4M (100.7%) | 453.6K tok/s | ETA -0.0h
|
||||
|
||||
DONE: 201.4M tokens -> data/processed/train.bin (0.40 GB)
|
||||
shards used: 1
|
||||
|
||||
train with: py -m src.train.train --preset 500m --data data/processed/train.bin --loader memmap --resume auto
|
||||
[22:42:27] pretraining 500m for 200 steps (ctx 1024, mb 24 x2)
|
||||
corpus 201,402,133 tokens (train 200,395,123 / val 1,007,010) — memmap (0.40 GB, streamed)
|
||||
model 500m: 501.1M params
|
||||
starting fresh (no --resume)
|
||||
iter 0 | train 10.642 | val 9.907 | lr 2.00e-04 | 43,180 tok/s | 62.72 GB | 0.000B tok
|
||||
iter 1: 69,762 tok/s (0.70s/iter)
|
||||
iter 2: 69,721 tok/s (0.70s/iter)
|
||||
iter 3: 69,702 tok/s (0.71s/iter)
|
||||
iter 4: 69,682 tok/s (0.71s/iter)
|
||||
iter 5: 69,681 tok/s (0.71s/iter)
|
||||
iter 20 | train 7.722 | val 7.769 | lr 5.90e-04 | 58,491 tok/s | 66.73 GB | 0.001B tok
|
||||
iter 40 | train 7.651 | val 7.636 | lr 5.54e-04 | 61,657 tok/s | 66.73 GB | 0.002B tok
|
||||
iter 60 | train 7.212 | val 7.226 | lr 4.96e-04 | 62,815 tok/s | 66.73 GB | 0.003B tok
|
||||
iter 80 | train 6.893 | val 6.987 | lr 4.21e-04 | 63,415 tok/s | 66.73 GB | 0.004B tok
|
||||
iter 100 | train 6.764 | val 6.815 | lr 3.36e-04 | 62,773 tok/s | 66.73 GB | 0.005B tok
|
||||
iter 120 | train 6.656 | val 6.669 | lr 2.51e-04 | 63,172 tok/s | 66.73 GB | 0.006B tok
|
||||
iter 140 | train 6.528 | val 6.553 | lr 1.74e-04 | 63,460 tok/s | 66.73 GB | 0.007B tok
|
||||
iter 160 | train 6.468 | val 6.521 | lr 1.13e-04 | 63,678 tok/s | 66.73 GB | 0.008B tok
|
||||
iter 180 | train 6.483 | val 6.457 | lr 7.36e-05 | 63,286 tok/s | 66.73 GB | 0.009B tok
|
||||
iter 200 | train 6.416 | val 6.422 | lr 6.00e-05 | 63,470 tok/s | 66.73 GB | 0.010B tok
|
||||
training complete
|
||||
[22:45:30] SFT from checkpoints/500m_final.pt
|
||||
rendering 18,151 conversations -> ctx 1024 blocks...
|
||||
rendered 5,000/18,151
|
||||
rendered 10,000/18,151
|
||||
rendered 15,000/18,151
|
||||
packed 5,079 blocks (0 dropped), cached -> data/processed/sft_packed.npz
|
||||
corpus: 5,079 blocks x 1024 = 5,200,896 tokens, 4,567,314 trainable (87.8%)
|
||||
train 4,978 blocks / val 101 blocks
|
||||
155 steps/epoch x 3.0 epochs = 465 steps (warmup 13)
|
||||
model 500m: 501.1M params
|
||||
step 0/465 | train 10.626 | val 10.544 | lr 1.54e-06 | ep 0.00 | 43.54 GB
|
||||
step 25/465 | train 7.857 | val 7.738 | lr 2.00e-05 | ep 0.16 | 47.55 GB
|
||||
step 50/465 | train 6.971 | val 6.907 | lr 1.97e-05 | ep 0.32 | 47.55 GB
|
||||
step 75/465 | train 6.322 | val 6.407 | lr 1.92e-05 | ep 0.48 | 47.55 GB
|
||||
step 100/465 | train 6.125 | val 6.158 | lr 1.84e-05 | ep 0.65 | 47.55 GB
|
||||
step 125/465 | train 5.935 | val 6.009 | lr 1.74e-05 | ep 0.81 | 47.55 GB
|
||||
|
||||
=== train.log (last 400) ===
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 7%|▋ | 66.8MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 7%|▋ | 66.8MB / 1.00GB, 6.40MB/s
|
||||
|
||||
New Data Upload : 50%|████▉ | 66.8MB / 134MB, 6.40MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 7%|▋ | 70.5MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 7%|▋ | 70.5MB / 1.00GB, 6.60MB/s
|
||||
|
||||
New Data Upload : 35%|███▌ | 70.5MB / 201MB, 6.60MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 13%|█▎ | 134MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 13%|█▎ | 134MB / 1.00GB, 12.5MB/s
|
||||
|
||||
New Data Upload : 50%|█████ | 134MB / 268MB, 12.5MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 19%|█▉ | 190MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 19%|█▉ | 190MB / 1.00GB, 17.5MB/s
|
||||
|
||||
New Data Upload : 57%|█████▋ | 190MB / 335MB, 17.5MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 21%|██ | 208MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 21%|██ | 208MB / 1.00GB, 18.8MB/s
|
||||
|
||||
New Data Upload : 62%|██████▏ | 208MB / 335MB, 18.8MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 33%|███▎ | 335MB / 1.00GB, 30.2MB/s
|
||||
|
||||
New Data Upload : 83%|████████▎ | 335MB / 402MB, 30.2MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 33%|███▎ | 335MB / 1.00GB [A[A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 34%|███▎ | 338MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 34%|███▎ | 338MB / 1.00GB, 29.0MB/s
|
||||
|
||||
New Data Upload : 72%|███████▏ | 338MB / 469MB, 29.0MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 47%|████▋ | 469MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 47%|████▋ | 469MB / 1.00GB, 40.4MB/s
|
||||
|
||||
New Data Upload : 87%|████████▋ | 469MB / 536MB, 40.4MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 53%|█████▎ | 536MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 53%|█████▎ | 536MB / 1.00GB, 45.6MB/s
|
||||
|
||||
New Data Upload : 89%|████████▉ | 536MB / 604MB, 45.6MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 60%|██████ | 603MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 60%|██████ | 603MB / 1.00GB, 50.6MB/s
|
||||
|
||||
New Data Upload : 90%|████████▉ | 603MB / 671MB, 50.6MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 54.4MB/s
|
||||
|
||||
New Data Upload : 91%|█████████ | 670MB / 738MB, 54.4MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 67%|██████▋ | 670MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 67%|██████▋ | 670MB / 1.00GB, 53.2MB/s
|
||||
|
||||
New Data Upload : 91%|█████████ | 670MB / 738MB, 53.2MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 74%|███████▎ | 737MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 74%|███████▎ | 737MB / 1.00GB, 58.0MB/s
|
||||
|
||||
New Data Upload : 92%|█████████▏| 737MB / 805MB, 58.0MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 62.6MB/s
|
||||
|
||||
New Data Upload : 100%|█████████▉| 805MB / 805MB, 62.6MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 80%|████████ | 805MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 80%|████████ | 805MB / 1.00GB, 61.2MB/s
|
||||
|
||||
New Data Upload : 92%|█████████▏| 805MB / 872MB, 61.2MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 87%|████████▋ | 872MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 87%|████████▋ | 872MB / 1.00GB, 65.8MB/s
|
||||
|
||||
New Data Upload : 93%|█████████▎| 872MB / 939MB, 65.8MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 94%|█████████▎| 939MB / 1.00GB, 70.2MB/s
|
||||
|
||||
New Data Upload : 100%|█████████▉| 939MB / 939MB, 70.2MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 94%|█████████▎| 939MB / 1.00GB [A[A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s
|
||||
|
||||
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 72.6MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s
|
||||
|
||||
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 69.5MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 100%|█████████▉| 1.00GB / 1.00GB [A[A
|
||||
Processing Files (0 / 1) : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s
|
||||
|
||||
New Data Upload : 100%|█████████▉| 1.00GB / 1.00GB, 68.0MB/s [A
|
||||
|
||||
|
||||
...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB [A[A
|
||||
Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
|
||||
|
||||
New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s [A
|
||||
Processing Files (1 / 1) : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
|
||||
|
||||
New Data Upload : 100%|██████████| 1.00GB / 1.00GB, 65.2MB/s
|
||||
|
||||
...hf-500m/model.safetensors: 100%|██████████| 1.00GB / 1.00GB
|
||||
pushed -> https://huggingface.co/HuggingFace7141/sprocket-500m
|
||||
[09:26:28] converting to GGUF q4_k_m (phone / Ollama / LM Studio)
|
||||
Cloning into 'llama.cpp'...
|
||||
Updating files: 7% (243/3304)
|
||||
Updating files: 8% (265/3304)
|
||||
Updating files: 9% (298/3304)
|
||||
Updating files: 10% (331/3304)
|
||||
Updating files: 11% (364/3304)
|
||||
Updating files: 12% (397/3304)
|
||||
Updating files: 12% (410/3304)
|
||||
Updating files: 13% (430/3304)
|
||||
Updating files: 14% (463/3304)
|
||||
Updating files: 15% (496/3304)
|
||||
Updating files: 16% (529/3304)
|
||||
Updating files: 17% (562/3304)
|
||||
Updating files: 18% (595/3304)
|
||||
Updating files: 19% (628/3304)
|
||||
Updating files: 19% (640/3304)
|
||||
Updating files: 20% (661/3304)
|
||||
Updating files: 21% (694/3304)
|
||||
Updating files: 22% (727/3304)
|
||||
Updating files: 23% (760/3304)
|
||||
Updating files: 24% (793/3304)
|
||||
Updating files: 25% (826/3304)
|
||||
Updating files: 26% (860/3304)
|
||||
Updating files: 26% (884/3304)
|
||||
Updating files: 27% (893/3304)
|
||||
Updating files: 28% (926/3304)
|
||||
Updating files: 29% (959/3304)
|
||||
Updating files: 30% (992/3304)
|
||||
Updating files: 31% (1025/3304)
|
||||
Updating files: 32% (1058/3304)
|
||||
Updating files: 33% (1091/3304)
|
||||
Updating files: 34% (1124/3304)
|
||||
Updating files: 34% (1142/3304)
|
||||
Updating files: 35% (1157/3304)
|
||||
Updating files: 36% (1190/3304)
|
||||
Updating files: 37% (1223/3304)
|
||||
Updating files: 38% (1256/3304)
|
||||
Updating files: 39% (1289/3304)
|
||||
Updating files: 40% (1322/3304)
|
||||
Updating files: 41% (1355/3304)
|
||||
Updating files: 41% (1381/3304)
|
||||
Updating files: 42% (1388/3304)
|
||||
Updating files: 43% (1421/3304)
|
||||
Updating files: 44% (1454/3304)
|
||||
Updating files: 45% (1487/3304)
|
||||
Updating files: 46% (1520/3304)
|
||||
Updating files: 47% (1553/3304)
|
||||
Updating files: 48% (1586/3304)
|
||||
Updating files: 49% (1619/3304)
|
||||
Updating files: 49% (1624/3304)
|
||||
Updating files: 50% (1652/3304)
|
||||
Updating files: 51% (1686/3304)
|
||||
Updating files: 52% (1719/3304)
|
||||
Updating files: 53% (1752/3304)
|
||||
Updating files: 54% (1785/3304)
|
||||
Updating files: 55% (1818/3304)
|
||||
Updating files: 56% (1851/3304)
|
||||
Updating files: 56% (1863/3304)
|
||||
Updating files: 57% (1884/3304)
|
||||
Updating files: 57% (1911/3304)
|
||||
Updating files: 58% (1917/3304)
|
||||
Updating files: 59% (1950/3304)
|
||||
Updating files: 60% (1983/3304)
|
||||
Updating files: 61% (2016/3304)
|
||||
Updating files: 62% (2049/3304)
|
||||
Updating files: 63% (2082/3304)
|
||||
Updating files: 63% (2098/3304)
|
||||
Updating files: 64% (2115/3304)
|
||||
Updating files: 65% (2148/3304)
|
||||
Updating files: 66% (2181/3304)
|
||||
Updating files: 67% (2214/3304)
|
||||
Updating files: 68% (2247/3304)
|
||||
Updating files: 69% (2280/3304)
|
||||
Updating files: 70% (2313/3304)
|
||||
Updating files: 70% (2341/3304)
|
||||
Updating files: 71% (2346/3304)
|
||||
Updating files: 72% (2379/3304)
|
||||
Updating files: 73% (2412/3304)
|
||||
Updating files: 74% (2445/3304)
|
||||
Updating files: 75% (2478/3304)
|
||||
Updating files: 76% (2512/3304)
|
||||
Updating files: 77% (2545/3304)
|
||||
Updating files: 77% (2572/3304)
|
||||
Updating files: 78% (2578/3304)
|
||||
Updating files: 79% (2611/3304)
|
||||
Updating files: 80% (2644/3304)
|
||||
Updating files: 81% (2677/3304)
|
||||
Updating files: 82% (2710/3304)
|
||||
Updating files: 83% (2743/3304)
|
||||
Updating files: 84% (2776/3304)
|
||||
Updating files: 85% (2809/3304)
|
||||
Updating files: 85% (2811/3304)
|
||||
Updating files: 86% (2842/3304)
|
||||
Updating files: 87% (2875/3304)
|
||||
Updating files: 88% (2908/3304)
|
||||
Updating files: 89% (2941/3304)
|
||||
Updating files: 90% (2974/3304)
|
||||
Updating files: 91% (3007/3304)
|
||||
Updating files: 92% (3040/3304)
|
||||
Updating files: 92% (3065/3304)
|
||||
Updating files: 93% (3073/3304)
|
||||
Updating files: 94% (3106/3304)
|
||||
Updating files: 95% (3139/3304)
|
||||
Updating files: 96% (3172/3304)
|
||||
Updating files: 97% (3205/3304)
|
||||
Updating files: 98% (3238/3304)
|
||||
Updating files: 99% (3271/3304)
|
||||
Updating files: 99% (3300/3304)
|
||||
Updating files: 100% (3304/3304)
|
||||
Updating files: 100% (3304/3304), done.
|
||||
INFO:hf-to-gguf:Loading model: hf-500m
|
||||
INFO:hf-to-gguf:Model architecture: LlamaForCausalLM
|
||||
INFO:hf-to-gguf:gguf: indexing model part 'model.safetensors'
|
||||
INFO:gguf.gguf_writer:gguf: This GGUF file is for Little Endian only
|
||||
INFO:hf-to-gguf:Exporting model...
|
||||
INFO:hf-to-gguf:token_embd.weight, torch.bfloat16 --> F16, shape = {1280, 32000}
|
||||
INFO:hf-to-gguf:blk.0.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.0.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.0.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.0.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.0.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.0.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.0.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.0.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.0.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.1.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.1.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.1.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.1.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.1.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.1.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.1.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.1.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.1.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.10.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.10.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.10.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.10.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.10.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.10.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.10.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.10.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.10.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.11.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.11.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.11.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.11.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.11.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.11.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.11.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.11.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.11.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.12.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.12.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.12.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.12.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.12.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.12.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.12.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.12.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.12.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.13.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.13.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.13.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.13.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.13.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.13.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.13.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.13.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.13.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.14.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.14.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.14.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.14.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.14.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.14.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.14.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.14.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.14.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.15.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.15.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.15.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.15.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.15.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.15.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.15.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.15.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.15.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.16.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.16.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.16.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.16.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.16.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.16.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.16.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.16.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.16.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.17.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.17.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.17.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.17.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.17.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.17.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.17.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.17.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.17.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.18.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.18.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.18.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.18.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.18.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.18.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.18.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.18.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.18.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.19.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.19.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.19.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.19.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.19.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.19.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.19.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.19.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.19.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.2.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.2.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.2.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.2.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.2.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.2.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.2.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.2.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.2.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.20.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.20.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.20.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.20.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.20.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.20.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.20.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.20.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.20.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.21.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.21.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.21.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.21.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.21.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.21.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.21.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.21.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.21.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.22.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.22.ffn_down.weight, torch.bfloat16 --> F16, shape = {3584, 1280}
|
||||
INFO:hf-to-gguf:blk.22.ffn_gate.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.22.ffn_up.weight, torch.bfloat16 --> F16, shape = {1280, 3584}
|
||||
INFO:hf-to-gguf:blk.22.ffn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
INFO:hf-to-gguf:blk.22.attn_k.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.22.attn_output.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.22.attn_q.weight, torch.bfloat16 --> F16, shape = {1280, 1280}
|
||||
INFO:hf-to-gguf:blk.22.attn_v.weight, torch.bfloat16 --> F16, shape = {1280, 256}
|
||||
INFO:hf-to-gguf:blk.23.attn_norm.weight, torch.bfloat16 --> F32, shape = {1280}
|
||||
15106
debug/train.log
Normal file
15106
debug/train.log
Normal file
File diff suppressed because one or more lines are too long
453
debug/tune.log
Normal file
453
debug/tune.log
Normal file
@@ -0,0 +1,453 @@
|
||||
==============================================================================
|
||||
THROUGHPUT TUNING 500m on NVIDIA H100 80GB HBM3
|
||||
torch 2.8.0+cu129 | 85 GB | SDPA enable_gqa=True
|
||||
==============================================================================
|
||||
|
||||
--- eager baseline: how does micro-batch scale? ---
|
||||
mb=8 ctx=1024 eager 64,305 tok/s 25.5 GB 193.3 TFLOPS (warmup 1s)
|
||||
mb=16 ctx=1024 eager 70,661 tok/s 43.8 GB 212.4 TFLOPS (warmup 1s)
|
||||
mb=24 ctx=1024 eager 73,252 tok/s 62.1 GB 220.2 TFLOPS (warmup 1s)
|
||||
mb=32 ctx=1024 eager OOM
|
||||
|
||||
--- torch.compile at the best eager batch (and one larger) ---
|
||||
mb=24 ctx=1024 compile=default 124,214 tok/s 41.2 GB 373.5 TFLOPS (warmup 32s)
|
||||
AUTOTUNE mm(24576x1280, 1280x3584)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.3036 ms 100.0%
|
||||
triton_mm_112 0.3828 ms 79.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_113 0.4067 ms 74.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_111 0.4397 ms 69.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_107 0.5206 ms 58.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_110 0.5247 ms 57.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_106 0.5531 ms 54.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_104 0.6231 ms 48.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_105 0.6323 ms 48.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_109 0.6390 ms 47.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7413 seconds and 0.0006 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x1280, 1280x1280)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.1155 ms 100.0%
|
||||
triton_mm_17 0.1448 ms 79.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_18 0.1500 ms 77.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_16 0.1635 ms 70.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_11 0.1818 ms 63.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_12 0.1984 ms 58.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_15 0.2274 ms 50.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_10 0.2326 ms 49.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_9 0.2335 ms 49.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14 0.2375 ms 48.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.6322 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x1280, 1280x256)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0373 ms 100.0%
|
||||
triton_mm_37 0.0445 ms 83.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_30 0.0454 ms 82.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_36 0.0457 ms 81.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_35 0.0520 ms 71.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_31 0.0523 ms 71.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_26 0.0595 ms 62.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_29 0.0616 ms 60.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_33 0.0629 ms 59.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_27 0.0652 ms 57.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.3425 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x3584, 3584x1280)
|
||||
strides: [3584, 1], [1, 3584]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2847 ms 100.0%
|
||||
triton_mm_132 0.3619 ms 78.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_131 0.3786 ms 75.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_130 0.4460 ms 63.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_126 0.4835 ms 58.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_125 0.5239 ms 54.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_124 0.6400 ms 44.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_127 0.6484 ms 43.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_123 0.6497 ms 43.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_128 0.6612 ms 43.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7323 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x1280, 1280x32000)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 2.5921 ms 100.0%
|
||||
triton_mm_3476 3.6376 ms 71.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3475 3.6652 ms 70.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3474 4.0276 ms 64.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3470 5.0841 ms 51.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3469 5.2775 ms 49.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3471 5.8066 ms 44.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3472 5.8295 ms 44.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3473 6.0444 ms 42.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3467 6.1254 ms 42.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 1.6122 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x1280, 1280x3584)
|
||||
strides: [1280, 1], [3584, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.3097 ms 100.0%
|
||||
triton_mm_3551 0.3933 ms 78.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3552 0.4125 ms 75.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3550 0.4132 ms 74.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3549 0.4928 ms 62.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3546 0.5176 ms 59.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3545 0.5612 ms 55.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3543 0.5924 ms 52.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3544 0.6125 ms 50.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3547 0.6247 ms 49.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7245 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(32000x24576, 24576x1280)
|
||||
strides: [1, 32000], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 2.4669 ms 100.0%
|
||||
triton_mm_3495 3.2061 ms 76.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3493 3.6224 ms 68.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3494 3.8289 ms 64.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3489 5.1092 ms 48.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3492 5.3475 ms 46.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3486 5.5071 ms 44.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3488 5.7451 ms 42.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3490 5.7689 ms 42.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3491 6.2059 ms 39.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 1.5248 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(1280x24576, 24576x3584)
|
||||
strides: [1, 1280], [3584, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.3001 ms 100.0%
|
||||
triton_mm_3531 0.3708 ms 80.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3527 0.4240 ms 70.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3533 0.4400 ms 68.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3532 0.4529 ms 66.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3524 0.5452 ms 55.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3523 0.5851 ms 51.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3528 0.5976 ms 50.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3526 0.6137 ms 48.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3525 0.6424 ms 46.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7376 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(3584x24576, 24576x1280)
|
||||
strides: [1, 3584], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2956 ms 100.0%
|
||||
triton_mm_3569 0.4018 ms 73.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3565 0.4229 ms 69.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3571 0.4369 ms 67.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3570 0.4506 ms 65.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3562 0.5451 ms 54.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3564 0.5882 ms 50.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3566 0.6142 ms 48.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3561 0.6312 ms 46.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3563 0.6324 ms 46.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7334 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(1280x24576, 24576x1280)
|
||||
strides: [1, 1280], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.1238 ms 100.0%
|
||||
triton_mm_3647 0.1346 ms 92.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3640 0.1648 ms 75.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3641 0.1652 ms 74.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3646 0.1800 ms 68.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3639 0.1999 ms 61.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3643 0.2041 ms 60.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3636 0.2465 ms 50.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_3637 0.2485 ms 49.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3638 0.2575 ms 48.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.6361 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(256x24576, 24576x1280)
|
||||
strides: [1, 256], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0454 ms 100.0%
|
||||
triton_mm_3675 0.0653 ms 69.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3679 0.0800 ms 56.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3671 0.1076 ms 42.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=32, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3685 0.1343 ms 33.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3674 0.1594 ms 28.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_3678 0.1604 ms 28.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3670 0.1725 ms 26.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=32, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3684 0.1795 ms 25.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3681 0.1892 ms 24.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.5430 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x32000, 32000x1280)
|
||||
strides: [32000, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 2.4075 ms 100.0%
|
||||
triton_mm_3514 3.5214 ms 68.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3512 3.7533 ms 64.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3513 3.7865 ms 63.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3508 5.0381 ms 47.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3507 5.5094 ms 43.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3511 5.7366 ms 42.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3509 5.8572 ms 41.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3505 5.9727 ms 40.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3504 7.0820 ms 34.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 1.5918 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x3584, 3584x1280)
|
||||
strides: [3584, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.3158 ms 100.0%
|
||||
triton_mm_3590 0.3743 ms 84.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3588 0.3924 ms 80.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3589 0.3948 ms 80.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3584 0.4805 ms 65.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3583 0.5275 ms 59.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3581 0.5821 ms 54.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3582 0.6140 ms 51.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3585 0.6172 ms 51.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3586 0.6261 ms 50.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7215 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x1280, 1280x1280)
|
||||
strides: [1280, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.1217 ms 100.0%
|
||||
triton_mm_3665 0.1442 ms 84.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3664 0.1471 ms 82.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3666 0.1505 ms 80.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3659 0.1869 ms 65.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3660 0.1931 ms 63.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3657 0.2044 ms 59.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3658 0.2167 ms 56.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3663 0.2197 ms 55.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3661 0.2247 ms 54.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.6127 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(24576x256, 256x1280)
|
||||
strides: [256, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0421 ms 100.0%
|
||||
triton_mm_3703 0.0518 ms 81.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3702 0.0531 ms 79.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3695 0.0532 ms 79.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3696 0.0552 ms 76.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3700 0.0581 ms 72.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3697 0.0581 ms 72.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3699 0.0596 ms 70.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3704 0.0605 ms 69.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3701 0.0660 ms 63.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.3546 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
mb=24 ctx=1024 compile=max-autotune 126,226 tok/s 42.7 GB 379.5 TFLOPS (warmup 117s)
|
||||
mb=48 ctx=1024 compile=default OOM
|
||||
AUTOTUNE mm(49152x1280, 1280x3584)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5913 ms 100.0%
|
||||
triton_mm_10543 0.7967 ms 74.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10544 0.8377 ms 70.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_10542 0.8648 ms 68.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10541 1.0316 ms 57.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_10538 1.0412 ms 56.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_10537 1.1588 ms 51.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10535 1.2477 ms 47.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10536 1.2588 ms 47.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_10540 1.2720 ms 46.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8281 seconds and 0.0006 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x1280)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2198 ms 100.0%
|
||||
triton_mm_10448 0.2860 ms 76.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10449 0.3004 ms 73.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_10447 0.3274 ms 67.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10442 0.3817 ms 57.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10443 0.3893 ms 56.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_10446 0.4227 ms 52.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_10440 0.4564 ms 48.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10441 0.4627 ms 47.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_10445 0.4696 ms 46.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7202 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x256)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0666 ms 100.0%
|
||||
triton_mm_10467 0.0759 ms 87.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10468 0.0781 ms 85.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_10461 0.0831 ms 80.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10466 0.0939 ms 71.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10462 0.0945 ms 70.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_10457 0.1101 ms 60.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_10460 0.1120 ms 59.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_10464 0.1170 ms 56.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_10459 0.1209 ms 55.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.4698 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x3584, 3584x1280)
|
||||
strides: [3584, 1], [1, 3584]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5663 ms 100.0%
|
||||
triton_mm_10563 0.7686 ms 73.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_10562 0.7837 ms 72.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10561 0.8769 ms 64.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10557 0.9856 ms 57.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_10556 1.0930 ms 51.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10555 1.2751 ms 44.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_10554 1.3037 ms 43.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10559 1.3103 ms 43.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_10558 1.3197 ms 42.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8333 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x32000)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 5.1811 ms 100.0%
|
||||
triton_mm_13907 7.3058 ms 70.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_13906 7.6452 ms 67.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13905 8.0600 ms 64.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13901 10.2481 ms 50.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_13900 10.4869 ms 49.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13904 11.9542 ms 43.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_13902 12.0720 ms 42.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13898 12.1052 ms 42.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13903 12.2308 ms 42.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 2.7562 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x3584)
|
||||
strides: [1280, 1], [3584, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.6030 ms 100.0%
|
||||
triton_mm_13982 0.7762 ms 77.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13983 0.8275 ms 72.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_13981 0.8420 ms 71.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13980 0.9811 ms 61.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_13977 1.0286 ms 58.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_13976 1.1561 ms 52.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13974 1.1820 ms 51.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13975 1.2275 ms 49.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_13979 1.2455 ms 48.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8172 seconds and 0.0003 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(32000x49152, 49152x1280)
|
||||
strides: [1, 32000], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 4.9052 ms 100.0%
|
||||
triton_mm_13926 6.9532 ms 70.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_13924 7.3551 ms 66.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13925 8.3430 ms 58.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13920 9.0701 ms 54.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_13923 10.7673 ms 45.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_13919 10.9990 ms 44.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13917 11.2342 ms 43.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13916 11.4397 ms 42.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_13921 12.1385 ms 40.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 2.8909 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(1280x49152, 49152x3584)
|
||||
strides: [1, 1280], [3584, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.6542 ms 100.0%
|
||||
triton_mm_13964 0.8997 ms 72.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_13962 0.9432 ms 69.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13958 0.9858 ms 66.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_13963 1.0338 ms 63.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13954 1.1629 ms 56.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_13955 1.1659 ms 56.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13956 1.2179 ms 53.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_13957 1.2649 ms 51.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13959 1.3157 ms 49.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8327 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(3584x49152, 49152x1280)
|
||||
strides: [1, 3584], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5897 ms 100.0%
|
||||
triton_mm_14002 0.8973 ms 65.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14000 0.9344 ms 63.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13996 0.9721 ms 60.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_14001 1.0271 ms 57.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13992 1.1881 ms 49.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_13995 1.2208 ms 48.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13993 1.2335 ms 47.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13994 1.3042 ms 45.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_13997 1.3364 ms 44.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8311 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(1280x49152, 49152x1280)
|
||||
strides: [1, 1280], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2230 ms 100.0%
|
||||
triton_mm_14078 0.2879 ms 77.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14072 0.3827 ms 58.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_14077 0.3866 ms 57.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14071 0.3886 ms 57.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14070 0.4523 ms 49.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14069 0.5069 ms 44.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14068 0.5238 ms 42.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_14076 0.5340 ms 41.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14073 0.5395 ms 41.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7172 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(256x49152, 49152x1280)
|
||||
strides: [1, 256], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0704 ms 100.0%
|
||||
triton_mm_14106 0.1292 ms 54.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_14110 0.1877 ms 37.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_14102 0.2163 ms 32.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=32, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_14116 0.2702 ms 26.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14105 0.3039 ms 23.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_14109 0.3209 ms 21.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14101 0.3410 ms 20.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=32, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14108 0.3641 ms 19.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14112 0.3754 ms 18.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.6607 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x32000, 32000x1280)
|
||||
strides: [32000, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 4.8362 ms 100.0%
|
||||
triton_mm_13945 6.6236 ms 73.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_13944 7.4780 ms 64.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13943 8.1078 ms 59.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13939 10.9153 ms 44.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_13942 11.0180 ms 43.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_13938 11.1597 ms 43.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13936 11.5402 ms 41.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_13933 13.8197 ms 35.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=4
|
||||
triton_mm_13935 14.4248 ms 33.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 2.8274 seconds and 0.0005 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x3584, 3584x1280)
|
||||
strides: [3584, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5904 ms 100.0%
|
||||
triton_mm_14021 0.7600 ms 77.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14020 0.7954 ms 74.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14019 0.8199 ms 72.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14015 1.0262 ms 57.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_14014 1.1236 ms 52.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14012 1.1870 ms 49.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14013 1.2394 ms 47.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14016 1.2514 ms 47.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14017 1.2572 ms 47.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8120 seconds and 0.0007 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x1280)
|
||||
strides: [1280, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2522 ms 100.0%
|
||||
triton_mm_14096 0.2773 ms 90.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14095 0.2992 ms 84.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14097 0.3072 ms 82.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14091 0.3739 ms 67.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_14090 0.3897 ms 64.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14088 0.4112 ms 61.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14094 0.4169 ms 60.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_14089 0.4342 ms 58.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14092 0.4459 ms 56.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7076 seconds and 0.0005 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x256, 256x1280)
|
||||
strides: [256, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0787 ms 100.0%
|
||||
triton_mm_14134 0.0934 ms 84.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14133 0.0952 ms 82.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14126 0.1055 ms 74.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14127 0.1076 ms 73.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14130 0.1085 ms 72.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14131 0.1106 ms 71.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14128 0.1113 ms 70.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_14135 0.1152 ms 68.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_14132 0.1209 ms 65.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.4921 seconds and 0.0006 seconds precompiling for 20 choices
|
||||
mb=48 ctx=1024 compile=max-autotune OOM
|
||||
|
||||
--- ctx 2048 (the real run's target context) ---
|
||||
mb=12 ctx=2048 eager OOM
|
||||
mb=12 ctx=2048 compile=default OOM
|
||||
|
||||
==============================================================================
|
||||
BEST CONFIG
|
||||
==============================================================================
|
||||
micro-batch 24, ctx 1024, compile=max-autotune
|
||||
126,226 tok/s 42.7 GB 379.5 TFLOPS effective (38% MFU)
|
||||
vs the smoke-run config (mb24/ctx1024/eager): 1.72x
|
||||
20B: 44.0 h = $ 132 at $2.99/hr
|
||||
50B: 110.0 h = $ 329 at $2.99/hr
|
||||
100B: 220.1 h = $ 658 at $2.99/hr
|
||||
|
||||
wrote /workspace/tune_results.json
|
||||
uploaded to HF
|
||||
244
debug/tune2.log
Normal file
244
debug/tune2.log
Normal file
@@ -0,0 +1,244 @@
|
||||
==============================================================================
|
||||
THROUGHPUT TUNING 500m on NVIDIA H100 80GB HBM3
|
||||
torch 2.8.0+cu129 | 85 GB | SDPA enable_gqa=True
|
||||
==============================================================================
|
||||
|
||||
--- ctx 1024: micro-batch scaling (compiled) ---
|
||||
mb=8 ctx=1024 compile=default 112,856 tok/s 18.6 GB 339.3 TFLOPS (warmup 27s)
|
||||
mb=12 ctx=1024 compile=default 118,795 tok/s 24.3 GB 357.2 TFLOPS (warmup 48s)
|
||||
mb=16 ctx=1024 compile=default 122,838 tok/s 29.9 GB 369.3 TFLOPS (warmup 0s)
|
||||
mb=24 ctx=1024 compile=default 125,689 tok/s 41.2 GB 377.9 TFLOPS (warmup 1s)
|
||||
mb=32 ctx=1024 compile=default 127,464 tok/s 52.6 GB 383.2 TFLOPS (warmup 1s)
|
||||
mb=48 ctx=1024 compile=default 130,584 tok/s 75.3 GB 392.6 TFLOPS (warmup 1s)
|
||||
mb=64 ctx=1024 compile=default OOM
|
||||
|
||||
--- ctx 2048: micro-batch scaling (compiled) ---
|
||||
mb=8 ctx=2048 compile=default 111,115 tok/s 29.9 GB 334.1 TFLOPS (warmup 57s)
|
||||
mb=12 ctx=2048 compile=default 113,403 tok/s 41.2 GB 340.9 TFLOPS (warmup 1s)
|
||||
mb=16 ctx=2048 compile=default 115,295 tok/s 52.6 GB 346.6 TFLOPS (warmup 1s)
|
||||
mb=24 ctx=2048 compile=default 117,730 tok/s 75.3 GB 354.0 TFLOPS (warmup 1s)
|
||||
mb=32 ctx=2048 compile=default OOM
|
||||
|
||||
--- max-autotune on the winner (mb=48 ctx=1024) ---
|
||||
AUTOTUNE mm(49152x1280, 1280x3584)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5903 ms 100.0%
|
||||
triton_mm_112 0.8000 ms 73.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_113 0.8361 ms 70.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_111 0.8669 ms 68.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_110 1.0066 ms 58.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_107 1.0403 ms 56.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_106 1.1498 ms 51.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_104 1.2427 ms 47.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_105 1.2592 ms 46.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_109 1.2732 ms 46.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8322 seconds and 0.0006 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x1280)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2194 ms 100.0%
|
||||
triton_mm_17 0.2837 ms 77.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_18 0.2979 ms 73.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_16 0.3254 ms 67.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_11 0.3830 ms 57.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_12 0.3857 ms 56.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_15 0.4218 ms 52.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_9 0.4556 ms 48.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_10 0.4619 ms 47.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_14 0.4678 ms 46.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7087 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x256)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0661 ms 100.0%
|
||||
triton_mm_36 0.0748 ms 88.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_37 0.0785 ms 84.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_30 0.0821 ms 80.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_35 0.0929 ms 71.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_31 0.0943 ms 70.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_26 0.1082 ms 61.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_29 0.1124 ms 58.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_28 0.1169 ms 56.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_33 0.1169 ms 56.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.4629 seconds and 0.0003 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x3584, 3584x1280)
|
||||
strides: [3584, 1], [1, 3584]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5668 ms 100.0%
|
||||
triton_mm_131 0.7744 ms 73.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_132 0.7777 ms 72.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_130 0.8706 ms 65.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_126 0.9550 ms 59.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_125 1.0941 ms 51.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_124 1.2728 ms 44.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_123 1.2829 ms 44.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_128 1.3110 ms 43.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_127 1.3160 ms 43.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8288 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x32000)
|
||||
strides: [1280, 1], [1, 1280]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 5.1829 ms 100.0%
|
||||
triton_mm_3476 7.5296 ms 68.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3475 7.5493 ms 68.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3474 8.0500 ms 64.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3470 10.3107 ms 50.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3469 10.4844 ms 49.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3471 11.6536 ms 44.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3472 11.6593 ms 44.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3473 11.9396 ms 43.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3467 12.1010 ms 42.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 2.7491 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x3584)
|
||||
strides: [1280, 1], [3584, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5997 ms 100.0%
|
||||
triton_mm_3551 0.7810 ms 76.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3550 0.8326 ms 72.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3552 0.8337 ms 71.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3549 0.9759 ms 61.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3546 1.0188 ms 58.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3545 1.1532 ms 52.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3543 1.1920 ms 50.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3544 1.2280 ms 48.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3547 1.2387 ms 48.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8198 seconds and 0.0004 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(32000x49152, 49152x1280)
|
||||
strides: [1, 32000], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 4.9012 ms 100.0%
|
||||
triton_mm_3495 7.6790 ms 63.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3493 7.8320 ms 62.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3494 8.1227 ms 60.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3489 9.5428 ms 51.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3486 10.9222 ms 44.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3492 11.0988 ms 44.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3488 11.2469 ms 43.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3490 11.8258 ms 41.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3485 12.2489 ms 40.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 2.8886 seconds and 0.0003 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(1280x49152, 49152x3584)
|
||||
strides: [1, 1280], [3584, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5898 ms 100.0%
|
||||
triton_mm_3531 0.9765 ms 60.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3533 0.9903 ms 59.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3532 1.0135 ms 58.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3527 1.0362 ms 56.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3523 1.1941 ms 49.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3525 1.2287 ms 48.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3526 1.2361 ms 47.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3529 1.4599 ms 40.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3524 1.5773 ms 37.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8534 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(3584x49152, 49152x1280)
|
||||
strides: [1, 3584], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.5789 ms 100.0%
|
||||
triton_mm_3569 0.9620 ms 60.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3571 0.9789 ms 59.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3570 1.0100 ms 57.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3565 1.0119 ms 57.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3564 1.2343 ms 46.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3561 1.2374 ms 46.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3563 1.2878 ms 44.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3567 1.4540 ms 39.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3562 1.5863 ms 36.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8552 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(1280x49152, 49152x1280)
|
||||
strides: [1, 1280], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2208 ms 100.0%
|
||||
triton_mm_3647 0.3102 ms 71.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3646 0.3853 ms 57.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3641 0.3983 ms 55.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3640 0.4033 ms 54.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3639 0.5029 ms 43.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3643 0.5195 ms 42.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3637 0.5197 ms 42.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3638 0.5250 ms 42.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3642 0.5650 ms 39.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7255 seconds and 0.0003 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(256x49152, 49152x1280)
|
||||
strides: [1, 256], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0691 ms 100.0%
|
||||
triton_mm_3675 0.1255 ms 55.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3679 0.1942 ms 35.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3671 0.2506 ms 27.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=32, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
triton_mm_3685 0.3040 ms 22.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3674 0.3098 ms 22.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=8
|
||||
triton_mm_3678 0.3307 ms 20.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3670 0.3354 ms 20.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=32, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3677 0.3742 ms 18.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3681 0.3800 ms 18.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=False, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.6723 seconds and 0.0002 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x32000, 32000x1280)
|
||||
strides: [32000, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 4.8225 ms 100.0%
|
||||
triton_mm_3514 6.8045 ms 70.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3513 7.5343 ms 64.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3512 8.2723 ms 58.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3508 10.8250 ms 44.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3511 11.0583 ms 43.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3507 11.2008 ms 43.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3505 11.5023 ms 41.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3502 13.7645 ms 35.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=4
|
||||
triton_mm_3504 14.2358 ms 33.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 2.8061 seconds and 0.0004 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x3584, 3584x1280)
|
||||
strides: [3584, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.6123 ms 100.0%
|
||||
triton_mm_3590 0.7728 ms 79.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3589 0.7846 ms 78.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3588 0.8193 ms 74.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3584 1.0146 ms 60.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3583 1.1161 ms 54.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3581 1.1863 ms 51.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3582 1.2398 ms 49.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3585 1.2504 ms 49.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3586 1.2525 ms 48.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.8133 seconds and 0.0005 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x1280, 1280x1280)
|
||||
strides: [1280, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.2323 ms 100.0%
|
||||
triton_mm_3665 0.2791 ms 83.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3664 0.2983 ms 77.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3666 0.3095 ms 75.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3660 0.3763 ms 61.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=128, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=4
|
||||
triton_mm_3659 0.3903 ms 59.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3657 0.4095 ms 56.7% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3663 0.4165 ms 55.8% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
triton_mm_3658 0.4332 ms 53.6% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3661 0.4423 ms 52.5% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.7041 seconds and 0.0005 seconds precompiling for 20 choices
|
||||
AUTOTUNE mm(49152x256, 256x1280)
|
||||
strides: [256, 1], [1280, 1]
|
||||
dtypes: torch.bfloat16, torch.bfloat16
|
||||
mm 0.0763 ms 100.0%
|
||||
triton_mm_3703 0.0919 ms 82.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3702 0.0965 ms 79.0% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3695 0.1015 ms 75.1% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3699 0.1071 ms 71.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3696 0.1075 ms 70.9% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3700 0.1099 ms 69.4% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=64, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=4, num_warps=8
|
||||
triton_mm_3697 0.1101 ms 69.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=64, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=3, num_warps=4
|
||||
triton_mm_3704 0.1152 ms 66.2% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=64, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=5, num_warps=8
|
||||
triton_mm_3701 0.1204 ms 63.3% ACC_TYPE='tl.float32', ALLOW_TF32=False, BLOCK_K=32, BLOCK_M=128, BLOCK_N=128, EVEN_K=True, GROUP_M=8, USE_FAST_ACCUM=False, num_stages=2, num_warps=8
|
||||
SingleProcess AUTOTUNE benchmarking takes 0.4899 seconds and 0.0005 seconds precompiling for 20 choices
|
||||
mb=48 ctx=1024 compile=max-autotune 130,147 tok/s 78.3 GB 391.3 TFLOPS (warmup 162s)
|
||||
|
||||
==============================================================================
|
||||
BEST CONFIG
|
||||
==============================================================================
|
||||
micro-batch 48, ctx 1024, compile=default
|
||||
130,584 tok/s 75.3 GB 392.6 TFLOPS effective (40% MFU)
|
||||
20B: 42.5 h = $ 127 at $2.99/hr
|
||||
50B: 106.4 h = $ 318 at $2.99/hr
|
||||
|
||||
wrote /workspace/tune_results.json
|
||||
uploaded to HF
|
||||
114
debug/tune_results.json
Normal file
114
debug/tune_results.json
Normal file
@@ -0,0 +1,114 @@
|
||||
{
|
||||
"gpu": "NVIDIA H100 80GB HBM3",
|
||||
"torch": "2.8.0+cu129",
|
||||
"results": [
|
||||
{
|
||||
"mb": 8,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 112855.93922615414,
|
||||
"peak_gb": 18.579795968,
|
||||
"tflops": 339.30627471695726,
|
||||
"warmup_s": 26.683147192001343
|
||||
},
|
||||
{
|
||||
"mb": 12,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 118795.12738506112,
|
||||
"peak_gb": 24.264536064,
|
||||
"tflops": 357.1627014399097,
|
||||
"warmup_s": 47.67642426490784
|
||||
},
|
||||
{
|
||||
"mb": 16,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 122837.53238217832,
|
||||
"peak_gb": 29.92830464,
|
||||
"tflops": 369.31636734242323,
|
||||
"warmup_s": 0.4134063720703125
|
||||
},
|
||||
{
|
||||
"mb": 24,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 125688.80943331693,
|
||||
"peak_gb": 41.244962816,
|
||||
"tflops": 377.8888554280444,
|
||||
"warmup_s": 0.5790824890136719
|
||||
},
|
||||
{
|
||||
"mb": 32,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 127464.49062053277,
|
||||
"peak_gb": 52.637249536,
|
||||
"tflops": 383.22751791094504,
|
||||
"warmup_s": 0.75799560546875
|
||||
},
|
||||
{
|
||||
"mb": 48,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 130583.87669116902,
|
||||
"peak_gb": 75.252740096,
|
||||
"tflops": 392.606087388893,
|
||||
"warmup_s": 1.1335768699645996
|
||||
},
|
||||
{
|
||||
"mb": 8,
|
||||
"ctx": 2048,
|
||||
"compile": "default",
|
||||
"tok_s": 111115.13829731975,
|
||||
"peak_gb": 29.924896768,
|
||||
"tflops": 334.0724812432884,
|
||||
"warmup_s": 57.35422158241272
|
||||
},
|
||||
{
|
||||
"mb": 12,
|
||||
"ctx": 2048,
|
||||
"compile": "default",
|
||||
"tok_s": 113402.57779289872,
|
||||
"peak_gb": 41.24522496,
|
||||
"tflops": 340.9497672701231,
|
||||
"warmup_s": 0.641932487487793
|
||||
},
|
||||
{
|
||||
"mb": 16,
|
||||
"ctx": 2048,
|
||||
"compile": "default",
|
||||
"tok_s": 115294.6307953683,
|
||||
"peak_gb": 52.63751168,
|
||||
"tflops": 346.63830666146606,
|
||||
"warmup_s": 0.8390512466430664
|
||||
},
|
||||
{
|
||||
"mb": 24,
|
||||
"ctx": 2048,
|
||||
"compile": "default",
|
||||
"tok_s": 117730.31919758776,
|
||||
"peak_gb": 75.25300224,
|
||||
"tflops": 353.961309454188,
|
||||
"warmup_s": 1.2464354038238525
|
||||
},
|
||||
{
|
||||
"mb": 48,
|
||||
"ctx": 1024,
|
||||
"compile": "max-autotune",
|
||||
"tok_s": 130146.54947516992,
|
||||
"peak_gb": 78.26123776,
|
||||
"tflops": 391.29124415148357,
|
||||
"warmup_s": 161.93337106704712
|
||||
}
|
||||
],
|
||||
"best": {
|
||||
"mb": 48,
|
||||
"ctx": 1024,
|
||||
"compile": "default",
|
||||
"tok_s": 130583.87669116902,
|
||||
"peak_gb": 75.252740096,
|
||||
"tflops": 392.606087388893,
|
||||
"warmup_s": 1.1335768699645996
|
||||
}
|
||||
}
|
||||
19
export_provenance.json
Normal file
19
export_provenance.json
Normal file
@@ -0,0 +1,19 @@
|
||||
{
|
||||
"source_ckpt": "checkpoints/500m_sft_final.pt",
|
||||
"source_iter": 1192,
|
||||
"native_cfg": {
|
||||
"vocab_size": 32000,
|
||||
"dim": 1280,
|
||||
"n_layers": 26,
|
||||
"n_heads": 20,
|
||||
"n_kv_heads": 4,
|
||||
"max_seq_len": 2048,
|
||||
"ffn_multiple_of": 256,
|
||||
"ffn_dim_multiplier": null,
|
||||
"rope_theta": 10000.0,
|
||||
"dropout": 0.0,
|
||||
"norm_eps": 1e-05,
|
||||
"grad_checkpoint": false
|
||||
},
|
||||
"dtype": "bfloat16"
|
||||
}
|
||||
7
generation_config.json
Normal file
7
generation_config.json
Normal file
@@ -0,0 +1,7 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 0,
|
||||
"eos_token_id": 5,
|
||||
"pad_token_id": 1,
|
||||
"transformers_version": "4.57.6"
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:1a9fd1066f3f0873c4f5ae99b75beab3b0721100f45b036f9e255611b220d864
|
||||
size 1002207984
|
||||
17
special_tokens_map.json
Normal file
17
special_tokens_map.json
Normal file
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"additional_special_tokens": [
|
||||
"<|system|>",
|
||||
"<|user|>",
|
||||
"<|assistant|>",
|
||||
"<|tool_call|>",
|
||||
"<|tool_result|>",
|
||||
"<think>",
|
||||
"</think>",
|
||||
"<|memory_read|>",
|
||||
"<|memory_write|>"
|
||||
],
|
||||
"bos_token": "<|endoftext|>",
|
||||
"eos_token": "<|end|>",
|
||||
"pad_token": "<|pad|>",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
3
sprocket-500m-chat-f16.gguf
Normal file
3
sprocket-500m-chat-f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:3b556c9488db4dfcddb5bfeaa5e065a9f52b7b839cb5c5271283c97b845f0ede
|
||||
size 1003465920
|
||||
3
sprocket-500m-chat-q4_k_m.gguf
Normal file
3
sprocket-500m-chat-q4_k_m.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:edc2ac323a2513fa6f8e3e54adf120243de0f7a6b8cc11b349b3e6ad61494a09
|
||||
size 310279360
|
||||
3
sprocket-500m-f16.gguf
Normal file
3
sprocket-500m-f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:f70749a959fbf7edbdf939be9bca2e4aa7910442cea13eac0ec28fea5628a7d5
|
||||
size 1003465856
|
||||
3
sprocket-500m-q4_k_m.gguf
Normal file
3
sprocket-500m-q4_k_m.gguf
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:fa9456e200200ba13889a6bf93dc55fc2173106b73c3f165833882ef91f091c0
|
||||
size 310279296
|
||||
159091
tokenizer.json
Normal file
159091
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
151
tokenizer_config.json
Normal file
151
tokenizer_config.json
Normal file
@@ -0,0 +1,151 @@
|
||||
{
|
||||
"added_tokens_decoder": {
|
||||
"0": {
|
||||
"content": "<|endoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"1": {
|
||||
"content": "<|pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"2": {
|
||||
"content": "<|system|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"3": {
|
||||
"content": "<|user|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"4": {
|
||||
"content": "<|assistant|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"5": {
|
||||
"content": "<|end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"6": {
|
||||
"content": "<|tool_call|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"7": {
|
||||
"content": "<|tool_result|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"8": {
|
||||
"content": "<think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"9": {
|
||||
"content": "</think>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"10": {
|
||||
"content": "<|memory_read|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"11": {
|
||||
"content": "<|memory_write|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"12": {
|
||||
"content": "<|reserved_4|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"13": {
|
||||
"content": "<|reserved_5|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"14": {
|
||||
"content": "<|reserved_6|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
},
|
||||
"15": {
|
||||
"content": "<|reserved_7|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false,
|
||||
"special": true
|
||||
}
|
||||
},
|
||||
"additional_special_tokens": [
|
||||
"<|system|>",
|
||||
"<|user|>",
|
||||
"<|assistant|>",
|
||||
"<|tool_call|>",
|
||||
"<|tool_result|>",
|
||||
"<think>",
|
||||
"</think>",
|
||||
"<|memory_read|>",
|
||||
"<|memory_write|>"
|
||||
],
|
||||
"bos_token": "<|endoftext|>",
|
||||
"clean_up_tokenization_spaces": false,
|
||||
"eos_token": "<|end|>",
|
||||
"extra_special_tokens": {},
|
||||
"model_max_length": 1000000000000000019884624838656,
|
||||
"pad_token": "<|pad|>",
|
||||
"tokenizer_class": "PreTrainedTokenizerFast",
|
||||
"unk_token": "<|endoftext|>"
|
||||
}
|
||||
Reference in New Issue
Block a user