ModelHub XC 30197bd489 初始化项目,由ModelHub XC社区提供模型
Model: zenlm/zen-5-flash-gguf
Source: Original Platform
2026-08-21 19:53:12 +08:00

license, base_model, library_name, pipeline_tag, language, tags
license base_model library_name pipeline_tag language tags
apache-2.0 Qwen/Qwen3-4B transformers text-generation
en
zh
zen5
zen5-flash
zenlm
hanzo
safetensors
instruct
agentic
zen-5

Zen5 Flash

Smallest and fastest tier in the Zen5 family. A dense 4B-parameter instruct model with sub-100ms time-to-first-token at 32K context, tuned for high-throughput routing and simple agent loops.

Repackaged from Qwen/Qwen3-4B (apache-2.0, Alibaba Qwen) — redistributed as safetensors from the abliterated huihui-ai/Qwen3-4B-abliterated variant. Not trained from scratch — a permissively-licensed redistribution for the OSS-clean Zen model line.

Part of the canonical Zen5 ladder:

SKU Hardware fit This repo
zen5-flash anything (4 GB VRAM) ← you are here
zen5-mini hosted only zen-5-mini-gguf
zen5 (default) 24 GB+ VRAM zen-5-gguf
zen5-pro Mac M4 Max / DGX Spark / H100 80GB zen-5-pro-gguf
zen5-max Mac Studio M3 Ultra 512GB / 8x H100 zen-5-max-gguf

Files

File Format
model-00001-of-00002.safetensors + model-00002-of-00002.safetensors sharded safetensors
tokenizer.json, tokenizer_config.json, special_tokens_map.json tokenizer
config.json, generation_config.json model config
chat_template.jinja chat template

Run

Hosted via the Hanzo gateway (api.hanzo.ai) as zen5-flash — see https://docs.hanzo.ai/zen.

Local with the zen5-engine or any transformers-compatible runtime:

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("zenlm/zen-5-flash-gguf")
model = AutoModelForCausalLM.from_pretrained("zenlm/zen-5-flash-gguf", device_map="auto")

License

apache-2.0. Upstream: Qwen/Qwen3-4B by Alibaba Qwen; abliterated variant by huihui-ai. This repository redistributes a derivative under the same license.

Description
Model synced from source: zenlm/zen-5-flash-gguf
Readme 15 MiB
Languages
Text 100%