Files
Nemotron-Mini-4B-Instruct-b…/README.md
ModelHub XC 272b0415c7 初始化项目,由ModelHub XC社区提供模型
Model: mlx-community/Nemotron-Mini-4B-Instruct-bf16-mlx
Source: Original Platform
2026-09-01 06:31:13 +08:00

1.3 KiB
Raw Blame History

language, license, license_name, license_link, tags, base_model
language license license_name license_link tags base_model
en
other nvidia-open-model-license https://developer.download.nvidia.com/licenses/nvidia-open-model-license-agreement-june-2024.pdf
mlx
llm
nemotron
apple-silicon
nvidia/Nemotron-Mini-4B-Instruct

mlx-community/Nemotron-Mini-4B-Instruct-bf16-mlx

This model was converted from nvidia/Nemotron-Mini-4B-Instruct to MLX format for use on Apple Silicon.

Quantization: No quantization full bfloat16

Usage

from mlx_lm import load, generate

model, tokenizer = load("mlx-community/{repo_name}")

prompt = (
    "<extra_id_0>System\\n"
    "You are a helpful, honest AI assistant.\\n\\n"
    "<extra_id_1>User\\n"
    "Who are you?\\n"
    "<extra_id_1>Assistant\\n"
)

print(generate(model, tokenizer, prompt, max_tokens=256))

Benchmark (Apple Silicon, single prompt, 23 tokens)

Variant tok/s
bf16 (this) 2.47
4-bit default 4.37
mxfp4-q4 4.56
nvfp4-q4 9.69
mixed-3-6 9.72

Original model

See nvidia/Nemotron-Mini-4B-Instruct for the original model card, license, and usage terms.