frameworks, license, license_name, license_link, pipeline_tag, tags, inference, base_model, library_name
frameworks license license_name license_link pipeline_tag tags inference base_model library_name
Pytorch
other glm-4 LICENSE text-generation
glm
edge
mlx
false THUDM/glm-edge-1.5b-chat mlx

GLM-Edge-1.5B-Chat (16-bit MLX Full-Precision)

This repository contains Zhipu AI's GLM-Edge-1.5B-Chat in its original unquantized 16-bit float (bfloat16) precision. It is compiled natively for Apple Silicon under the MLX framework.

16-bit precision operates with 0% quantization loss. It preserves the exact mathematical weights of the original model.

Performance Benchmarks

  • Inference Speed: ~32.50 tokens per second (base M1 Apple Silicon)
  • VRAM Footprint: ~2.94 GB
  • Memory Efficiency: Highly optimized for M-series unified memory architecture. Runs smoothly on 8GB machines.

Installation

Install the MLX LM package.

pip install mlx-lm

Usage

Command Line Interface

Chat with the model in your terminal.

mlx_lm.chat --model SirSahOl/glm-edge-1.5b-chat-mlx-16bit

Python API

Load and generate text programmatically.

from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/glm-edge-1.5b-chat-mlx-16bit")

messages = [{"role": "user", "content": "Explain quantum superposition."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

response = generate(model, tokenizer, prompt=prompt, verbose=True)

Multi-Quantization Comparison

Evaluate your hardware budget and choose the optimal precision:

Variant Disk Size VRAM Footprint M1 Speed Key Advantage
4-bit MLX ~800 MB ~0.87 GB ~72.5 tokens/sec Maximum speed, lowest RAM.
8-bit MLX ~1.56 GB ~1.56 GB ~48.5 tokens/sec Lossless balance, highly stable reasoning.
16-bit MLX (This Repo) ~2.94 GB ~3.00 GB ~32.5 tokens/sec Raw full-precision, absolute peak quality.

Limitations

  • Largest footprint of the three variants. Use only if raw precision and unquantized quality are critical for your tasks.

LM Studio Configuration (Universal Preset Fix)

If you load this model in LM Studio, you must configure custom Stop Strings to prevent the model from entering an infinite self-dialogue loop.

You can create a custom prompt preset to configure all settings automatically. Create a JSON file named GLM-Edge.json inside your LM Studio config directory:

  • macOS / Linux: ~/.lmstudio/config-presets/GLM-Edge.json
  • Windows: %USERPROFILE%\.lmstudio\config-presets\GLM-Edge.json

Add the following JSON content:

{
  "name": "GLM-Edge",
  "inference_params": {
    "pre_prompt": "You are a helpful, direct, and honest assistant.",
    "input_prefix": "<|user|>\n",
    "input_suffix": "\n<|assistant|>\n",
    "pre_prompt_prefix": "<|system|>\n",
    "pre_prompt_suffix": "\n",
    "antiprompt": [
      "<|user|>",
      "<|observation|>",
      "<|endoftext|>"
    ],
    "stopStrings": [
      "<|user|>",
      "<|observation|>",
      "<|endoftext|>"
    ],
    "temperature": 0.7,
    "max_tokens": 2048
  }
}

Restart LM Studio, open a Chat session, and select "GLM-Edge" from the Prompt Template dropdown.

Option B: Manual Configuration

Alternatively, configure the settings manually in the Advanced Configuration sidebar:

  1. Stop Strings (Antiprompts / stopStrings): Add <|user|>, <|observation|>, and <|endoftext|>
  2. Prompt Formatting:
    • User Prefix: <|user|>\n
    • Assistant Suffix: \n<|assistant|>\n
    • System Prefix: <|system|>\n
    • System Suffix: \n
Description
Model synced from source: SirSahOl/glm-edge-1.5b-chat-mlx-16bit
Readme 1.3 MiB
Languages
Jinja 100%