--- frameworks: - Pytorch license: other license_name: glm-4 license_link: LICENSE pipeline_tag: text-generation tags: - glm - edge - mlx inference: false base_model: THUDM/glm-edge-1.5b-chat library_name: mlx --- # GLM-Edge-1.5B-Chat (16-bit MLX Full-Precision) This repository contains Zhipu AI's GLM-Edge-1.5B-Chat in its original unquantized 16-bit float (bfloat16) precision. It is compiled natively for Apple Silicon under the MLX framework. 16-bit precision operates with 0% quantization loss. It preserves the exact mathematical weights of the original model. ## Performance Benchmarks * **Inference Speed**: ~32.50 tokens per second (base M1 Apple Silicon) * **VRAM Footprint**: ~2.94 GB * **Memory Efficiency**: Highly optimized for M-series unified memory architecture. Runs smoothly on 8GB machines. ## Installation Install the MLX LM package. ```bash pip install mlx-lm ``` ## Usage ### Command Line Interface Chat with the model in your terminal. ```bash mlx_lm.chat --model SirSahOl/glm-edge-1.5b-chat-mlx-16bit ``` ### Python API Load and generate text programmatically. ```python from mlx_lm import load, generate model, tokenizer = load("SirSahOl/glm-edge-1.5b-chat-mlx-16bit") messages = [{"role": "user", "content": "Explain quantum superposition."}] prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) response = generate(model, tokenizer, prompt=prompt, verbose=True) ``` ## Multi-Quantization Comparison Evaluate your hardware budget and choose the optimal precision: | Variant | Disk Size | VRAM Footprint | M1 Speed | Key Advantage | | :--- | :--- | :--- | :--- | :--- | | **[4-bit MLX](https://huggingface.co/SirSahOl/glm-edge-1.5b-chat-mlx-4bit)** | ~800 MB | ~0.87 GB | ~72.5 tokens/sec | Maximum speed, lowest RAM. | | **[8-bit MLX](https://huggingface.co/SirSahOl/glm-edge-1.5b-chat-mlx-8bit)** | ~1.56 GB | ~1.56 GB | ~48.5 tokens/sec | Lossless balance, highly stable reasoning. | | **16-bit MLX** (This Repo) | ~2.94 GB | ~3.00 GB | **~32.5 tokens/sec** | Raw full-precision, absolute peak quality. | ## Limitations * Largest footprint of the three variants. Use only if raw precision and unquantized quality are critical for your tasks. ## LM Studio Configuration (Universal Preset Fix) If you load this model in LM Studio, you must configure custom Stop Strings to prevent the model from entering an infinite self-dialogue loop. ### Option A: Automatic Preset (Recommended) You can create a custom prompt preset to configure all settings automatically. Create a JSON file named `GLM-Edge.json` inside your LM Studio config directory: * **macOS / Linux**: `~/.lmstudio/config-presets/GLM-Edge.json` * **Windows**: `%USERPROFILE%\.lmstudio\config-presets\GLM-Edge.json` Add the following JSON content: ```json { "name": "GLM-Edge", "inference_params": { "pre_prompt": "You are a helpful, direct, and honest assistant.", "input_prefix": "<|user|>\n", "input_suffix": "\n<|assistant|>\n", "pre_prompt_prefix": "<|system|>\n", "pre_prompt_suffix": "\n", "antiprompt": [ "<|user|>", "<|observation|>", "<|endoftext|>" ], "stopStrings": [ "<|user|>", "<|observation|>", "<|endoftext|>" ], "temperature": 0.7, "max_tokens": 2048 } } ``` Restart LM Studio, open a Chat session, and select **"GLM-Edge"** from the Prompt Template dropdown. ### Option B: Manual Configuration Alternatively, configure the settings manually in the **Advanced Configuration** sidebar: 1. **Stop Strings (Antiprompts / stopStrings)**: Add `<|user|>`, `<|observation|>`, and `<|endoftext|>` 2. **Prompt Formatting**: * **User Prefix**: `<|user|>\n` * **Assistant Suffix**: `\n<|assistant|>\n` * **System Prefix**: `<|system|>\n` * **System Suffix**: `\n`