Files

114 lines
3.7 KiB
Markdown
Raw Permalink Normal View History

---
frameworks:
- Pytorch
license: other
license_name: glm-4
license_link: LICENSE
pipeline_tag: text-generation
tags:
- glm
- edge
- mlx
inference: false
base_model: THUDM/glm-edge-1.5b-chat
library_name: mlx
---
# GLM-Edge-1.5B-Chat (16-bit MLX Full-Precision)
This repository contains Zhipu AI's GLM-Edge-1.5B-Chat in its original unquantized 16-bit float (bfloat16) precision. It is compiled natively for Apple Silicon under the MLX framework.
16-bit precision operates with 0% quantization loss. It preserves the exact mathematical weights of the original model.
## Performance Benchmarks
* **Inference Speed**: ~32.50 tokens per second (base M1 Apple Silicon)
* **VRAM Footprint**: ~2.94 GB
* **Memory Efficiency**: Highly optimized for M-series unified memory architecture. Runs smoothly on 8GB machines.
## Installation
Install the MLX LM package.
```bash
pip install mlx-lm
```
## Usage
### Command Line Interface
Chat with the model in your terminal.
```bash
mlx_lm.chat --model SirSahOl/glm-edge-1.5b-chat-mlx-16bit
```
### Python API
Load and generate text programmatically.
```python
from mlx_lm import load, generate
model, tokenizer = load("SirSahOl/glm-edge-1.5b-chat-mlx-16bit")
messages = [{"role": "user", "content": "Explain quantum superposition."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
response = generate(model, tokenizer, prompt=prompt, verbose=True)
```
## Multi-Quantization Comparison
Evaluate your hardware budget and choose the optimal precision:
| Variant | Disk Size | VRAM Footprint | M1 Speed | Key Advantage |
| :--- | :--- | :--- | :--- | :--- |
| **[4-bit MLX](https://huggingface.co/SirSahOl/glm-edge-1.5b-chat-mlx-4bit)** | ~800 MB | ~0.87 GB | ~72.5 tokens/sec | Maximum speed, lowest RAM. |
| **[8-bit MLX](https://huggingface.co/SirSahOl/glm-edge-1.5b-chat-mlx-8bit)** | ~1.56 GB | ~1.56 GB | ~48.5 tokens/sec | Lossless balance, highly stable reasoning. |
| **16-bit MLX** (This Repo) | ~2.94 GB | ~3.00 GB | **~32.5 tokens/sec** | Raw full-precision, absolute peak quality. |
## Limitations
* Largest footprint of the three variants. Use only if raw precision and unquantized quality are critical for your tasks.
## LM Studio Configuration (Universal Preset Fix)
If you load this model in LM Studio, you must configure custom Stop Strings to prevent the model from entering an infinite self-dialogue loop.
### Option A: Automatic Preset (Recommended)
You can create a custom prompt preset to configure all settings automatically. Create a JSON file named `GLM-Edge.json` inside your LM Studio config directory:
* **macOS / Linux**: `~/.lmstudio/config-presets/GLM-Edge.json`
* **Windows**: `%USERPROFILE%\.lmstudio\config-presets\GLM-Edge.json`
Add the following JSON content:
```json
{
"name": "GLM-Edge",
"inference_params": {
"pre_prompt": "You are a helpful, direct, and honest assistant.",
"input_prefix": "<|user|>\n",
"input_suffix": "\n<|assistant|>\n",
"pre_prompt_prefix": "<|system|>\n",
"pre_prompt_suffix": "\n",
"antiprompt": [
"<|user|>",
"<|observation|>",
"<|endoftext|>"
],
"stopStrings": [
"<|user|>",
"<|observation|>",
"<|endoftext|>"
],
"temperature": 0.7,
"max_tokens": 2048
}
}
```
Restart LM Studio, open a Chat session, and select **"GLM-Edge"** from the Prompt Template dropdown.
### Option B: Manual Configuration
Alternatively, configure the settings manually in the **Advanced Configuration** sidebar:
1. **Stop Strings (Antiprompts / stopStrings)**: Add `<|user|>`, `<|observation|>`, and `<|endoftext|>`
2. **Prompt Formatting**:
* **User Prefix**: `<|user|>\n`
* **Assistant Suffix**: `\n<|assistant|>\n`
* **System Prefix**: `<|system|>\n`
* **System Suffix**: `\n`