114 lines
3.7 KiB
Markdown
114 lines
3.7 KiB
Markdown
|
|
---
|
||
|
|
frameworks:
|
||
|
|
- Pytorch
|
||
|
|
license: other
|
||
|
|
license_name: glm-4
|
||
|
|
license_link: LICENSE
|
||
|
|
pipeline_tag: text-generation
|
||
|
|
tags:
|
||
|
|
- glm
|
||
|
|
- edge
|
||
|
|
- mlx
|
||
|
|
inference: false
|
||
|
|
base_model: THUDM/glm-edge-1.5b-chat
|
||
|
|
library_name: mlx
|
||
|
|
---
|
||
|
|
|
||
|
|
# GLM-Edge-1.5B-Chat (16-bit MLX Full-Precision)
|
||
|
|
|
||
|
|
This repository contains Zhipu AI's GLM-Edge-1.5B-Chat in its original unquantized 16-bit float (bfloat16) precision. It is compiled natively for Apple Silicon under the MLX framework.
|
||
|
|
|
||
|
|
16-bit precision operates with 0% quantization loss. It preserves the exact mathematical weights of the original model.
|
||
|
|
|
||
|
|
## Performance Benchmarks
|
||
|
|
* **Inference Speed**: ~32.50 tokens per second (base M1 Apple Silicon)
|
||
|
|
* **VRAM Footprint**: ~2.94 GB
|
||
|
|
* **Memory Efficiency**: Highly optimized for M-series unified memory architecture. Runs smoothly on 8GB machines.
|
||
|
|
|
||
|
|
## Installation
|
||
|
|
Install the MLX LM package.
|
||
|
|
```bash
|
||
|
|
pip install mlx-lm
|
||
|
|
```
|
||
|
|
|
||
|
|
## Usage
|
||
|
|
|
||
|
|
### Command Line Interface
|
||
|
|
Chat with the model in your terminal.
|
||
|
|
```bash
|
||
|
|
mlx_lm.chat --model SirSahOl/glm-edge-1.5b-chat-mlx-16bit
|
||
|
|
```
|
||
|
|
|
||
|
|
### Python API
|
||
|
|
Load and generate text programmatically.
|
||
|
|
```python
|
||
|
|
from mlx_lm import load, generate
|
||
|
|
|
||
|
|
model, tokenizer = load("SirSahOl/glm-edge-1.5b-chat-mlx-16bit")
|
||
|
|
|
||
|
|
messages = [{"role": "user", "content": "Explain quantum superposition."}]
|
||
|
|
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
||
|
|
|
||
|
|
response = generate(model, tokenizer, prompt=prompt, verbose=True)
|
||
|
|
```
|
||
|
|
|
||
|
|
## Multi-Quantization Comparison
|
||
|
|
Evaluate your hardware budget and choose the optimal precision:
|
||
|
|
|
||
|
|
| Variant | Disk Size | VRAM Footprint | M1 Speed | Key Advantage |
|
||
|
|
| :--- | :--- | :--- | :--- | :--- |
|
||
|
|
| **[4-bit MLX](https://huggingface.co/SirSahOl/glm-edge-1.5b-chat-mlx-4bit)** | ~800 MB | ~0.87 GB | ~72.5 tokens/sec | Maximum speed, lowest RAM. |
|
||
|
|
| **[8-bit MLX](https://huggingface.co/SirSahOl/glm-edge-1.5b-chat-mlx-8bit)** | ~1.56 GB | ~1.56 GB | ~48.5 tokens/sec | Lossless balance, highly stable reasoning. |
|
||
|
|
| **16-bit MLX** (This Repo) | ~2.94 GB | ~3.00 GB | **~32.5 tokens/sec** | Raw full-precision, absolute peak quality. |
|
||
|
|
|
||
|
|
## Limitations
|
||
|
|
* Largest footprint of the three variants. Use only if raw precision and unquantized quality are critical for your tasks.
|
||
|
|
|
||
|
|
## LM Studio Configuration (Universal Preset Fix)
|
||
|
|
If you load this model in LM Studio, you must configure custom Stop Strings to prevent the model from entering an infinite self-dialogue loop.
|
||
|
|
|
||
|
|
### Option A: Automatic Preset (Recommended)
|
||
|
|
You can create a custom prompt preset to configure all settings automatically. Create a JSON file named `GLM-Edge.json` inside your LM Studio config directory:
|
||
|
|
* **macOS / Linux**: `~/.lmstudio/config-presets/GLM-Edge.json`
|
||
|
|
* **Windows**: `%USERPROFILE%\.lmstudio\config-presets\GLM-Edge.json`
|
||
|
|
|
||
|
|
Add the following JSON content:
|
||
|
|
```json
|
||
|
|
{
|
||
|
|
"name": "GLM-Edge",
|
||
|
|
"inference_params": {
|
||
|
|
"pre_prompt": "You are a helpful, direct, and honest assistant.",
|
||
|
|
"input_prefix": "<|user|>\n",
|
||
|
|
"input_suffix": "\n<|assistant|>\n",
|
||
|
|
"pre_prompt_prefix": "<|system|>\n",
|
||
|
|
"pre_prompt_suffix": "\n",
|
||
|
|
"antiprompt": [
|
||
|
|
"<|user|>",
|
||
|
|
"<|observation|>",
|
||
|
|
"<|endoftext|>"
|
||
|
|
],
|
||
|
|
"stopStrings": [
|
||
|
|
"<|user|>",
|
||
|
|
"<|observation|>",
|
||
|
|
"<|endoftext|>"
|
||
|
|
],
|
||
|
|
"temperature": 0.7,
|
||
|
|
"max_tokens": 2048
|
||
|
|
}
|
||
|
|
}
|
||
|
|
```
|
||
|
|
Restart LM Studio, open a Chat session, and select **"GLM-Edge"** from the Prompt Template dropdown.
|
||
|
|
|
||
|
|
### Option B: Manual Configuration
|
||
|
|
Alternatively, configure the settings manually in the **Advanced Configuration** sidebar:
|
||
|
|
1. **Stop Strings (Antiprompts / stopStrings)**: Add `<|user|>`, `<|observation|>`, and `<|endoftext|>`
|
||
|
|
2. **Prompt Formatting**:
|
||
|
|
* **User Prefix**: `<|user|>\n`
|
||
|
|
* **Assistant Suffix**: `\n<|assistant|>\n`
|
||
|
|
* **System Prefix**: `<|system|>\n`
|
||
|
|
* **System Suffix**: `\n`
|
||
|
|
|
||
|
|
|
||
|
|
|
||
|
|
|