218 lines
7.8 KiB
Markdown
218 lines
7.8 KiB
Markdown
---
|
||
language:
|
||
- en
|
||
- zh
|
||
- ja
|
||
- ko
|
||
- fr
|
||
- es
|
||
- pt
|
||
- de
|
||
- it
|
||
- ru
|
||
- ar
|
||
- vi
|
||
- th
|
||
tags:
|
||
- text-generation
|
||
- ternary
|
||
- bitnet
|
||
- 1.58bit
|
||
- cpu
|
||
- gguf
|
||
- qwen2.5
|
||
- deepseek
|
||
- efficient
|
||
- low-memory
|
||
- jirack
|
||
- web-ui
|
||
- routing
|
||
- tool-call
|
||
- robotics
|
||
license: mit
|
||
---
|
||
# JiRack Ultra 1B (CPU)
|
||
A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new **Routing**, **Tool call**, and **Robotics** tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations.
|
||
- JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative.
|
||
- Subscription: **$1 per month per user** (updated license for non-company use).
|
||
- Corp Subscription: **$3 per month per user** (updated license for company use).
|
||
- It works without subscription but send message about subscription
|
||
|
||
# Ollama production support
|
||
- We are working to support JiRack on Ollama for production systems also
|
||
- added Jirack chat without reasoning feature https://ollama.com/cmsmanhattan
|
||
- Follow fresh Ollama platform updates
|
||
|
||
# JiRack sevice options
|
||
- Current quantizations were done from the FP16 model, but the model allows for more compression thanks to its ternary architecture.
|
||
- If you need to do ternary compression, please write to me and I'll perform QAT from your dataset, tailored specifically to your task.
|
||
- Plus double QAT via ONNX QAT.
|
||
- Adapt train process to avoid catastrophic forgetting with NDA
|
||
- Adapt train process to avoid fast plato in training with NDA
|
||
- Convert model to TQ2_0 with support AVX2 and AVX-512 CPU instructions for high performance on CPU
|
||
- Adapts to agentic or instruct models for tool calling, using the JiRak tokenizer to enable high-quality tool calling on small models — built as a domain-specific tool expert.
|
||
- Deployment and scale
|
||
|
||
|
||
|
||
# Spring Boot AI tool calls examples for JiRack Ultra series
|
||
- Tool call library on java for Enterprise https://github.com/alibaba/spring-ai-alibaba
|
||
|
||
# GoEx AI tool calls examples for JiRack Ultra series
|
||
- Tool call library on python https://github.com/ShishirPatil/gorilla
|
||
|
||
# JiRack Ultra 1 tool calls to boost tool call quality
|
||
- Use JiRack Precision tokenzer tags for tool calls with ToolBench https://github.com/OpenBMB/ToolBench
|
||
- https://huggingface.co/xalss/Qwen2-7B-Instruct-glaive-function-calling
|
||
- https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1
|
||
- Add JiRack tool call tags in the dataset and modify tool call processor if needed
|
||
|
||
|
||
|
||
|
||
# JiRack RoboTech
|
||
- Advanced Tokenizer with Robotics & Routing & Tool calls Tokenizer and other
|
||
- [CMSManhattan/JiRackPrecisionTokenizer](https://huggingface.co/CMSManhattan/JiRackPrecisionTokenizer)
|
||
|
||
|
||
|
||
## Available Variants
|
||
| Tag | Quant | Size | Approx. RAM | Description |
|
||
|-----|-------|------|-------------|-------------|
|
||
| `cmsmanhattan/jirack-ultra-1b-cpu:latest` | Full | 0.55 GB | ~1.8 GB | Full ternary reference |
|
||
| `cmsmanhattan/jirack-ultra-1b-cpu-q4:latest` | Q4_K_M | 0.38 GB | ~1.4 GB | Recommended balance |
|
||
| `cmsmanhattan/jirack-ultra-1b-cpu-q3:latest` | Q3_K_M | 0.31 GB | ~1.2 GB | Good quality / size trade-off |
|
||
| `cmsmanhattan/jirack-ultra-1b-cpu-q2:latest` | Q2_K | 0.24 GB | ~1.0 GB | Maximum compression |
|
||
## Quick Start
|
||
### Run with Docker
|
||
**Default CPU (Q4 recommended)**
|
||
```bash
|
||
docker run -d \
|
||
--name jirack_ultra_1b \
|
||
-p 7869:7869 \
|
||
--cpus=16 \
|
||
-e THREADS=16 \
|
||
-e THREADS_BATCH=16 \
|
||
--restart unless-stopped \
|
||
cmsmanhattan/jirack-ultra-1b-cpu-q4:latest
|
||
```
|
||
**Q3**
|
||
```bash
|
||
docker run -d \
|
||
--name jirack_ultra_1b \
|
||
-p 7869:7869 \
|
||
--cpus=16 \
|
||
-e THREADS=16 \
|
||
-e THREADS_BATCH=16 \
|
||
--restart unless-stopped \
|
||
cmsmanhattan/jirack-ultra-1b-cpu-q3:latest
|
||
```
|
||
**Q2 (lowest memory)**
|
||
```bash
|
||
docker run -d \
|
||
--name jirack_ultra_1b \
|
||
-p 7869:7869 \
|
||
--cpus=16 \
|
||
-e THREADS=16 \
|
||
-e THREADS_BATCH=16 \
|
||
--restart unless-stopped \
|
||
cmsmanhattan/jirack-ultra-1b-cpu-q2:latest
|
||
```
|
||
**Full precision**
|
||
```bash
|
||
docker run -d \
|
||
--name jirack_ultra_1b \
|
||
-p 7869:7869 \
|
||
--cpus=16 \
|
||
-e THREADS=16 \
|
||
-e THREADS_BATCH=16 \
|
||
--restart unless-stopped \
|
||
cmsmanhattan/jirack-ultra-1b-cpu:latest
|
||
```
|
||
**Multi CPU**
|
||
```bash
|
||
docker run -d \
|
||
--name jirack_ultra_1b \
|
||
-p 7869:7869 \
|
||
--cpus=16 \
|
||
-e THREADS=16 \
|
||
-e THREADS_BATCH=16 \
|
||
--restart unless-stopped \
|
||
--memory=4g \
|
||
--cpus=4 \
|
||
cmsmanhattan/jirack-ultra-1b-cpu-q4:latest
|
||
```
|
||
### Docker Compose Example
|
||
```yaml
|
||
services:
|
||
jirack:
|
||
image: cmsmanhattan/jirack-ultra-1b-cpu-q4:latest
|
||
container_name: jirack_ultra_1b
|
||
ports:
|
||
- "7869:7869"
|
||
volumes:
|
||
- .:/app
|
||
- ./web:/app/web
|
||
environment:
|
||
- MAX_TOKENS=2048
|
||
- TEMPERATURE=0.7
|
||
- TOP_P=0.9
|
||
- DEFAULT_STREAM=False
|
||
- INTRA_THREADS=4
|
||
- USE_ENV_ALLOCATOR=1
|
||
- THREADS=16
|
||
- THREADS_BATCH=16
|
||
deploy:
|
||
resources:
|
||
limits:
|
||
memory: 4g
|
||
```
|
||
## Access the UI
|
||
Once the container is running, open your browser and navigate to:
|
||
`http://localhost:7869`
|
||
This opens the JiRack UI — a clean web interface.
|
||
## Changing the Port
|
||
The listening port can be easily modified directly from the **Settings** panel within the JiRack UI.
|
||
## Licensing
|
||
- The JiRack Ultra 1B model is provided under a commercial license ($12 per user per year).
|
||
- All JiRack UI clients are provided under a commercial license.
|
||
- However, the UI clients can be used for free when running together with the official JiRack Docker containers, as long as they are not redistributed separately.
|
||
For commercial licensing, cluster deployment, or enterprise use of JiRack models, please contact us.
|
||
- **JiRack MS Windows 11 Desktop Client (with Ollama API):**
|
||
https://huggingface.co/kgrabko/JiRackTernary_1b/resolve/main/jirack-chat.zip
|
||
- **Live email chat with the model:** support@cmsmanhattan.com
|
||
## Hardware Recommendations
|
||
### Recommended Hardware for JiRack Ultra 1B (single Docker container)
|
||
| Use Case | CPU | RAM | Recommended Quant | Expected Speed | Recommendation |
|
||
|-------------------|------------------------------|----------|-------------------|---------------------|----------------|
|
||
| Recommended | Ryzen 5 / Intel i5 | 4–8 GB | Q4_K_M | Excellent interactive | Best choice |
|
||
| High Performance | Ryzen 7 / Intel i7 | 8–16 GB | Full / Q4 | Excellent | Excellent |
|
||
| Low Memory | Modern 4+ core CPU | 2–4 GB | Q3_K_M or Q2_K | Usable | Acceptable |
|
||
| Edge / Minimal | Laptop / SBC CPU | 2 GB | Q2_K | Acceptable | Budget option |
|
||
## Important Memory Notes
|
||
Even though the quantized 1B models are very small, we recommend the following for best experience:
|
||
- Q4_K_M: 2–4 GB system RAM minimum
|
||
- Q3_K_M / Q2_K: 1.5–3 GB system RAM
|
||
- Full precision: 3–4 GB+ system RAM recommended
|
||
Reasons for extra headroom:
|
||
- KV-cache consumption during generation
|
||
- Runtime overhead and temporary buffers
|
||
- System stability and avoiding out-of-memory errors
|
||
- Room for larger context windows
|
||
**Minimum recommended (Q4):** 2–3 GB system RAM
|
||
**Ideal:** 4–8 GB system RAM
|
||
I added the default model in full precision. This serves as the base for quantization, allowing us to find the optimal balance between model size and performance.
|
||
## Architecture Notes
|
||
- **Refactored with BitNet features**: Native BitLinear ternary path (b1.58-style) with λ-warmup STE
|
||
- **Updated tokenizer**: Extended with new special tags for **Routing**, **Tool call**, and **Robotics**
|
||
- Base: Redesigned Llama-3.2-1B style (Hidden 2048, Intermediate 8192, 16 layers, GQA 32/8, vocab 128256)
|
||
- RoPE θ = 10000, RMSNorm ε = 1e-6
|
||
- Ready-to-run GGUF quantizations (Q2_K, Q3_K_M, Q4_K_M)
|
||
## 📧 Contact & Licensing
|
||
For joint venture opportunities, hardware integration, or licensing inquiries:
|
||
- **Email:** grabko@cmsmanhattan.com
|
||
- **Phone:** +1 (516) 777-0945
|
||
- **Location:** New York, USA
|
||
|
||
## License
|
||
MIT License |