初始化项目,由ModelHub XC社区提供模型
Model: BikoRiko/Qwen2.5-1.5B-1.3M-Stretched Source: Original Platform
This commit is contained in:
78
README.md
Normal file
78
README.md
Normal file
@@ -0,0 +1,78 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
base_model: Qwen/Qwen2.5-1.5B-Instruct
|
||||
tags:
|
||||
- text-generation
|
||||
- lmsys
|
||||
- long-context
|
||||
- rope-scaling
|
||||
- qwen
|
||||
- 1M-context
|
||||
library_name: transformers
|
||||
---
|
||||
|
||||
# Qwen2.5-1.5B-1.3M-Stretched 🚀
|
||||
|
||||
This is a custom, un-quantized, long-context variant of **Qwen2.5-1.5B-Instruct**. The model's native context window has been dynamically stretched from its original 32,768 tokens all the way out to **1300000 tokens (1.3m Context)** using advanced Dynamic-NTK RoPE scaling.
|
||||
|
||||
By scaling the rotary position embedding frequencies natively, the model preserves its low-context processing crispness while dramatically expanding its structural attention horizon—allowing it to ingest multiple entire codebases or full-length novels at once.
|
||||
|
||||
---
|
||||
|
||||
## Model Highlights 🌟
|
||||
|
||||
* **Massive 1.3M Context:** Capable of loading and processing over 1.3 million tokens in a single active session sequence.
|
||||
* **Un-Quantized High Fidelity:** Saved directly in native `BFloat16` precision. Zero quantization artifacts, zero precision degradation.
|
||||
* **Compact Footprint:** At only 1.5 Billion parameters, the base weights take up ~3GB, making it incredibly lightweight compared to massive 70B+ proprietary long-context models.
|
||||
|
||||
---
|
||||
|
||||
## Technical Configuration 🔧
|
||||
|
||||
The original RoPE frequency base was stretched by implementing a dynamic scaling factor calculated perfectly to expand the 32k window to 1.3M without breaking the model's perplexity metrics:
|
||||
|
||||
* **`max_position_embeddings`**: `2400000`
|
||||
* **`rope_scaling.type`**: `dynamic`
|
||||
* **`rope_scaling.factor`**: `40.0`
|
||||
|
||||
---
|
||||
|
||||
## Hardware Requirements & VRAM Optimization ⚠️
|
||||
|
||||
While the model weights only occupy **~3 GB** of memory, processing ultra-long sequences creates a massive **KV Cache** footprint that scales linearly.
|
||||
|
||||
### VRAM Scaling Expectations:
|
||||
* **~32k context:** ~4 GB total VRAM
|
||||
* **~500k context:** ~18 GB total VRAM
|
||||
* **~1.3M context:** ~36 GB to 40 GB total VRAM (Fits comfortably on an **NVIDIA A100 80GB**)
|
||||
|
||||
> **Deployment Note:** To test or run inference near the full 1.3M limit, an enterprise GPU cluster node (such as an **NVIDIA A100 40GB/80GB** or **H100**) is highly recommended. Ensure PyTorch SDPA is active to optimize memory layout overheads.
|
||||
> **Awnser me** If you need more Ai models with a bigger context ask me becouse i have acces to b200 and more gpus with over 256gb ram so do not be afriad to ask me Thank you have a nice day.
|
||||
---
|
||||
|
||||
## Quickstart Usage Local Setup 💻
|
||||
|
||||
You can run this model natively using the Hugging Face `transformers` library without needing to pass manual scaling parameters—the 1.3M configuration is fully baked into the `config.json` file.
|
||||
|
||||
```python
|
||||
import torch
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_id = "BikoRiko/Qwen2.5-1.5B-1.3M-Stretched"
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
model_id,
|
||||
torch_dtype=torch.bfloat16,
|
||||
device_map="cuda",
|
||||
attn_implementation="sdpa" # Essential for large KV-Cache stability
|
||||
)
|
||||
|
||||
# Put your massive text file here (Up to 1.3 Million Tokens!)
|
||||
prompt = "Your massive book or code content here..."
|
||||
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
|
||||
|
||||
with torch.no_grad():
|
||||
outputs = model.generate(**inputs, max_new_tokens=50)
|
||||
|
||||
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
||||
Reference in New Issue
Block a user