Files
ModelHub XC c60e11fd41 初始化项目,由ModelHub XC社区提供模型
Model: BikoRiko/Qwen2.5-1.5B-1.3M-Stretched
Source: Original Platform
2026-08-15 14:15:18 +08:00

78 lines
3.3 KiB
Markdown

---
license: apache-2.0
base_model: Qwen/Qwen2.5-1.5B-Instruct
tags:
- text-generation
- lmsys
- long-context
- rope-scaling
- qwen
- 1M-context
library_name: transformers
---
# Qwen2.5-1.5B-1.3M-Stretched 🚀
This is a custom, un-quantized, long-context variant of **Qwen2.5-1.5B-Instruct**. The model's native context window has been dynamically stretched from its original 32,768 tokens all the way out to **1300000 tokens (1.3m Context)** using advanced Dynamic-NTK RoPE scaling.
By scaling the rotary position embedding frequencies natively, the model preserves its low-context processing crispness while dramatically expanding its structural attention horizon—allowing it to ingest multiple entire codebases or full-length novels at once.
---
## Model Highlights 🌟
* **Massive 1.3M Context:** Capable of loading and processing over 1.3 million tokens in a single active session sequence.
* **Un-Quantized High Fidelity:** Saved directly in native `BFloat16` precision. Zero quantization artifacts, zero precision degradation.
* **Compact Footprint:** At only 1.5 Billion parameters, the base weights take up ~3GB, making it incredibly lightweight compared to massive 70B+ proprietary long-context models.
---
## Technical Configuration 🔧
The original RoPE frequency base was stretched by implementing a dynamic scaling factor calculated perfectly to expand the 32k window to 1.3M without breaking the model's perplexity metrics:
* **`max_position_embeddings`**: `2400000`
* **`rope_scaling.type`**: `dynamic`
* **`rope_scaling.factor`**: `40.0`
---
## Hardware Requirements & VRAM Optimization ⚠️
While the model weights only occupy **~3 GB** of memory, processing ultra-long sequences creates a massive **KV Cache** footprint that scales linearly.
### VRAM Scaling Expectations:
* **~32k context:** ~4 GB total VRAM
* **~500k context:** ~18 GB total VRAM
* **~1.3M context:** ~36 GB to 40 GB total VRAM (Fits comfortably on an **NVIDIA A100 80GB**)
> **Deployment Note:** To test or run inference near the full 1.3M limit, an enterprise GPU cluster node (such as an **NVIDIA A100 40GB/80GB** or **H100**) is highly recommended. Ensure PyTorch SDPA is active to optimize memory layout overheads.
> **Awnser me** If you need more Ai models with a bigger context ask me becouse i have acces to b200 and more gpus with over 256gb ram so do not be afriad to ask me Thank you have a nice day.
---
## Quickstart Usage Local Setup 💻
You can run this model natively using the Hugging Face `transformers` library without needing to pass manual scaling parameters—the 1.3M configuration is fully baked into the `config.json` file.
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BikoRiko/Qwen2.5-1.5B-1.3M-Stretched"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="cuda",
attn_implementation="sdpa" # Essential for large KV-Cache stability
)
# Put your massive text file here (Up to 1.3 Million Tokens!)
prompt = "Your massive book or code content here..."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))