Model: BikoRiko/Qwen2.5-1.5B-1.3M-Stretched Source: Original Platform
license, base_model, tags, library_name
| license | base_model | tags | library_name | ||||||
|---|---|---|---|---|---|---|---|---|---|
| apache-2.0 | Qwen/Qwen2.5-1.5B-Instruct |
|
transformers |
Qwen2.5-1.5B-1.3M-Stretched 🚀
This is a custom, un-quantized, long-context variant of Qwen2.5-1.5B-Instruct. The model's native context window has been dynamically stretched from its original 32,768 tokens all the way out to 1300000 tokens (1.3m Context) using advanced Dynamic-NTK RoPE scaling.
By scaling the rotary position embedding frequencies natively, the model preserves its low-context processing crispness while dramatically expanding its structural attention horizon—allowing it to ingest multiple entire codebases or full-length novels at once.
Model Highlights 🌟
- Massive 1.3M Context: Capable of loading and processing over 1.3 million tokens in a single active session sequence.
- Un-Quantized High Fidelity: Saved directly in native
BFloat16precision. Zero quantization artifacts, zero precision degradation. - Compact Footprint: At only 1.5 Billion parameters, the base weights take up ~3GB, making it incredibly lightweight compared to massive 70B+ proprietary long-context models.
Technical Configuration 🔧
The original RoPE frequency base was stretched by implementing a dynamic scaling factor calculated perfectly to expand the 32k window to 1.3M without breaking the model's perplexity metrics:
max_position_embeddings:2400000rope_scaling.type:dynamicrope_scaling.factor:40.0
Hardware Requirements & VRAM Optimization ⚠️
While the model weights only occupy ~3 GB of memory, processing ultra-long sequences creates a massive KV Cache footprint that scales linearly.
VRAM Scaling Expectations:
- ~32k context: ~4 GB total VRAM
- ~500k context: ~18 GB total VRAM
- ~1.3M context: ~36 GB to 40 GB total VRAM (Fits comfortably on an NVIDIA A100 80GB)
Deployment Note: To test or run inference near the full 1.3M limit, an enterprise GPU cluster node (such as an NVIDIA A100 40GB/80GB or H100) is highly recommended. Ensure PyTorch SDPA is active to optimize memory layout overheads. Awnser me If you need more Ai models with a bigger context ask me becouse i have acces to b200 and more gpus with over 256gb ram so do not be afriad to ask me Thank you have a nice day.
Quickstart Usage Local Setup 💻
You can run this model natively using the Hugging Face transformers library without needing to pass manual scaling parameters—the 1.3M configuration is fully baked into the config.json file.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BikoRiko/Qwen2.5-1.5B-1.3M-Stretched"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="cuda",
attn_implementation="sdpa" # Essential for large KV-Cache stability
)
# Put your massive text file here (Up to 1.3 Million Tokens!)
prompt = "Your massive book or code content here..."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))