Files
ModelHub XC c60e11fd41 初始化项目,由ModelHub XC社区提供模型
Model: BikoRiko/Qwen2.5-1.5B-1.3M-Stretched
Source: Original Platform
2026-08-15 14:15:18 +08:00

3.3 KiB

license, base_model, tags, library_name
license base_model tags library_name
apache-2.0 Qwen/Qwen2.5-1.5B-Instruct
text-generation
lmsys
long-context
rope-scaling
qwen
1M-context
transformers

Qwen2.5-1.5B-1.3M-Stretched 🚀

This is a custom, un-quantized, long-context variant of Qwen2.5-1.5B-Instruct. The model's native context window has been dynamically stretched from its original 32,768 tokens all the way out to 1300000 tokens (1.3m Context) using advanced Dynamic-NTK RoPE scaling.

By scaling the rotary position embedding frequencies natively, the model preserves its low-context processing crispness while dramatically expanding its structural attention horizon—allowing it to ingest multiple entire codebases or full-length novels at once.


Model Highlights 🌟

  • Massive 1.3M Context: Capable of loading and processing over 1.3 million tokens in a single active session sequence.
  • Un-Quantized High Fidelity: Saved directly in native BFloat16 precision. Zero quantization artifacts, zero precision degradation.
  • Compact Footprint: At only 1.5 Billion parameters, the base weights take up ~3GB, making it incredibly lightweight compared to massive 70B+ proprietary long-context models.

Technical Configuration 🔧

The original RoPE frequency base was stretched by implementing a dynamic scaling factor calculated perfectly to expand the 32k window to 1.3M without breaking the model's perplexity metrics:

  • max_position_embeddings: 2400000
  • rope_scaling.type: dynamic
  • rope_scaling.factor: 40.0

Hardware Requirements & VRAM Optimization ⚠️

While the model weights only occupy ~3 GB of memory, processing ultra-long sequences creates a massive KV Cache footprint that scales linearly.

VRAM Scaling Expectations:

  • ~32k context: ~4 GB total VRAM
  • ~500k context: ~18 GB total VRAM
  • ~1.3M context: ~36 GB to 40 GB total VRAM (Fits comfortably on an NVIDIA A100 80GB)

Deployment Note: To test or run inference near the full 1.3M limit, an enterprise GPU cluster node (such as an NVIDIA A100 40GB/80GB or H100) is highly recommended. Ensure PyTorch SDPA is active to optimize memory layout overheads. Awnser me If you need more Ai models with a bigger context ask me becouse i have acces to b200 and more gpus with over 256gb ram so do not be afriad to ask me Thank you have a nice day.


Quickstart Usage Local Setup 💻

You can run this model natively using the Hugging Face transformers library without needing to pass manual scaling parameters—the 1.3M configuration is fully baked into the config.json file.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "BikoRiko/Qwen2.5-1.5B-1.3M-Stretched"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="cuda",
    attn_implementation="sdpa" # Essential for large KV-Cache stability
)

# Put your massive text file here (Up to 1.3 Million Tokens!)
prompt = "Your massive book or code content here..." 
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=50)
    
print(tokenizer.decode(outputs[0], skip_special_tokens=True))