Logo
Explore Help
Register Sign In
dylanyunlong/project_6
1
0
Fork 0
You've already forked project_6
Code Issues Pull Requests Actions Projects Releases Wiki Activity
Files
8c1955dc920dd38da32a8cb52de553da7cfc8043
project_6/vllm/model_executor/models/phi3.py

18 lines
364 B
Python
Raw Normal View History

[DEPLOY] Complete submission: baseline + all optimizations Adds ALL files needed for Dockerfile build: - qwen3_6_scripts/ (baseline patches + our optimizations) - vllm/ (full vllm package) - paged_attention_v2_pytorch.py (V2 with single-bmm optimization) - Dockerfile + computility-run.yaml Our optimizations vs baseline: 1. paged_attn.py: pre-gathered context KV (eliminates 194 gather calls), Triton try/fallback, V2 heuristic, threshold 32K→64K 2. paged_attention_v2_pytorch.py: fills NotImplementedError, single-bmm Phase 1 (195 launches → 3) 3. patch_enable_triton.py: HAS_TRITON=True with safety fallback 4. patch_triton_tuning.py: BLOCK=64, NUM_WARPS=4 for BI-V100 5. computility-run.yaml: gpu-memory-utilization 0.9→0.95, max-num-batched-tokens 8192→16384 This repo can now be submitted to dev.modelhub.org.cn as-is.
2026-07-30 16:06:20 +00:00
# coding=utf-8
# Adapted from llama.py
"""Inference-only Phi3 model code inherit from Llama.py"""
from vllm.model_executor.models.llama import LlamaForCausalLM
class Phi3ForCausalLM(LlamaForCausalLM):
packed_modules_mapping = {
"qkv_proj": [
"qkv_proj",
],
"gate_up_proj": [
"gate_up_proj",
],
}
Reference in New Issue Copy Permalink
Powered by Gitea Version: 1.24.3 Page: 78ms Template: 1ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API