初始化项目,由ModelHub XC社区提供模型

Model: eivintobias/heartly-qwen-code
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-09-04 04:10:21 +08:00
commit dd9551524a
12 changed files with 973 additions and 0 deletions

36
.gitattributes vendored Normal file
View File

@@ -0,0 +1,36 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
tokenizer.json filter=lfs diff=lfs merge=lfs -text

116
README.md Normal file
View File

@@ -0,0 +1,116 @@
---
language:
- en
license: mit
base_model: Qwen/Qwen2.5-Coder-1.5B
library_name: transformers
model_type: qwen2
tags:
- code
- hallucination-reduction
- heartly
- decide-verify-stop
- boundary-head
- pytorch
- text-generation
---
# Heartly Qwen-Code v3
A 1.5B coding LLM with the **Heartly** hallucination-reduction architecture,
fine-tuned from **Qwen2.5-Coder-1.5B** with the conversational Stage-5 SFT recipe
(Fix14: natural phrasing, single refusal, persona — 5,200 samples in
`heartly-qwen-code/sft_dataset_code_v3.jsonl`).
v3 builds on the same v1/v2 Stage 14 numbers (grammar adoption 100%, boundary-head
AUROC 1.000, critic AUROC 1.000) — same Qwen2.5-Coder-1.5B base, now trained for
multi-turn conversational code chat. See [`HF_MODEL_CARD.md`](HF_MODEL_CARD.md) for
the Stage 12 probe/critic results carried over from the identical architecture.
## Output grammar
```
thinking [reasoning] response<decide>speak|stop</decide><verify>known|unknown</verify> [answer] <stop>
```
Only `[answer]` should reach the user.
## Usage
### 0. Recommended — chat via the GitHub server (strips the grammar for you)
This model emits the Heartly grammar as ordinary multi-token text (the tags are
**not** tokenizer special tokens), so some front-ends (e.g. LM Studio) may decode
them mangled. The **`server.py`** FastAPI loader on GitHub loads this model and
runs every reply through **`reply_formatter.py`**, which canonicalises the tags
and returns only the clean answer.
```bash
pip install -r requirements.txt # fastapi + uvicorn + transformers + torch
python server.py --model eivintobias/heartly-qwen-code --port 8000
curl -X POST http://127.0.0.1:8000/chat \
-H "Content-Type: application/json" \
-d '{"prompt":"Write a function that reverses a string"}'
```
Response: `{"model":"eivintobias/heartly-qwen-code","raw":"...<decide>...","reply":"<clean answer>"}`.
Quick browser test (no curl): open `http://127.0.0.1:8000/``server.py` serves an
HTML chat UI at `GET /`. The first message lazy-loads the model; code answers render
with real line breaks, and the Heartly grammar is stripped by the reply formatter.
Quick offline test (no server): `python chat_smoke.py "Write a function that sorts a list"`.
📦 **Model card source:** this file (`HF_MODEL_CARD_v3.md`). When uploaded to
HuggingFace, copy it to `README.md` on the hub repo.
### 1. Transformers
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("eivintobias/heartly-qwen-code")
model = AutoModelForCausalLM.from_pretrained(
"eivintobias/heartly-qwen-code", torch_dtype=torch.float32, device_map="cpu"
)
model.eval()
ids = tok.encode("User: Write a function that reverses a string\nAssistant: ", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=256, pad_token_id=tok.eos_token_id, do_sample=False)
raw = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=False)
# Strip the grammar -> clean answer:
from reply_formatter import format_reply
print(format_reply(raw))
```
> `reply_formatter.py` (grammar strip) and `server.py` are bundled in this repo (HF clone = flat layout; GitHub = `heartly-qwen-code/`). Clone it so `from reply_formatter import format_reply` resolves before the offline example.
## Files in this repo
| File | Description |
|------|-------------|
| `config.json` | Qwen2ForCausalLM (28 layers, d=1536) + `heartly_stop_token_id=9495` |
| `generation_config.json` | default generate params |
| `chat_template.jinja` | standard Qwen chat template |
| `tokenizer.json` / `tokenizer_config.json` | Qwen BPE tokenizer |
| `model.safetensors` | v3 fine-tuned weights (**full fine-tune**, not a LoRA adapter) |
| `server.py` | FastAPI server: lazy-loads the model; serves `/`, `/health`, `/chat`; strips Heartly grammar via reply_formatter |
| `reply_formatter.py` | strips `thinking` / `<decide>` / `<verify>` / `<stop>` -> clean answer; unescapes code newlines |
| `chat_smoke.py` | offline load + chat smoke test (no server) |
| `requirements.txt` | torch, transformers, fastapi, uvicorn, sentencepiece, datasets, scikit-learn, numpy, accelerate, huggingface_hub |
## Training
- **Base:** Qwen/Qwen2.5-Coder-1.5B
- **Method:** full fine-tune (fp16), max-length 512, 2 epochs, freeze bottom 12 layers
- **Dataset:** `sft_dataset_code_v3.jsonl` (5,200 conversational Heartly samples)
- **GPU:** 1× RTX 3090 (24GB)
## License
MIT — built on Qwen2.5-Coder (Apache 2.0).
## Links
- [GitHub (server + tools)](https://github.com/eivintobias/heartly/tree/master/heartly-qwen-code)
- [v1/v2 model card (LoRA)](HF_MODEL_CARD.md)
- [Heartly RWKV7 model](https://huggingface.co/eivintobias/heartly-rwkv7-1.5b)

68
chat_smoke.py Normal file
View File

@@ -0,0 +1,68 @@
#!/usr/bin/env python3
"""chat_smoke.py - quick offline test of heartly-qwen-code-v3.
Loads the model once and runs a prompt through the SAME loader + reply_formatter
that server.py uses, printing RAW (grammar) and REPLY (clean). No HTTP server.
Usage:
python chat_smoke.py "Write a function that reverses a string"
python chat_smoke.py --mode debug --prompt "Explain a closure"
echo "What is a closure?" | python chat_smoke.py
"""
from __future__ import annotations
import argparse
import sys
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from reply_formatter import format_reply
def main():
p = argparse.ArgumentParser(description="Smoke-test heartly-qwen-code-v3")
p.add_argument("prompt_pos", nargs="?", default=None,
help="question to ask (positional, e.g. chat_smoke.py 'reverse a list')")
p.add_argument("--prompt", "-p", dest="prompt_opt", default=None,
help="question to ask (alternative to positional)")
p.add_argument("--model", default="heartly-qwen-code-v3")
p.add_argument("--mode", default="chat", choices=["chat", "debug", "raw"])
p.add_argument("--max-new-tokens", type=int, default=200)
a = p.parse_args()
# Precedence: --prompt > positional > stdin > hard-coded default.
if a.prompt_opt is not None:
prompt = a.prompt_opt
elif a.prompt_pos is not None:
prompt = a.prompt_pos
else:
prompt = sys.stdin.read().strip()
if not prompt:
prompt = "Write a function that reverses a string" # safety default
print(f"Loading {a.model} ...", file=sys.stderr)
tok = AutoTokenizer.from_pretrained(a.model)
model = AutoModelForCausalLM.from_pretrained(
a.model, torch_dtype=torch.float32, device_map="cpu"
).eval()
ids = tok.encode(f"User: {prompt}\nAssistant: ", return_tensors="pt")
with torch.no_grad():
out = model.generate(
ids,
max_new_tokens=a.max_new_tokens,
pad_token_id=tok.eos_token_id,
do_sample=False,
)
raw = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=False)
reply = format_reply(raw, mode=a.mode)
print("\n=== RAW ===")
print(raw)
print("\n=== REPLY (" + a.mode + ") ===")
print(reply)
if __name__ == "__main__":
main()

54
chat_template.jinja Normal file
View File

@@ -0,0 +1,54 @@
{%- if tools %}
{{- '<|im_start|>system\n' }}
{%- if messages[0]['role'] == 'system' %}
{{- messages[0]['content'] }}
{%- else %}
{{- 'You are Qwen, created by Alibaba Cloud. You are a helpful assistant.' }}
{%- endif %}
{{- "\n\n# Tools\n\nYou may call one or more functions to assist with the user query.\n\nYou are provided with function signatures within <tools></tools> XML tags:\n<tools>" }}
{%- for tool in tools %}
{{- "\n" }}
{{- tool | tojson }}
{%- endfor %}
{{- "\n</tools>\n\nFor each function call, return a json object with function name and arguments within <tool_call></tool_call> XML tags:\n<tool_call>\n{\"name\": <function-name>, \"arguments\": <args-json-object>}\n</tool_call><|im_end|>\n" }}
{%- else %}
{%- if messages[0]['role'] == 'system' %}
{{- '<|im_start|>system\n' + messages[0]['content'] + '<|im_end|>\n' }}
{%- else %}
{{- '<|im_start|>system\nYou are Qwen, created by Alibaba Cloud. You are a helpful assistant.<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- for message in messages %}
{%- if (message.role == "user") or (message.role == "system" and not loop.first) or (message.role == "assistant" and not message.tool_calls) %}
{{- '<|im_start|>' + message.role + '\n' + message.content + '<|im_end|>' + '\n' }}
{%- elif message.role == "assistant" %}
{{- '<|im_start|>' + message.role }}
{%- if message.content %}
{{- '\n' + message.content }}
{%- endif %}
{%- for tool_call in message.tool_calls %}
{%- if tool_call.function is defined %}
{%- set tool_call = tool_call.function %}
{%- endif %}
{{- '\n<tool_call>\n{"name": "' }}
{{- tool_call.name }}
{{- '", "arguments": ' }}
{{- tool_call.arguments | tojson }}
{{- '}\n</tool_call>' }}
{%- endfor %}
{{- '<|im_end|>\n' }}
{%- elif message.role == "tool" %}
{%- if (loop.index0 == 0) or (messages[loop.index0 - 1].role != "tool") %}
{{- '<|im_start|>user' }}
{%- endif %}
{{- '\n<tool_response>\n' }}
{{- message.content }}
{{- '\n</tool_response>' }}
{%- if loop.last or (messages[loop.index0 + 1].role != "tool") %}
{{- '<|im_end|>\n' }}
{%- endif %}
{%- endif %}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|im_start|>assistant\n' }}
{%- endif %}

62
config.json Normal file
View File

@@ -0,0 +1,62 @@
{
"architectures": [
"Qwen2ForCausalLM"
],
"attention_dropout": 0.0,
"bos_token_id": 151643,
"dtype": "bfloat16",
"eos_token_id": 151643,
"hidden_act": "silu",
"hidden_size": 1536,
"initializer_range": 0.02,
"intermediate_size": 8960,
"layer_types": [
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention",
"full_attention"
],
"max_position_embeddings": 32768,
"max_window_layers": 28,
"model_type": "qwen2",
"num_attention_heads": 12,
"num_hidden_layers": 28,
"num_key_value_heads": 2,
"pad_token_id": null,
"rms_norm_eps": 1e-06,
"rope_parameters": {
"rope_theta": 1000000.0,
"rope_type": "default"
},
"sliding_window": null,
"tie_word_embeddings": true,
"transformers_version": "5.14.1",
"use_cache": false,
"use_sliding_window": false,
"vocab_size": 151936,
"heartly_stop_token_id": 9495
}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token_id": 151643,
"eos_token_id": 151643,
"max_new_tokens": 2048,
"transformers_version": "5.14.1"
}

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:23faa74e0f892a87cfd8a8fda7f3c8c777465d50fa41c91b18d715c829de940d
size 3087467144

452
reply_formatter.py Normal file
View File

@@ -0,0 +1,452 @@
#!/usr/bin/env python3
"""
reply_formatter.py -- extract the user-visible answer from a Heartly
Qwen-Code (v3) grammar output.
The v3 model is trained to emit:
thinking {reasoning} response<decide>speak|stop</decide><verify>known|unknown</verify> {answer} <stop>
(or `<decide>stop</decide>` for the silence case).
The user should only ever see {answer}, never the scaffolding. In practice the
model -- especially through a GUI like LM Studio -- produces *noisy* variants:
the tags are NOT registered as tokenizer special tokens, so Qwen's BPE shatters
`<decide>` into subwords that can re-decode mangled (e.g. ``<deside>``), the
`<stop>` end-marker can come back truncated (no closing `>`), and verify values
can carry junk (`<verify>known, true</verify>`). Qwen chat-template tokens also
leak as plain text (`<stop> <begin> <sep>`).
This parser is tolerant: it canonicalises the tags, extracts the structured
fields, and has an aggressive fallback so NO grammar token ever reaches the user.
Modes:
"chat" -- only the clean answer (default)
"debug" -- answer + [decide=X verify=Y] metadata per turn
"raw" -- the original raw output, untouched
Usage:
from reply_formatter import format_reply, clean_reply
shown = format_reply(raw_model_output)
"""
from __future__ import annotations
import re
from dataclasses import dataclass
# ---------------------------------------------------------------------------
# 1. Junk tokens -- decoded-to-plain-text Qwen special tokens that are never
# answer content. Stripped first so they cannot confuse grammar parsing.
# ---------------------------------------------------------------------------
_JUNK_TOKENS = (
"<|im_start|>", "<|im_end|>", "im_start", "im_end",
"<|object_ref_start|>", "<|object_ref_end|>",
"<|box_start|>", "<|box_end|>",
"<|quad_start|>", "<|quad_end|>",
"<|vision_start|>", "<|vision_end|>", "<|vision_pad|>", "<|image_pad|>",
"<|endoftext|>", "endoftext",
"<|tool_call_begin|>", "<|tool_call_end|>", "<|tool_call_argument_begin|>",
"<|tool_call_argument_end|>", "<|tool_call_argument_name|>", "<|tool_call_argument|>",
"<|tool_calls_section_begin|>", "<|tool_calls_section_end|>",
"<tool_calls>", "</tool_calls>", "tool_call", "tool_calls",
"<begin>", "<sep>",
)
# ---------------------------------------------------------------------------
# 2. Meta-commentary patterns -- the model's self-talk about what it knows /
# doesn't know. If any bleeds into the answer zone, strip it. Sourced from
# the reasoning / refusal / silence templates in render_code_sft_v3.py.
# (The think block is removed entirely in the common case, so this is a
# safety net, not the primary mechanism.)
# ---------------------------------------------------------------------------
_META_PATTERNS = [
r"\bI (?:do )?know (?:how )?to do this\b[^.]*\.",
r"\bI (?:do )?know how to write this\b[^.]*\.",
r"\bI (?:do )?know this function well\b[^.]*\.",
r"\bI (?:can|will) (?:write|produce|implement|respond|answer|write clean code)\b[^.]*\.",
r"\bThis is a standard programming task\b[^.]*\.",
r"\bI recognise this programming problem\b[^.]*\.",
r"\bI (?:recognise|recognize) this (?:programming|pattern|algorithm)\b[^.]*\.",
r"\bStandard problem\b[^.]*\.",
r"\bClear task\b[^.]*\.",
r"\bI've seen this pattern before\b[^.]*\.",
r"\bThis is straightforward\b[^.]*\.",
r"\bI (?:do )?not know this API or library\b[^.]*\.",
r"\bI should not invent a solution\b[^.]*\.",
r"\bI have no knowledge of this framework\b[^.]*\.",
r"\bThe honest response is\b[^.]*\.",
r"\bI cannot verify the correct implementation\b[^.]*\.",
r"\bMaking something up would be worse than admitting it\b[^.]*\.",
r"\bI'll say so rather than produce fake code\b[^.]*\.",
r"\bThe input is empty or not a real question\b[^.]*\.",
r"\bNo meaningful request was made\b[^.]*\.",
r"\bSpeaking would add nothing\b[^.]*\.",
r"\bSocial turn\b[^.]*\.",
r"\bGreeting[ —-]\s*respond\b[^.]*\.",
r"\bNot a factual question\b[^.]*\.",
r"\bCasual conversation\b[^.]*\.",
r"\bFollow-up question\b[^.]*\.",
r"\bThey need motivation\b[^.]*\.",
r"\bEmotional context\b[^.]*\.",
r"\bMeta-conversation about how we work together\b[^.]*\.",
r"\bBe warm\b[^.]*\.",
r"\bThey (?:want|are|need|'re|'ve|gave) ",
r"\bThey're (?:just saying|opening up) ",
r"\bThey gave more detail\b[^.]*\.",
]
_META_RE = re.compile("|".join(_META_PATTERNS), re.IGNORECASE)
# Stray control words left behind after a mangled tag is stripped.
# Any leftover angle-bracket control construct, tolerant of truncation/mangles
# (e.g. `</decide` with no closing `>`, or `<deside speak>`).
_TAG_RE = re.compile(r"</?\s*[a-z_][^>]*>?", re.IGNORECASE)
# Stray control words left behind after a mangled tag is stripped (used inline
# in _clean_text as leading/trailing guards).
_CONTROL_WORDS = r"(?:speak|stop|silent|done|response|known|unknown|precise)"
# Collapse runs of whitespace.
_MULTI_SPACE_RE = re.compile(r"\s{2,}")
# ---------------------------------------------------------------------------
# 3. Normalization -- turn messy real-world tokens into canonical tags so the
# structured parser can run cleanly.
# ---------------------------------------------------------------------------
def _normalize(raw: str) -> str:
t = raw or ""
# 3a. Drop Qwen / decoded-token litter.
for junk in _JUNK_TOKENS:
t = t.replace(junk, "")
# 3b. Decide tags. Canonical is ``<decide>`` / ``</decide>`` but tokenization
# can drop a letter (<deside>) or add spaces, and the close can be
# truncated (``</decide``). ``de[cs]ide`` catches both spellings.
t = re.sub(
r"<\s*/?\s*de[cs]ide\b[^>]*>?>?",
lambda m: "</decide>" if "/" in m.group(0) else "<decide>",
t,
flags=re.IGNORECASE,
)
# 3c. Verify tags -- same tolerance.
t = re.sub(
r"<\s*/?\s*verify\b[^>]*>?>?",
lambda m: "</verify>" if "/" in m.group(0) else "<verify>",
t,
flags=re.IGNORECASE,
)
# 3d. Stop tags -- ``<stop>``, ``<stop (truncated`` (no close), and a bare
# ``<stop`` at end-of-output. The optional ``>?>`` handles all of these.
t = re.sub(r"<\s*stop\b[^>]*>?>?", "<stop>", t, flags=re.IGNORECASE)
# 3e. Strip a leading noise run (``<stop> <begin> <sep>`` ...) that some
# outputs carry before any thinking/decide content. This <stop> is junk,
# not the Heartly end-of-turn marker (which is mid/trailing).
t = re.sub(
r"^\s*((?:<stop>|<begin>|<sep>)\s*)*", "", t, flags=re.IGNORECASE
).lstrip()
return t
# ---------------------------------------------------------------------------
# 4. Value normalization helpers.
# ---------------------------------------------------------------------------
_DECIDE_VAL_RE = re.compile(r"\b(speak|stop|silent)\b", re.IGNORECASE)
def _decide_value(raw_val: str) -> str:
"""Reduce a decide tag body to canonical ``speak`` / ``stop``."""
m = _DECIDE_VAL_RE.search(raw_val or "")
if m:
v = m.group(1).lower()
return "stop" if v == "silent" else v
return "" # unparseable
_VERIFY_VAL_RE = re.compile(r"\b(known|unknown)\b", re.IGNORECASE)
def _verify_value(raw_val: str) -> str:
m = _VERIFY_VAL_RE.search(raw_val or "")
return m.group(1).lower() if m else ""
# ---------------------------------------------------------------------------
# 5. Parsed turn + segment parser.
# ---------------------------------------------------------------------------
@dataclass
class ParsedTurn:
reasoning: str = ""
decide: str = "" # speak | stop | ""
verify: str = "" # known | unknown | ""
answer: str = ""
# ``thinking {reasoning} response<decide>...`` -- the word "response" is
# OPTIONAL (the model drops it sometimes; seen in the 2026-08-01 log). The
# lookahead ``(?=<decide>|\\Z)`` bounds the reasoning block at the decide tag
# (or end of string) so the (non-greedy) body doesn't greedily swallow the
# answer, and it does NOT consume the decide opener (leaving it for the decide
# parser to read the decide value).
_THINK_RE = re.compile(
r"\bthinking\b\s*(.*?)(\s*\bresponse\b)?\s*(?=<decide>|\Z)",
re.DOTALL | re.IGNORECASE,
)
# After normalization, tags are canonical.
_DECIDE_CLOSE_RE = re.compile(r"<decide>(.*?)</decide>", re.DOTALL | re.IGNORECASE)
_VERIFY_CLOSE_RE = re.compile(r"<verify>(.*?)</verify>", re.DOTALL | re.IGNORECASE)
# Tolerant open-tag grabs (for mangled/missing close).
_DECIDE_OPEN_RE = re.compile(r"<decide\b([^<]*)", re.IGNORECASE)
_VERIFY_OPEN_RE = re.compile(r"<verify\b([^<]*)", re.IGNORECASE)
def _parse_segment(segment: str) -> ParsedTurn:
"""Parse one turn (already split on ``<stop>``)."""
seg = segment.strip()
if not seg:
return ParsedTurn()
# --- thinking block (word "response" optional) ---
reasoning = ""
m = _THINK_RE.search(seg)
if m:
reasoning = (m.group(1) or "").strip()
seg = seg[m.end():]
# --- decide ---
decide = ""
m = _DECIDE_CLOSE_RE.search(seg)
if m:
decide = _decide_value(m.group(1))
seg = seg[m.end():]
else:
mo = _DECIDE_OPEN_RE.search(seg)
if mo:
decide = _decide_value(mo.group(1))
seg = seg[mo.end():]
# --- verify ---
verify = ""
m = _VERIFY_CLOSE_RE.search(seg)
if m:
verify = _verify_value(m.group(1))
seg = seg[m.end():]
else:
mo = _VERIFY_OPEN_RE.search(seg)
if mo:
verify = _verify_value(mo.group(1))
seg = seg[mo.end():]
# --- answer zone: everything left, up to any stray stop ---
answer = seg
sm = re.search(r"<stop", answer, re.IGNORECASE)
if sm:
answer = answer[: sm.start()]
# Default an unparseable decide to "speak" (the model is answering).
if not decide:
decide = "speak"
return ParsedTurn(
reasoning=reasoning,
decide=decide,
verify=verify,
answer=answer,
)
def parse_reply(raw: str) -> list[ParsedTurn]:
"""Parse raw model output into one or more turns (multi-turn aware)."""
text = _normalize(raw)
turns: list[ParsedTurn] = []
for segment in re.split(r"<stop>", text):
seg = segment.strip()
if not seg:
continue
# A real turn must carry a decide marker; stray fence/backtick noise
# emitted between repeated <stop> markers is skipped.
if not _DECIDE_OPEN_RE.search(seg):
continue
turns.append(_parse_segment(seg))
if not turns:
# No structured grammar found at all -- treat the whole output as one
# best-effort turn.
turns.append(_parse_segment(text))
return turns
# ---------------------------------------------------------------------------
# 6. Answer sanitizers.
# ---------------------------------------------------------------------------
def _strip_meta(text: str) -> str:
"""Remove the model's reasoning self-talk from a piece of text."""
return _META_RE.sub("", text).strip()
def _dedupe(text: str) -> str:
"""Drop verbatim-repeated sentences, keep first occurrence."""
out, seen = [], set()
for sent in re.split(r"(?<=[.!?])\s+", text):
s = sent.strip()
if not s:
continue
key = s.lower()
if key not in seen:
seen.add(key)
out.append(s)
return " ".join(out)
def _clean_text(text: str) -> str:
"""Final pass on an answer zone: strip residual tags/junk, meta, dedupe.
Whitespace collapse and sentence de-duplication run on the PROSE regions
only -- fenced code blocks (``` ``` ```) are preserved verbatim, because
their indentation is significant and must not be collapsed to a single
space.
"""
t = text or ""
# Qwen litter first (never answer content).
for junk in _JUNK_TOKENS:
t = t.replace(junk, "")
# Fence out code blocks so their whitespace survives intact. re.split with
# a capturing group yields [prose, code, prose, code, ...] -- code blocks
# land at odd indices and are emitted untouched.
pieces = re.split(r"(```[^\n]*\n.*?```)", t, flags=re.DOTALL)
out = []
for i, chunk in enumerate(pieces):
if not chunk:
continue
if i % 2 == 1: # captured fenced block -> keep verbatim
# The model emits literal backslash-n / backslash-t as plain text in
# code (training artifact). Restore real newlines/tabs so multi-line
# code renders instead of collapsing to a single line.
chunk = chunk.replace(chr(92) + "n", chr(10)).replace(chr(92) + "t", chr(9))
out.append(chunk)
continue
# prose region
chunk = _TAG_RE.sub(" ", chunk)
chunk = re.sub(rf"^\s*{_CONTROL_WORDS}\b\s*", "", chunk, flags=re.IGNORECASE)
chunk = re.sub(rf"\s*{_CONTROL_WORDS}\s*$", "", chunk, flags=re.IGNORECASE)
chunk = _strip_meta(chunk)
chunk = _MULTI_SPACE_RE.sub(" ", chunk).strip()
chunk = _dedupe(chunk)
if chunk:
out.append(chunk)
return " ".join(out).strip()
def _answer_from_reasoning(reasoning: str) -> str:
"""Last-resort: if the answer zone is empty, try to harvest a real answer
from the reasoning block (e.g. 'The answer is X' phrasing). Returns "" when
the reasoning only holds placeholder / meta chatter -- in that case the
caller maps the empty answer to the silence sentinel ("...")."""
if not reasoning:
return ""
m = re.search(
r"(?:the answer is|answer is|it's|it is)\s*:?\s*(.+?)(?:\.\s*$|\.\s*<|$)",
reasoning,
re.IGNORECASE | re.DOTALL,
)
if m:
return _clean_text(m.group(1))
return ""
# ---------------------------------------------------------------------------
# 7. Resolution -- decide what the user actually sees.
# ---------------------------------------------------------------------------
def _resolve_chat(turns: list[ParsedTurn]) -> str:
parts = []
for t in turns:
if t.decide == "stop":
parts.append("...")
continue
ans = _clean_text(t.answer)
if not ans and t.verify == "unknown":
parts.append(ans or "I don't have that information.")
continue
if not ans:
ans = _answer_from_reasoning(t.reasoning)
parts.append(ans if ans else "...")
return "\n".join(parts)
def _resolve_debug(turns):
parts = []
for t in turns:
ans = _clean_text(t.answer)
if not ans:
ans = _answer_from_reasoning(t.reasoning)
meta = " ".join(
f"[{k}={v}]" for k, v in (("decide", t.decide), ("verify", t.verify)) if v
)
shown = f"{ans} {meta}".strip() if ans else (meta or "(empty)")
parts.append(shown)
return "\n".join(parts)
# 6b. Last-resort legacy cleaner -- used only if structured parsing leaves a
# grammar token in the chat output. Aggressively strips every angle-bracket
# construct + control word, guaranteeing a clean result.
def _legacy_clean(text):
t = _normalize(text)
t = re.split(r"<stop", t, maxsplit=1)[0]
dm = _DECIDE_OPEN_RE.search(t)
if dm:
t = t[dm.start():]
else:
m = _THINK_RE.search(t)
if m:
t = t[m.end():] if m.end() <= len(t) else ""
t = _DECIDE_CLOSE_RE.sub(" ", t)
t = _VERIFY_CLOSE_RE.sub(" ", t)
t = _TAG_RE.sub(" ", t)
t = re.sub(rf"^\s*{_CONTROL_WORDS}\b\s*", "", t, flags=re.IGNORECASE)
t = re.sub(rf"\s*{_CONTROL_WORDS}\b\s*$", "", t, flags=re.IGNORECASE)
return _clean_text(t)
# 8. Public API
def format_reply(raw, mode="chat"):
"""Format a raw Heartly model output for display.
mode: "chat" (default), "debug", or "raw".
"""
if mode == "raw":
return raw if raw else ""
if mode not in ("chat", "debug"):
mode = "chat"
if not raw or not raw.strip():
return ""
turns = parse_reply(raw)
if mode == "debug":
return _resolve_debug(turns)
out = _resolve_chat(turns)
# Leak guarantee: if any grammar token survived, fall back hard.
if out == "" and raw.strip():
out = _legacy_clean(raw)
elif out != "" and _TAG_RE.search(out):
out = _legacy_clean(raw) or out
# Safety net: the model intermittently emits code (especially un-fenced) as
# literal backslash-n / backslash-t text. Render those as real newlines/tabs
# so answers never collapse to a single line in the browser or CLI.
out = out.replace(chr(92) + "n", chr(10)).replace(chr(92) + "t", chr(9))
return out
def clean_reply(text):
"""Legacy-compatible drop-in replacement for the old clean_reply()."""
return format_reply(text, mode="chat")
if __name__ == "__main__":
import sys
sample = sys.stdin.read() if not sys.argv[1:] else open(sys.argv[1]).read()
print(format_reply(sample, mode="debug"))

10
requirements.txt Normal file
View File

@@ -0,0 +1,10 @@
torch>=2.0.0
transformers>=4.37.0
datasets>=2.14.0
scikit-learn>=1.0
sentencepiece>=0.1.99
accelerate>=0.25.0
numpy>=1.24
fastapi>=0.110.0
uvicorn>=0.29.0
huggingface_hub>=0.25

133
server.py Normal file
View File

@@ -0,0 +1,133 @@
#!/usr/bin/env python3
"""server.py - HTTP server for heartly-qwen-code-v3.
Serves a browser chat UI at / plus JSON endpoints /health and /chat.
The model's <decide>/<verify>/<stop>/<thinking> scaffolding is stripped by
reply_formatter (same-directory module) before the answer reaches the user.
"""
from __future__ import annotations
import argparse
import os
import threading
from typing import Optional
import torch
from fastapi import FastAPI
from fastapi.responses import HTMLResponse
from pydantic import BaseModel
from transformers import AutoModelForCausalLM, AutoTokenizer
from reply_formatter import format_reply
class ChatRequest(BaseModel):
prompt: str
max_new_tokens: int = 512
temperature: float = 0.7
top_p: float = 0.9
do_sample: bool = True
mode: str = "chat" # chat | debug | raw
class _State:
model: Optional[object] = None
tokenizer: Optional[object] = None
model_name: str = os.environ.get("HEARTLY_MODEL", "eivintobias/heartly-qwen-code")
lock: threading.Lock = threading.Lock()
def _format_prompt(prompt: str) -> str:
return f"User: {prompt}\nAssistant: "
def _load() -> None:
if _State.model is not None:
return
with _State.lock:
if _State.model is None:
dtype = torch.float16 if torch.cuda.is_available() else torch.float32
device = "cuda" if torch.cuda.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(_State.model_name)
model = AutoModelForCausalLM.from_pretrained(_State.model_name, torch_dtype=dtype, device_map=device)
model.eval()
_State.tokenizer, _State.model = tok, model
app = FastAPI(title="Heartly Qwen-Code v3", version="3.0")
CHAT_HTML = """<!doctype html>
<html lang="en">
<head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
<title>Heartly Qwen-Code v3</title>
<style>
html,body{margin:0;height:100%;font-family:-apple-system,BlinkMacSystemFont,Segoe UI,Roboto,Helvetica,Arial,sans-serif;background:#0b0f14;color:#e6e8ec}
#wrap{max-width:820px;height:100vh;margin:0 auto;display:flex;flex-direction:column}
#msgs{flex:1;overflow-y:auto;padding:18px 16px 14px;display:flex;flex-direction:column;gap:12px}
.msg{max-width:82%;white-space:pre-wrap;line-height:1.45;padding:10px 14px;border-radius:14px;font-size:14px}
.user{background:#1f2430;margin-left:auto;border-radius:16px 4px 16px 16px}
.bot{background:#171b23;margin-right:auto;border-radius:4px 16px 16px 16px}
.bot.placeholder{opacity:.55}
#form{display:flex;gap:8px;padding:12px;background:#0f131a;border-top:1px solid #1d222c}
#input{flex:1;background:#171b23;border:1px solid #2a2f3c;border-radius:10px;color:#e6e8ec;padding:10px 12px;font-size:14px;outline:none}
#input::placeholder{color:#7a8190}
button{background:#2f6fec;border:none;color:#fff;border-radius:10px;padding:10px 16px;cursor:pointer;font-size:14px}
button:disabled{opacity:.5;cursor:not-allowed}
</style></head><body>
<div id="wrap"><div id="msgs"></div>
<form id="form" autocomplete="off">
<input id="input" placeholder="Ask the Heartly Qwen-Code v3 model..." autofocus>
<button id="send">Send</button>
</form></div>
<script>
const msgs=document.getElementById('msgs'),form=document.getElementById('form'),input=document.getElementById('input'),btn=document.getElementById('send');
function addMsg(c,t,ph){const d=document.createElement('div');d.className='msg '+c;if(ph)d.classList.add('placeholder');d.textContent=t;msgs.appendChild(d);msgs.scrollTop=msgs.scrollHeight;return d;}
function busy(b){btn.disabled=b;input.disabled=b;}
form.onsubmit=function(e){e.preventDefault();const p=input.value.trim();if(!p||btn.disabled)return;
addMsg('user',p);const bot=addMsg('bot','thinking...',true);busy(true);input.value='';
fetch('/chat',{method:'POST',headers:{'Content-Type':'application/json'},body:JSON.stringify({prompt:p,max_new_tokens:256,temperature:0.3,top_p:0.9,mode:'chat'})})
.then(r=>r.json()).then(d=>{bot.classList.remove('placeholder');bot.textContent=d.reply||'(no reply)';})
.catch(err=>{bot.classList.remove('placeholder');bot.textContent='error: '+err;}).finally(()=>{busy(false);input.focus();});};
</script></body></html>
"""
@app.get("/", response_class=HTMLResponse)
async def chat_page():
"""Browser chat UI - open the tab and start typing."""
return CHAT_HTML
@app.get("/health")
@torch.no_grad()
async def health():
return {"status": "ready" if _State.model is not None else "loading (loads on first /chat)", "model": _State.model_name}
@app.post("/chat")
@torch.no_grad()
async def chat(req: ChatRequest):
_load()
model, tok = _State.model, _State.tokenizer
device = next(model.parameters()).device
ids = tok.encode(_format_prompt(req.prompt), return_tensors="pt").to(device)
out = model.generate(ids, max_new_tokens=req.max_new_tokens, temperature=req.temperature, top_p=req.top_p, do_sample=req.do_sample, pad_token_id=tok.eos_token_id)
raw = tok.decode(out[0][ids.shape[1]:], skip_special_tokens=False)
reply = format_reply(raw, mode=req.mode)
return {"model": _State.model_name, "raw": raw, "reply": reply}
def main():
p = argparse.ArgumentParser(description="Heartly Qwen-Code v3 HTTP server")
p.add_argument("--model", default=os.environ.get("HEARTLY_MODEL", "eivintobias/heartly-qwen-code"))
p.add_argument("--host", default="127.0.0.1")
p.add_argument("--port", type=int, default=8000)
a = p.parse_args()
_State.model_name = a.model
import uvicorn
uvicorn.run("server:app", host=a.host, port=a.port)
if __name__ == "__main__":
main()

3
tokenizer.json Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3fd169731d2cbde95e10bf356d66d5997fd885dd8dbb6fb4684da3f23b2585d8
size 11421892

30
tokenizer_config.json Normal file
View File

@@ -0,0 +1,30 @@
{
"add_prefix_space": false,
"backend": "tokenizers",
"bos_token": null,
"clean_up_tokenization_spaces": false,
"eos_token": "<|endoftext|>",
"errors": "replace",
"extra_special_tokens": [
"<|im_start|>",
"<|im_end|>",
"<|object_ref_start|>",
"<|object_ref_end|>",
"<|box_start|>",
"<|box_end|>",
"<|quad_start|>",
"<|quad_end|>",
"<|vision_start|>",
"<|vision_end|>",
"<|vision_pad|>",
"<|image_pad|>",
"<|video_pad|>"
],
"is_local": false,
"local_files_only": false,
"model_max_length": 32768,
"pad_token": "<|endoftext|>",
"split_special_tokens": false,
"tokenizer_class": "Qwen2Tokenizer",
"unk_token": null
}