--- base_model: Qwen/Qwen3-1.7B license: apache-2.0 pipeline_tag: text-generation library_name: transformers tags: - qwen3 - cogs - json - structured-output - rag - mlx --- # qwen3-1.7b-cogs-ask **Merged** full model (Qwen/Qwen3-1.7B + LoRA, fused, bf16) — the student model for the [cogs](https://github.com/trunksio/cogs) `ask` pipeline. Standalone adapter: [`lewisdog/qwen3-1.7b-cogs-ask-lora`](https://huggingface.co/lewisdog/qwen3-1.7b-cogs-ask-lora). Sibling model for the ingest pipeline: [`lewisdog/qwen3-1.7b-cogs-ingest`](https://huggingface.co/lewisdog/qwen3-1.7b-cogs-ingest). Two tasks, both emitting compact JSON exactly as the cogs runtime parses it: | task | output JSON | |--------------|-------------------------------------------------------------------| | `decompose` | `{"subquestions": [...]}` — question → 1-4 sub-questions | | `synthesize` | `{"answer": "... [note-id] ...", "citations": [...], "abstained": bool}` — grounded answer over retrieved notes; **abstains** on out-of-domain questions | ## Eval (full 63-record valid set, temp 0, repetition_penalty 1.1) - JSON parse: **100%** · decompose non-empty: **100%** - synthesize citation-validity: **77%** (→ **~90%** with the prefix-repair below) - abstention preserved: **100%** · false-abstention: **4%** ## Apple silicon (MLX / omlx) ```sh python3 -m mlx_lm.convert --hf-path lewisdog/qwen3-1.7b-cogs-ask -q --q-bits 8 --mlx-path qwen3-cogs-ask-mlx # NOTE: if transformers writes v5-format configs, backfill rope_theta / rope_scaling / # torch_dtype from the base Qwen3-1.7B config and swap in the original Qwen3-1.7B # tokenizer files before converting. Wire as [llm.ask] in cogs.toml. ``` ## ⚠️ Serving notes 1. **Decode at temp 0 + `repetition_penalty=1.1`** (pure greedy runs away on long list outputs), `enable_thinking=False`, stop on `<|im_end|>`. 2. **Repair citation prefixes.** Most citation errors are a dropped/wrong namespace prefix (`enisa-…` → `sources-enisa-…`). Normalize against the notes actually in the prompt and drop anything unverifiable: ```python PFX = ("sources-", "source-", "concepts-", "concept-", "entities-", "entity-") core = lambda x: next((x[len(p):] for p in PFX if x.startswith(p)), x) def repair(cite, note_ids): if cite in note_ids: return cite return {core(n): n for n in note_ids}.get(core(cite)) # None => drop it ``` 3. Abstention is reliable — trust it to decline out-of-domain questions. ## Training LoRA r=32/α=64, 5 epochs, eff. batch 16, max_seq 8192, lr 1e-4 cosine, bf16, on a DGX Spark (GB10). Data: `cogs distill` ask pairs (475 base + abstentions oversampled 3×). Eval loss 1.906 → 0.758; eval token-acc 0.836. Full write-up in the repo `RESULTS.md`. - PEFT 0.19.1 · TRL 1.7.1 · Transformers 5.13.0 · PyTorch 2.12.1+cu130