--- license: gemma base_model: perfectPresentation/rcrc-gemma-1b-cpt datasets: - perfectPresentation/rcrc-qa-v5 language: - ar - en tags: - chat - rcrc - arabic - gemma - gemma3 - closed-book library_name: transformers --- # rcrc-chat-v5-gemma-1b-cpt-sft — Closed-Book RCRC Chatbot A 1B closed-book chatbot for the Royal Commission for Riyadh City (RCRC). Trained as **Path B** of the v5 comparison: continued pre-training on raw KB text, then chat SFT on Qwen3-235B-synthesized QA pairs. Closed-book here means the model answers from baked-in knowledge — there is no retrieval at inference. ## Pipeline ``` google/gemma-3-1b-pt ↓ CPT: 3 epochs on cleaned RCRC + Hanifa raw text perfectPresentation/rcrc-gemma-1b-cpt ↓ SFT: 3 epochs on rcrc-qa-v5 (16,761 Qwen-synthesized QA pairs) perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft ← this repo ``` ## Training data [`perfectPresentation/rcrc-qa-v5`](https://huggingface.co/datasets/perfectPresentation/rcrc-qa-v5): 16,426 train + 335 validation single-turn (system, user, assistant) pairs synthesized by `Qwen/Qwen3-235B-A22B-Instruct-2507` from the cleaned RCRC website + Hanifa Urban Code chunks. ## Training recipe | | | |---|---| | Base | `perfectPresentation/rcrc-gemma-1b-cpt` | | Epochs | 3 | | LR | 2e-5, cosine, 5% warmup | | Effective batch | 16 (per_device 4 × grad_accum 4) | | Max seq length | 1024, packing enabled | | Final eval loss | 0.71 | | Final eval token-accuracy | 83.5% | | Hardware | HF Jobs · 1× L4 (~1.6 h) | ## Use ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch tok = AutoTokenizer.from_pretrained("perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft") model = AutoModelForCausalLM.from_pretrained( "perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft", dtype=torch.bfloat16, ) messages = [ {"role": "system", "content": "أنت مساعد للهيئة الملكية لمدينة الرياض. تجيب على استفسارات المستخدمين " "عن خدمات وبرامج ومشاريع وأنظمة الهيئة بدقة وأدب."}, {"role": "user", "content": "ما هو الكود العمراني لوادي حنيفة؟"}, ] prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tok(prompt, return_tensors="pt").to(model.device) out = model.generate( **inputs, max_new_tokens=400, do_sample=True, temperature=0.5, top_p=0.9, repetition_penalty=1.15, no_repeat_ngram_size=6, ) print(tok.decode(out[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)) ``` ## Honest evaluation: closed-book vs RAG On a 50-question internal eval set (RCRC + Hanifa + edge cases), this closed-book model was compared to a sibling RAG pipeline (the v3-rag checkpoint reading retrieved chunks from `rcrc-rag-index-v2`): | | accuracy | relevance | clarity | wins | |---|---:|---:|---:|---:| | RAG (v3-rag + index v2) | 3.92 | 4.08 | 4.42 | 34/50 | | **This model (closed-book)** | **2.64** | **3.36** | **3.56** | **13/50** | | Ties | | | | 3/50 | **Where closed-book is competitive**: dialect responses (Najdi/Hijazi), free-form opinion-style queries. **Where RAG dominates**: factual specifics from the Hanifa Urban Code, project details, organizational facts, numerical specs. For accuracy-sensitive deployments on the RCRC + Hanifa corpus, a RAG pipeline at the same parameter count outperforms this closed-book model by ~1.3 points on average. This model is provided for completeness of the v5 study and for offline / no-retrieval scenarios. ## Limitations - 1B-scale closed-book recall is brittle on specifics (numbers, exact procedure steps). Verify against rcrc.gov.sa. - Training data was Qwen-synthesized; some questions may carry the synthesizer's biases. - No multi-turn conversational SFT — single-turn QA only. ## GGUF builds (llama.cpp / Ollama) Quantized GGUF files live at the repo root. | File | Quant | Approx size | |---|---|---| | `rcrc-v5-gemma-1b-cpt-sft-F16.gguf` | F16 | ~2.0 GB | | `rcrc-v5-gemma-1b-cpt-sft-Q8_0.gguf` | Q8_0 | ~1.0 GB | | `rcrc-v5-gemma-1b-cpt-sft-Q5_K_M.gguf` | Q5_K_M | ~720 MB | | `rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf` | Q4_K_M | ~620 MB | ### Ollama (one-liner) ```bash ollama run huggingface.co/perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft:Q4_K_M # or :Q8_0, :Q5_K_M, :F16 ``` If you hit a host-redirect error (`hf.co` → `huggingface.co`), upgrade Ollama to a recent version, or use the `huggingface.co/...` URL above. ### Manual Modelfile route ```bash hf download perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf Modelfile --local-dir ./model cd model ollama create rcrc-v5-gemma-1b-cpt-sft -f Modelfile ollama run rcrc-v5-gemma-1b-cpt-sft ``` ### llama.cpp ```bash hf download perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf --local-dir . ./llama-cli -m rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf -cnv ``` ## GGUF builds (llama.cpp / Ollama) Quantized GGUF files live under `gguf/`. They are built directly from the safetensors above with `llama.cpp convert_hf_to_gguf.py` followed by `llama-quantize`. | File | Quant | Approx size | |---|---|---| | `gguf/rcrc-v5-gemma-1b-cpt-sft-F16.gguf` | F16 | ~2.0 GB | | `gguf/rcrc-v5-gemma-1b-cpt-sft-Q8_0.gguf` | Q8_0 | ~1.0 GB | | `gguf/rcrc-v5-gemma-1b-cpt-sft-Q5_K_M.gguf` | Q5_K_M | ~720 MB | | `gguf/rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf` | Q4_K_M | ~620 MB | ### llama.cpp ```bash hf download perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft gguf/rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf --local-dir . ./llama-cli -m gguf/rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf -cnv ``` ### Ollama ```bash hf download perfectPresentation/rcrc-chat-v5-gemma-1b-cpt-sft gguf/rcrc-v5-gemma-1b-cpt-sft-Q4_K_M.gguf gguf/Modelfile --local-dir ./model cd model/gguf ollama create rcrc-v5-gemma-1b-cpt-sft -f Modelfile ollama run rcrc-v5-gemma-1b-cpt-sft ```