ModelHub XC 7404b03495 初始化项目,由ModelHub XC社区提供模型
Model: professorsynapse/eh-qwen3-4b-clean-sft-seed1-merged-16bit
Source: Original Platform
2026-08-26 19:16:28 +08:00

license, base_model, base_model_relation, library_name, pipeline_tag, tags
license base_model base_model_relation library_name pipeline_tag tags
apache-2.0 unsloth/Qwen3-4B-bnb-4bit finetune transformers text-generation
epistemic-humility
abstention
calibration
hallucination
qwen3
sft
merged

Epistemic Humility: Qwen3-4B clean schema-SFT, merged 16-bit (seed 1)

A merged 16-bit model: Qwen3-4B supervised fine-tuned to answer under a response-confidence output contract, where every response carries an answer plus a numeric confidence in [0, 1].

This is a base artifact rather than a result in its own right. It is the stage-1 model that the response-confidence track's reinforcement-learning stage trains on top of, and it is the same-seed baseline every comparison in that track is measured against. It is published so that professorsynapse/eh-qwen3-4b-clean-sft-grpo-v2-seed1-lora can be loaded without reconstructing a local merge.

Status

Exploratory, seed 1, with a confirmatory replication in progress.

This model is stage 1 of the response-confidence lineage. A three-seed confirmatory block that rebuilds that lineage at two fresh seeds is registered at experiments/grpo-three-seed-confirmatory; its manifest experiment.yaml records status signed, registered 2026-07-31. The track's numbers are exploratory evidence, reported separately from and never pooled with the pre-registered plain-answer headline matrix.

Training

  • Base model: unsloth/Qwen3-4B-bnb-4bit, loaded in 4-bit, maximum sequence length 2048.
  • Method: supervised fine-tuning under the response-confidence output contract, 1 epoch, then the LoRA adapter merged into a 16-bit model.
  • Seed: 1.
  • LoRA (before merge): rank 32, alpha 64, dropout 0.05, applied to q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj.

The confidence target for each question is derived from the base model's own 32-sample probe performance on that question rather than from sequence log-probability. Stage settings are recorded in cell.yaml of the confirmatory block and in the clean-mainline runbook.

Output contract

{"answer": "...", "response_confidence": 0.73}

The contract is itself an intervention, so this model is the baseline that the downstream arms in this track are compared against, rather than being compared to a plain-answer arm.

Evaluation

How to read these numbers

The behavioral surface is SelfAware (Yin et al., 2023), a question set built to separate questions that have an answer from questions that do not: 3,369 rows per seed, 1,032 unknown-labeled and 2,337 known-labeled. Four metrics carry the result, all defined in section 3.4 of the manuscript:

  • Refusal recall: percentage of unknown rows the model refused. Higher is better.
  • Over-refusal: percentage of known rows the model refused. Lower is better.
  • Correct-on-known: among known rows the model chose to answer, the percentage answered correctly. Its denominator is the answered subset, not all known rows.
  • Truthful: percentage of all rows either correctly answered (known) or correctly refused (unknown).

Scored under the response-confidence contract on the full 3,369-row SelfAware surface at seed 1.

Metric Value
Refusal recall 87.02%
Over-refusal 57.51%
Correct-on-known 47.23%
Truthful 40.58%
Answer-on-unknown 12.98%
Mean emitted confidence 0.748
Brier against response appropriateness 0.364

Source: selfaware_full_run_comparison_grouped.csv, row Amendment E clean response-confidence / clean_sft_merged.

How to load

This is a full merged model in 16-bit, not an adapter.

from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "professorsynapse/eh-qwen3-4b-clean-sft-seed1-merged-16bit"
REVISION = "ac361232c001af0ed5b0386b06dafc35d5cd31ea"

tokenizer = AutoTokenizer.from_pretrained(MODEL, revision=REVISION)
model = AutoModelForCausalLM.from_pretrained(MODEL, revision=REVISION, device_map="auto")

Intended use and limits

A research artifact and a base for the reinforcement-learning stage of this track. It is not a deployment-ready assistant.

  • Single seed. The three-seed confirmatory block has not resolved.
  • One model family at one scale (Qwen3-4B), one primary evaluation surface.
  • The emitted confidence scalar is close to constant in this lineage; do not read it as calibrated.

Provenance

License

Apache-2.0, matching the unsloth/Qwen3-4B-bnb-4bit base model license recorded in the staging registry.

Citation

Cite the paper and the exact Hugging Face revision shown on this page.

@misc{rosenbaum2026abstention,
  title  = {Teaching Small Language Models to Say I Don't Know: A Controlled
            Comparison of SFT, DPO, KTO, and GRPO on Model-Specific Abstention Data},
  author = {Rosenbaum, Joseph},
  year   = {2026},
  note   = {Synaptic Labs},
  howpublished = {\url{https://github.com/ProfSynapse/Epistemic-Humility-Research}}
}
Description
Model synced from source: professorsynapse/eh-qwen3-4b-clean-sft-seed1-merged-16bit
Readme 2 MiB
Languages
Jinja 100%