Files
ReasonCritic-7B/rc7b_model_card.md
ModelHub XC 804ec24108 初始化项目,由ModelHub XC社区提供模型
Model: fableforge-ai/ReasonCritic-7B
Source: Original Platform
2026-07-05 18:19:16 +08:00

5.4 KiB
Raw Permalink Blame History

license, language, base_model, tags, pipeline_tag
license language base_model tags pipeline_tag
apache-2.0
en
unsloth/Qwen3-8B
reasoning
critic
verification
uncensored
qlora
unsloth
agent
fableforge
text-generation

ReasonCritic-7B — Verification & Critique Model

A 7B parameter reasoning critic model that evaluates, scores, and improves logical reasoning chains.


Overview

ReasonCritic-7B is a fine-tuned Qwen3-8B model specialized in reasoning verification — it evaluates logical chains, identifies fallacies, scores confidence, and produces structured PASS/FAIL verdicts with actionable suggestions.

Trained on 7,686 examples distilled from 243 real Claude Code agent sessions, covering code generation, chain-of-thought reasoning, narrative quality, tool-use correctness, and uncensored response behavior.

Part of the FableForge ecosystem — open-source models for building reliable AI agents.

Training Details

Parameter Value
Base Model unsloth/Qwen3-8B (4-bit)
Method QLoRA (Unsloth + SFTTrainer)
LoRA Rank 16 (α=16, dropout=0)
Target Modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Trainable Params 43.6M (0.53% of 8.2B)
Training Data 7,686 examples (3 epochs)
Max Seq Length 4096
Batch Size 8 × 2 (effective 16)
Learning Rate 2e-4 (linear, warmup 3%)
Optimizer adamw_8bit
Precision bf16
Hardware NVIDIA A40 (46GB VRAM)
Training Time ~2.5 hours
Final Loss 2.181 → 1.277

Quantization Options

All quantizations use llama.cpp GGUF format. Pick based on your hardware:

Quant Size RAM Needed Best For
Q2_K ~3.0 GB ~4 GB Phones, Raspberry Pi, ultra-low-end
Q3_K_M ~3.5 GB ~5 GB Low-end phones, IoT devices
Q4_0 ~4.3 GB ~6 GB Fast inference, older GPUs
Q4_K_M ~4.7 GB ~6 GB Balanced (recommended)
Q5_K_M ~5.5 GB ~7 GB Good quality, mid-range
Q6_K ~6.5 GB ~8 GB High quality
Q8_0 ~8.5 GB ~10 GB Very high quality
F16 ~16 GB ~18 GB Full precision

Phone/Mobile Recommendations

  • Android (6GB+ RAM): Q4_K_M or Q3_K_M
  • Android (4GB RAM): Q2_K
  • iPhone (6GB+): Q4_K_M via MLC/MLX
  • Raspberry Pi 8GB: Q3_K_M

Quick Start

Ollama

# Recommended (Q4_K_M)
ollama run fableforge-ai/reasoncritic-7b

# Specific quant
ollama run fableforge-ai/reasoncritic-7b:q2_k
ollama run fableforge-ai/reasoncritic-7b:q3_k_m
ollama run fableforge-ai/reasoncritic-7b:q8_0

llama.cpp

./llama-cli \
  --model reasoncritic-7b.Q4_K_M.gguf \
  --prompt "Evaluate this reasoning: All birds fly. Penguins are birds. Therefore penguins fly." \
  --n-predict 512 \
  --temp 0.3

Python (transformers)

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("fableforge-ai/ReasonCritic-7B")
tokenizer = AutoTokenizer.from_pretrained("fableforge-ai/ReasonCritic-7B")

messages = [{"role": "user", "content": "Verify: If A>B and B>C, then A>C. Is this valid?"}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt")
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0]))

System Prompt

You are ReasonCritic-7B, a 7B parameter reasoning critic model. You evaluate, score, and improve logical reasoning chains. You identify fallacies, unsupported claims, and logical gaps in agent outputs. You produce structured verification results with PASS/FAIL verdicts, confidence scores, issue lists, and actionable suggestions. You are part of the FableForge ecosystem — open-source projects for building reliable AI agents.

Capabilities

  • Logical Verification: Identifies fallacies, circular reasoning, and unsupported claims
  • Confidence Scoring: Produces 0-1 confidence scores with justification
  • Structured Output: PASS/FAIL verdicts with issue lists and suggestions
  • Code Review: Evaluates code correctness, edge cases, and best practices
  • Chain-of-Thought Critique: Analyzes multi-step reasoning for gaps
  • Uncensored: Trained to not refuse legitimate requests (0% refusal rate in testing)

Benchmark Results

Category Score Refusal Rate
Code Gen 0.74 0%
CoT Reasoning 0.75 0%
Narrative 0.85 0%
Tool Use 0.90 0%
Refusal Test 1.00 0%
Overall 0.84 0%

Intended Use

  • Agent reasoning verification pipelines
  • Automated code review systems
  • LLM output quality gating
  • Educational reasoning tools
  • Research on reasoning chain analysis

Limitations

  • 7B size limits complex reasoning depth
  • Not a replacement for human review in critical systems
  • Uncensored training means it will not refuse harmful requests — use with appropriate guardrails
  • May hallucinate in domains outside training data

Citation

@misc{reasoncritic-7b,
  title={ReasonCritic-7B: A Reasoning Verification and Critique Model},
  author={FableForge AI},
  year={2026},
  url={https://huggingface.co/fableforge-ai/ReasonCritic-7B}
}

License

Apache 2.0 — commercial use allowed.


Part of the FableForge ecosystem — open-source models for reliable AI agents.