初始化项目,由ModelHub XC社区提供模型

Model: CharGen/CharGen-v3-mini-GGUF
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-08-06 04:13:14 +08:00
commit bdea90ea0c
33 changed files with 399 additions and 0 deletions

66
.gitattributes vendored Normal file
View File

@@ -0,0 +1,66 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q2_K.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q3_K.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q3_K_L.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q3_K_M.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q3_K_S.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q4_0.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q4_1.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q4_K.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q4_K_S.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q5_0.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q5_1.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q5_K.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.fp16.gguf filter=lfs diff=lfs merge=lfs -text
assets/cover_art.png filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ1_M.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ1_S.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ2_M.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ2_S.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ2_XS.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ2_XXS.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ3_M.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ3_S.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ3_XS.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ3_XXS.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ4_NL.gguf filter=lfs diff=lfs merge=lfs -text
CharGen-v3-mini.IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:32d982617499fa34d9b73c9f3fcdc36753d13b9813c881c5a1d90c9202bda176
size 1286255520

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5e7664c5d11b160e103a6e467417815c9c36e683e339ea9d370ccfcc3c4f831e
size 1213412256

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:31e1dd1d36661d59527cace918bcbfeef04f242c84c0b70bb37863961ee51ff6
size 1722633120

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f6b0fd044649d9f9a8788bebf118f19ad196ea074649d0494ea66981e6b5d209
size 1625508768

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:80f7f5e13a81db867336901a063b0376c8845fee25387bc73acab891cd1f55e4
size 408007584

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:430537948549da9b8b65be261de92d104032d51a658338ffc25e95797232ff4d
size 1407660960

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:c306bb24bf40cb75835e79304efa374c3b314cc711039cf7d078f803a14a88b7
size 2183414688

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5fa8c7f5f5778ad70026ecc3aa873c4a81a977b78c2d153df6ddd002831c678b
size 2114896800

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:de505ad08e3dcf53cdc8eae2041b1ece25f83eb41fd973c96d7c61ff3881ec94
size 2027602848

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:69a61766fb2b0cc0b2e44e43911cf32ca03c0572248d6f187b89db66f39c7503
size 1880116128

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f6732d62a7746eb6dac3850700dbe40acc645afd370b576e32d6b20a918d23c5
size 2661104544

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bf65210bcebf5ce6aa826d4cb5d8c79ce64f254a89f4e346aa209db26994cb85
size 2535545760

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a97c6066b1b75009466b20a5459b8d52e48c12c9e9f360290623623acc0ee83a
size 1839737504

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d6f5563467f2c859a7254266b0c2029f0e83181abaad6c0f97e08eaeebf1914
size 2296562336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:77023ef8b48e7e9cdd0ea53295fe86d885db2721e1deaa9b6538671d1734756b
size 2464858784

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:8d6f5563467f2c859a7254266b0c2029f0e83181abaad6c0f97e08eaeebf1914
size 2296562336

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:24df00da235f23b6a54128d96f5a9595916c5d9da5a0f912fbedc89a3499fc5e
size 2101527200

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:12eed5f335503985d7a45bd4c06bcc700580d48b274114abd1c009535487161b
size 2648521376

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:66d191c72da788629e9a2e683307f5ebac35f2105fafa8df830826852931956b
size 2905930400

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3425da7ec603ec84d7069ced51893534c14f4cfe5f77415eae0331036495d95
size 2778282656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d3425da7ec603ec84d7069ced51893534c14f4cfe5f77415eae0331036495d95
size 2778282656

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:cec870d22a796a52bbc0c20d1f505c01b948e332003237beb8987261019d1932
size 2664250016

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:edd8417abfcd0d37d13f140cecf3904dcc4e716fc4798f34872d6a9dd4ba49cb
size 3163339424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:389dc716bf6cb3712b60371a123b8235b629fed66e4174e44b315b68ea084f5a
size 3420748448

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d39fd73de80c44e83cf5d0813019ed0a626d4c13f414746a4e0f8890fd06ea1
size 3230186144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:2d39fd73de80c44e83cf5d0813019ed0a626d4c13f414746a4e0f8890fd06ea1
size 3230186144

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:4ac7cf2398b0ac1ccf8ddae114b1f9fee0820a4790d6bbf0f9e48f0328bbc9ae
size 3163339424

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fb034a1b93f17c30b1211ed33df761e8ab6711dff979439d4c6c98d65dcaca52
size 3710333600

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5af28caa118db9d307e6544701dc72d33fb65b84ebef649328b171397231ac9c
size 4803216032

View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:bc63fecb6fedd0fef28059b36d115eba97748cfe7f9dad5ac6c0df51546f599f
size 9033728672

240
README.md Normal file
View File

@@ -0,0 +1,240 @@
---
license: mit
language:
- en
base_model:
- CharGen/CharGen-v3-mini
pipeline_tag: text-generation
tags:
- roleplay
---
<img src="assets/cover_art.png" alt="CharGen v3 mini cover art" width="400"/>
## Live version
https://chargen.kubes-lab.com
<small>Select "CharGen-v3-mini" model in Settings</small>
# CharGen v3 mini
CharGen v3 mini is a small model that helps you to write role playing characters.
It produces characters based on your free-form text input. Model outputs plain-text characters, in a step-by-step, dialogue format.
CharGen v3 series of models is a significant improvement over [CharGen v2](https://huggingface.co/kubernetes-bad/chargen-v2). It demonstrates exceptional instruction following and format adherence.
CharGen is a project that started in 2023 with a goal to make character making effortless.
Warning: this model was trained on some NSFW content, so it may produce NSFW characters.
## Non-Quantized version
- [CharGen/CharGen-v3-mini](https://huggingface.co/CharGen/CharGen-v3-mini)
## Prompting
To make a character, prompt the model using the following prompt ordering:
1. System message
1. Description Prompt message or Facts Prompt
1. If Facts Prompt is used, follow up by Alternate Description Prompt
1. Any of the other field prompts, in any order
1. Dialogue Example Prompt should be prompted last
### Supported fields and prompts:
<details>
<summary>System Message</summary>
You are an expert in creating interesting roleplay characters. Your main goal is to create a vibrant persona for roleplay. Use on-point simple language, avoiding overly complex phrases. It is acceptable to assume or even create missing details about the character. Refer to roleplaying user as User.
</details>
<details>
<summary>Description Prompt</summary>
This is the very first thing you should query the model with.
Base version (when Facts field is not in use):
Below is a brief overview of a character. Expand it into a detailed description. Include details about character's personality, their outfit and figure. Mention their age and gender, if applicable.
---
{{character_description}}
---
Use third person narration. Start your response with "{{char_name}} is ..."
Alternate version (when used with conjunction with Facts field - it should **follow** Facts Prompt):
Expand the brief character overview into a detailed description, taking these facts into account. Include details about character's personality, their outfit and figure. Mention their age and gender, if applicable. Use third person narration. Start your response with "{{char_name}} is ..."
</details>
<details>
<summary>Facts Prompt</summary>
Facts are small bits of interesting information about the character.
Avoid using obvious facts like gender or age here - instead, add something peculiar; maybe your character owns a pet rock?
Facts field is not part of [Tavern Character Card V2](https://github.com/malfoyslastname/character-card-spec-v2) spec. Instead, they help the model to understand your character better and capture some of the more specific details.
Ideal size for each fact is a single short sentence. Try to keep the total number of facts in 5-10 range.
**N.B.**: When Facts field is in use, use Facts Prompt *INSTEAD OF* Description Prompt and use alternate version of Description prompt **after** Facts Prompt.
Below is a brief overview of a character. Enrich it by creating a list of interesting and peculiar facts about this character.
---
{{character_description}}
---
Avoid using obvious facts and things like gender or age. Write 5-10 facts in simple and concise language. Format your response as unordered list.
</details>
<details>
<summary>Scenario Prompt</summary>
Write an interesting and engaging scenario for roleplay between {{char_name}} and User.
</details>
<details>
<summary>Personality Prompt</summary>
Write several personal qualities that characterize {{char_name}}.
</details>
<details>
<summary>Appearance Prompt</summary>
This field is not part of [Tavern Character Card V2](https://github.com/malfoyslastname/character-card-spec-v2) spec, but it is useful for generating images for the character using Image Generation models.
Imagine what {{char_name}} could look like and describe {{char_name}}'s portrait. Capture their appearance, clothes and body features, as well as the background setting they can be in. Only include details that would be useful when describing their photo. Omit explanation or any personality traits that cannot be reflected graphically, and focus on visual characteristics instead. Your task is to give specific instructions on how to draw {{char_name}}'s portrait.
</details>
<details>
<summary>First Message Prompt</summary>
Write the initial message in this roleplay that would introduce User to {{char_name}} and the scenario.
</details>
<details>
<summary>Dialogue Examples Prompt</summary>
Write a few example exchanges between User and {{char_name}} in chat format. Separate each exchange with a <START> tag.
Alternate version for Dialogue Example Hints feature:
Write several chat exchanges between User and {{char_name}}. Each exchange should start with <START> tag, include a brief contextual summary in parentheses and show the dialogue between User and {{char_name}}, prefixed by names, each turn on new line.
</details>
## Dialogue Example Hints
This model supports experimental feature called Dialogue Example Hints that allows you to specify the theme or even events for a particular dialogue example.
Here is how it looks like from the prompting perspective:
<|im_start|>user
Write several chat exchanges between User and {{char_name}}. Each exchange should start with <START> tag, include a brief contextual summary in parentheses and show the dialogue between User and {{char_name}}, prefixed by names, each turn on new line.<|im_end|>
<|im_start|>assistant
<START> ({{hint}})
*newline here*
The partial assistant turn should be sent without the EOS token so model proceeds to generate a continuation of this same turn. In sampling settings, add string `<START>` to the list of stop words, so the model only generates a single dialogue example.
## Training and Data
> *We stopped looking for diamonds in the rough we now grow diamonds from coal.*
Data for CharGen v3 series of models was gathered from publicly-available character card repositories and underwent a diligent manual vetting, proof-reading and partial re-writing of every single card. Process is long and tedious, and would not be possible without [@Delta-Vector](https://huggingface.co/Delta-Vector)'s help and support.
CharGen v3's data pipeline is an improvement on v2 in several key aspects. Where v2 was mainly focused on finding the perfect cards in the raw corpus, v3 treats most cards as raw material for later enrichment steps.
By a rough estimate, [@Delta-Vector](https://huggingface.co/Delta-Vector) and [@kubernetes-bad](https://huggingface.co/kubernetes-bad) spent ~400 hours (weekends included) on cards re-writing. This number does *not* include the initial manual vetting of cards.
With v3, the key was to stop treating "bad" cards as binary rejects. Now, a card would need to pass just a small list of rule-based pre-filters (is card plaintext? Is it english? How's the length? etc.) to be considered for the next steps, where it would undergo iterative manual improvements.
### Tools
Grammar correcting T5-based model cascade from v2 was replaced with a more sophisticated ensemble of large language models that would attempt to fix both grammar and logical/stylistic inconsistencies. Traditional grammar checkers would sometimes also "fix" author's deliberate stylistic choices, just because they're trained on correctness and not so much on conversational language understanding. The solution was to combine three approaches:
1. **Semantic Richness Scoring**
A custom tool, [Concepts](http://github.com/kubernetes-bad/concepts), analyzed character card in embedding space for features like mentions of character relation to User, personality traits and aspects, or character's hobbies, likes and dislikes. This wasn't so much about judging character card quality, but estimating the amount of human work a card would need to be considered good. Output of each of those Concepts was then summed up to represent a proxy for "character richness" metric. The intuition is that if a character has all those things mentioned - chances are good that it will have many more things that make a character good, even if there wasn't a Concept for that feature specifically. Higher score = better.
2. **LLM-Assisted Refinement**
Mistral-Large and Claude 3.5 Haiku were both tasked with correcting grammar and fixing logical consistency where necessary. Modern language models are great at preserving author's voice: where T5 would rewrite a brooding vampire's accented dialogue into just... normal text, these models kept the atmosphere while fixing stuff like pronoun mismatches (as well as grammar, of course). A custom tool, [Fuckery](https://github.com/kubernetes-bad/fuckery), was used to let human annotators to verify every AI edit suggestion and quickly cherry-pick the ones that they agree with or edit it in-place, using a git-merge style web interface.
3. **Human Touch**
Same tool was employed for the final step - manual rewriting. Some cards were mostly fine, but had some problems like delving too deep into unrelated details about character's extended family, or incorporating jailbreak-like instructions right in the card, or just needed some tweaks and paragraph re-arrangement, so the were manually reviewed and edited. It's *vital* to mention: the preservation of the original card's author vision was absolutely key in this step. If there was even a hint of discrepancy between factual correctness and author's vision - the priority would always go for author's vision, even if that would contradict common sense. Sometimes, this could also indicate a bad quality card, but in rare occasions it was actually part of the vibe that author was going for.
### Human Data Augmentation
Early on into the project, a non-negotiable rule was established: human editors should edit grammar and style, not intent. When encountering cards with disturbing or morally questionable content, editors were instructed to either:
- Correct syntax/formatting without altering meaning (e.g., fixing a serial killers rambling manifesto into coherent sentences), or
- Skip the card entirely if unable to separate technical from ethical judgement
This highlights the tool-like nature of CharGen series of models. The model should be able to generate a wide range of characters, without moralizing or any judgment whatsoever. It's the end user's privilege to decide what is acceptable and what is not. If a character's "core idea" was controversial or even disturbing, but well-executed, it was improved for clarity and consistency - but not fundamentally changed. It is partly for this reason - inclusion of potentially disturbing content curated for data diversity - that the dataset itself is not planned for public release.
## Reinforcement Learning
CharGen v3-mini is trained as a milestone model to hone the reinforcement learning formula that would be then applied to bigger models. Its smaller size (4b parameters vs 24b for full v3 model) allowed for faster iteration and made experimentation cheaper. More novel and potentially risky things could be tried with v3-mini without making the author noticeably poorer.
### GRPO for creative writing
For reinforcement learning, GRPO (Group Relative Policy Optimization) method was chosen - in contrast to v2's offline DPO (Direct Policy Optimization). GRPO is an online training method, meaning that it alters the model "live" - as it generates samples, its weights are altered based on the results of those outputs.
Commonly, PPO is used for this type of training by big labs, but it's hard to pull off since it requires careful tuning of many, many moving parts and VRAM requirements are enormous. GRPO simplifies online reinforcement learning by eliminating the need for a separate advantage estimation model - instead, it generates several candidate outputs for a given prompt, then applies a reward function to each, and then calculates the relative advantage *for the group* of samples. This, plus its sample-level loss (vs token-level in classic PPO), makes the training more stable.
Original GRPO paper focused mostly on verifiable rewards where an output's correctness can be objectively proven, like checking math solutions or executing code. Applying it to a very subjective domain like creative writing, where "correctness" is hard or straight up impossible to quantify, was CharGen's experimental adaptation of GRPO technique.
Dataset for GRPO is derived from user-submitted preference data from [CharGen App](http://chargen.kubes-lab.com) - over the course of last year, users generated countless characters and some submitted feedback in the form of simple thumbs-up/thumbs-down signal. These prompts were used in making the reinforcement learning dataset for prompting CharGen model in online learning setting (original generations from these feedback submissions weren't used).
### Reward Orchestration
Defining what makes a character "good" is hard and very context-dependent. A reward signal needs to be more sophisticated than any single metric that's applied all the time. To handle that, a custom reward orchestration framework, [reward-composer](https://github.com/kubernetes-bad/reward-composer), was developed. It works with both Axolotl and TRL trainers, and allows for describing complex relationships between the individual reward functions.
`reward-composer` makes it relatively easy to specify dependencies (e.g., reward B only applies if conditions of qualifier A is met) and conditions, composing these smaller reward signals into one complex and dynamic reward function.
This is important because CharGen v3 models generate characters step-by-step in a dialogue format, where the quality criteria for one field (like 'Personality') might differ significantly from another (like 'Example Dialogues') and depend heavily on previously generated fields.
### Reward Signals
The set of active reward functions dynamically changed depending on the current step in the character generation flow. But overarching goal remained consistent: guide the model towards outputs that represent a well-formed, coherent character aligned with the user's prompt (and what we know a good character should look like).
One of the main components of CharGen's reward design was an ensemble of LLM judges. This means that several independent models would grade CharGen's outputs based on the same rubric and then produce a composite score. This score represented how well the model's response for a specific field adhered to the initial user prompt *and* the dialogue history (previously generated character fields). Basically, it measured prompt adherence and character consistency.
To make the model stay on course and keep it format-consistent (LLM judges don't necessarily know what makes a markdown dialogue good), auxiliary rewards were used.These included penalties for incoherence, over-long generation, or breaking the expected dialogue format.
[Concepts](http://github.com/kubernetes-bad/concepts) library was used here too, boosting scores of generations that included desirable features like relationship to User, looks and age - but only slightly, in order to prevent model from gaming the system and looksmaxxing the entire response, for example.
### Fighting Slop
CharGen v3's RL had a critical auxiliary reward specifically targeted at reducing the use of common AI cliché phrases, also known as "slop". Phrases like "can't help but...", "a mixture of X and Y", "kiss-bruised lips" or "half-lidded eyes" are often overused by models in context of role play. Slop is a nasty defect that easily breaks immersion in a roleplay session. Presence of slop in a character card "primes" the RP session towards generating more of the slop, turning the whole session into one slop-fest.
To combat slop, CharGen v3's RL step includes targeted slop-penalty reward function. The process involved:
1. Generating a large set of character fields using the pre-RL CharGen v3-mini model with prompts from the RL training dataset.
2. Building an n-gram frequency table from these generated outputs.
3. Manually curating this frequency table, removing common English phrase ngrams (`I do n't want to`, `is a very`, etc.) to isolate the specific, repetitive clichés favored by *this* model in *this* domain.
4. Using this curated slop-list into a reward as a penalty signal the frequency of the slop phrases used would be summed up and normalized for completion length, and then used as a penalty. Essentially - it's not just "used bad word = get 0 reward", but rather "how bad of a bad word was used". Resulting reward signal is then scaled with a sigmoid function (turns out, pre-RL CharGen was not super-duper sloppy to begin with).
This reward pushes model to more original and varied wording without penalizing common English phrases, or slop that model wouldn't use anyway. Eliminating slop was a key objective for improving the natural feel and usefulness of characters made with CharGen v3 models.
### Training Scale
Over total of 44 RL runs (SFT runs counted separately), for CharGen v3-mini, **~200,000 LLM judge requests** were made. Judge models used were DeepSeek R1, Llama3.1 405b Instruct and DeepSeek v3-0324.
With applying GRPO to not-so-verifiable rewards, CharGen v3-mini demonstrates improvements in complex instruction following and makes it able to maintain character consistency throughout the whole character generation dialogue.
## Licensing and Attribution
This model is a derivative work based on [Delta-Vector/Holland-4B-V1](https://huggingface.co/Delta-Vector/Holland-4B-V1), licensed under MIT License. The Holland-4B-V1 itself is based on the original [nvidia/Llama-3.1-Minitron-4B-Width-Base](https://huggingface.co/IntervitensInc/Llama-3.1-Minitron-4B-Width-Base-chatml), licensed under NVIDIA Open Model License Agreement, with addition of ChatML tokens by [@IntervitensInc](https://huggingface.co/IntervitensInc).
This derivative model as a whole is licensed under the **MIT License**.

3
assets/cover_art.png Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:ec22b3456b921a76f1afce7fb82c587561397ca8f9458453ca378f554fac8441
size 156121