初始化项目,由ModelHub XC社区提供模型
Model: nightmedia/Qwen3-4B-Agent-Eva Source: Original Platform
This commit is contained in:
78
README.md
Normal file
78
README.md
Normal file
@@ -0,0 +1,78 @@
|
||||
---
|
||||
license: apache-2.0
|
||||
language:
|
||||
- en
|
||||
- zh
|
||||
base_model:
|
||||
- janhq/Jan-v1-2509
|
||||
- Gen-Verse/Qwen3-4B-RA-SFT
|
||||
- TeichAI/Qwen3-4B-Instruct-2507-Polaris-Alpha-Distill
|
||||
- TeichAI/Qwen3-4B-Thinking-2507-Gemini-2.5-Flash-Distill
|
||||
- DavidAU/Qwen3-4B-Apollo-V0.1-4B-Thinking-Heretic-Abliterated
|
||||
- nightmedia/Qwen3-4B-Agent
|
||||
- FutureMa/Eva-4B
|
||||
pipeline_tag: text-generation
|
||||
library_name: transformers
|
||||
tags:
|
||||
- coding
|
||||
- research
|
||||
- deep thinking
|
||||
- 1M context
|
||||
- 256k context
|
||||
- Qwen3
|
||||
- All use cases
|
||||
- creative
|
||||
- creative writing
|
||||
- fiction writing
|
||||
- plot generation
|
||||
- sub-plot generation
|
||||
- story generation
|
||||
- scene continue
|
||||
- storytelling
|
||||
- fiction story
|
||||
- science fiction
|
||||
- all genres
|
||||
- story
|
||||
- writing
|
||||
- vivid prosing
|
||||
- vivid writing
|
||||
- fiction
|
||||
- roleplaying
|
||||
- bfloat16
|
||||
- finetune
|
||||
- mergekit
|
||||
- merge
|
||||
- mlx
|
||||
---
|
||||
# Qwen3-4B-Agent-Eva-qx86-hi-mlx
|
||||
|
||||
This is a model merge between Qwen3-4B-Agent and FutureMa/Eva-4B.
|
||||
|
||||
Brainwaves of qx86-hi quants of the parent models and Qwen3-4B-Agent-Eva
|
||||
```brainwave
|
||||
Agent 0.603,0.817,0.838,0.743,0.426,0.780,0.708
|
||||
Eva-4B 0.539,0.747,0.864,0.606,0.412,0.751,0.605
|
||||
|
||||
Qwen3-4B-Agent-Eva
|
||||
bf16 0.565,0.779,0.872,0.700,0.418,0.776,0.653
|
||||
qx86-hi 0.568,0.775,0.872,0.699,0.418,0.777,0.654
|
||||
```
|
||||
The Agent base is abliterated and contains only the essential models to top 0.6/0.8 arc--this level of performance is usually found in much larger models, if at all. Agent has an unforgiving nature, and here with the Eva merge, it will look for accountability :)
|
||||
|
||||
The qx86-hi quant performs at the same level with full precision in this model, and lower than the Agent base because Eva came with lower metrics to begin with.
|
||||
|
||||
The Element models are profiled to act as agents on the Star Trek DS9 station, in a roleplay scenario.
|
||||
|
||||
The models can be used for regular tasks as well.
|
||||
|
||||
Each comes with different skills. I found FutureMa/Eva-4B recently with an interesting model card:
|
||||
|
||||
> Eva-4B is a 4B-parameter model for detecting evasive answers in earnings call Q&A.
|
||||
|
||||
In Element8-Eva, that would be Quark. Element8 is a very rich merge, with lower metrics than Agent.
|
||||
|
||||
Like I mentioned on the Element8-Eva model card, the FutureMa/Eva-4B was simply included for conversational skills.
|
||||
|
||||
Without them, who do you get, Worf?
|
||||
|
||||
-G
|
||||
Reference in New Issue
Block a user