127 lines
4.2 KiB
Markdown
127 lines
4.2 KiB
Markdown
---
|
||
license: apache-2.0
|
||
base_model: marketeam/Qwen-Marketing
|
||
library_name: transformers
|
||
pipeline_tag: text-generation
|
||
language:
|
||
- en
|
||
tags:
|
||
- marketing
|
||
- content-generation
|
||
- social-media
|
||
- structured-output
|
||
- json
|
||
- qlora
|
||
---
|
||
|
||
# Qwen-Marketing-S1 — reliable JSON social-media content generation
|
||
|
||
**Qwen-Marketing-S1** is an 8B instruction model specialised for one job and does
|
||
it near-perfectly: turn a **brand profile + campaign brief** into a **clean,
|
||
schema-valid JSON array of 6–8 platform-ready social posts** — each with a
|
||
caption, content type, hashtags, an image (media) prompt, and short reasoning.
|
||
|
||
It emits **bare JSON directly** — no markdown fences, no chain-of-thought
|
||
leakage, no malformed structure — so it drops into a production pipeline with no
|
||
post-processing.
|
||
|
||
## Why it exists
|
||
|
||
General instruction models are strong writers but unreliable at *structured
|
||
output*: a few percent of the time they wrap JSON in fences, leak reasoning, or
|
||
return the wrong shape/count — which breaks any automated pipeline. This model
|
||
closes that reliability gap for the marketing-content-plan task.
|
||
|
||
## Results (100 held-out scenarios, judge-independent hard metrics)
|
||
|
||
| Metric | Baseline | **Qwen-Marketing-S1** |
|
||
|---|---|---|
|
||
| **Aggregate score** | 0.954 | **0.999** |
|
||
| Parses as JSON | 0.95 | **1.00** |
|
||
| Correct array shape | 0.95 | **1.00** |
|
||
| Post count 6–8 | 0.93 | **1.00** |
|
||
| Posts with all keys | 0.94 | **1.00** |
|
||
| Valid platforms | 0.95 | **1.00** |
|
||
| Valid content types | 0.95 | **1.00** |
|
||
| Hashtags 5–15 | 0.94 | **0.99** |
|
||
| Caption within platform limit | 0.94 | **1.00** |
|
||
|
||
Every gate reaches 100%: the outputs the baseline broke (invalid JSON on ~5%,
|
||
wrong post count on ~7%) drop to **zero**.
|
||
|
||
## Output schema
|
||
|
||
Each element of the returned array:
|
||
|
||
```json
|
||
{
|
||
"platform": "instagram",
|
||
"content_type": "carousel",
|
||
"caption": "...",
|
||
"hashtags": ["...", "..."],
|
||
"media_prompt": "a prompt for an image model",
|
||
"reasoning": "why this post fits the brief"
|
||
}
|
||
```
|
||
|
||
Valid `platform`: instagram, twitter/x, linkedin, facebook, tiktok.
|
||
Valid `content_type`: text, image, video, carousel, reel.
|
||
|
||
## Usage
|
||
|
||
```python
|
||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||
import torch
|
||
|
||
model_id = "AbdulrahmanOmar/qwen-marketing-s1"
|
||
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
||
model = AutoModelForCausalLM.from_pretrained(
|
||
model_id, torch_dtype=torch.bfloat16, device_map="auto"
|
||
)
|
||
|
||
messages = [
|
||
{"role": "system", "content": "You are a senior content creator within an AI marketing platform. Output the deliverable immediately as one valid JSON array and nothing else."},
|
||
{"role": "user", "content": "Create a social media content calendar for a specialty coffee roaster launching a summer single-origin Ethiopian bean. Platforms: instagram, tiktok. Create 6-8 posts."},
|
||
]
|
||
inputs = tokenizer.apply_chat_template(
|
||
messages, add_generation_prompt=True, return_tensors="pt"
|
||
).to(model.device)
|
||
out = model.generate(inputs, max_new_tokens=2048, do_sample=False)
|
||
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
|
||
```
|
||
|
||
Trained with `enable_thinking=False`; keep thinking mode off at inference.
|
||
|
||
## How it was trained
|
||
|
||
- **Knowledge distillation → QLoRA SFT → merged** into a standalone model.
|
||
- **Teacher:** a 14B instruction model (`Qwen/Qwen2.5-14B-Instruct-AWQ`) generated
|
||
content plans over 1,400 synthetic brand/campaign scenarios; only schema-valid
|
||
outputs were kept as training targets (1,321 pairs; 94.4% valid).
|
||
- **Fine-tuning:** QLoRA (4-bit NF4 base), LoRA on all attention + MLP
|
||
projections, **assistant-only loss**, then merged to fp16 for release.
|
||
|
||
| Setting | Value |
|
||
|---|---|
|
||
| LoRA rank / alpha / dropout | 16 / 32 / 0.05 |
|
||
| Epochs | 3 |
|
||
| Learning rate / schedule | 2e-4 / cosine, 3% warmup |
|
||
| Max sequence length | 2560 |
|
||
| Effective batch size | 16 |
|
||
| Optimizer | paged AdamW 8-bit |
|
||
| Training examples | 1,321 |
|
||
|
||
## Limitations
|
||
|
||
- Purpose-built for the JSON content-plan schema above — not a general chat model.
|
||
- English marketing scenarios only; evaluated in-distribution on synthetic data.
|
||
- Can still produce off-brand copy or unverified claims — keep a human in the loop.
|
||
|
||
## License
|
||
|
||
Apache-2.0.
|
||
|
||
## Base model
|
||
|
||
Fine-tuned from `marketeam/Qwen-Marketing` (Apache-2.0).
|