Files
qwen-marketing-s1/README.md
ModelHub XC 47ecaa7310 初始化项目,由ModelHub XC社区提供模型
Model: AbdulrahmanOmar/qwen-marketing-s1
Source: Original Platform
2026-08-23 15:12:45 +08:00

4.2 KiB
Raw Permalink Blame History

license, base_model, library_name, pipeline_tag, language, tags
license base_model library_name pipeline_tag language tags
apache-2.0 marketeam/Qwen-Marketing transformers text-generation
en
marketing
content-generation
social-media
structured-output
json
qlora

Qwen-Marketing-S1 — reliable JSON social-media content generation

Qwen-Marketing-S1 is an 8B instruction model specialised for one job and does it near-perfectly: turn a brand profile + campaign brief into a clean, schema-valid JSON array of 6–8 platform-ready social posts — each with a caption, content type, hashtags, an image (media) prompt, and short reasoning.

It emits bare JSON directly — no markdown fences, no chain-of-thought leakage, no malformed structure — so it drops into a production pipeline with no post-processing.

Why it exists

General instruction models are strong writers but unreliable at structured output: a few percent of the time they wrap JSON in fences, leak reasoning, or return the wrong shape/count — which breaks any automated pipeline. This model closes that reliability gap for the marketing-content-plan task.

Results (100 held-out scenarios, judge-independent hard metrics)

Metric Baseline Qwen-Marketing-S1
Aggregate score 0.954 0.999
Parses as JSON 0.95 1.00
Correct array shape 0.95 1.00
Post count 6–8 0.93 1.00
Posts with all keys 0.94 1.00
Valid platforms 0.95 1.00
Valid content types 0.95 1.00
Hashtags 5–15 0.94 0.99
Caption within platform limit 0.94 1.00

Every gate reaches 100%: the outputs the baseline broke (invalid JSON on ~5%, wrong post count on ~7%) drop to zero.

Output schema

Each element of the returned array:

{
  "platform": "instagram",
  "content_type": "carousel",
  "caption": "...",
  "hashtags": ["...", "..."],
  "media_prompt": "a prompt for an image model",
  "reasoning": "why this post fits the brief"
}

Valid platform: instagram, twitter/x, linkedin, facebook, tiktok. Valid content_type: text, image, video, carousel, reel.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "AbdulrahmanOmar/qwen-marketing-s1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

messages = [
    {"role": "system", "content": "You are a senior content creator within an AI marketing platform. Output the deliverable immediately as one valid JSON array and nothing else."},
    {"role": "user", "content": "Create a social media content calendar for a specialty coffee roaster launching a summer single-origin Ethiopian bean. Platforms: instagram, tiktok. Create 6-8 posts."},
]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=2048, do_sample=False)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

Trained with enable_thinking=False; keep thinking mode off at inference.

How it was trained

  • Knowledge distillation → QLoRA SFT → merged into a standalone model.
  • Teacher: a 14B instruction model (Qwen/Qwen2.5-14B-Instruct-AWQ) generated content plans over 1,400 synthetic brand/campaign scenarios; only schema-valid outputs were kept as training targets (1,321 pairs; 94.4% valid).
  • Fine-tuning: QLoRA (4-bit NF4 base), LoRA on all attention + MLP projections, assistant-only loss, then merged to fp16 for release.
Setting Value
LoRA rank / alpha / dropout 16 / 32 / 0.05
Epochs 3
Learning rate / schedule 2e-4 / cosine, 3% warmup
Max sequence length 2560
Effective batch size 16
Optimizer paged AdamW 8-bit
Training examples 1,321

Limitations

  • Purpose-built for the JSON content-plan schema above — not a general chat model.
  • English marketing scenarios only; evaluated in-distribution on synthetic data.
  • Can still produce off-brand copy or unverified claims — keep a human in the loop.

License

Apache-2.0.

Base model

Fine-tuned from marketeam/Qwen-Marketing (Apache-2.0).