--- license: apache-2.0 base_model: Tongyi-MAI/Z-Image-Turbo tags: - image-generation - prompt-engineering - qwen - photography new_version: BennyDaBall/Qwen3-4b-Z-Image-Engineer-V4 --- # ๐Ÿš€ Z-Engineer V2.5 (4B) **Follow me on X [@BennyDaBall_OG](https://x.com/BennyDaBall_OG) !** **The "Z-Engineer" is back โ€” longer, deeper, and smarter.** This is **Z-Engineer V2.5**, a specialized 4B parameter model fine-tuned on the **Qwen 3** architecture. It serves as a dedicated **Creative Director** for your image generation workflow, capable of extrapolating complex, cohesive visual narratives from minimal seed concepts. It doesn't just describe a scene; it engineers the light, lens, and atmosphere necessary to render it. ## ๐Ÿง  What is this? Z-Engineer V2.5 is a merged LoRA fine-tuned version of high-performance text encoder from `Tongyi-MAI/Z-Image-Turbo`. It has been trained to specifically understand the nuances of **AI Image Generation** (Z-Image-Turbo, Flux2 Klein). It excels at: * **Expanding Concepts**: Turn "dog on a bike" into a cinematic narrative. * **Technical Precision**: It understands lenses (35mm vs 85mm), lighting (rembrandt, volumetric), and film stocks. * **Stylistic Consistency**: It avoids the robotic "AI feel" and writes with a distinct, creative voice. ## ๐Ÿ”‘ Key Use Cases * **โœจ Prompt Enhancement**: A lightweight, low-VRAM solution to create, edit, and enrich simple image ideas into detailed narratives. * **๐Ÿ”Œ Z-Image Turbo Encoder**: Fully backwards compatible as a drop-in CLIP text encoder for Z-Image Turbo workflows, producing varied and unique results from the same seed. * **๐Ÿ›ก๏ธ Local & Private**: Runs entirely on your machine. No API fees, no data logging, no censorship. * **โšก Hybrid Power**: Use it to expand a prompt, then use the model itself as the encoder for the generation stage. ## ๐Ÿ“‰ Key Improvements * **Base Model Upgrade**: Switched from standard Qwen3 Instruct to the native text encoder from **Z-Image-Turbo** for perfect alignment. * **All-Layer Training**: Unlike typical lightweight LoRAs, I trained adapters on **all 36 layers** of the model, ensuring deep behavioral alignment. * **Massive Iteration Count**: Trained for **10,000 iterations** to fully saturate the weights with the dataset concepts. ## ๐Ÿ“Š CLIP Model Comparison Z-Engineer V2.5 can be used as a **drop-in CLIP text encoder** for Z-Image-Turbo workflows. Here's how it compares to previous versions and the base model: | Model | Result | | :--- | :--- | | **Z-Engineer V2.5** | โœ… Clean, natural output with excellent detail and coherence. | | **Z-Engineer V2** | โœ… Good quality, but V2.5 shows improved texture and lighting. | | **Z-Engineer V1** | โŒ **Broken**: Produces severe visual artifacts and distortions. | | **Base Qwen3 4B** | โš ๏ธ Functional but generic; lacks the specialized prompt understanding. | ### Visual Comparison ![CLIP Comparison 1](z-image-gguf-comparison_00023_.png) ![CLIP Comparison 2](z-image-gguf-comparison_00026_.png) *Note: V1 exhibits catastrophic artifacts (bottom-left in each grid) due to training instabilities. V2.5 (top-left) consistently produces the cleanest, most natural results.* ## ๐Ÿ”Œ ComfyUI Integration (Recommended) I have released a custom node for seamless integration with ComfyUI! * **Features**: Optimized for local OpenAI API compatible backends (LM Studio, Ollama, etc.). * **Get it here**: [ComfyUI-Z-Engineer](https://github.com/BennyDaBall930/ComfyUI-Z-Engineer) ## ๐Ÿ’ป Training Facts I believe in open science. Here is exactly how this was built: * **Hardware**: Trained locally on a Mac with **48GB Unified Memory** (Apple Silicon). * **Framework**: **MLX** (Apple's native machine learning framework). * **Dataset**: Generated locally using [Qwen3 VL 30B A3B Instruct](https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct) * **Size**: ~34,678 high-quality examples. * **Content**: A curated mix of "Prompt Enhancement" pairs, teaching the model how to take a seed idea and "engineer" it into a final prompt. * **Hyperparameters**: * **Iterations**: 10,000 * **Batch Size**: 4 * **LoRA Layers**: 36 (All Linear Layers) * **Learning Rate**: 1e-5 ## ๐Ÿ“ฆ GGUF & Quantization I provide a full suite of GGUF quantizations for use with `llama.cpp`, Ollama, and LM Studio. | Quantization | Size | Use Case | | :--- | :--- | :--- | | **Q4_K_S** | **2.2 GB** | ๐Ÿ”ป Max Compression | | **Q4_K_M** | **2.3 GB** | โšก๏ธ Fast / Mobile / Edge | | **Q5_K_M** | **2.7 GB** | โš–๏ธ **Recommended Balance** | | **Q6_K** | **3.1 GB** | ๐Ÿ’Ž High Quality | | **Q8_0** | **4.0 GB** | ๐ŸŽฌ Near-Lossless | | **F16** | **7.5 GB** | ๐Ÿงช Reference / Conversion | ## โš ๏ธ Disclaimer This model generates text for image prompts. While I have filtered the dataset, users should use their best judgment. I am not responsible for the content you generate. **Follow me on X [@BennyDaBall_OG](https://x.com/BennyDaBall_OG) !**