--- license: mit base_model: google/gemma-2b tags: - text-generation-inference - transformers - gemma - mini-gemma - agentic-ai model_type: gemma pipeline_tag: text-generation --- # Mini-Gemma Custom Model This repository contains a custom domain-specialized fine-tune of the Gemma architecture, optimized for specific text distributions and patterns. The model was trained using the Hugging Face `Trainer` on an accelerated NVIDIA GPU cluster. ## 📊 Training Performance & Metrics The model successfully converged over its training run with highly stable gradients: * **Total Training Steps:** 20,000 * **Final Total Train Loss:** `3.478` * **Final Step Loss:** `2.988` * **Gradient Norm Stability:** Stable at `~1.12` * **Training Status:** Complete / Fully Converged ## 🚀 Quick Start & Usage You can easily load and run this model locally using the Transformers library: ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM, pipeline model_id = "agentbyumer/mini-gemma" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.float16, device_map="auto" ) generator = pipeline("text-generation", model=model, tokenizer=tokenizer) prompt = "Your specialized prompt here" outputs = generator( prompt, max_new_tokens=150, do_sample=True, temperature=0.7, return_full_text=False ) print(outputs[0]['generated_text']) ``` ## 📜 License This project is licensed under the permissive MIT License. See the accompanying [LICENSE](./LICENSE) file for full details.