--- base_model: unsloth/phi-3-mini-4k-instruct-bnb-4bit tags: - text-generation-inference - transformers - unsloth - phi3 - finance - lora - qlora - finetuned license: apache-2.0 language: - en pipeline_tag: text-generation --- # 💹 Phi-3-mini Finance LoRA (fp16) A domain-specialized version of Microsoft's **Phi-3-mini-4k-instruct**, fine-tuned on financial Q&A data using **LoRA + 4-bit NF4 quantization** — trained entirely on a free Google Colab T4 GPU. [](https://github.com/unslothai/unsloth) --- ## 📋 Model Details | Field | Details | |---|---| | **Developed by** | [somrajmondal](https://huggingface.co/somrajmondal) | | **Base model** | unsloth/phi-3-mini-4k-instruct-bnb-4bit | | **Model type** | Causal Language Model (Phi-3 architecture) | | **Parameters** | 3.8B total / 29.8M trainable (0.78%) | | **Language** | English | | **License** | Apache 2.0 | | **Fine-tuning method** | LoRA (Low-Rank Adaptation) | | **Quantization** | 4-bit NF4 during training, merged to fp16 | --- ## 🎯 What This Model Does This model is fine-tuned to answer **finance and investment questions** clearly and accurately. It was trained on the `gbharti/finance-alpaca` dataset covering topics like: - Stock market concepts (P/E ratio, dividends, market cap) - Investment strategies (ETFs, mutual funds, dollar-cost averaging) - Fixed income (bonds, yields, interest rates) - Personal finance (compound interest, savings, budgeting) - Financial planning and portfolio diversification --- ## 🏋️ Training Details | Setting | Value | |---|---| | **Dataset** | [gbharti/finance-alpaca](https://huggingface.co/datasets/gbharti/finance-alpaca) | | **Training rows** | 5,000 | | **Epochs** | 2 | | **Total steps** | 1,250 | | **Batch size** | 2 (effective: 8 with grad accumulation) | | **Learning rate** | 2e-4 (cosine scheduler) | | **Optimizer** | AdamW 8-bit | | **LoRA rank (r)** | 16 | | **LoRA alpha** | 16 | | **LoRA dropout** | 0.05 | | **LoRA target modules** | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj | | **Max sequence length** | 1024 | | **Final training loss** | 2.18 | | **Peak VRAM used** | 3.69 GB | | **Training hardware** | Google Colab Free T4 (15GB VRAM) | | **Training time** | ~2 hours 20 minutes | | **Framework** | Unsloth + HuggingFace TRL | --- ## 🚀 How to Use ### Quick Start (Transformers) ```python from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_id = "somrajmondal/phi3-mini-finance-lora-fp16" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.float16, device_map="auto", ) question = "What is compound interest and why is it important?" prompt = f"""<|user|> {question} <|end|> <|assistant|> """ inputs = tokenizer(prompt, return_tensors="pt").to(model.device) outputs = model.generate( **inputs, max_new_tokens=300, do_sample=False, repetition_penalty=1.3, eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.eos_token_id, ) response = tokenizer.decode( outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True ) print(response) ``` ### With Unsloth (Faster Inference) ```python from unsloth import FastLanguageModel import torch model, tokenizer = FastLanguageModel.from_pretrained( model_name = "somrajmondal/phi3-mini-finance-lora-fp16", max_seq_length = 1024, dtype = None, load_in_4bit = True, # set False for full fp16 ) FastLanguageModel.for_inference(model) question = "What is the difference between a stock and a bond?" prompt = f"""<|user|> {question} <|end|> <|assistant|> """ inputs = tokenizer(prompt, return_tensors="pt").to("cuda") outputs = model.generate( **inputs, max_new_tokens=300, do_sample=False, repetition_penalty=1.3, eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.eos_token_id, ) response = tokenizer.decode( outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True ) print(response) ``` --- ## 💬 Prompt Format This model uses the **Phi-3 chat template**. Always wrap your input like this: ``` <|user|> Your finance question here <|end|> <|assistant|> ``` --- ## 📊 Example Outputs **Q: What is a P/E ratio?** > The price-to-earnings (P/E) ratio measures a company's current share price relative to its earnings per share. A high P/E suggests investors expect future growth, while a low P/E may indicate an undervalued stock or slower expected growth. **Q: What is dollar cost averaging?** > Dollar cost averaging is an investment strategy where you invest a fixed amount of money at regular intervals, regardless of market conditions. This reduces the impact of volatility and removes the need to time the market. --- ## ⚠️ Limitations - Trained on only 5,000 rows — may lack depth on niche financial topics - Not suitable for real financial advice — always consult a professional - May occasionally produce incomplete answers on complex multi-part questions - Training loss of 2.18 indicates room for improvement with more epochs/data --- ## 🔧 Recommended Inference Settings ```python # For factual / accurate answers (recommended) do_sample = False repetition_penalty = 1.3 max_new_tokens = 300 # For more creative / detailed answers do_sample = True temperature = 0.3 top_p = 0.9 ``` --- ## 📦 Training Framework - [Unsloth](https://github.com/unslothai/unsloth) — 2x faster training, 60% less VRAM - [HuggingFace TRL](https://github.com/huggingface/trl) — SFTTrainer - [PEFT](https://github.com/huggingface/peft) — LoRA adapters - [BitsAndBytes](https://github.com/TimDettmers/bitsandbytes) — 4-bit NF4 quantization --- ## 📄 Citation If you use this model, please cite the base model and dataset: ```bibtex @misc{phi3-mini-finance-lora, author = {somrajmondal}, title = {Phi-3-mini Finance LoRA}, year = {2025}, publisher = {HuggingFace}, url = {https://huggingface.co/somrajmondal/phi3-mini-finance-lora-fp16} } ```