license, datasets, language, base_model, pipeline_tag, tags
license
datasets
language
base_model
pipeline_tag
tags
apache-2.0
bitext/Bitext-customer-support-llm-chatbot-training-dataset
unsloth/llama-3-8b-bnb-4bit
text-generation
text-generation-inference
transformers
unsloth
llama
gguf
Customer-Support-Bot
Customer Support Chatbot with LLaMA 3.1
An end-to-end customer support chatbot solution powered by fine-tuned LLaMA 3.1 8B model, deployed using Flask, Docker, and AWS ECS.
Overview
This project implements a sophisticated customer support chatbot leveraging the LLaMA 3.1 8B model fine-tuned on customer support conversations. The solution uses LoRA fine-tuning and various quantization techniques for optimized inference, deployed as a containerized application on AWS ECS with Fargate.
Features
Fine-tuned LLaMA 3.1 Model : Customized for customer support using the Bitext customer support dataset
Optimized Inference : Implements 4-bit, 8-bit, and 16-bit quantization
Containerized Deployment : Docker-based deployment for consistency and scalability
Cloud Infrastructure : Hosted on AWS ECS with Fargate for serverless container management
CI/CD Pipeline : Automated deployment using AWS CodePipeline
Monitoring : Comprehensive logging and monitoring via AWS CloudWatch
Model Details
The fine-tuned model is hosted on Hugging Face:
Tech Stack
Backend : Flask API
Model Serving : Ollama
Containerization : Docker
Cloud Services :
AWS ECS (Fargate)
AWS CodePipeline
AWS CloudWatch
Model Training : LoRA, Quantization
Screenshots
Chatbot Interface
AWS CloudWatch Monitoring
Docker Logs
AWS Deployment
Push Docker image to Amazon ECR
Configure AWS ECS Task Definition
Set up AWS CodePipeline for CI/CD
Configure CloudWatch monitoring
Uploaded model
Developed by: praneethposina
License: apache-2.0
Finetuned from model : unsloth/llama-3-8b-bnb-4bit