--- license: apache-2.0 base_model: Qwen/Qwen2.5-7B-Instruct library_name: llama.cpp pipeline_tag: text-generation language: - de - en tags: - gguf - qwen2.5 - history - archaeology - ancient-texts - religious-studies - biblical-studies - research-assistant - local-ai - german - english - lm-studio --- # Chaperon-Sophia-v2 An experimental research assistant for archaeology, history, biblical studies, and ancient texts. Developed by **Chaperon GmbH** and **Lucas Bischof**, Sophia-v2 demonstrates a complete end-to-end AI workflow including multimodal data processing, document extraction, dataset generation, LoRA fine-tuning, GGUF deployment, and local AI integration. ## Overview Sophia-v2 is a custom fine-tuned variant of **Qwen2.5-7B-Instruct** designed to explore historical, textual, archaeological, and religious source material through natural conversation. The project serves both as a technical showcase of Chaperon GmbH's AI capabilities and as a practical proof-of-concept for domain-specific model adaptation. ## Project Goals This project demonstrates the ability to: - Process large collections of PDFs - Extract text from historical documents - Convert audio and video into training datasets - Process image-based source material - Generate structured instruction datasets - Fine-tune Large Language Models using LoRA - Merge and deploy custom models - Publish GGUF models for local inference ## Training Sources The training process included the preparation and processing of curated material related to: - CDLI Cuneiform Collections - ETCSL Translations - Ancient Mesopotamian Literature - Biblical Texts - Apocryphal Literature - Gnostic Literature - Gospel of Mary - Pistis Sophia - Comparative Religion - Historical Source Material - Archaeological Research The focus of Sophia-v2 is source-oriented exploration and discussion rather than authoritative interpretation. ## Intended Audience Sophia-v2 may be useful for: - Archaeology enthusiasts - History researchers - Biblical scholars - Comparative religion communities - Independent researchers - Educational projects - Digital humanities initiatives ## Available Models ### Sophia_v2_q5_K_M.gguf Recommended version for most users. Provides an excellent balance between quality, performance, and memory usage. ### Sophia_v2_f16.gguf Reference-quality version. Provides maximum fidelity while requiring significantly more storage and memory. ## Technical Specifications Base Model: - Qwen/Qwen2.5-7B-Instruct Training Method: - Supervised Fine Tuning (SFT) - LoRA Adaptation - Model Merge - GGUF Conversion Deployment Targets: - LM Studio - llama.cpp - Open WebUI - KoboldCpp ## Limitations Sophia-v2 is a language model and may: - Generate incorrect information - Misinterpret source material - Produce inaccurate citations - Present plausible but incorrect conclusions Users should verify important claims using primary and scholarly sources. ## About Chaperon GmbH Sophia-v2 serves as a public demonstration of the capabilities of Chaperon GmbH in: - Multimodal data processing - Knowledge extraction - AI training pipelines - Custom model development - Local AI deployment - Enterprise AI integration ## Credits Created by: **Lucas Bischof** **Chaperon GmbH** Base Model: **Qwen/Qwen2.5-7B-Instruct** GGUF Conversion: **llama.cpp** >>>>>>> 038febb (Improve model card)