初始化项目,由ModelHub XC社区提供模型
Model: yasserrmd/GeoScholar-QA-1.2B Source: Original Platform
This commit is contained in:
37
.gitattributes
vendored
Normal file
37
.gitattributes
vendored
Normal file
@@ -0,0 +1,37 @@
|
||||
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||
*.model filter=lfs diff=lfs merge=lfs -text
|
||||
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||
banner[[:space:]](2).png filter=lfs diff=lfs merge=lfs -text
|
||||
banner.png filter=lfs diff=lfs merge=lfs -text
|
||||
126
README.md
Normal file
126
README.md
Normal file
@@ -0,0 +1,126 @@
|
||||
---
|
||||
license: cc-by-4.0
|
||||
language:
|
||||
- en
|
||||
tags:
|
||||
- geoscience
|
||||
- question-answering
|
||||
- supervised-fine-tuning
|
||||
- academic
|
||||
- earth-sciences
|
||||
- remote-sensing
|
||||
- gis
|
||||
- text-generation-inference
|
||||
- transformers
|
||||
- unsloth
|
||||
- lfm2
|
||||
datasets:
|
||||
- GeoGPT-Research-Project/GeoGPT-QA
|
||||
base_model: LiquidAI/LFM2-1.2B
|
||||
---
|
||||
|
||||
# GeoScholar-QA
|
||||
|
||||
<img src="banner.png" />
|
||||
|
||||
[](https://creativecommons.org/licenses/by/4.0/)
|
||||
[](https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA)
|
||||
|
||||
## 📖 Overview
|
||||
|
||||
GeoScholar-QA is a large language model fine-tuned for academic question answering in the field of geoscience. It is built upon the Liquid AI LFM2 base model and has been trained using the Unsloth framework with the GeoGPT-QA dataset. [4] This model's primary strength lies in its ability to explain concepts and theories within the earth sciences. It is not designed to provide consistently accurate statistics or citations.
|
||||
|
||||
---
|
||||
|
||||
## 📚 Dataset: GeoGPT-QA
|
||||
|
||||
* **Dataset name:** GeoGPT-QA
|
||||
* **Publisher / Project:** [GeoGPT-Research-Project](https://huggingface.co/GeoGPT-Research-Project)
|
||||
* **Size:** Approximately 41,400 rows in the `train` split. [4]
|
||||
* **Format:** The dataset is tabular (originally CSV, automatically converted to Parquet) and includes fields such as `question`, `answer`, `title`, `authors`, `doi`, `journal`, `volume`, `pages`, and `license`. [4]
|
||||
* **Language:** English [4]
|
||||
* **License:** CC-BY 4.0 — You are permitted to share and adapt the dataset, but you must provide attribution and indicate if any changes were made. [4]
|
||||
|
||||
---
|
||||
|
||||
## 🔧 Model Details & Training
|
||||
|
||||
* **Base model:** [Liquid AI LFM2](https://huggingface.co/LiquidAI/LFM2-350M) is a hybrid model designed for on-device deployment, offering a balance of quality, speed, and memory efficiency. [8]
|
||||
* **Fine-tuning framework:** [Unsloth](https://github.com/unslothai/unsloth) is a framework designed to speed up and optimize the fine-tuning of large language models, making it more accessible on limited hardware. [10, 11]
|
||||
* **Training data:** [GeoGPT-QA](https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA) [4]
|
||||
* **Objective:** The model was trained through Supervised Fine-Tuning (SFT) using question-answer pairs. The focus of this training was on geoscience theory, conceptual knowledge, and explanations. [19]
|
||||
* **Effective batch size:** A low per-device batch size was utilized along with gradient accumulation to prevent out-of-memory (OOM) errors during training.
|
||||
* **Training progress:** (Example) Approximately 3000 steps, covering about 58% of the dataset. *Please adjust this based on your final training run.*
|
||||
|
||||
---
|
||||
|
||||
## 🧑🏫 Intended Use
|
||||
|
||||
GeoScholar-QA is intended for the following applications:
|
||||
|
||||
* Providing academic explanations in various fields of geoscience, such as plate tectonics, hydrology, and geomorphology.
|
||||
* Serving as a teaching and learning aid for students and educators.
|
||||
* Enhancing conceptual and theoretical understanding of earth science principles.
|
||||
|
||||
**This model is not intended for:**
|
||||
|
||||
* High-risk or decision-making tasks that require precise numerical data or statistics.
|
||||
* Reliance on the generated citations or study results without independent verification.
|
||||
|
||||
---
|
||||
|
||||
## ⚠️ Limitations
|
||||
|
||||
* The model may generate inaccurate numbers, study names, datasets, or locations.
|
||||
* Answers in applied or technical contexts may be overgeneralized or vague.
|
||||
* It should not be used as a substitute for verification by a domain expert, especially in research or policy-making settings.
|
||||
|
||||
---
|
||||
|
||||
## ✅ Example Usage
|
||||
|
||||
```python
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
|
||||
model_name = "yasserrmd/GeoScholar-QA-1.2B"
|
||||
|
||||
tokenizer = AutoTokenizer.from_pretrained(model_name)
|
||||
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
|
||||
|
||||
messages = [
|
||||
{"role": "user", "content": "How do plate tectonics explain the formation of volcanoes along subduction zones?"}
|
||||
]
|
||||
|
||||
inputs = tokenizer.apply_chat_template(
|
||||
messages,
|
||||
add_generation_prompt=True,
|
||||
return_tensors="pt"
|
||||
).to("cuda")
|
||||
|
||||
outputs = model.generate(
|
||||
**inputs,
|
||||
max_new_tokens=256,
|
||||
temperature=0.3,
|
||||
repetition_penalty=1.05
|
||||
)
|
||||
|
||||
print(tokenizer.decode(outputs, skip_special_tokens=True))
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📝 License & Attribution
|
||||
|
||||
* This model is trained using the **GeoGPT-QA** dataset, which is licensed under [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/). [4]
|
||||
* You must give appropriate credit to the *GeoGPT Research Project* and include a link to the [dataset](https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA).
|
||||
* If you adapt or build upon this model, you must indicate that changes have been made.
|
||||
|
||||
---
|
||||
|
||||
## 🔭 Tags / Metadata for Hugging Face Model Card
|
||||
|
||||
* **Model type:** Text generation / QA
|
||||
* **Domain:** Geoscience, Earth Sciences
|
||||
* **Base model:** Liquid AI LFM2
|
||||
* **Training method:** SFT (Supervised Fine-Tuning)
|
||||
* **License:** CC-BY 4.0
|
||||
3
banner.png
Normal file
3
banner.png
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:723c834439e6809c7533f685c43b2051b372ccdb7ee65c5193731678291fcfb5
|
||||
size 1945670
|
||||
4
chat_template.jinja
Normal file
4
chat_template.jinja
Normal file
@@ -0,0 +1,4 @@
|
||||
{{bos_token}}{% for message in messages %}{{'<|im_start|>' + message['role'] + '
|
||||
' + message['content'] + '<|im_end|>' + '
|
||||
'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
|
||||
' }}{% endif %}
|
||||
59
config.json
Normal file
59
config.json
Normal file
@@ -0,0 +1,59 @@
|
||||
{
|
||||
"architectures": [
|
||||
"Lfm2ForCausalLM"
|
||||
],
|
||||
"block_auto_adjust_ff_dim": true,
|
||||
"block_dim": 2048,
|
||||
"block_ff_dim": 12288,
|
||||
"block_ffn_dim_multiplier": 1.0,
|
||||
"block_mlp_init_scale": 1.0,
|
||||
"block_multiple_of": 256,
|
||||
"block_norm_eps": 1e-05,
|
||||
"block_out_init_scale": 1.0,
|
||||
"block_use_swiglu": true,
|
||||
"block_use_xavier_init": true,
|
||||
"bos_token_id": 1,
|
||||
"conv_L_cache": 3,
|
||||
"conv_bias": false,
|
||||
"conv_dim": 2048,
|
||||
"conv_dim_out": 2048,
|
||||
"conv_use_xavier_init": true,
|
||||
"dtype": "float16",
|
||||
"eos_token_id": 7,
|
||||
"hidden_size": 2048,
|
||||
"initializer_range": 0.02,
|
||||
"intermediate_size": 12288,
|
||||
"layer_types": [
|
||||
"conv",
|
||||
"conv",
|
||||
"full_attention",
|
||||
"conv",
|
||||
"conv",
|
||||
"full_attention",
|
||||
"conv",
|
||||
"conv",
|
||||
"full_attention",
|
||||
"conv",
|
||||
"full_attention",
|
||||
"conv",
|
||||
"full_attention",
|
||||
"conv",
|
||||
"full_attention",
|
||||
"conv"
|
||||
],
|
||||
"max_position_embeddings": 32768,
|
||||
"model_type": "lfm2",
|
||||
"norm_eps": 1e-05,
|
||||
"num_attention_heads": 32,
|
||||
"num_heads": 32,
|
||||
"num_hidden_layers": 16,
|
||||
"num_key_value_heads": 8,
|
||||
"pad_token_id": 0,
|
||||
"rope_theta": 1000000.0,
|
||||
"transformers_version": "4.56.1",
|
||||
"unsloth_fixed": true,
|
||||
"unsloth_version": "2025.9.4",
|
||||
"use_cache": true,
|
||||
"use_pos_enc": true,
|
||||
"vocab_size": 65536
|
||||
}
|
||||
10
generation_config.json
Normal file
10
generation_config.json
Normal file
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"_from_model_config": true,
|
||||
"bos_token_id": 1,
|
||||
"eos_token_id": [
|
||||
7
|
||||
],
|
||||
"max_length": 32768,
|
||||
"pad_token_id": 0,
|
||||
"transformers_version": "4.56.1"
|
||||
}
|
||||
3
model.safetensors
Normal file
3
model.safetensors
Normal file
@@ -0,0 +1,3 @@
|
||||
version https://git-lfs.github.com/spec/v1
|
||||
oid sha256:4ce28a8f2fd957ef7bd7b053080f92464104927f1ef737019a1302b10ac328f6
|
||||
size 2340697784
|
||||
23
special_tokens_map.json
Normal file
23
special_tokens_map.json
Normal file
@@ -0,0 +1,23 @@
|
||||
{
|
||||
"bos_token": {
|
||||
"content": "<|startoftext|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"eos_token": {
|
||||
"content": "<|im_end|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
},
|
||||
"pad_token": {
|
||||
"content": "<|pad|>",
|
||||
"lstrip": false,
|
||||
"normalized": false,
|
||||
"rstrip": false,
|
||||
"single_word": false
|
||||
}
|
||||
}
|
||||
323812
tokenizer.json
Normal file
323812
tokenizer.json
Normal file
File diff suppressed because it is too large
Load Diff
4076
tokenizer_config.json
Normal file
4076
tokenizer_config.json
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user