168 lines
6.0 KiB
Markdown
168 lines
6.0 KiB
Markdown
|
|
---
|
|||
|
|
license: llama3.1
|
|||
|
|
base_model:
|
|||
|
|
- meta-llama/Llama-3.1-8B-Instruct
|
|||
|
|
tags:
|
|||
|
|
- satcom
|
|||
|
|
- satellite
|
|||
|
|
- satellite communications
|
|||
|
|
- space
|
|||
|
|
- esa
|
|||
|
|
- artes
|
|||
|
|
- llm
|
|||
|
|
- fine-tuned
|
|||
|
|
- question-answering
|
|||
|
|
language:
|
|||
|
|
- en
|
|||
|
|
model_name: esa-sceva/llama3-satcom-8b
|
|||
|
|
pipeline_tag: text-generation
|
|||
|
|
library_name: transformers
|
|||
|
|
datasets:
|
|||
|
|
- esa-sceva/satcom-qa
|
|||
|
|
- esa-sceva/satcom-mcqa
|
|||
|
|
- esa-sceva/satcom-synth-qa
|
|||
|
|
- esa-sceva/satcom-synth-qa-cot
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# LLama3 SatCom 8B
|
|||
|
|
|
|||
|
|
**LLama3 SatCom 8B** is a fine-tuned open Large Language Model (LLM) developed under the **ESA ARTES** programme as part of the **SatcomLLM / SCEVA (SatCom Expert Virtual Assistant)** project.
|
|||
|
|
It is designed to support **satellite communications (SatCom)** experts, engineers, and mission planners through domain-specialised reasoning, question answering, and document-based assistance.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Model Description
|
|||
|
|
|
|||
|
|
- **Base model:** `meta-llama/Llama-3.1-8B-Instruct`
|
|||
|
|
- **Fine-tuning type:** Instruction fine-tuning (IFT)
|
|||
|
|
- **Training data:** Domain-specific question–answer datasets (manual, synthetic, and multiple-choice)
|
|||
|
|
- **Architecture:** Decoder-only transformer, 8 billion parameters
|
|||
|
|
- **Languages:** English
|
|||
|
|
- **License:** LLama-3.1 Communiy License Agreement
|
|||
|
|
|
|||
|
|
The model has been fine-tuned on curated SatCom-related corpora to enhance its understanding of technical language, protocols, and reasoning processes common to satellite communications, including 5G/6G non-terrestrial networks, link budget evaluation, and mission engineering tasks.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Training Datasets
|
|||
|
|
|
|||
|
|
| Dataset | Description |
|
|||
|
|
| --------------------------------- | -------------------------------------------------------------------------------------- |
|
|||
|
|
| **esa-sceva/satcom-synth-qa** | Synthetic QA data generated via agentic pipelines using large teacher models |
|
|||
|
|
| **esa-sceva/satcom-synth-qa-cot** | Chain-of-thought annotated QA used to improve reasoning depth and factual traceability |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
|
|||
|
|
## Intended Use
|
|||
|
|
|
|||
|
|
**Primary use cases:**
|
|||
|
|
- Technical Q&A and reasoning on SatCom systems
|
|||
|
|
- Support for link budget and RF engineering questions
|
|||
|
|
- Guidance for 5G/6G NTN (Non-Terrestrial Network) operations
|
|||
|
|
- Mission design, planning, and anomaly detection support
|
|||
|
|
- Educational and research use within the SatCom sector
|
|||
|
|
|
|||
|
|
**Intended users:**
|
|||
|
|
- ESA engineers and project officers
|
|||
|
|
- SatCom and aerospace researchers
|
|||
|
|
- SMEs and technical operators in satellite communication
|
|||
|
|
- Academic and educational users
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Limitations
|
|||
|
|
|
|||
|
|
- The model does **not access real-time mission data** or proprietary ESA documents.
|
|||
|
|
- Answers are based on training data and may require expert validation for operational use.
|
|||
|
|
- It should not be relied upon for flight-critical or safety-critical decisions.
|
|||
|
|
- Limited context window (base 8B configuration) may constrain long-document reasoning.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Technical Details
|
|||
|
|
|
|||
|
|
| Parameter | Value |
|
|||
|
|
|------------|--------|
|
|||
|
|
| Base Model | Llama 3.1 8B Instruct |
|
|||
|
|
| Parameters | 8 billion |
|
|||
|
|
| Context length | 8k tokens |
|
|||
|
|
| Precision | bfloat16 / fp16 |
|
|||
|
|
| Framework | Lit-GPT (Lightning AI) |
|
|||
|
|
| Training infra | EuroHPC MareNostrum5 + AWS EC2 |
|
|||
|
|
| Optimisation | LoRA fine-tuning, cosine LR schedule |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## Evaluation
|
|||
|
|
|
|||
|
|
### Evaluation Datasets
|
|||
|
|
|
|||
|
|
The model was benchmarked on both general-purpose and domain-specific QA tasks. Regarding Satcom-specific datasets:
|
|||
|
|
|
|||
|
|
| Dataset | Subset | Description |
|
|||
|
|
| ------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------- |
|
|||
|
|
| **esa-sceva/satcom-qa** | **Open SatCom QA** | Conceptual and reasoning-based questions on SatCom workflows, regulations, and mission/system design |
|
|||
|
|
| | **Math SatCom QA** | Quantitative and formula-based questions derived from system design and orbital mechanics topics |
|
|||
|
|
| **esa-sceva/satcom-mcqa** | **Open MCQA** | Conceptual multiple-choice questions on RF systems, communication protocols, and architecture |
|
|||
|
|
| | **Math MCQA** | Numerical and link-budget-focused multiple-choice questions testing applied calculations |
|
|||
|
|
|
|||
|
|
### Results
|
|||
|
|
|
|||
|
|
<table>
|
|||
|
|
<thead>
|
|||
|
|
<tr>
|
|||
|
|
<th rowspan="2">Model</th>
|
|||
|
|
<th colspan="3">MCQA (Accuracy)</th>
|
|||
|
|
<th colspan="3">Satcom-QA</th>
|
|||
|
|
<th colspan="3">EVE-QA</th>
|
|||
|
|
</tr>
|
|||
|
|
<tr>
|
|||
|
|
<th>Satcom</th>
|
|||
|
|
<th>EVE</th>
|
|||
|
|
<th>TeleQnA</th>
|
|||
|
|
<th>Norm.</th>
|
|||
|
|
<th>Bin.</th>
|
|||
|
|
<th>WR</th>
|
|||
|
|
<th>Norm.</th>
|
|||
|
|
<th>Bin.</th>
|
|||
|
|
<th>WR</th>
|
|||
|
|
</tr>
|
|||
|
|
</thead>
|
|||
|
|
<tbody>
|
|||
|
|
<tr>
|
|||
|
|
<td>Llama-3.1-8B-Instruct</td>
|
|||
|
|
<td>78.59</td>
|
|||
|
|
<td><strong>81.35</strong></td>
|
|||
|
|
<td>68.40</td>
|
|||
|
|
<td>75.44</td>
|
|||
|
|
<td>70.40</td>
|
|||
|
|
<td>—</td>
|
|||
|
|
<td>61.47</td>
|
|||
|
|
<td>65.20</td>
|
|||
|
|
<td>—</td>
|
|||
|
|
</tr>
|
|||
|
|
<tr style="border-top: 1px dashed #999;">
|
|||
|
|
<td><code>llama3-satcom-8b</code></td>
|
|||
|
|
<td><strong>80.15</strong></td>
|
|||
|
|
<td>80.95</td>
|
|||
|
|
<td><strong>68.80</strong></td>
|
|||
|
|
<td><strong>77.75</strong></td>
|
|||
|
|
<td><strong>73.62</strong></td>
|
|||
|
|
<td><strong>51.49</strong></td>
|
|||
|
|
<td><strong>62.90</strong></td>
|
|||
|
|
<td><strong>68.12</strong></td>
|
|||
|
|
<td><strong>51.41</strong></td>
|
|||
|
|
</tr>
|
|||
|
|
</tbody>
|
|||
|
|
</table>
|
|||
|
|
|
|||
|
|
**Table:** Evaluation results (%).
|
|||
|
|
- **Norm.** denotes the normalized score obtained by averaging 1–5 ratings from a panel of LLM judges (Qwen3, gpt-4.1-mini, Mistral-Large-2512, and DeepSeek-V3.2) and scaling to [0,1].
|
|||
|
|
- **Bin.** is the binary accuracy computed from correctness judgments.
|
|||
|
|
- **WR** (Adjusted Win Rate) is defined as (wins + 0.5 × ties) / total, based on pairwise comparisons with randomized answer order.
|
|||
|
|
Multiple-choice performance is measured using standard accuracy.
|
|||
|
|
All results are averaged over 3–5 runs; standard deviation ≤ 0.25 pp for open-ended QA and ≤ 0.10 pp for MCQA.
|
|||
|
|
|
|||
|
|
---
|