87 lines
3.2 KiB
Markdown
87 lines
3.2 KiB
Markdown
---
|
||
base_model: unsloth/Meta-Llama-3.1-8B-bnb-4bit
|
||
library_name: peft
|
||
pipeline_tag: text-generation
|
||
tags:
|
||
- base_model:adapter:unsloth/Meta-Llama-3.1-8B-bnb-4bit
|
||
- lora
|
||
- sft
|
||
- transformers
|
||
- trl
|
||
- unsloth
|
||
---
|
||
|
||
# Model Card
|
||
|
||
|
||
## Model Description
|
||
|
||
A French grammar correction model designed primarily for learners of French as a second language (FSL/FLE). It corrects grammar, spelling, syntax, punctuation, and stylistic issues while preserving the original meaning and tone as much as possible.
|
||
|
||
The model is especially effective for:
|
||
|
||
- French learners and students
|
||
- Academic and professional writing
|
||
- Language practice and self-correction
|
||
- Improving fluency and sentence naturalness
|
||
|
||
It can also perform high-quality translation from English to French, making it useful both as a grammar corrector and as a lightweight bilingual writing assistant.
|
||
|
||
Optimized for local inference with LM Studio and compatible with GGUF quantizations for efficient CPU or GPU deployment.
|
||
|
||
|
||
- **Developed by:** Adel Jebali, Concordia University
|
||
- **Funded by:** SSHRC
|
||
- **Language(s) (NLP):** French
|
||
- **License:** Apache 2.0
|
||
- **Finetuned:** from Llama 3.1-8B
|
||
|
||
## Uses
|
||
|
||
Correct you French written texts! It is not a chat model.
|
||
|
||
|
||
## Bias, Risks, and Limitations
|
||
|
||
This AI LLM is not 100% bullet-proof. Errors are still possible.
|
||
|
||
|
||
## How to use
|
||
|
||
Works best with [LM Studio](https://lmstudio.ai) and is available in three quantization formats:
|
||
|
||
* **4-bit — Q4_K_M — 4.92 GB**
|
||
Best choice for low-memory systems and entry-level hardware. Recommended for Macs with only 8 GB of unified memory or older GPUs. Offers the fastest loading times and lowest VRAM usage, with a small trade-off in output quality and coherence.
|
||
|
||
* **6-bit — Q6_K — 6.6 GB**
|
||
Excellent balance between quality, speed, and memory consumption. A strong default option for most users with 12–16 GB of RAM or mid-range GPUs. In many cases, it delivers quality close to 8-bit while remaining significantly lighter.
|
||
|
||
* **8-bit — Q8_0 — 8.54 GB**
|
||
Highest quality and most faithful outputs among the available quantizations. Recommended if your hardware can handle it, especially with a dedicated GPU or Apple Silicon Mac with sufficient unified memory. Produces more stable generations, better grammatical consistency, and fewer hallucinations.
|
||
|
||
## Recommendations
|
||
|
||
* **8 GB Macs:** use **Q4_K_M**
|
||
* **16 GB systems:** use **Q6_K** for the best balance
|
||
* **24 GB+ RAM or modern GPU:** use **Q8_0** for maximum quality
|
||
|
||
For optimal performance in [LM Studio](https://lmstudio.ai):
|
||
|
||
* Enable **GPU offloading** when available
|
||
* Increase the **context length** only if needed, since larger contexts consume more memory
|
||
* On Apple Silicon Macs, Metal acceleration significantly improves inference speed
|
||
* Temperature 0 (or 0.1)
|
||
* Min p 0
|
||
* Top k 0
|
||
* Top p 1
|
||
|
||
If your priority is:
|
||
|
||
* **Maximum speed / lowest memory usage → Q4_K_M**
|
||
* **Best balance → Q6_K**
|
||
* **Best overall quality → Q8_0**
|
||
|
||
|
||
# Paper
|
||
A. Jebali, "Developing a Grammatical Error Correction System for French Second Language Written Texts," 2025 5th International Conference on Electrical, Computer and Energy Technologies (ICECET), Paris, France, 2025, pp. 1-6, doi: 10.1109/ICECET63943.2025.11472110.
|