Files
llama3.1-gec-strict/README.md
ModelHub XC 4cab16e4c7 初始化项目,由ModelHub XC社区提供模型
Model: adeljebali/llama3.1-gec-strict
Source: Original Platform
2026-09-10 04:32:15 +08:00

3.2 KiB
Raw Permalink Blame History

base_model, library_name, pipeline_tag, tags
base_model library_name pipeline_tag tags
unsloth/Meta-Llama-3.1-8B-bnb-4bit peft text-generation
base_model:adapter:unsloth/Meta-Llama-3.1-8B-bnb-4bit
lora
sft
transformers
trl
unsloth

Model Card

Model Description

A French grammar correction model designed primarily for learners of French as a second language (FSL/FLE). It corrects grammar, spelling, syntax, punctuation, and stylistic issues while preserving the original meaning and tone as much as possible.

The model is especially effective for:

  • French learners and students
  • Academic and professional writing
  • Language practice and self-correction
  • Improving fluency and sentence naturalness

It can also perform high-quality translation from English to French, making it useful both as a grammar corrector and as a lightweight bilingual writing assistant.

Optimized for local inference with LM Studio and compatible with GGUF quantizations for efficient CPU or GPU deployment.

  • Developed by: Adel Jebali, Concordia University
  • Funded by: SSHRC
  • Language(s) (NLP): French
  • License: Apache 2.0
  • Finetuned: from Llama 3.1-8B

Uses

Correct you French written texts! It is not a chat model.

Bias, Risks, and Limitations

This AI LLM is not 100% bullet-proof. Errors are still possible.

How to use

Works best with LM Studio and is available in three quantization formats:

  • 4-bit — Q4_K_M — 4.92 GB Best choice for low-memory systems and entry-level hardware. Recommended for Macs with only 8 GB of unified memory or older GPUs. Offers the fastest loading times and lowest VRAM usage, with a small trade-off in output quality and coherence.

  • 6-bit — Q6_K — 6.6 GB Excellent balance between quality, speed, and memory consumption. A strong default option for most users with 12–16 GB of RAM or mid-range GPUs. In many cases, it delivers quality close to 8-bit while remaining significantly lighter.

  • 8-bit — Q8_0 — 8.54 GB Highest quality and most faithful outputs among the available quantizations. Recommended if your hardware can handle it, especially with a dedicated GPU or Apple Silicon Mac with sufficient unified memory. Produces more stable generations, better grammatical consistency, and fewer hallucinations.

Recommendations

  • 8 GB Macs: use Q4_K_M
  • 16 GB systems: use Q6_K for the best balance
  • 24 GB+ RAM or modern GPU: use Q8_0 for maximum quality

For optimal performance in LM Studio:

  • Enable GPU offloading when available
  • Increase the context length only if needed, since larger contexts consume more memory
  • On Apple Silicon Macs, Metal acceleration significantly improves inference speed
  • Temperature 0 (or 0.1)
  • Min p 0
  • Top k 0
  • Top p 1

If your priority is:

  • Maximum speed / lowest memory usage → Q4_K_M
  • Best balance → Q6_K
  • Best overall quality → Q8_0

Paper

A. Jebali, "Developing a Grammatical Error Correction System for French Second Language Written Texts," 2025 5th International Conference on Electrical, Computer and Energy Technologies (ICECET), Paris, France, 2025, pp. 1-6, doi: 10.1109/ICECET63943.2025.11472110.