88 lines
2.8 KiB
Markdown
88 lines
2.8 KiB
Markdown
---
|
|
license: apache-2.0
|
|
language:
|
|
- en
|
|
library_name: transformers
|
|
pipeline_tag: text-generation
|
|
datasets:
|
|
- HuggingFaceFW/fineweb-edu
|
|
- mlfoundations/dclm-baseline-1.0-parquet
|
|
- HuggingFaceTB/smol-smoltalk
|
|
tags:
|
|
- boris
|
|
- nmai
|
|
- gpt2
|
|
- 75M
|
|
- instruct
|
|
base_model:
|
|
- KSP-NMAI/Boris-1.3-75M
|
|
---
|
|
|
|

|
|
|
|
# Boris-1.3-75M-Instruct
|
|
|
|
Boris-1.3-75M-Instruct is a 75 million-parameter language model created by New
|
|
Millennium Artificial Intelligence (NMAI).
|
|
|
|
This is an **instruction-tuned model**. It follows instructions and holds a
|
|
conversation. For the base (pretrained-only) version, see
|
|
[KSP-NMAI/Boris-1.3-75M](https://huggingface.co/KSP-NMAI/Boris-1.3-75M).
|
|
|
|
## Usage
|
|
|
|
```python
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
|
|
tok = AutoTokenizer.from_pretrained("KSP-NMAI/Boris-1.3-75M-Instruct")
|
|
model = AutoModelForCausalLM.from_pretrained("KSP-NMAI/Boris-1.3-75M-Instruct")
|
|
|
|
messages = [{"role": "user", "content": "What's a good way to start learning C?"}]
|
|
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
|
|
out = model.generate(ids, max_new_tokens=120, do_sample=True, top_p=0.95)
|
|
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
|
|
```
|
|
|
|
## Details
|
|
|
|
| | |
|
|
|---|---|
|
|
| Architecture | GPT-2 (pre-LN, learned positional embeddings, tied embeddings) |
|
|
| Layers / heads / d_model | 12 / 9 / 576 |
|
|
| Context length | 1024 |
|
|
| Vocab | 50304 (GPT-NeoX-20B BPE, padded) |
|
|
| Tokenizer | `EleutherAI/gpt-neox-20b` |
|
|
| Precision | trained in bf16 autocast with fp32 master weights |
|
|
|
|
## Fine-tuning
|
|
|
|
Fine-tuned on [smol-smoltalk](https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk) and [OpenAssistant](https://huggingface.co/datasets/OpenAssistant/oasst1) for *279,815* examples over *8:19:30* on one RTX 3060.
|
|
|
|
| | |
|
|
|---|---|
|
|
| Final loss | *1.4162* |
|
|
| Final grad norm | *1.185* |
|
|
| Final learning rate | *1.00e-05* |
|
|
|
|

|
|
|
|
## Limitations
|
|
|
|
A model of this size will produce text that is frequently inaccurate,
|
|
inconsistent, or offensive. It has received no alignment or safety tuning and
|
|
should not be used for factual reference or deployed without supervision.
|
|
|
|
## Copyright & License
|
|
|
|
*Copyright 2026 Joseph Jones*
|
|
|
|
This project and all associated files (the "Work") are licensed under the Apache
|
|
License, Version 2.0 (the "License"); you may not use this project except in
|
|
compliance with the License. You may obtain a copy of the License at:
|
|
|
|
http://www.apache.org/licenses/LICENSE-2.0
|
|
|
|
Unless required by applicable law or agreed to in writing, software distributed
|
|
under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR
|
|
CONDITIONS OF ANY KIND, either express or implied. See the License for the
|
|
specific language governing permissions and limitations under the License. |