初始化项目,由ModelHub XC社区提供模型

Model: OpceanAI/Yuuki-best
Source: Original Platform
This commit is contained in:
ModelHub XC
2026-07-19 00:18:14 +08:00
commit 285572f504
26 changed files with 302822 additions and 0 deletions

37
.gitattributes vendored Normal file
View File

@@ -0,0 +1,37 @@
*.7z filter=lfs diff=lfs merge=lfs -text
*.arrow filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.bz2 filter=lfs diff=lfs merge=lfs -text
*.ckpt filter=lfs diff=lfs merge=lfs -text
*.ftz filter=lfs diff=lfs merge=lfs -text
*.gz filter=lfs diff=lfs merge=lfs -text
*.h5 filter=lfs diff=lfs merge=lfs -text
*.joblib filter=lfs diff=lfs merge=lfs -text
*.lfs.* filter=lfs diff=lfs merge=lfs -text
*.mlmodel filter=lfs diff=lfs merge=lfs -text
*.model filter=lfs diff=lfs merge=lfs -text
*.msgpack filter=lfs diff=lfs merge=lfs -text
*.npy filter=lfs diff=lfs merge=lfs -text
*.npz filter=lfs diff=lfs merge=lfs -text
*.onnx filter=lfs diff=lfs merge=lfs -text
*.ot filter=lfs diff=lfs merge=lfs -text
*.parquet filter=lfs diff=lfs merge=lfs -text
*.pb filter=lfs diff=lfs merge=lfs -text
*.pickle filter=lfs diff=lfs merge=lfs -text
*.pkl filter=lfs diff=lfs merge=lfs -text
*.pt filter=lfs diff=lfs merge=lfs -text
*.pth filter=lfs diff=lfs merge=lfs -text
*.rar filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
*.tar.* filter=lfs diff=lfs merge=lfs -text
*.tar filter=lfs diff=lfs merge=lfs -text
*.tflite filter=lfs diff=lfs merge=lfs -text
*.tgz filter=lfs diff=lfs merge=lfs -text
*.wasm filter=lfs diff=lfs merge=lfs -text
*.xz filter=lfs diff=lfs merge=lfs -text
*.zip filter=lfs diff=lfs merge=lfs -text
*.zst filter=lfs diff=lfs merge=lfs -text
*tfevents* filter=lfs diff=lfs merge=lfs -text
*.gguf filter=lfs diff=lfs merge=lfs -text
*.llamafile filter=lfs diff=lfs merge=lfs -text

137
LICENSE Normal file
View File

@@ -0,0 +1,137 @@
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
Copyright 2026 OpceanAI
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License.
Subject to the terms and conditions of this License, each Contributor
hereby grants to You a perpetual, worldwide, non-exclusive, no-charge,
royalty-free, irrevocable copyright license to reproduce, prepare
Derivative Works of, publicly display, publicly perform, sublicense,
and distribute the Work and such Derivative Works in Source or Object
form.
3. Grant of Patent License.
Subject to the terms and conditions of this License, each Contributor
hereby grants to You a perpetual, worldwide, non-exclusive, no-charge,
royalty-free, irrevocable (except as stated in this section) patent
license to make, have made, use, offer to sell, sell, import, and
otherwise transfer the Work, where such license applies only to those
patent claims licensable by such Contributor that are necessarily
infringed by their Contribution(s) alone or by combination of their
Contribution(s) with the Work to which such Contribution(s) was
submitted. If You institute patent litigation against any entity
alleging that the Work or a Contribution constitutes patent
infringement, then any patent licenses granted under this License
shall terminate as of the date such litigation is filed.
4. Redistribution.
You may reproduce and distribute copies of the Work or Derivative
Works thereof in any medium, with or without modifications, provided
that You meet the following conditions:
(a) You must give recipients a copy of this License; and
(b) You must cause modified files to carry prominent notices stating
that You changed the files; and
(c) You must retain all copyright, patent, trademark, and attribution
notices; and
(d) Any NOTICE file must be included if present.
5. Submission of Contributions.
Unless You explicitly state otherwise, any Contribution submitted
shall be under the terms of this License.
6. Trademarks.
This License does not grant permission to use the trade names,
trademarks, or service marks of the Licensor.
7. Disclaimer of Warranty.
The Work is provided on an "AS IS" BASIS, WITHOUT WARRANTIES OR
CONDITIONS OF ANY KIND.
8. Limitation of Liability.
In no event shall any Contributor be liable for damages arising from
the use of the Work.
9. Accepting Warranty or Additional Liability.
You may offer support or warranty only on Your own behalf.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
Copyright 2026 OpceanAI
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

783
README.md Normal file
View File

@@ -0,0 +1,783 @@
---
license: apache-2.0
datasets:
- bigcode/the-stack
- OpceanAI/Yuuki-dataset
language:
- en
- es
base_model:
- openai-community/gpt2
pipeline_tag: text-generation
library_name: pytorch
tags:
- code
- transformers
metrics:
- perplexity
- code_eval
new_version: OpceanAI/Yuuki-v0.1
inference: true
widget:
- text: "def fibonacci(n):"
example_title: "Python Function"
- text: "module Main where"
example_title: "Agda Module"
- text: "int main() {"
example_title: "C Program"
---
<div align="center">
<br>
<img src="https://img.shields.io/badge/%E2%9C%A6-YUUKI--BEST-000000?style=for-the-badge&labelColor=000000" alt="Yuuki Best" height="50">
<br><br>
# The Best Checkpoint of the $0 Phone-Trained LLM
**Strongest initial model trained entirely on a smartphone.**<br>
**GPT-2 architecture. Checkpoint 2000. 146% improvement over checkpoint 1400.**
<br>
<a href="#features"><img src="https://img.shields.io/badge/FEATURES-000000?style=for-the-badge" alt="Features"></a>
&nbsp;&nbsp;
<a href="https://huggingface.co/spaces/OpceanAI/Yuuki"><img src="https://img.shields.io/badge/LIVE_DEMO-000000?style=for-the-badge" alt="Demo"></a>
&nbsp;&nbsp;
<a href="https://github.com/sponsors/aguitauwu"><img src="https://img.shields.io/badge/SPONSOR-000000?style=for-the-badge" alt="Sponsor"></a>
<br><br>
[![License](https://img.shields.io/badge/Apache_2.0-222222?style=flat-square&logo=opensourceinitiative&logoColor=white)](LICENSE)
&nbsp;
[![Model](https://img.shields.io/badge/GPT--2-222222?style=flat-square&logo=openai&logoColor=white)](https://huggingface.co/openai-community/gpt2)
&nbsp;
[![Dataset](https://img.shields.io/badge/The_Stack-222222?style=flat-square&logo=databricks&logoColor=white)](https://huggingface.co/datasets/bigcode/the-stack)
&nbsp;
[![PyTorch](https://img.shields.io/badge/PyTorch-222222?style=flat-square&logo=pytorch&logoColor=white)](https://pytorch.org/)
&nbsp;
[![Transformers](https://img.shields.io/badge/Transformers-222222?style=flat-square&logo=huggingface&logoColor=white)](https://huggingface.co/docs/transformers)
<br>
---
<br>
</div>
## ⚠️ Important Disclaimer
**Yuuki-best** is the **strongest checkpoint** of the Yuuki project at checkpoint 2000 (5.3% training progress). This is an **early-stage research snapshot**, not a production-ready model.
- 🔬 **Research project** - Exploring mobile-based LLM training
- 📱 **Single-person effort** - Trained entirely on a smartphone
- 📄 **Research paper coming** - Full methodology and findings
- 🚧 **Early development** - Performance will improve significantly in v0.1
<br>
---
<br>
<div align="center">
## What is Yuuki-best?
</div>
<br>
**Yuuki-best** represents the **best checkpoint** (step 2000) of the Yuuki code generation model — a multilingual LLM trained entirely on a **Redmi 12 smartphone** with **zero cloud budget**. This checkpoint demonstrates major qualitative improvements over earlier versions, with a **146% average score increase** and clear evidence of real language learning.
The model is based on **GPT-2 architecture** (82M parameters) and has been trained on **The Stack** dataset with 75,000 code examples. This checkpoint shows:
-**Functional training pipeline** - Proven to work on mobile CPU
-**Real language learning** - Generates actual Agda imports and structures
-**Structured code outputs** - Syntactic scaffolding emerging
-**Measurable progress** - 146% improvement in just 1.6% more training
<br>
---
<br>
<div align="center">
## Features
</div>
<br>
<table>
<tr>
<td width="50%" valign="top">
**Best Initial Checkpoint**
Checkpoint 2000 represents the strongest model snapshot so far, with clear improvements in code structure, language awareness, and quality scores compared to earlier checkpoints.
<br>
**Real Language Learning**
Generates genuine Agda imports (Cubical, Data.Nat, Function) and shows early understanding of language-specific tokens and patterns across multiple programming languages.
<br>
**Transparent Evaluation**
Unfiltered generation samples showing both successes and limitations. Honest assessment of current capabilities at ~5% training progress with clear quality metrics.
<br>
**146% Quality Improvement**
Average evaluation score increased from 10/100 (checkpoint 1400) to 24.6/100 (checkpoint 2000) despite only 1.6% additional training — demonstrating rapid early learning.
</td>
<td width="50%" valign="top">
**Zero-Budget Training**
Trained on a $150 Android phone with no cloud compute, no GPU acceleration, and no data center infrastructure. Proof that AI training is accessible to everyone.
<br>
**Multiple Quantized Formats**
Available in GGUF format with multiple quantization levels (Q4_0, Q4_K_M, Q5_K_M, Q8_0, F32) for efficient CPU and mobile inference.
<br>
**Early Language Specialization**
Due to alphabetical dataset ordering, shows strongest performance in Agda (55/100), with C, Assembly, and other languages progressively learning.
<br>
**Part of Complete Ecosystem**
Integrated with CLI tools (yuy, yuy-chat), web interfaces (Yuuki-chat, Yuuki-web), and comprehensive documentation for easy deployment.
</td>
</tr>
</table>
<br>
---
<br>
<div align="center">
## Checkpoint Comparison
</div>
<br>
### Performance Metrics
| Metric | Checkpoint 1400 | Checkpoint 2000 | Improvement |
|:-------|:----------------|:----------------|:------------|
| **Training Progress** | 1,400 / 37,500 (3.7%) | 2,000 / 37,500 (5.3%) | +1.6% |
| **Average Loss** | 1.70 2.23 | 1.69 2.31 | Similar |
| **Training Speed** | ~100 sec/step | ~86 sec/step | **14% faster** |
| **Model Size** | 988 MB | 988 MB | Same |
| **Evaluated Languages** | 5 languages | 5 languages | Same |
<br>
### Language Evaluation Scores
<table>
<tr>
<td width="50%" valign="top">
**Checkpoint 1400**
| Language | Score |
|:---------|:------|
| Agda | 20 / 100 |
| C | 8 / 100 |
| Assembly | 2 / 100 |
| **Average** | **10 / 100** |
</td>
<td width="50%" valign="top">
**Checkpoint 2000 (Yuuki-best)**
| Language | Score |
|:---------|:------|
| Agda | 55 / 100 |
| C | 20 / 100 |
| Assembly | 15 / 100 |
| **Average** | **24.6 / 100** |
</td>
</tr>
</table>
<br>
**📈 +146% average improvement** despite minimal additional training (≈1.6% more steps).
<br>
---
<br>
<div align="center">
## Code Quality Evolution
</div>
<br>
### 🔻 Checkpoint 1400 — Early, Noisy, Unstable
<table>
<tr>
<td width="50%" valign="top">
**Agda Test**
```agda
BLAS : Bad memory unallocation! : 50 0x72ec90a000
module Main where
open import Data.PropositionalEquality using (PropositionalEquality;
open import Data.Nat
data import Data.PropositionalEquality using (Propos;
open-Equality -PropositionalEquality) using (Prop-} (Prop-Prop-ElAdvertisements)
```
❌ Severe corruption, mixed tokens, runtime noise
</td>
<td width="50%" valign="top">
**C Test**
```c
int main() {
prostitution of '/*' (prostitution of '*'
prost = $0)
```
❌ No syntactic coherence, random token associations
</td>
</tr>
<tr>
<td width="50%" valign="top">
**Assembly Test**
```asm
mov eax,
the International Press S.
```
❌ No instruction continuation
</td>
<td width="50%" valign="top">
**Python Test**
```python
def hello():
open program
open import HUnit.H.P.B.P.P-E
```
❌ Python not learned, heavy cross-language contamination
</td>
</tr>
</table>
<br>
### 🔺 Checkpoint 2000 — Structured, Language-Aware (Best)
<table>
<tr>
<td width="50%" valign="top">
**Agda Test**
```agda
module Main where
open import Function
--
open import Data.Nat
open import Function
open import Data.Nat
open import Data.Unit
open import Data.Nat
open import Data.Nat.Dec
open import Data.Properties.Nat
-- Dec
open import Data.Nat
open import Data.Properties
open import Data.Base
open import Data.Nat.Properties
open import Data.Unary
open import data
```
**Strong familiarity with real Agda modules and imports**
</td>
<td width="50%" valign="top">
**C Test** ⚠️
```c
int main() {
no-print "no"
no-print "no"
print "no-print "
print "no"
print "no"
print "no"
fault-print "no"
}
run run
```
**Syntax unstable, but structure and intent present**
</td>
</tr>
<tr>
<td width="50%" valign="top">
**Assembly Test** 🌱
```asm
mov eax, "v2.2"
ov eax, "v1"
ov eax, "v2.2"
ov eax, "v2._v1"
ov eax, "v3"
```
**Token-level familiarity with assembly-like patterns**
</td>
<td width="50%" valign="top">
**JavaScript Test** ⚠️
```javascript
function test() {
for (i=0;i=0; i++, 1) {
(a=0; i++) {
.-1 . A
```
**Logic broken, but syntactic scaffolding emerging**
</td>
</tr>
</table>
<br>
---
<br>
<div align="center">
## Interpretation
</div>
<br>
These outputs confirm that Yuuki at checkpoint 2000:
<table>
<tr>
<td width="50%" valign="top">
**What It Learned**
- 📚 **Real libraries** - Actual Agda modules (Data.Nat, Cubical, Function)
- 🏗️ **Early structure** - Syntactic scaffolding before correctness
- 🎯 **Language tokens** - Recognizes language-specific patterns
- 📈 **Gradual improvement** - Natural learning progression
</td>
<td width="50%" valign="top">
**What It Shows**
-**Dataset ordering effects** - Early languages learned first
-**Real progress** - 146% quality improvement
-**Functional pipeline** - Training works on mobile CPU
-**Healthy behavior** - Expected at ~5% training progress
</td>
</tr>
</table>
<br>
**This behavior is expected and healthy at ~5% total training.**
<br>
---
<br>
<div align="center">
## Key Takeaway
</div>
<br>
Between **3.7% → 5.3%** training progress, Yuuki shows:
-**Major qualitative gains** - 146% improvement in evaluation scores
-**Clear specialization trends** - Strong Agda performance emerging
-**Rapid early learning** - Despite CPU-only constraints
This validates the project's core claim:
> **Progress is real, measurable, and reproducible — even at $0 cost.**
<br>
---
<br>
<div align="center">
## Available Formats
</div>
<br>
### GGUF Quantized Models
Optimized for CPU inference with llama.cpp and Ollama.
| Format | Size | Use Case | Quality |
|:-------|:-----|:---------|:--------|
| **yuuki-best-f32.gguf** | ~328 MB | Full precision baseline | Best |
| **yuuki-best-q8_0.gguf** | ~87 MB | High quality, smaller size | Excellent |
| **yuuki-best-q5_k_m.gguf** | ~56 MB | Balanced quality/size | Very Good |
| **yuuki-best-q4_k_m.gguf** | ~47 MB | Good quality, efficient | Good |
| **yuuki-best-q4_0.gguf** | ~46 MB | Most efficient, fast | Good |
<br>
---
<br>
<div align="center">
## Usage
</div>
<br>
### With Transformers (PyTorch)
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load model
model = AutoModelForCausalLM.from_pretrained("OpceanAI/Yuuki-best")
tokenizer = AutoTokenizer.from_pretrained("OpceanAI/Yuuki-best")
# Generate code
prompt = "module Main where"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=100, temperature=0.7)
code = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(code)
```
<br>
### With llama.cpp (GGUF)
```bash
# Run inference with quantized model
./llama.cpp/main -m yuuki-best-q4_k_m.gguf \
-p "module Main where" \
-n 50 \
-t 4 \
--temp 0.7
```
<br>
### With Ollama
```bash
# Create Modelfile
cat > Modelfile << EOF
FROM ./yuuki-best-q4_k_m.gguf
TEMPLATE """{{ .Prompt }}"""
PARAMETER temperature 0.7
PARAMETER top_p 0.9
EOF
# Import and run
ollama create yuuki-best -f Modelfile
ollama run yuuki-best "module Main where"
```
<br>
---
<br>
<div align="center">
## Training Configuration
</div>
<br>
<table>
<tr>
<td width="50%" valign="top">
**Hardware**
| Component | Specification |
|:----------|:--------------|
| Device | Redmi 12 (Android phone) |
| CPU | Snapdragon 685 (8-core ARM) |
| RAM | 6 GB |
| Storage | 128 GB |
| Training Mode | CPU only |
| Cost | **$0.00** |
</td>
<td width="50%" valign="top">
**Model Parameters**
| Parameter | Value |
|:----------|:------|
| Base Model | GPT-2 (82M parameters) |
| Dataset | The Stack + Yuuki-dataset |
| Checkpoint | 2000 / 37,500 steps |
| Progress | 5.3% |
| Training Speed | ~86 sec/step |
| Loss Range | 1.69 2.31 |
</td>
</tr>
</table>
<br>
---
<br>
<div align="center">
## Philosophy
</div>
<br>
> **"Progress is real, measurable, and reproducible — even at $0 cost."**
This checkpoint proves:
-**LLM training works on mobile** - Real learning on consumer hardware
-**Quality improves rapidly** - 146% gain with minimal training
-**$0 budget is viable** - No cloud compute needed
-**Anyone can contribute** - Breaking barriers to AI development
<br>
---
<br>
<div align="center">
## Related Projects
</div>
<br>
| Project | Description |
|:--------|:------------|
| [Yuuki-v0.1](https://huggingface.co/OpceanAI/Yuuki-v0.1) | Latest release version (coming soon) |
| [Yuuki-3.7](https://huggingface.co/OpceanAI/Yuuki-3.7) | Intermediate checkpoint model |
| [yuy](https://github.com/YuuKi-OS/yuy) | CLI for downloading, managing, and running Yuuki models |
| [yuy-chat](https://github.com/YuuKi-OS/yuy-chat) | TUI chat interface for local AI conversations |
| [Yuuki-chat](https://github.com/YuuKi-OS/Yuuki-chat) | Web-based chat interface with research modes |
| [Yuuki-web](https://github.com/YuuKi-OS/Yuuki-web) | Official landing page and project showcase |
| [yuuki-training](https://github.com/YuuKi-OS/yuuki-training) | Training code and scripts |
| [Yuuki Space](https://huggingface.co/spaces/OpceanAI/Yuuki) | Web-based interactive demo |
<br>
---
<br>
<div align="center">
## Links
</div>
<br>
<div align="center">
[![Model Weights](https://img.shields.io/badge/Model_Weights-Hugging_Face-ffd21e?style=for-the-badge&logo=huggingface&logoColor=black)](https://huggingface.co/OpceanAI/Yuuki-best)
&nbsp;
[![Live Demo](https://img.shields.io/badge/Live_Demo-Spaces-ffd21e?style=for-the-badge&logo=huggingface&logoColor=black)](https://huggingface.co/spaces/OpceanAI/Yuuki)
&nbsp;
[![Training Code](https://img.shields.io/badge/Training_Code-GitHub-181717?style=for-the-badge&logo=github&logoColor=white)](https://github.com/YuuKi-OS/yuuki-training)
<br>
[![YUY CLI](https://img.shields.io/badge/Yuy_CLI-GitHub-181717?style=for-the-badge&logo=github&logoColor=white)](https://github.com/YuuKi-OS/yuy)
&nbsp;
[![YUY Chat](https://img.shields.io/badge/Yuy_Chat-GitHub-181717?style=for-the-badge&logo=github&logoColor=white)](https://github.com/YuuKi-OS/yuy-chat)
&nbsp;
[![Sponsor](https://img.shields.io/badge/Sponsor-GitHub_Sponsors-ea4aaa?style=for-the-badge&logo=githubsponsors&logoColor=white)](https://github.com/sponsors/aguitauwu)
</div>
<br>
---
<br>
<div align="center">
## Community
</div>
<br>
Join the Yuuki community:
- 💬 [Discord Server](https://discord.gg/j8zV2u8k) - Chat with other users and contributors
- 🐦 [Twitter Updates](https://twitter.com/aguitauwu) - Follow development progress
- 📺 [GitHub](https://github.com/aguitauwu) - Star repos and contribute
- 💖 [GitHub Sponsors](https://github.com/sponsors/aguitauwu) - Support the project
- 🦙 .[ollama](https://ollama.com/aguitachan3/yuuki-best)
<br>
---
<br>
<div align="center">
## Acknowledgments
</div>
<br>
- **HuggingFace** - Infrastructure and transformers library
- **BigCode** - The Stack dataset
- **The ML community** - For inspiration and support
- **Everyone following along** - Your interest makes this worthwhile
<br>
---
<br>
<div align="center">
## License
</div>
<br>
```
Apache License 2.0
Copyright (c) 2026 Yuuki Project
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
```
**You can use Yuuki commercially, modify it, distribute it. Just give credit.**
<br>
---
<br>
<div align="center">
**Built with patience, a phone, and zero budget.**
<br>
[![Yuuki Project](https://img.shields.io/badge/Yuuki_Project-2026-000000?style=for-the-badge)](https://huggingface.co/OpceanAI)
<br>
*The best checkpoint so far. More improvements coming soon.* 🌸
</div>

45
config.json Normal file
View File

@@ -0,0 +1,45 @@
{
"_num_labels": 1,
"activation_function": "gelu_new",
"architectures": [
"GPT2LMHeadModel"
],
"attn_pdrop": 0.1,
"bos_token_id": 50256,
"dtype": "float32",
"embd_pdrop": 0.1,
"eos_token_id": 50256,
"id2label": {
"0": "LABEL_0"
},
"initializer_range": 0.02,
"label2id": {
"LABEL_0": 0
},
"layer_norm_epsilon": 1e-05,
"model_type": "gpt2",
"n_ctx": 1024,
"n_embd": 768,
"n_head": 12,
"n_inner": null,
"n_layer": 6,
"n_positions": 1024,
"reorder_and_upcast_attn": false,
"resid_pdrop": 0.1,
"scale_attn_by_inverse_layer_idx": false,
"scale_attn_weights": true,
"summary_activation": null,
"summary_first_dropout": 0.1,
"summary_proj_to_labels": true,
"summary_type": "cls_index",
"summary_use_proj": true,
"task_specific_params": {
"text-generation": {
"do_sample": true,
"max_length": 50
}
},
"transformers_version": "4.57.3",
"use_cache": true,
"vocab_size": 50257
}

6
generation_config.json Normal file
View File

@@ -0,0 +1,6 @@
{
"_from_model_config": true,
"bos_token_id": 50256,
"eos_token_id": 50256,
"transformers_version": "4.57.3"
}

50001
merges.txt Normal file

File diff suppressed because it is too large Load Diff

3
model.safetensors Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:5c96d284b18a64cce26aacb4b1a632fdee492d5befc907e30f0af74c00d060ca
size 327657928

3
optimizer.pt Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:3c0c195a1e422cdf40f0111f3d439123654ab525403ce7044c6304d325ad3364
size 655362763

3
rng_state.pth Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:d6d709785c99d1ad065c4bc294f83293a500ac714a6810ca85d36d3ac5930bcb
size 14455

3
scheduler.pt Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:f9d6bc69318fecc0eb83dc1f3ff910ae9f261fdcef983cd951757193ca67b223
size 1465

6
special_tokens_map.json Normal file
View File

@@ -0,0 +1,6 @@
{
"bos_token": "<|endoftext|>",
"eos_token": "<|endoftext|>",
"pad_token": "<|endoftext|>",
"unk_token": "<|endoftext|>"
}

250306
tokenizer.json Normal file

File diff suppressed because it is too large Load Diff

21
tokenizer_config.json Normal file
View File

@@ -0,0 +1,21 @@
{
"add_prefix_space": false,
"added_tokens_decoder": {
"50256": {
"content": "<|endoftext|>",
"lstrip": false,
"normalized": true,
"rstrip": false,
"single_word": false,
"special": true
}
},
"bos_token": "<|endoftext|>",
"clean_up_tokenization_spaces": false,
"eos_token": "<|endoftext|>",
"extra_special_tokens": {},
"model_max_length": 1024,
"pad_token": "<|endoftext|>",
"tokenizer_class": "GPT2Tokenizer",
"unk_token": "<|endoftext|>"
}

1434
trainer_state.json Normal file

File diff suppressed because it is too large Load Diff

3
training_args.bin Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:51f2309beb94ee35ddf1e12568a4bc05482fa1aa824622ff30b84e67703c36d6
size 5841

1
vocab.json Normal file

File diff suppressed because one or more lines are too long

3
yuuki-best-f32.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd10400e41b71b4dd554706e9f028d12258a1751af76db42d11b03b89868adb7
size 329422368

3
yuuki-best-f32.llamafile Executable file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:fd10400e41b71b4dd554706e9f028d12258a1751af76db42d11b03b89868adb7
size 329422368

3
yuuki-best-q4_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:634d2b32b500399ca7844b24cb3b4e7d52407aa300f506735b91b675a3c8c94d
size 60713344

3
yuuki-best-q4_0.llamafile Executable file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:634d2b32b500399ca7844b24cb3b4e7d52407aa300f506735b91b675a3c8c94d
size 60713344

3
yuuki-best-q4_k_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a87d704e516a9e68f1b231bc5412c1b5ab579199375c07135328913be7d05cd0
size 63256960

3
yuuki-best-q4_k_m.llamafile Executable file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:a87d704e516a9e68f1b231bc5412c1b5ab579199375c07135328913be7d05cd0
size 63256960

3
yuuki-best-q5_k_m.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:51e2d24fa211cba8b2b9b36f8a528d752e3a2dfb0ff0d79adb0dbf53770fedef
size 68058496

3
yuuki-best-q5_k_m.llamafile Executable file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:51e2d24fa211cba8b2b9b36f8a528d752e3a2dfb0ff0d79adb0dbf53770fedef
size 68058496

3
yuuki-best-q8_0.gguf Normal file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:611d329e0b05b1420cc93fb3a01bd7376a61c048c33fb3c92de90efc9040913f
size 91294816

3
yuuki-best-q8_0.llamafile Executable file
View File

@@ -0,0 +1,3 @@
version https://git-lfs.github.com/spec/v1
oid sha256:611d329e0b05b1420cc93fb3a01bd7376a61c048c33fb3c92de90efc9040913f
size 91294816