--- license: apache-2.0 base_model: Qwen/Qwen2.5-7B-Instruct tags: - text-generation - shell - bash - code - qlora - unsloth language: - en datasets: - AuricErgeson/text-to-shell-dataset pipeline_tag: text-generation --- # ShellWhiz-7B A fine-tune of Qwen2.5-7B-Instruct that turns a plain-English request into a shell command. Type what you want to do, get back `find`, `grep`, `docker`, `git`, or whatever fits. ## Why this exists I wanted something I could actually type "show me the 5 biggest files in this folder" into and get a working `du`/`sort`/`head` pipeline back, instead of half-remembering the flags myself. There are commercial tools that do this (Warp, some IDE plugins), but I couldn't find a small open model that just did the one thing well, so I built one. ## What it's good at Trained on 697 natural-language-to-shell-command pairs covering: - File and directory operations (`find`, `cp`, `mv`, `rm`, `chmod`, `du`) - Text processing (`grep`, `sed`, `awk`, `sort`, `cut`) - Git workflows - Docker and docker-compose - Process management (`ps`, `kill`, `systemctl`) - Networking (`curl`, `ssh`, `scp`, `ping`) - Archiving and package management (`tar`, `zip`, `apt`, `pip`, `npm`) ## Examples These are from the actual post-training sanity check, not cherry-picked from the training set: ``` > find all python files modified in the last 24 hours find -name '*.py' -mtime -1 > show me the 5 largest files in this directory find . -type f -exec du -h {} + | sort -rh | head -5 > list all running docker containers docker ps ``` The `` placeholder is intentional. The model was trained to use placeholders where a real path would depend on context, rather than guessing one. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_name = "AuricErgeson/shellwhiz-7b" tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16, device_map="auto") system_msg = "You are a helpful assistant that converts natural language requests into precise shell commands. Respond with ONLY the shell command, no explanation." messages = [ {"role": "system", "content": system_msg}, {"role": "user", "content": "find all files larger than 100MB"}, ] inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device) outputs = model.generate(inputs, max_new_tokens=100, temperature=0.1) print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)) ``` A GGUF (q4_k_m) build is also available in this repo if you want to run it locally through Ollama or llama.cpp. ## Training details - **Base model:** Qwen2.5-7B-Instruct - **Method:** QLoRA (4-bit), rank 16, alpha 16, no dropout - **Trainable parameters:** 40,370,176 of 7,655,986,688 (0.53 percent) - **Hardware:** single T4 GPU, Google Colab free tier - **Epochs:** 3, 264 total steps, effective batch size 8 - **Training time:** about 17 minutes Loss dropped from 2.96 at step 10 to 0.18 by the end of training and flattened out around step 190, with no spikes or divergence: | Step | Loss | |------|------| | 10 | 2.960 | | 50 | 0.391 | | 100 | 0.268 | | 150 | 0.258 | | 200 | 0.188 | | 260 | 0.188 | ## Evaluation I ran the model against 105 held-out prompts it never saw during training, phrased differently from the training set on purpose to test generalization rather than recall. Each output was judged by Claude against a known-correct reference command, allowing for different-but-equivalent approaches (there's rarely only one right way to write a shell command). | Verdict | Count | Percent | |---------|-------|---------| | Correct | 58 | 55.2% | | Partial (right idea, has a bug) | 26 | 24.8% | | Wrong | 21 | 20.0% | Syntax validity (does `bash -n` parse it without error) came out at 100/105, or 95.2 percent. The five syntax failures were almost all cases where the model left a bracketed placeholder like `` or `` in a spot where bash needs an actual token, which reads as a formatting habit rather than the model not understanding the command it's building. The wrong and partial cases cluster into a few recognizable patterns, worth knowing before you rely on this for anything important: - **Hallucinated flags.** A couple of failures invented flags that don't exist on the real tool (`docker images --sort`, `pkill --exclude`). These would fail immediately with an error, so at least they're not silently wrong. - **Negation and inversion.** When a prompt asks for something to be turned *off* or a filter to be the *inverse* of the obvious reading, the model sometimes gets the polarity backwards (e.g. batch mode requested off, model turns it on). - **Dropped constraints on multi-part requests.** Given an instruction with two or three requirements stacked together (filter by host AND connection state, auto-remove AND port map), the model sometimes satisfies one and quietly drops another. - **Leftover placeholders when a real value was given.** A few outputs used `:` style placeholders even when the prompt spelled out a concrete value like `nginx:latest`. None of this is surprising for 700 training examples on a 7B model, but it's worth knowing which categories to double check rather than trusting blindly. ## Dataset The training data is published separately at [AuricErgeson/text-to-shell-dataset](https://huggingface.co/datasets/AuricErgeson/text-to-shell-dataset), generated synthetically and deduplicated on instruction text.