初始化项目,由ModelHub XC社区提供模型
Model: selorahomes/Selora-AI Source: Original Platform
This commit is contained in:
82
.gitattributes
vendored
Normal file
82
.gitattributes
vendored
Normal file
@@ -0,0 +1,82 @@
|
|||||||
|
*.7z filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.arrow filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bin filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ftz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.gz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.h5 filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.joblib filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.model filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npy filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.npz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.onnx filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.ot filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.parquet filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pb filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pickle filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pkl filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pt filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.pth filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.rar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
||||||
|
saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar.* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tar filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tflite filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.tgz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.wasm filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.xz filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zip filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*.zst filter=lfs diff=lfs merge=lfs -text
|
||||||
|
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen25_15b_answer.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen25_15b_automation.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen25_15b_base.Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen25_15b_clarification.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen25_15b_command.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_answer.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_automation.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_base.IQ4_XS.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_base.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_clarification.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_command.lora.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v042-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v042-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v042-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v042-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v043-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v043-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v043-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v043-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v043-recipe.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v046-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
qwen3_17b_base.Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v047-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v047-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v047-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v047-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v048-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v048-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v048-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v048-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-v048-utilities.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-answer.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-automation.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-clarification.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-command.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-utilities.f16.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
ollama-single-q6/selora-qwen-0.4.8-q6.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
ollama/selora-qwen.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
ollama/selora-ollama.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
|
selora-ollama.Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
|
||||||
34
Modelfile.answer
Normal file
34
Modelfile.answer
Normal file
@@ -0,0 +1,34 @@
|
|||||||
|
# selora-qwen-answer (Selora AI v0.4.8) — Ollama recipe.
|
||||||
|
# Put this file in the same directory as the downloaded GGUFs, then:
|
||||||
|
# ollama create selora-qwen-answer -f Modelfile.answer
|
||||||
|
FROM ./qwen3_17b_base.Q6_K.gguf
|
||||||
|
ADAPTER ./selora-answer.f16.gguf
|
||||||
|
|
||||||
|
TEMPLATE """{{ if .System }}<|im_start|>system
|
||||||
|
{{ .System }}<|im_end|>
|
||||||
|
{{ end }}{{ if .Prompt }}<|im_start|>user
|
||||||
|
/no_think {{ .Prompt }}<|im_end|>
|
||||||
|
{{ end }}<|im_start|>assistant
|
||||||
|
"""
|
||||||
|
|
||||||
|
SYSTEM """You are Selora AI's answer specialist for Home Assistant.
|
||||||
|
|
||||||
|
Given a user question and the AVAILABLE ENTITIES list, respond with ONE JSON object only:
|
||||||
|
{"r":"<response with {entity_id} placeholders where state is needed>","q":["<entity_id>",...]}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- r: response template. Use {entity_id} placeholders for any state references; the consumer substitutes live state. Keep r short — 1-2 sentences max.
|
||||||
|
- q: array of entity_ids to look up. Omit when no live state is needed.
|
||||||
|
- Either field can be omitted if not used, but never both.
|
||||||
|
- Only reference entity_ids that appear in AVAILABLE ENTITIES below.
|
||||||
|
- Never invent state values; always template them via {entity_id}.
|
||||||
|
- If the question is outside the home's scope, return {"r":"I can only answer questions about your home."}.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
"""
|
||||||
|
|
||||||
|
PARAMETER temperature 0.0
|
||||||
|
PARAMETER repeat_penalty 1.0
|
||||||
|
PARAMETER repeat_last_n 256
|
||||||
|
PARAMETER stop "<|im_end|>"
|
||||||
|
PARAMETER stop "<|endoftext|>"
|
||||||
47
Modelfile.automation
Normal file
47
Modelfile.automation
Normal file
@@ -0,0 +1,47 @@
|
|||||||
|
# selora-qwen-automation (Selora AI v0.4.8) — Ollama recipe.
|
||||||
|
# Put this file in the same directory as the downloaded GGUFs, then:
|
||||||
|
# ollama create selora-qwen-automation -f Modelfile.automation
|
||||||
|
FROM ./qwen3_17b_base.Q6_K.gguf
|
||||||
|
ADAPTER ./selora-automation.f16.gguf
|
||||||
|
|
||||||
|
TEMPLATE """{{ if .System }}<|im_start|>system
|
||||||
|
{{ .System }}<|im_end|>
|
||||||
|
{{ end }}{{ if .Prompt }}<|im_start|>user
|
||||||
|
/no_think {{ .Prompt }}<|im_end|>
|
||||||
|
{{ end }}<|im_start|>assistant
|
||||||
|
"""
|
||||||
|
|
||||||
|
SYSTEM """You are Selora AI, an automation architect for Home Assistant. The user wants a recurring rule, schedule, or multi-step sequence saved as an automation, OR a reusable parameterized blueprint.
|
||||||
|
|
||||||
|
Choose the output format from the request:
|
||||||
|
|
||||||
|
A) BLUEPRINT request — when the user asks to "create a blueprint automation", gives a "## Detailed Description" with an input table, or otherwise wants a REUSABLE/PARAMETERIZED automation with named inputs. Respond with markdown containing a single yaml block. The response MUST start with ```yaml and end with ``` because it is parsed by code. Inside, emit a Home Assistant automation BLUEPRINT:
|
||||||
|
- top-level `blueprint:` with `name:`, `description:`, `domain: automation`, and `input:` containing EVERY input named in the request (use the exact input keys from the request's input table).
|
||||||
|
- each input has a `name:` and a `selector:` (entity/number/duration/media/etc.); add `default:` for timeouts/durations/levels.
|
||||||
|
- wire inputs into triggers/conditions/actions with `!input <input_name>` references — never hardcode entity_ids in a blueprint.
|
||||||
|
- use `mode: single` and `max_exceeded: silent`. States are quoted strings. Times/durations are "HH:MM:SS".
|
||||||
|
- Output ONLY the ```yaml block (inline comments allowed).
|
||||||
|
|
||||||
|
B) CONCRETE request — a one-off command/schedule grounded in this user's AVAILABLE ENTITIES. Return ONE JSON object:
|
||||||
|
{"intent":"automation","response":"<1-2 sentence explanation>","description":"<2-3 sentences>","automation":{"alias":"<max 4 words>","description":"<...>","triggers":[<one-or-more>],"conditions":[<optional>],"actions":[<one-or-more>]}}
|
||||||
|
Or, if no entity matches, the clarification shape:
|
||||||
|
{"intent":"clarification","response":"<ONE specific follow-up question naming candidate devices from AVAILABLE ENTITIES>"}
|
||||||
|
|
||||||
|
BLUEPRINT RULES (format A):
|
||||||
|
- The ```yaml fence is mandatory and literal — lowercase `yaml`, not `yml`/`YAML`/bare ```.
|
||||||
|
- `blueprint.domain` is ALWAYS `automation`.
|
||||||
|
- Triggers inside a blueprint use `platform:` (e.g. `- platform: state`, `- platform: numeric_state`, `- platform: time`, `- platform: sun`).
|
||||||
|
- `!input` may reference a scalar (`entity_id: !input door_sensor`), a whole `target:` (`target: !input light_switch`), or a whole `data:` (`data: !input alert_media`).
|
||||||
|
- Do NOT add a required input the caller will not supply; any input beyond the requested set MUST have a `default:`.
|
||||||
|
|
||||||
|
CONCRETE RULES (format B):
|
||||||
|
- Use HA 2024+ plural keys: 'triggers', 'actions', 'conditions'. Service calls use the 'service' key.
|
||||||
|
- State 'to'/'from' MUST be strings. Times "HH:MM:SS". Durations "HH:MM:SS" or {"hours":N,...}.
|
||||||
|
- EVERY entity_id MUST appear VERBATIM in AVAILABLE ENTITIES; never invent placeholder names.
|
||||||
|
"""
|
||||||
|
|
||||||
|
PARAMETER temperature 0.0
|
||||||
|
PARAMETER repeat_penalty 1.0
|
||||||
|
PARAMETER repeat_last_n 256
|
||||||
|
PARAMETER stop "<|im_end|>"
|
||||||
|
PARAMETER stop "<|endoftext|>"
|
||||||
33
Modelfile.clarification
Normal file
33
Modelfile.clarification
Normal file
@@ -0,0 +1,33 @@
|
|||||||
|
# selora-qwen-clarification (Selora AI v0.4.8) — Ollama recipe.
|
||||||
|
# Put this file in the same directory as the downloaded GGUFs, then:
|
||||||
|
# ollama create selora-qwen-clarification -f Modelfile.clarification
|
||||||
|
FROM ./qwen3_17b_base.Q6_K.gguf
|
||||||
|
ADAPTER ./selora-clarification.f16.gguf
|
||||||
|
|
||||||
|
TEMPLATE """{{ if .System }}<|im_start|>system
|
||||||
|
{{ .System }}<|im_end|>
|
||||||
|
{{ end }}{{ if .Prompt }}<|im_start|>user
|
||||||
|
/no_think {{ .Prompt }}<|im_end|>
|
||||||
|
{{ end }}<|im_start|>assistant
|
||||||
|
"""
|
||||||
|
|
||||||
|
SYSTEM """You are Selora AI's clarification specialist for Home Assistant.
|
||||||
|
|
||||||
|
When the user's request is ambiguous, respond with ONE JSON object only:
|
||||||
|
{"q":"<question text>","o":["<option1>","<option2>",...]}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- q: short, specific clarifying question. 1 sentence max.
|
||||||
|
- o: optional array of suggested answers. Omit the o key when free-form input is appropriate.
|
||||||
|
- Reference entity aliases from AVAILABLE ENTITIES when the ambiguity is about which entity.
|
||||||
|
- Don't ask multiple questions in one turn — pick the single most important blocker.
|
||||||
|
- Don't restate the user's full request; ask the one thing you need.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
"""
|
||||||
|
|
||||||
|
PARAMETER temperature 0.0
|
||||||
|
PARAMETER repeat_penalty 1.0
|
||||||
|
PARAMETER repeat_last_n 256
|
||||||
|
PARAMETER stop "<|im_end|>"
|
||||||
|
PARAMETER stop "<|endoftext|>"
|
||||||
35
Modelfile.command
Normal file
35
Modelfile.command
Normal file
@@ -0,0 +1,35 @@
|
|||||||
|
# selora-qwen-command (Selora AI v0.4.8) — Ollama recipe.
|
||||||
|
# Put this file in the same directory as the downloaded GGUFs, then:
|
||||||
|
# ollama create selora-qwen-command -f Modelfile.command
|
||||||
|
FROM ./qwen3_17b_base.Q6_K.gguf
|
||||||
|
ADAPTER ./selora-command.f16.gguf
|
||||||
|
|
||||||
|
TEMPLATE """{{ if .System }}<|im_start|>system
|
||||||
|
{{ .System }}<|im_end|>
|
||||||
|
{{ end }}{{ if .Prompt }}<|im_start|>user
|
||||||
|
/no_think {{ .Prompt }}<|im_end|>
|
||||||
|
{{ end }}<|im_start|>assistant
|
||||||
|
"""
|
||||||
|
|
||||||
|
SYSTEM """You are Selora AI's command specialist for Home Assistant.
|
||||||
|
|
||||||
|
Given a user command and the AVAILABLE ENTITIES list, respond with ONE JSON object only:
|
||||||
|
{"c":[{"s":"<service>","e":"<entity_id>","d":{<optional params>}}],"r":"<short confirmation>"}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- c: ordered array of one or more service calls. Calls execute in array order.
|
||||||
|
- s: HA service in "domain.action" form (e.g. "light.turn_on", "lock.lock", "media_player.play_media", "scene.turn_on").
|
||||||
|
- e: canonical entity_id from AVAILABLE ENTITIES. Never use the human alias — always the entity_id.
|
||||||
|
- d: service parameters object. Omit the d key entirely when there are no params (do not include "d":{}).
|
||||||
|
- r: ≤ 1 sentence past-tense confirmation describing what got done (e.g. "Kitchen light on.").
|
||||||
|
- The service domain (before the dot) must match the entity_id's domain. light.turn_on goes with light.* entities, lock.lock goes with lock.* entities, etc.
|
||||||
|
- For multi-target requests, produce one c entry per (service, entity_id) pair.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
"""
|
||||||
|
|
||||||
|
PARAMETER temperature 0.0
|
||||||
|
PARAMETER repeat_penalty 1.0
|
||||||
|
PARAMETER repeat_last_n 256
|
||||||
|
PARAMETER stop "<|im_end|>"
|
||||||
|
PARAMETER stop "<|endoftext|>"
|
||||||
36
Modelfile.utilities
Normal file
36
Modelfile.utilities
Normal file
@@ -0,0 +1,36 @@
|
|||||||
|
# selora-qwen-utilities (Selora AI v0.4.8) — Ollama recipe.
|
||||||
|
# Put this file in the same directory as the downloaded GGUFs, then:
|
||||||
|
# ollama create selora-qwen-utilities -f Modelfile.utilities
|
||||||
|
FROM ./qwen3_17b_base.Q6_K.gguf
|
||||||
|
ADAPTER ./selora-utilities.f16.gguf
|
||||||
|
|
||||||
|
TEMPLATE """{{ if .System }}<|im_start|>system
|
||||||
|
{{ .System }}<|im_end|>
|
||||||
|
{{ end }}{{ if .Prompt }}<|im_start|>user
|
||||||
|
/no_think {{ .Prompt }}<|im_end|>
|
||||||
|
{{ end }}<|im_start|>assistant
|
||||||
|
"""
|
||||||
|
|
||||||
|
SYSTEM """You are Selora AI's utilities specialist for Home Assistant.
|
||||||
|
|
||||||
|
You handle docs-grounded help: maintenance (pending updates, version conflicts), troubleshooting (why a device is unavailable / offline / not responding), and setup help (how to add or configure an integration). You are given the user question, the AVAILABLE ENTITIES list, and a RELEVANT DOCS list of retrieved Home Assistant documentation chunks.
|
||||||
|
|
||||||
|
Respond with ONE JSON object only:
|
||||||
|
{"r":"<advice with {entity_id} placeholders where live state is needed>","q":["<entity_id>",...],"src":["<doc_chunk_id>",...]}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- r: advice prose grounded in RELEVANT DOCS. Use {entity_id} placeholders for any state references (e.g. which integration has an update, which device is unavailable); the consumer substitutes live state. Keep it focused — a short paragraph.
|
||||||
|
- q: array of entity_ids whose live state the advice depends on (update.* entities, problem/connectivity binary_sensors, the unavailable entity). Omit or use [] when no live state is needed (e.g. setup help for a device that does not exist yet).
|
||||||
|
- src: array of doc chunk ids from RELEVANT DOCS that the advice is drawn from. This is the citation provenance — only cite chunks that appear in RELEVANT DOCS.
|
||||||
|
- Only reference entity_ids that appear in AVAILABLE ENTITIES below.
|
||||||
|
- Ground the fix/steps in the retrieved docs, never in the entity state alone — the state is the signal, the docs supply the explanation.
|
||||||
|
- Never invent state values; always template them via {entity_id}.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
"""
|
||||||
|
|
||||||
|
PARAMETER temperature 0.0
|
||||||
|
PARAMETER repeat_penalty 1.0
|
||||||
|
PARAMETER repeat_last_n 256
|
||||||
|
PARAMETER stop "<|im_end|>"
|
||||||
|
PARAMETER stop "<|endoftext|>"
|
||||||
583
README.md
Normal file
583
README.md
Normal file
@@ -0,0 +1,583 @@
|
|||||||
|
---
|
||||||
|
license: apache-2.0
|
||||||
|
base_model: Qwen/Qwen3-1.7B
|
||||||
|
tags:
|
||||||
|
- text-generation
|
||||||
|
- qwen
|
||||||
|
- qwen3
|
||||||
|
- lora
|
||||||
|
- gguf
|
||||||
|
- home-assistant
|
||||||
|
- home-automation
|
||||||
|
- smart-home
|
||||||
|
- iot
|
||||||
|
- instruction-tuned
|
||||||
|
- tool-use
|
||||||
|
- ollama
|
||||||
|
language:
|
||||||
|
- en
|
||||||
|
library_name: transformers
|
||||||
|
pipeline_tag: text-generation
|
||||||
|
---
|
||||||
|
|
||||||
|
**Selora Homes:** [selorahomes.com](https://selorahomes.com)
|
||||||
|
**Selora AI Home Assistant Integration:** [github.com/SeloraHomes/ha-selora-ai](https://github.com/SeloraHomes/ha-selora-ai)
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# Selora AI
|
||||||
|
|
||||||
|
Selora AI is an instruction-tuned language model for
|
||||||
|
[**Home Assistant**](https://www.home-assistant.io/), the open-source **smart home**
|
||||||
|
platform. It ships in two forms: **five small, single-purpose LoRA specialists** — the
|
||||||
|
reference setup the integration hot-swaps between — and a single general model,
|
||||||
|
**`selora-ollama`**, that folds all five into one for Ollama. Each specialist emits
|
||||||
|
one strict, compact JSON **"slim envelope"** for its intent, and
|
||||||
|
the [Selora AI integration](https://github.com/SeloraHomes/ha-selora-ai) executes that
|
||||||
|
envelope against your home — running service calls, resolving live state, or saving an
|
||||||
|
automation (see [Specialists](#specialists) for the exact shapes). The five specialists:
|
||||||
|
|
||||||
|
- **command** — control devices ("turn off the kitchen lights")
|
||||||
|
- **automation** — author home automations and blueprints
|
||||||
|
- **answer** — answer questions about the home's live state
|
||||||
|
- **clarification** — ask a follow-up when a request is ambiguous
|
||||||
|
- **utilities** — docs-grounded maintenance, troubleshooting, and setup help
|
||||||
|
|
||||||
|
The five specialists are LoRA adapters fine-tuned on a shared
|
||||||
|
[Qwen3 1.7B](https://huggingface.co/Qwen/Qwen3-1.7B) base.
|
||||||
|
|
||||||
|
**Base:** Qwen3 1.7B · **Format:** GGUF Q6_K base + 5 per-specialist LoRA adapters (F16) · **License:** Apache-2.0
|
||||||
|
|
||||||
|
Selora AI powers the [Selora AI Home Assistant integration](https://github.com/SeloraHomes/ha-selora-ai)
|
||||||
|
and runs locally on Apple Silicon, Linux, or Windows via
|
||||||
|
[llama-server](#llama-server-home-assistant-integration-runtime) or
|
||||||
|
[Ollama](#ollama-evaluate-the-model-before-integrating), or in the cloud via
|
||||||
|
[vLLM](#vllm-cloud) — built for self-hosted **IoT** deployments that stay private and
|
||||||
|
offline-first.
|
||||||
|
|
||||||
|
## Use cases
|
||||||
|
|
||||||
|
- **Chat control of smart-home devices** — "turn off the kitchen
|
||||||
|
lights", "set the thermostat to 68", "open the garage door" — resolved against
|
||||||
|
live Home Assistant entity state.
|
||||||
|
- **Natural-language home automation creation** — describe an automation in
|
||||||
|
plain English ("when the front door opens after 10pm, turn on the porch
|
||||||
|
light") and Selora returns valid Home Assistant YAML as a draft you review
|
||||||
|
before it's saved.
|
||||||
|
- **Scene and routine orchestration** — chain actions across multiple entities
|
||||||
|
("good night" → lock doors, dim bedroom lights, set thermostat) without
|
||||||
|
hand-writing scripts.
|
||||||
|
- **Q&A about your home** — "is the laundry running?", "what's the temperature
|
||||||
|
upstairs?" — the `answer` adapter returns the entities to check plus a
|
||||||
|
response template, and the integration fills in the live state.
|
||||||
|
- **Docs-grounded help** — "why is the living-room sensor unavailable?", "how do
|
||||||
|
I add the Hue integration?", "which integrations have pending updates?" —
|
||||||
|
answered from retrieved Home Assistant documentation, with live state pulled in
|
||||||
|
via entity placeholders and citations back to the source docs.
|
||||||
|
- **Privacy-first home assistant** — runs entirely on local hardware
|
||||||
|
(Raspberry Pi 5, Mac mini, NUC-class boxes) with no cloud dependency, so
|
||||||
|
device commands and home telemetry never leave the LAN.
|
||||||
|
|
||||||
|
## Specialists
|
||||||
|
|
||||||
|
Every specialist LoRA emits a compact **"slim envelope"** — a small JSON object whose
|
||||||
|
high-frequency keys are single characters to save tokens — and that envelope is an
|
||||||
|
**intermediate representation, not a finished action**. The Selora AI integration's
|
||||||
|
`_convert_slim_shape` parser turns it into real Home Assistant activity: run a
|
||||||
|
downloaded GGUF on its own and you get the envelope back; it takes the integration (or
|
||||||
|
an equivalent parser) to execute service calls, resolve live state, build YAML, or
|
||||||
|
attach citations. By design the model **never fabricates device state** — it emits the
|
||||||
|
calls to make, the entities to look up, the question to ask, or the automation to
|
||||||
|
build, and the integration completes the loop against live HA.
|
||||||
|
|
||||||
|
| Adapter | Emits (raw model output) | What the integration does with it |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `command` | `{"c":[{"s":"<service>","e":"<entity_id>","d":{…}}],"r":"<confirmation>"}` | Executes each `c` as an HA service call, in array order, then shows `r`. The model does **not** call services itself. |
|
||||||
|
| `answer` | `{"r":"<text with {entity_id} placeholders>","q":["<entity_id>",…]}` | Looks up the `q` entities' live state and substitutes it into the `{entity_id}` placeholders in `r`. The model templates state; it never reads it. |
|
||||||
|
| `clarification` | `{"q":"<question>","o":["<option>",…]}` | Surfaces the question and the optional quick-reply `o` options; the user's next reply is the action. |
|
||||||
|
| `utilities` | `{"r":"<advice with {entity_id} placeholders>","q":["<entity_id>",…],"src":["<doc_chunk_id>",…]}` | Substitutes live `q` state into the advice and surfaces the `src` citations. Advice is grounded in a `RELEVANT DOCS` block injected at inference. |
|
||||||
|
| `automation` | Blueprint request → a fenced `yaml` HA blueprint (`blueprint:` / `input:` / `!input`). Concrete request → `{"intent":"automation","response":…,"automation":{"alias":…,"triggers":[…],"conditions":[…],"actions":[…]}}` | Uses the blueprint YAML (nearly) as-is, or maps the concrete `automation` object JSON→YAML into a saved HA automation. |
|
||||||
|
|
||||||
|
Key map, applied once on the HA side: `r`=response, `q`=query / entity list, `c`=calls,
|
||||||
|
`s`=service, `e`=entity_id, `d`=data (service params), `o`=options.
|
||||||
|
|
||||||
|
The integration's `selora_local` provider classifies each request to one specialist (a
|
||||||
|
regex pre-classifier), then activates the matching adapter — `llama-server`'s
|
||||||
|
`/lora-adapters` hot-swap on the production hub, or vLLM `--enable-lora`.
|
||||||
|
|
||||||
|
### Example envelopes
|
||||||
|
|
||||||
|
Real output for each specialist — the raw model emission; the integration then executes
|
||||||
|
or resolves it:
|
||||||
|
|
||||||
|
```jsonc
|
||||||
|
// command — "Turn off the kitchen light and lock the front door"
|
||||||
|
{"c":[{"s":"light.turn_off","e":"light.kitchen"},{"s":"lock.lock","e":"lock.front_door"}],"r":"Kitchen light off and front door locked."}
|
||||||
|
|
||||||
|
// answer — "Is the bedroom light on?"
|
||||||
|
{"q":["light.bedroom"],"r":"The bedroom light is {light.bedroom}."}
|
||||||
|
|
||||||
|
// clarification — "Turn on a light" (when several exist)
|
||||||
|
{"q":"Which light — kitchen, bedroom, or living room?","o":["kitchen","bedroom","living room"]}
|
||||||
|
|
||||||
|
// utilities — "Why does the Hue bridge show an update?"
|
||||||
|
{"r":"The {update.philips_hue} update is pending; install it from Settings → Devices & Services → Updates.","q":["update.philips_hue"],"src":["home-assistant/updates#install"]}
|
||||||
|
|
||||||
|
// automation (concrete) — "Lock the front door at 10pm if it's still unlocked"
|
||||||
|
{"intent":"automation","response":"Created the 10pm auto-lock.","automation":{"alias":"Nightly Auto-Lock","triggers":[{"trigger":"time","at":"22:00:00"}],"conditions":[{"condition":"state","entity_id":"lock.front_door","state":"unlocked"}],"actions":[{"service":"lock.lock","entity_id":"lock.front_door"}]}}
|
||||||
|
```
|
||||||
|
|
||||||
|
The `automation` adapter's *blueprint* mode instead emits a fenced `yaml` block — a
|
||||||
|
complete HA blueprint with `blueprint:`, `input:`, and `!input` wiring — rather than the
|
||||||
|
concrete JSON shown above.
|
||||||
|
|
||||||
|
### Why split into specialist LoRAs
|
||||||
|
|
||||||
|
The split is what makes high task accuracy attainable on small (1.7B) weights, and it
|
||||||
|
is the reason the model behaves well on the evaluation surfaces below:
|
||||||
|
|
||||||
|
- **Contract-first, by construction.** The slim-envelope schema was designed up front and
|
||||||
|
baked into every training example, so each specialist learns one clean, deterministic
|
||||||
|
output shape — exactly what the integration needs to parse, and what lets the router fire
|
||||||
|
the right adapter.
|
||||||
|
- **One job per adapter.** Each LoRA learns a single tightly-constrained output
|
||||||
|
contract (its slim envelope) rather than a general chat distribution, so it can
|
||||||
|
saturate that one task instead of trading off across many.
|
||||||
|
- **Intent, not execution.** Execution, live-state lookup, and YAML construction are
|
||||||
|
the integration's responsibility. The model never invents state or device values — it
|
||||||
|
templates them — which removes a whole class of hallucination failures.
|
||||||
|
- **Slim envelopes are short.** Single-char-keyed JSON means fewer tokens to mispredict
|
||||||
|
and faster generation, and the rigid shape is trivial for the integration to parse
|
||||||
|
and validate.
|
||||||
|
- **Shared base, hot-swapped adapters.** One Q6_K base (~1.6 GB) loads once; the right
|
||||||
|
~10–40 MB adapter is activated per request, so five specialists cost roughly the
|
||||||
|
memory of a single model and every request is served by the adapter trained for it.
|
||||||
|
- **Multi-state entity context (v0.4.8).** Per-entity attribute tails in the
|
||||||
|
`AVAILABLE ENTITIES` block give each specialist richer single-turn grounding.
|
||||||
|
|
||||||
|
## Results
|
||||||
|
|
||||||
|
Evaluated on the open
|
||||||
|
[allenporter/home-assistant-datasets](https://github.com/allenporter/home-assistant-datasets)
|
||||||
|
suite — the basis of the Home Assistant LLM leaderboard (v0.4.8, temperature 0):
|
||||||
|
|
||||||
|
Each runtime is scored two ways — the model's **Raw** envelope graded directly, and **Raw + Integration** through the Selora integration's routing + repair.
|
||||||
|
|
||||||
|
### llama.cpp — 5 specialists (hot-swap)
|
||||||
|
|
||||||
|
**Raw**
|
||||||
|
|
||||||
|
| Surface | Score | Pass / total | 95% CI | Scored on |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| assist | 29.5% | 28 / 95 | ±9.2 | raw model |
|
||||||
|
| assist-mini | 74.0% | 37 / 50 | ±12.2 | raw model |
|
||||||
|
| questions | 26.3% | 10 / 38 | ±14.0 | raw model |
|
||||||
|
| automations | 66.7% | 40 / 60 | ±11.9 | raw model |
|
||||||
|
|
||||||
|
**Raw + Integration**
|
||||||
|
|
||||||
|
| Surface | Score | Pass / total | 95% CI | Scored on |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| assist | 91.7% † | 422 / 460 | ±2.5 | raw model + integration |
|
||||||
|
| assist-mini | 100% † | 196 / 196 | ±1.9 | raw model + integration |
|
||||||
|
| questions | 97.3% † | 360 / 370 | ±1.7 | raw model + integration |
|
||||||
|
| automations | 25.0% | 15 / 60 | ±11.0 | raw model + integration |
|
||||||
|
|
||||||
|
† Integration mini/assist/questions is the shared deterministic-repair layer (backend-independent — measured on the unified path); only automations is measured per-backend.
|
||||||
|
|
||||||
|
### ollama — `selora-ollama` (single merged model)
|
||||||
|
|
||||||
|
**Raw**
|
||||||
|
|
||||||
|
| Surface | Score | Pass / total | 95% CI | Scored on |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| assist | 45.3% | 43 / 95 | ±10.0 | raw model |
|
||||||
|
| assist-mini | 70.0% | 35 / 50 | ±12.7 | raw model |
|
||||||
|
| questions | 31.6% | 12 / 38 | ±14.8 | raw model |
|
||||||
|
| automations | 50.0% | 30 / 60 | ±12.7 | raw model |
|
||||||
|
|
||||||
|
**Raw + Integration**
|
||||||
|
|
||||||
|
| Surface | Score | Pass / total | 95% CI | Scored on |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| assist | 91.7% † | 422 / 460 | ±2.5 | raw model + integration |
|
||||||
|
| assist-mini | 100% † | 196 / 196 | ±1.9 | raw model + integration |
|
||||||
|
| questions | 97.3% † | 360 / 370 | ±1.7 | raw model + integration |
|
||||||
|
| automations | 16.7% | 10 / 60 | ±9.4 | raw model + integration |
|
||||||
|
|
||||||
|
† Same shared repair layer. Merging the 5 specialists into one costs peak per-task accuracy on raw — clearest on automations (67% → 50%).
|
||||||
|
|
||||||
|
See [Evaluation](#evaluation) for what each surface measures — the LoRA adapter vs the integration vs the raw model.
|
||||||
|
|
||||||
|
## Quick start
|
||||||
|
|
||||||
|
You have a choice in how you start with Selora AI:
|
||||||
|
|
||||||
|
- **Ready to deploy with Home Assistant?** Use [llama-server](#llama-server-home-assistant-integration-runtime) — the runtime the HA integration is built around.
|
||||||
|
- **Want to evaluate the model first?** Use [Ollama](#ollama-evaluate-the-model-before-integrating) — try each specialist on your machine, smoke-test the LoRAs on your hardware, decide if Selora AI is right for you before committing to the full Home Assistant integration. Fastest taste is a single line, no download: `ollama run hf.co/selorahomes/Selora-AI:selora-ollama.Q6_K.gguf`.
|
||||||
|
- **Serving in the cloud?** Use [vLLM](#vllm-cloud).
|
||||||
|
|
||||||
|
### llama-server (Home Assistant integration runtime)
|
||||||
|
|
||||||
|
The reference runtime — what the model was trained against and what the Home Assistant integration uses. `llama-server`'s `/lora-adapters` endpoint is the in-process LoRA hot-swap that lets the integration pick a specialist per turn without reloading the base.
|
||||||
|
|
||||||
|
Download the base and all five LoRA files into a single directory, then:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
llama-server \
|
||||||
|
--model qwen3_17b_base.Q6_K.gguf \
|
||||||
|
--lora-init-without-apply \
|
||||||
|
--lora selora-command.f16.gguf \
|
||||||
|
--lora selora-automation.f16.gguf \
|
||||||
|
--lora selora-answer.f16.gguf \
|
||||||
|
--lora selora-clarification.f16.gguf \
|
||||||
|
--lora selora-utilities.f16.gguf \
|
||||||
|
--ctx-size 8192
|
||||||
|
```
|
||||||
|
|
||||||
|
POST to `/lora-adapters` to switch the active LoRA before each
|
||||||
|
`/v1/chat/completions` call. Build instructions for `llama-server` are in the [llama.cpp build guide](https://github.com/ggerganov/llama.cpp/blob/master/docs/build.md).
|
||||||
|
|
||||||
|
### Ollama (evaluate the model before integrating)
|
||||||
|
|
||||||
|
Ollama lets you try Selora AI on your machine and validate the LoRAs work before setting up the full Home Assistant integration. Useful for kicking the tyres on each specialist, smoke-testing the model on your hardware, or driving it from a script.
|
||||||
|
|
||||||
|
Selora requires **Ollama 0.30 or later** (for LoRA inference) installed locally. Pick whichever fits your machine:
|
||||||
|
|
||||||
|
- **macOS / Linux / Windows:** [official installer](https://ollama.com/download) (single download per platform)
|
||||||
|
- **macOS via [Homebrew](https://brew.sh):** `brew install ollama`
|
||||||
|
- **Linux via shell:** `curl -fsSL https://ollama.com/install.sh | sh`
|
||||||
|
- **Windows via [Winget](https://learn.microsoft.com/windows/package-manager/winget/):** `winget install Ollama.Ollama`
|
||||||
|
|
||||||
|
Download the base, the LoRAs, and the Modelfiles from this repo into one
|
||||||
|
directory, then from that directory:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ollama create selora-qwen-command -f Modelfile.command
|
||||||
|
ollama create selora-qwen-automation -f Modelfile.automation
|
||||||
|
ollama create selora-qwen-answer -f Modelfile.answer
|
||||||
|
ollama create selora-qwen-clarification -f Modelfile.clarification
|
||||||
|
ollama create selora-qwen-utilities -f Modelfile.utilities
|
||||||
|
```
|
||||||
|
|
||||||
|
Each Modelfile pins the per-specialist system prompt and generation parameters,
|
||||||
|
so no extra configuration is needed. The Q6_K base is stored once in Ollama's
|
||||||
|
blob store and shared across all the specialists; only the ~10–40 MB LoRA
|
||||||
|
adapter is added per slot — but `ollama list` will show one named entry per
|
||||||
|
specialist.
|
||||||
|
|
||||||
|
Ollama 0.30+ does **not** support in-process LoRA hot-swap, so each specialist runs as its own named model. This path is best for direct chat or scripting use; for the Home Assistant integration use `llama-server` above.
|
||||||
|
|
||||||
|
#### Single general model — `selora-ollama`
|
||||||
|
|
||||||
|
If you'd rather not manage five separate models, **`selora-ollama`** folds all five specialists into one self-routing model — one model, all five intents, no hot-swap. The merged Q6_K weights (~1.3 GB) ship at the repo root with the unified router system prompt, chat template, and generation parameters baked into the GGUF, so Ollama pulls and runs it straight from Hugging Face — no manual download, no Modelfile step:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ollama run hf.co/selorahomes/Selora-AI:selora-ollama.Q6_K.gguf
|
||||||
|
```
|
||||||
|
|
||||||
|
Prefer to build it locally — to tweak the system prompt or parameters first? Download the [`ollama/`](ollama/) folder and create it from the bundled Modelfile instead:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ollama create selora-ollama -f Modelfile
|
||||||
|
```
|
||||||
|
|
||||||
|
Either path gives you the same model. Per-surface scores for both runtimes are in the [Results](#results) section above.
|
||||||
|
|
||||||
|
### vLLM (cloud)
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python -m vllm.entrypoints.openai.api_server \
|
||||||
|
--model ./qwen3_17b_hf \
|
||||||
|
--enable-lora --max-loras 5 --max-lora-rank 32 \
|
||||||
|
--lora-modules \
|
||||||
|
selora-command=/path/to/peft/command \
|
||||||
|
selora-automation=/path/to/peft/automation \
|
||||||
|
selora-answer=/path/to/peft/answer \
|
||||||
|
selora-clarification=/path/to/peft/clarification \
|
||||||
|
selora-utilities=/path/to/peft/utilities
|
||||||
|
```
|
||||||
|
|
||||||
|
vLLM activates the matching LoRA based on the request's `model` field;
|
||||||
|
no extra routing layer needed.
|
||||||
|
|
||||||
|
## Getting started in Home Assistant
|
||||||
|
|
||||||
|
A walk-through from zero to "Selora AI is answering me in Home Assistant." If you already have HA running and just want to plug in the model, skip to step 4.
|
||||||
|
|
||||||
|
### 1. Create a Selora Homes Connect account
|
||||||
|
|
||||||
|
Sign up at [selorahomes.com/connect](https://selorahomes.com/connect). The account ties your local install to:
|
||||||
|
|
||||||
|
- Cloud-side OAuth flows (needed by integrations that require external authentication — e.g. some appliance providers)
|
||||||
|
- Optional remote-access tunnels so you can reach your home from outside the LAN
|
||||||
|
- Configuration sync between multiple HA installs in the same household
|
||||||
|
|
||||||
|
The local model runs without an account — Connect is for cloud-bridged features and remote access. If you only want offline-only local AI, you can skip this step and revisit later.
|
||||||
|
|
||||||
|
### 2. Set up Home Assistant
|
||||||
|
|
||||||
|
Install HA on a Pi, NUC, NAS, or x86 server using the [official installation guide](https://www.home-assistant.io/installation/). HA OS is the recommended path for new users; Docker is fine for power users.
|
||||||
|
|
||||||
|
Confirm you can reach the HA web UI at `http://homeassistant.local:8123` before continuing.
|
||||||
|
|
||||||
|
### 3. Install the Selora AI integration
|
||||||
|
|
||||||
|
The custom component lives at [`github.com/SeloraHomes/ha-selora-ai`](https://github.com/SeloraHomes/ha-selora-ai). Two install paths:
|
||||||
|
|
||||||
|
**Via HACS (recommended).** HACS — the Home Assistant Community Store — handles updates automatically.
|
||||||
|
|
||||||
|
1. Install HACS itself if you don't have it: [HACS install guide](https://hacs.xyz/docs/setup/download)
|
||||||
|
2. In HA: **HACS → Integrations → ⋮ → Custom repositories**
|
||||||
|
3. Add `https://github.com/SeloraHomes/ha-selora-ai` as type **Integration**
|
||||||
|
4. Search for **Selora AI**, click **Install**, restart Home Assistant
|
||||||
|
|
||||||
|
**Manual install.** Clone directly into HA's custom_components folder:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /config/custom_components
|
||||||
|
git clone https://github.com/SeloraHomes/ha-selora-ai.git selora_ai
|
||||||
|
# Restart Home Assistant
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. Download the model files
|
||||||
|
|
||||||
|
From this HuggingFace repo, get:
|
||||||
|
|
||||||
|
- `qwen3_17b_base.Q6_K.gguf` (the shared base, ~1.6 GB)
|
||||||
|
- `selora-command.f16.gguf`
|
||||||
|
- `selora-automation.f16.gguf`
|
||||||
|
- `selora-answer.f16.gguf`
|
||||||
|
- `selora-clarification.f16.gguf`
|
||||||
|
- `selora-utilities.f16.gguf`
|
||||||
|
- The `Modelfile.*` files (for Ollama users; skip for `llama-server` users)
|
||||||
|
|
||||||
|
Put them all in a single directory on the machine that'll run the model. Many users put this on the same box as HA; others run it on a dedicated GPU machine and point HA at it over the LAN.
|
||||||
|
|
||||||
|
### 5. Run the model locally
|
||||||
|
|
||||||
|
Pick one runtime — both are covered in the **Quick start** section above:
|
||||||
|
|
||||||
|
- **Ollama 0.30+** — simpler if you already use Ollama. One model per specialist; the HA integration treats each as a separate provider.
|
||||||
|
- **`llama-server`** — the reference runtime, full LoRA hot-swap support. Best for the HA integration because it lets the integration pick the right specialist per turn.
|
||||||
|
|
||||||
|
Either way, the model needs to be reachable from wherever HA is running. Confirm with `curl http://<host>:8080/v1/models` (llama-server) or `ollama list` (Ollama).
|
||||||
|
|
||||||
|
### 6. Connect HA to Selora AI Local
|
||||||
|
|
||||||
|
In Home Assistant: **Settings → Devices & Services → Add Integration → Selora AI**. From the provider dropdown, pick **Selora AI Local**.
|
||||||
|
|
||||||
|
The integration auto-discovers a running `llama-server` (or Ollama) on the standard ports. If discovery fails, enter the host manually in the config flow.
|
||||||
|
|
||||||
|
### 7. Verify it works
|
||||||
|
|
||||||
|
Type one of these into the Selora AI chat panel that appears after setup:
|
||||||
|
|
||||||
|
- `turn on the kitchen light` — should flip a light
|
||||||
|
- `what lights are on?` — should list them
|
||||||
|
- `create an automation that turns on the porch light at sunset` — should produce an automation card
|
||||||
|
- `turn on a light` — should ask which one (if you have several)
|
||||||
|
- `why is the living-room sensor unavailable?` — should give docs-grounded troubleshooting steps
|
||||||
|
|
||||||
|
If they all work, you're done. If any fail, see Troubleshooting at the bottom of this page.
|
||||||
|
|
||||||
|
## What's new in v0.4.8
|
||||||
|
|
||||||
|
### New `utilities` specialist (slot 4)
|
||||||
|
|
||||||
|
A fifth specialist handles docs-grounded help — maintenance (pending updates,
|
||||||
|
version conflicts), troubleshooting (why a device is unavailable / offline /
|
||||||
|
not responding), and setup help (how to add or configure an integration). It is
|
||||||
|
given the user question, the `AVAILABLE ENTITIES` list, and a `RELEVANT DOCS`
|
||||||
|
list of retrieved Home Assistant documentation chunks, and returns a compact
|
||||||
|
envelope:
|
||||||
|
|
||||||
|
```
|
||||||
|
{"r":"<advice with {entity_id} placeholders>","q":["<entity_id>",…],"src":["<doc_chunk_id>",…]}
|
||||||
|
```
|
||||||
|
|
||||||
|
`r` is advice grounded in the retrieved docs with `{entity_id}` placeholders for
|
||||||
|
any live-state references; `q` lists the entities whose live state the advice
|
||||||
|
depends on (so the consumer substitutes current values); `src` cites the doc
|
||||||
|
chunks the advice is drawn from. The fix is grounded in the docs, never in entity
|
||||||
|
state alone — the state is the signal, the docs supply the explanation.
|
||||||
|
|
||||||
|
### Slot order
|
||||||
|
|
||||||
|
The specialist slot contract is now `0=command, 1=automation, 2=answer,
|
||||||
|
3=clarification, 4=utilities`. Backends that hot-swap LoRAs (llama-server's
|
||||||
|
`/lora-adapters`, vLLM `--enable-lora`) load all five at startup.
|
||||||
|
|
||||||
|
### Entity-block format reconciled with the integration
|
||||||
|
|
||||||
|
`format_entities_block` in `scripts/gen_utils.py` emits the exact per-line shape produced by `_format_entity_line` in `custom_components/selora_ai/llm_client/sanitize.py`:
|
||||||
|
|
||||||
|
```
|
||||||
|
AVAILABLE ENTITIES:
|
||||||
|
- entity_id=light.kitchen; state=off; friendly_name=Kitchen Lights
|
||||||
|
- entity_id=sensor.sun; state=below_horizon; friendly_name=Sun
|
||||||
|
```
|
||||||
|
|
||||||
|
This keeps train-vs-inference entity-context blocks in lock-step so the model stays in-distribution.
|
||||||
|
|
||||||
|
### Inference
|
||||||
|
|
||||||
|
The manifest carries `runtime.cache_prompt = true` so the hub starts
|
||||||
|
`llama-server` with system-prompt KV caching enabled, amortizing the
|
||||||
|
per-specialist system prompt across requests.
|
||||||
|
|
||||||
|
## Generation parameters
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"temperature": 0.0,
|
||||||
|
"repeat_penalty": 1.0,
|
||||||
|
"repeat_last_n": 256,
|
||||||
|
"max_tokens": 384,
|
||||||
|
"stop": ["<|im_end|>", "<|endoftext|>"]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Bump `max_tokens` to 1536 for automation requests (longer JSON output).
|
||||||
|
|
||||||
|
## Training
|
||||||
|
|
||||||
|
Base: [Qwen3 1.7B](https://huggingface.co/Qwen/Qwen3-1.7B) fine-tuned
|
||||||
|
with [Apple mlx-lm](https://github.com/ml-explore/mlx-examples). Each
|
||||||
|
specialist has its own LoRA (rank 8–24, scale 20) trained on a curated
|
||||||
|
HA-domain corpus (forum threads, HA docs, synthetic command /
|
||||||
|
automation pairs). System prompts trained per-specialist; see
|
||||||
|
[`prompts/`](prompts/). The `answer` adapter is trained to emit the slim
|
||||||
|
`{r,q}` shape — a `q` list of entities to look up alongside its
|
||||||
|
`{entity_id}`-templated response — which the integration resolves against
|
||||||
|
live state; see `prompts/answer_system_prompt.txt` and the
|
||||||
|
`Modelfile.answer` SYSTEM block. The `utilities` adapter is trained on
|
||||||
|
docs-grounded maintenance / troubleshooting / setup examples with
|
||||||
|
retrieved-doc context and citation provenance.
|
||||||
|
|
||||||
|
### Training examples
|
||||||
|
|
||||||
|
Every specialist trains on the same wrapper — its system prompt, then a user turn carrying
|
||||||
|
`USER REQUEST` and the `AVAILABLE ENTITIES` block (plus `RELEVANT DOCS` for `utilities`) —
|
||||||
|
with the target slim envelope as the completion. One real pair per specialist (entity / doc
|
||||||
|
lists trimmed; requests and target envelopes verbatim):
|
||||||
|
|
||||||
|
```text
|
||||||
|
# command
|
||||||
|
USER REQUEST: resume the living room TV
|
||||||
|
AVAILABLE ENTITIES:
|
||||||
|
- entity_id=media_player.living_room_tv; state=paused; friendly_name="living room TV"
|
||||||
|
→ {"c":[{"s":"media_player.media_play","e":"media_player.living_room_tv"}],"r":"Playing on living room TV."}
|
||||||
|
|
||||||
|
# answer
|
||||||
|
USER REQUEST: are hallway light and coffee maker on
|
||||||
|
AVAILABLE ENTITIES:
|
||||||
|
- entity_id=light.hallway; state=off; friendly_name="hallway light"
|
||||||
|
- entity_id=switch.coffee_maker; state=on; friendly_name="coffee maker"
|
||||||
|
→ {"q":["light.hallway","switch.coffee_maker"],"r":"Hallway light: {light.hallway}, coffee maker: {switch.coffee_maker}."}
|
||||||
|
|
||||||
|
# clarification
|
||||||
|
USER REQUEST: can you do something
|
||||||
|
AVAILABLE ENTITIES:
|
||||||
|
- entity_id=light.entry; state=off; friendly_name="entry light"
|
||||||
|
- entity_id=media_player.living_room_tv; state=idle; friendly_name="living room TV"
|
||||||
|
→ {"q":"Could you tell me what you'd like to do?"}
|
||||||
|
|
||||||
|
# automation
|
||||||
|
USER REQUEST: activate movie night every day at 7pm
|
||||||
|
AVAILABLE ENTITIES:
|
||||||
|
- entity_id=scene.movie_night; friendly_name="movie night"
|
||||||
|
→ {"intent":"automation","response":"At **7pm** every day, this activates **movie night**. Want me to skip weekends?","description":"Activate scene scene.movie_night at 7pm.","automation":{"alias":"Auto-Movie Night","description":"Activate movie night at 7pm daily.","triggers":[{"trigger":"time","at":"19:00:00"}],"conditions":[],"actions":[{"service":"scene.turn_on","target":{"entity_id":"scene.movie_night"},"data":{}}]}}
|
||||||
|
|
||||||
|
# utilities
|
||||||
|
USER REQUEST: guide me through adding HACS
|
||||||
|
RELEVANT DOCS:
|
||||||
|
- [hacs] Setting up the HACS integration (https://hacs.xyz/docs/use/configuration/basic/) — "Follow these steps to set up the HACS integration and authenticate it with GitHub…"
|
||||||
|
AVAILABLE ENTITIES:
|
||||||
|
- entity_id=light.entry; state=off; friendly_name="entry light"
|
||||||
|
→ {"r":"To add HACS, go to Settings > Devices & Services > Add Integration and search for it, then follow the config flow. The integration's documentation page covers the prerequisites and step-by-step setup.","src":["hacs:https://hacs.xyz/docs/use/configuration/basic/#2","hacs:https://hacs.xyz/docs/use/configuration/options/#2"]}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Evaluation
|
||||||
|
|
||||||
|
Selora AI is evaluated on the open
|
||||||
|
[allenporter/home-assistant-datasets](https://github.com/allenporter/home-assistant-datasets)
|
||||||
|
suite — the harness behind the Home Assistant LLM leaderboard, so its results are
|
||||||
|
comparable to other models' conversation agents. The headline numbers are in
|
||||||
|
[Results](#results) at the top of the card; what matters here is **what each surface
|
||||||
|
actually grades** — each scores a different slice of the stack:
|
||||||
|
|
||||||
|
- **The LoRA adapter** — each specialist is a small LoRA **trained on the slim-envelope
|
||||||
|
schema**: one strict, compact output contract per intent (the service calls to run, the
|
||||||
|
entities to read, the question to ask, or the blueprint to build), designed up front and
|
||||||
|
baked into every training example. Learning that single fixed shape per intent is what
|
||||||
|
lets the adapter emit it reliably — and a malformed envelope is a miss on every surface,
|
||||||
|
so this training is what every number ultimately grades.
|
||||||
|
- **The integration** is in the loop for **assist**, **assist-mini**, and **questions**.
|
||||||
|
The adapter emits the envelope; `_convert_slim_shape` runs the `command` calls as real
|
||||||
|
Home Assistant service calls and resolves the `answer` entity list into live-state
|
||||||
|
answers; the metric then grades the resulting device state / produced answer. So those
|
||||||
|
three are **end-to-end** scores — the adapter **plus** the integration's parser and
|
||||||
|
routing, exactly how a user experiences it.
|
||||||
|
- **The raw model** is graded directly for **automations**. The `automation` adapter's
|
||||||
|
blueprint YAML *is* the scored artifact, so the integration is bypassed — routing it
|
||||||
|
through would wrap the YAML in a prose summary, the wrong surface for authoring. The
|
||||||
|
66.7% is the bare adapter; the open misses are `light_on_door`, where it emits a
|
||||||
|
`2 * 60` duration the Home Assistant loader rejects.
|
||||||
|
|
||||||
|
## Files in this bundle
|
||||||
|
|
||||||
|
| Artifact | Purpose | Distribution |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| `qwen3_17b_base.Q6_K.gguf` | Q6_K base for Ollama / llama.cpp | Hugging Face, ollama.com |
|
||||||
|
| `selora-{intent}.f16.gguf` (×5) | Specialist LoRA adapters | Hugging Face, ollama.com |
|
||||||
|
| `Modelfile.{intent}` | Ollama recipes (base + LoRA + system prompt) | this repo, ollama.com |
|
||||||
|
| `prompts/{intent}_system_prompt.txt` | Plain-text trained prompts (reference / testing) | this repo |
|
||||||
|
| `selora-ollama.Q6_K.gguf` | Single general model at repo root — all 5 intents merged, Q6_K (~1.3 GB), system prompt / template / params baked in for `ollama run hf.co/selorahomes/Selora-AI:selora-ollama.Q6_K.gguf` | Hugging Face |
|
||||||
|
| `ollama/selora-ollama.gguf` + `ollama/Modelfile` | Same general model as a local `ollama create` bundle (Modelfile pins the router prompt + params) | Hugging Face |
|
||||||
|
|
||||||
|
The full-precision (f16) base and HF safetensors set used by vLLM /
|
||||||
|
TGI / SageMaker live separately in the cloud bundle and are not yet
|
||||||
|
mirrored to Hugging Face.
|
||||||
|
|
||||||
|
## First-run verification
|
||||||
|
|
||||||
|
Five prompts — one per specialist — let you confirm every slot loaded cleanly. Type them into HA's Selora AI panel (or hit the `selora_ai/chat_stream` WebSocket directly):
|
||||||
|
|
||||||
|
| Prompt | Specialist | Expected behaviour |
|
||||||
|
|---|---|---|
|
||||||
|
| `turn on the kitchen light` | command | Light flips on; response: `"Kitchen light on."` |
|
||||||
|
| `what lights are on?` | answer | List of currently-on lights with `[[entities:...]]` markers |
|
||||||
|
| `create an automation that turns on the porch light at sunset` | automation | Automation card with `trigger: sun, event: sunset` and the porch light target |
|
||||||
|
| `turn on a light` (with multiple lights present) | clarification | Asks which one and offers options |
|
||||||
|
| `why is the living-room sensor unavailable?` | utilities | Docs-grounded troubleshooting steps citing source docs |
|
||||||
|
|
||||||
|
A clean run on all five = LoRAs loaded, classifier routing correctly, and the v0.4.8 training format reaching the model. If any prompt returns garbage or empty output, check Troubleshooting below.
|
||||||
|
|
||||||
|
## Troubleshooting
|
||||||
|
|
||||||
|
| Symptom | Likely cause | Fix |
|
||||||
|
|---|---|---|
|
||||||
|
| `Selora AI Local` not in provider dropdown | Probe couldn't reach any host candidate | Verify `curl http://localhost:8080/v1/models` works on the HA host. Add the host manually in config flow if HA can't reach localhost (common on HA OS) |
|
||||||
|
| Chat returns empty / repeats one token | `repeat_penalty != 1.0` somewhere | Confirm llama-server is started without an override, or that the Modelfile's `PARAMETER repeat_penalty 1.0` line wasn't edited out |
|
||||||
|
| Wrong specialist responds (e.g. answer for a command) | Hot-swap call hasn't fired | Check HA logs for `Activating LoRA slot N`; if absent, the integration didn't classify the prompt as that intent — file an issue with the prompt text |
|
||||||
|
| Model invents entity_ids that don't exist | AVAILABLE ENTITIES block not being sent | The integration sends this automatically; if you're hitting the model directly, mirror the integration's `_format_entity_line` output exactly (see "Entity-block format reconciled with the integration" above) |
|
||||||
|
| `ollama run` works but HA can't reach it | Ollama default `localhost:11434`, llama-server `0.0.0.0:8080` — different ports | Either point the integration at port `11434` (Ollama path) or run llama-server explicitly. The integration probes `:8080` first |
|
||||||
|
| Pipeline hangs for 30s on automation prompts | Pre-v0.4.7 build of the integration | Update the integration to current `main` |
|
||||||
|
|
||||||
|
For deeper issues, the integration's debug log (`logger: custom_components.selora_ai: debug` in `configuration.yaml`) prints the full classifier decision, the request payload sent to `llama-server`, and the raw model response — enough to diagnose any reproducible case.
|
||||||
|
|
||||||
|
## Citation
|
||||||
|
|
||||||
|
```bibtex
|
||||||
|
@misc{selora-ai-2026,
|
||||||
|
title = {Selora AI: Qwen3 1.7B + LoRA Specialists for Home Assistant},
|
||||||
|
author = {{Selora Homes}},
|
||||||
|
year = {2026},
|
||||||
|
url = {https://huggingface.co/selorahomes/Selora-AI}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## License
|
||||||
|
|
||||||
|
Apache-2.0
|
||||||
105
manifest.json
Normal file
105
manifest.json
Normal file
@@ -0,0 +1,105 @@
|
|||||||
|
{
|
||||||
|
"name": "selora-ai-local",
|
||||||
|
"version": "0.4.8",
|
||||||
|
"description": "Selora AI v0.4.8 \u2014 Qwen3-1.7B Q6_K base + 5 LoRA specialists with slim action-then-confirm output schemas. Adds a utilities specialist (slot 4) alongside command, automation, answer, and clarification. Multi-state entity context (per-entity attribute tails in AVAILABLE ENTITIES) for richer single-turn grounding. Inference: cache_prompt enabled to amortize system-prompt KV cache across requests.",
|
||||||
|
"base_model": {
|
||||||
|
"id": "Qwen/Qwen3-1.7B",
|
||||||
|
"format": "gguf",
|
||||||
|
"dtype": "Q6_K",
|
||||||
|
"filename": "qwen3_17b_base.Q6_K.gguf",
|
||||||
|
"size_bytes": 1673006880,
|
||||||
|
"sha256": "a00bbdb411872149d73e1a0683b9b8a9f13cf74f98ba70ff8e8e430d9a093179"
|
||||||
|
},
|
||||||
|
"loras": [
|
||||||
|
{
|
||||||
|
"slot": 0,
|
||||||
|
"name": "command",
|
||||||
|
"filename": "selora-command.f16.gguf",
|
||||||
|
"size_bytes": 19938528,
|
||||||
|
"sha256": "de8871f56d04cdfb0caf634e843f5d1ce9d7eead3b76ccd68b9813499035eba1"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"slot": 1,
|
||||||
|
"name": "automation",
|
||||||
|
"filename": "selora-automation.f16.gguf",
|
||||||
|
"size_bytes": 37374880,
|
||||||
|
"sha256": "e6d8b3b9cd7dc05b3de017cb43915586f15bca797985b50eb112ba585e2e25c9"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"slot": 2,
|
||||||
|
"name": "answer",
|
||||||
|
"filename": "selora-answer.f16.gguf",
|
||||||
|
"size_bytes": 14957792,
|
||||||
|
"sha256": "ab3342bb35c1c97d995121b3548eb8c8406a975d293c9f5ea4cdc5974fdd16a4"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"slot": 3,
|
||||||
|
"name": "clarification",
|
||||||
|
"filename": "selora-clarification.f16.gguf",
|
||||||
|
"size_bytes": 9977056,
|
||||||
|
"sha256": "16d2a2f852ca6b4e73f03caa235a52dcd49804c8a3acae5420914dc7a878d610"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"slot": 4,
|
||||||
|
"name": "utilities",
|
||||||
|
"filename": "selora-utilities.f16.gguf",
|
||||||
|
"size_bytes": 19938528,
|
||||||
|
"sha256": "d1e47028cbc3ad81252853bc363921764456d50a1329d41e4ded3f2b7a9b3af9"
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"system_prompts": {
|
||||||
|
"command": {
|
||||||
|
"filename": "command_system_prompt.txt",
|
||||||
|
"size_bytes": 1071,
|
||||||
|
"sha256": "9921c6fef09c6ebad4a2ed4fad1dbe7e76efe0bfe4e532bf7c7fe096864de6a4"
|
||||||
|
},
|
||||||
|
"automation": {
|
||||||
|
"filename": "automation_system_prompt.txt",
|
||||||
|
"size_bytes": 2711,
|
||||||
|
"sha256": "04e2d8231e91d00afba0c50964a0647b036303bd4c4dbc9e7c58dd3434d2b682"
|
||||||
|
},
|
||||||
|
"answer": {
|
||||||
|
"filename": "answer_system_prompt.txt",
|
||||||
|
"size_bytes": 856,
|
||||||
|
"sha256": "ec4c2dfb6bcd378e65f891a15d9066d9f0c295a1ac3fc9dbc01cb01a9c0d6cb2"
|
||||||
|
},
|
||||||
|
"clarification": {
|
||||||
|
"filename": "clarification_system_prompt.txt",
|
||||||
|
"size_bytes": 683,
|
||||||
|
"sha256": "c6833a17147574946a7176447a88d65e687bc393e62db1aaa89c57d1fdf9a3ac"
|
||||||
|
},
|
||||||
|
"utilities": {
|
||||||
|
"filename": "utilities_system_prompt.txt",
|
||||||
|
"size_bytes": 1588,
|
||||||
|
"sha256": "ae1155b644b529ba63d9441b2abc347fd6f4e4b3d4bbb25b323509707df90d36"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"runtime": {
|
||||||
|
"cache_prompt": true,
|
||||||
|
"ctx_size": 4096
|
||||||
|
},
|
||||||
|
"training": {
|
||||||
|
"framework": "mlx-lm",
|
||||||
|
"base_model_repo": "Qwen/Qwen3-1.7B",
|
||||||
|
"optimizer": "adam",
|
||||||
|
"learning_rate": 0.0001,
|
||||||
|
"batch_size": 4,
|
||||||
|
"max_seq_length": 4096,
|
||||||
|
"english_only": true,
|
||||||
|
"data_source": "synthetic \u2014 slim schemas in slim_schemas.md, generated by scripts/gen_{intent}.py from 10 curated home specs + procedural variants; service_matrix.py covers 49 (domain, service) pairs. tools.home_specs.diversify_states() injects multi-state attributes per training example.",
|
||||||
|
"iterations_per_specialist": {
|
||||||
|
"command": 750,
|
||||||
|
"answer": 600,
|
||||||
|
"clarification": 450,
|
||||||
|
"automation": 1050,
|
||||||
|
"utilities": 600
|
||||||
|
},
|
||||||
|
"examples_per_specialist": {
|
||||||
|
"command": 8800,
|
||||||
|
"answer": 6600,
|
||||||
|
"clarification": 3300,
|
||||||
|
"automation": 6600,
|
||||||
|
"utilities": 6600
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
34
ollama/Modelfile
Normal file
34
ollama/Modelfile
Normal file
@@ -0,0 +1,34 @@
|
|||||||
|
FROM ./selora-ollama.gguf
|
||||||
|
TEMPLATE """{{ if .System }}<|im_start|>system
|
||||||
|
{{ .System }}<|im_end|>
|
||||||
|
{{ end }}{{ if .Prompt }}<|im_start|>user
|
||||||
|
/no_think {{ .Prompt }}<|im_end|>
|
||||||
|
{{ end }}<|im_start|>assistant
|
||||||
|
"""
|
||||||
|
SYSTEM """You are Selora AI for Home Assistant. You are given a USER REQUEST, the AVAILABLE ENTITIES list, EXISTING AUTOMATIONS, and (when relevant) RELEVANT DOCS. Decide which ONE of five response types the request needs, then reply with ONLY that type's JSON object — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
|
||||||
|
ROUTING — decide act-vs-ask first:
|
||||||
|
- PREFER TO ACT. If a target entity and an action are identifiable from AVAILABLE ENTITIES, emit a command — even when the request carries a number (brightness %, temperature, etc.). A numeric value is NOT a reason to clarify.
|
||||||
|
- Only CLARIFY when there is no actionable target or verb ("help", "do something", "set the temperature" with no device named) OR several entities are equally valid with no sensible default.
|
||||||
|
- UTILITIES is ONLY docs-grounded help about Home Assistant itself: a pending update, why a device is unavailable/offline, or how to set up/configure an integration. A bare "help" with no HA topic is CLARIFICATION, never utilities. Never invent state or entities.
|
||||||
|
|
||||||
|
THE FIVE TYPES:
|
||||||
|
1) command — control a device now (turn on/off, set, lock/unlock, open/close, play/pause, dim).
|
||||||
|
{"c":[{"s":"<domain.service>","e":"<entity_id>","d":{<params>}}],"r":"<one short past-tense confirmation>"}
|
||||||
|
One c entry per (service, entity_id); s is "domain.action" and its domain must match e's domain; omit d when there are no params.
|
||||||
|
2) automation — save a recurring rule, schedule, or multi-step sequence (cues: every, when, at <time>, if, whenever).
|
||||||
|
{"intent":"automation","response":"<short reply>","description":"<one line>","automation":{"alias":...,"triggers":[...],"conditions":[...],"actions":[...]}}
|
||||||
|
3) answer — a question about the home's current state.
|
||||||
|
{"r":"<answer, using {entity_id} placeholders wherever live state is needed>","q":["<entity_id>",...]}
|
||||||
|
4) clarification — the request is too vague to act on.
|
||||||
|
{"q":"<one short clarifying question>","o":["<option>",...]}
|
||||||
|
Here q is a QUESTION STRING (not an entity list); o offers optional quick-reply choices.
|
||||||
|
5) utilities — docs-grounded Home Assistant help (see ROUTING).
|
||||||
|
{"r":"<advice, with {entity_id} placeholders for live state>","q":["<entity_id>",...],"src":["<doc_chunk_id>",...]}
|
||||||
|
|
||||||
|
Use canonical entity_ids from AVAILABLE ENTITIES, never the human alias. Reference only entity_ids that appear there. Output exactly ONE JSON object of ONE type. JSON only."""
|
||||||
|
PARAMETER temperature 0.0
|
||||||
|
PARAMETER repeat_penalty 1.0
|
||||||
|
PARAMETER repeat_last_n 256
|
||||||
|
PARAMETER stop "<|im_end|>"
|
||||||
|
PARAMETER stop "<|endoftext|>"
|
||||||
20
ollama/README.md
Normal file
20
ollama/README.md
Normal file
@@ -0,0 +1,20 @@
|
|||||||
|
# Selora AI — Ollama (general self-routing model)
|
||||||
|
|
||||||
|
A single Qwen3-1.7B model that self-routes each Home Assistant request into command / automation / answer / clarification / utilities and emits one compact JSON "slim envelope" — one general model, no per-specialist hot-swapping. Q6_K, ~1.3 GB.
|
||||||
|
|
||||||
|
## Use with Ollama
|
||||||
|
|
||||||
|
Fastest path — pull and run straight from Hugging Face, no download or Modelfile step (the copy at the repo root has the system prompt, template, and params baked into the GGUF):
|
||||||
|
|
||||||
|
```
|
||||||
|
ollama run hf.co/selorahomes/Selora-AI:selora-ollama.Q6_K.gguf
|
||||||
|
```
|
||||||
|
|
||||||
|
Or build it locally from this folder — download the folder, then from it:
|
||||||
|
|
||||||
|
```
|
||||||
|
ollama create selora-ollama -f Modelfile
|
||||||
|
ollama run selora-ollama
|
||||||
|
```
|
||||||
|
|
||||||
|
The Modelfile bundles the unified router system prompt and chat template (`/no_think`, temperature 0). Run it standalone to see the raw envelope; pair with the [Selora AI integration](https://github.com/SeloraHomes/ha-selora-ai) to execute it against your home.
|
||||||
3
ollama/selora-ollama.gguf
Normal file
3
ollama/selora-ollama.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:226a85efd16ff4237277552611aa9e8f55e51ea34f4016c3bc556d14b8204288
|
||||||
|
size 1417754752
|
||||||
22
ollama/system_prompt.txt
Normal file
22
ollama/system_prompt.txt
Normal file
@@ -0,0 +1,22 @@
|
|||||||
|
You are Selora AI for Home Assistant. You are given a USER REQUEST, the AVAILABLE ENTITIES list, EXISTING AUTOMATIONS, and (when relevant) RELEVANT DOCS. Decide which ONE of five response types the request needs, then reply with ONLY that type's JSON object — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
|
||||||
|
ROUTING — decide act-vs-ask first:
|
||||||
|
- PREFER TO ACT. If a target entity and an action are identifiable from AVAILABLE ENTITIES, emit a command — even when the request carries a number (brightness %, temperature, etc.). A numeric value is NOT a reason to clarify.
|
||||||
|
- Only CLARIFY when there is no actionable target or verb ("help", "do something", "set the temperature" with no device named) OR several entities are equally valid with no sensible default.
|
||||||
|
- UTILITIES is ONLY docs-grounded help about Home Assistant itself: a pending update, why a device is unavailable/offline, or how to set up/configure an integration. A bare "help" with no HA topic is CLARIFICATION, never utilities. Never invent state or entities.
|
||||||
|
|
||||||
|
THE FIVE TYPES:
|
||||||
|
1) command — control a device now (turn on/off, set, lock/unlock, open/close, play/pause, dim).
|
||||||
|
{"c":[{"s":"<domain.service>","e":"<entity_id>","d":{<params>}}],"r":"<one short past-tense confirmation>"}
|
||||||
|
One c entry per (service, entity_id); s is "domain.action" and its domain must match e's domain; omit d when there are no params.
|
||||||
|
2) automation — save a recurring rule, schedule, or multi-step sequence (cues: every, when, at <time>, if, whenever).
|
||||||
|
{"intent":"automation","response":"<short reply>","description":"<one line>","automation":{"alias":...,"triggers":[...],"conditions":[...],"actions":[...]}}
|
||||||
|
3) answer — a question about the home's current state.
|
||||||
|
{"r":"<answer, using {entity_id} placeholders wherever live state is needed>","q":["<entity_id>",...]}
|
||||||
|
4) clarification — the request is too vague to act on.
|
||||||
|
{"q":"<one short clarifying question>","o":["<option>",...]}
|
||||||
|
Here q is a QUESTION STRING (not an entity list); o offers optional quick-reply choices.
|
||||||
|
5) utilities — docs-grounded Home Assistant help (see ROUTING).
|
||||||
|
{"r":"<advice, with {entity_id} placeholders for live state>","q":["<entity_id>",...],"src":["<doc_chunk_id>",...]}
|
||||||
|
|
||||||
|
Use canonical entity_ids from AVAILABLE ENTITIES, never the human alias. Reference only entity_ids that appear there. Output exactly ONE JSON object of ONE type. JSON only.
|
||||||
9
params
Normal file
9
params
Normal file
@@ -0,0 +1,9 @@
|
|||||||
|
{
|
||||||
|
"temperature": 0.0,
|
||||||
|
"repeat_penalty": 1.0,
|
||||||
|
"repeat_last_n": 256,
|
||||||
|
"stop": [
|
||||||
|
"<|im_end|>",
|
||||||
|
"<|endoftext|>"
|
||||||
|
]
|
||||||
|
}
|
||||||
14
prompts/answer_system_prompt.txt
Normal file
14
prompts/answer_system_prompt.txt
Normal file
@@ -0,0 +1,14 @@
|
|||||||
|
You are Selora AI's answer specialist for Home Assistant.
|
||||||
|
|
||||||
|
Given a user question and the AVAILABLE ENTITIES list, respond with ONE JSON object only:
|
||||||
|
{"r":"<response with {entity_id} placeholders where state is needed>","q":["<entity_id>",...]}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- r: response template. Use {entity_id} placeholders for any state references; the consumer substitutes live state. Keep r short — 1-2 sentences max.
|
||||||
|
- q: array of entity_ids to look up. Omit when no live state is needed.
|
||||||
|
- Either field can be omitted if not used, but never both.
|
||||||
|
- Only reference entity_ids that appear in AVAILABLE ENTITIES below.
|
||||||
|
- Never invent state values; always template them via {entity_id}.
|
||||||
|
- If the question is outside the home's scope, return {"r":"I can only answer questions about your home."}.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
27
prompts/automation_system_prompt.txt
Normal file
27
prompts/automation_system_prompt.txt
Normal file
@@ -0,0 +1,27 @@
|
|||||||
|
You are Selora AI, an automation architect for Home Assistant. The user wants a recurring rule, schedule, or multi-step sequence saved as an automation, OR a reusable parameterized blueprint.
|
||||||
|
|
||||||
|
Choose the output format from the request:
|
||||||
|
|
||||||
|
A) BLUEPRINT request — when the user asks to "create a blueprint automation", gives a "## Detailed Description" with an input table, or otherwise wants a REUSABLE/PARAMETERIZED automation with named inputs. Respond with markdown containing a single yaml block. The response MUST start with ```yaml and end with ``` because it is parsed by code. Inside, emit a Home Assistant automation BLUEPRINT:
|
||||||
|
- top-level `blueprint:` with `name:`, `description:`, `domain: automation`, and `input:` containing EVERY input named in the request (use the exact input keys from the request's input table).
|
||||||
|
- each input has a `name:` and a `selector:` (entity/number/duration/media/etc.); add `default:` for timeouts/durations/levels.
|
||||||
|
- wire inputs into triggers/conditions/actions with `!input <input_name>` references — never hardcode entity_ids in a blueprint.
|
||||||
|
- use `mode: single` and `max_exceeded: silent`. States are quoted strings. Times/durations are "HH:MM:SS".
|
||||||
|
- Output ONLY the ```yaml block (inline comments allowed).
|
||||||
|
|
||||||
|
B) CONCRETE request — a one-off command/schedule grounded in this user's AVAILABLE ENTITIES. Return ONE JSON object:
|
||||||
|
{"intent":"automation","response":"<1-2 sentence explanation>","description":"<2-3 sentences>","automation":{"alias":"<max 4 words>","description":"<...>","triggers":[<one-or-more>],"conditions":[<optional>],"actions":[<one-or-more>]}}
|
||||||
|
Or, if no entity matches, the clarification shape:
|
||||||
|
{"intent":"clarification","response":"<ONE specific follow-up question naming candidate devices from AVAILABLE ENTITIES>"}
|
||||||
|
|
||||||
|
BLUEPRINT RULES (format A):
|
||||||
|
- The ```yaml fence is mandatory and literal — lowercase `yaml`, not `yml`/`YAML`/bare ```.
|
||||||
|
- `blueprint.domain` is ALWAYS `automation`.
|
||||||
|
- Triggers inside a blueprint use `platform:` (e.g. `- platform: state`, `- platform: numeric_state`, `- platform: time`, `- platform: sun`).
|
||||||
|
- `!input` may reference a scalar (`entity_id: !input door_sensor`), a whole `target:` (`target: !input light_switch`), or a whole `data:` (`data: !input alert_media`).
|
||||||
|
- Do NOT add a required input the caller will not supply; any input beyond the requested set MUST have a `default:`.
|
||||||
|
|
||||||
|
CONCRETE RULES (format B):
|
||||||
|
- Use HA 2024+ plural keys: 'triggers', 'actions', 'conditions'. Service calls use the 'service' key.
|
||||||
|
- State 'to'/'from' MUST be strings. Times "HH:MM:SS". Durations "HH:MM:SS" or {"hours":N,...}.
|
||||||
|
- EVERY entity_id MUST appear VERBATIM in AVAILABLE ENTITIES; never invent placeholder names.
|
||||||
13
prompts/clarification_system_prompt.txt
Normal file
13
prompts/clarification_system_prompt.txt
Normal file
@@ -0,0 +1,13 @@
|
|||||||
|
You are Selora AI's clarification specialist for Home Assistant.
|
||||||
|
|
||||||
|
When the user's request is ambiguous, respond with ONE JSON object only:
|
||||||
|
{"q":"<question text>","o":["<option1>","<option2>",...]}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- q: short, specific clarifying question. 1 sentence max.
|
||||||
|
- o: optional array of suggested answers. Omit the o key when free-form input is appropriate.
|
||||||
|
- Reference entity aliases from AVAILABLE ENTITIES when the ambiguity is about which entity.
|
||||||
|
- Don't ask multiple questions in one turn — pick the single most important blocker.
|
||||||
|
- Don't restate the user's full request; ask the one thing you need.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
15
prompts/command_system_prompt.txt
Normal file
15
prompts/command_system_prompt.txt
Normal file
@@ -0,0 +1,15 @@
|
|||||||
|
You are Selora AI's command specialist for Home Assistant.
|
||||||
|
|
||||||
|
Given a user command and the AVAILABLE ENTITIES list, respond with ONE JSON object only:
|
||||||
|
{"c":[{"s":"<service>","e":"<entity_id>","d":{<optional params>}}],"r":"<short confirmation>"}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- c: ordered array of one or more service calls. Calls execute in array order.
|
||||||
|
- s: HA service in "domain.action" form (e.g. "light.turn_on", "lock.lock", "media_player.play_media", "scene.turn_on").
|
||||||
|
- e: canonical entity_id from AVAILABLE ENTITIES. Never use the human alias — always the entity_id.
|
||||||
|
- d: service parameters object. Omit the d key entirely when there are no params (do not include "d":{}).
|
||||||
|
- r: ≤ 1 sentence past-tense confirmation describing what got done (e.g. "Kitchen light on.").
|
||||||
|
- The service domain (before the dot) must match the entity_id's domain. light.turn_on goes with light.* entities, lock.lock goes with lock.* entities, etc.
|
||||||
|
- For multi-target requests, produce one c entry per (service, entity_id) pair.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
16
prompts/utilities_system_prompt.txt
Normal file
16
prompts/utilities_system_prompt.txt
Normal file
@@ -0,0 +1,16 @@
|
|||||||
|
You are Selora AI's utilities specialist for Home Assistant.
|
||||||
|
|
||||||
|
You handle docs-grounded help: maintenance (pending updates, version conflicts), troubleshooting (why a device is unavailable / offline / not responding), and setup help (how to add or configure an integration). You are given the user question, the AVAILABLE ENTITIES list, and a RELEVANT DOCS list of retrieved Home Assistant documentation chunks.
|
||||||
|
|
||||||
|
Respond with ONE JSON object only:
|
||||||
|
{"r":"<advice with {entity_id} placeholders where live state is needed>","q":["<entity_id>",...],"src":["<doc_chunk_id>",...]}
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
- r: advice prose grounded in RELEVANT DOCS. Use {entity_id} placeholders for any state references (e.g. which integration has an update, which device is unavailable); the consumer substitutes live state. Keep it focused — a short paragraph.
|
||||||
|
- q: array of entity_ids whose live state the advice depends on (update.* entities, problem/connectivity binary_sensors, the unavailable entity). Omit or use [] when no live state is needed (e.g. setup help for a device that does not exist yet).
|
||||||
|
- src: array of doc chunk ids from RELEVANT DOCS that the advice is drawn from. This is the citation provenance — only cite chunks that appear in RELEVANT DOCS.
|
||||||
|
- Only reference entity_ids that appear in AVAILABLE ENTITIES below.
|
||||||
|
- Ground the fix/steps in the retrieved docs, never in the entity state alone — the state is the signal, the docs supply the explanation.
|
||||||
|
- Never invent state values; always template them via {entity_id}.
|
||||||
|
|
||||||
|
Output JSON only — no narration, no markdown fences, no chain-of-thought.
|
||||||
3
qwen3_17b_base.Q6_K.gguf
Normal file
3
qwen3_17b_base.Q6_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:a00bbdb411872149d73e1a0683b9b8a9f13cf74f98ba70ff8e8e430d9a093179
|
||||||
|
size 1673006880
|
||||||
3
selora-answer.f16.gguf
Normal file
3
selora-answer.f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:ab3342bb35c1c97d995121b3548eb8c8406a975d293c9f5ea4cdc5974fdd16a4
|
||||||
|
size 14957792
|
||||||
3
selora-automation.f16.gguf
Normal file
3
selora-automation.f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:e6d8b3b9cd7dc05b3de017cb43915586f15bca797985b50eb112ba585e2e25c9
|
||||||
|
size 37374880
|
||||||
3
selora-clarification.f16.gguf
Normal file
3
selora-clarification.f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:16d2a2f852ca6b4e73f03caa235a52dcd49804c8a3acae5420914dc7a878d610
|
||||||
|
size 9977056
|
||||||
3
selora-command.f16.gguf
Normal file
3
selora-command.f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:de8871f56d04cdfb0caf634e843f5d1ce9d7eead3b76ccd68b9813499035eba1
|
||||||
|
size 19938528
|
||||||
3
selora-ollama.Q6_K.gguf
Normal file
3
selora-ollama.Q6_K.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:226a85efd16ff4237277552611aa9e8f55e51ea34f4016c3bc556d14b8204288
|
||||||
|
size 1417754752
|
||||||
3
selora-utilities.f16.gguf
Normal file
3
selora-utilities.f16.gguf
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
version https://git-lfs.github.com/spec/v1
|
||||||
|
oid sha256:d1e47028cbc3ad81252853bc363921764456d50a1329d41e4ded3f2b7a9b3af9
|
||||||
|
size 19938528
|
||||||
22
system
Normal file
22
system
Normal file
@@ -0,0 +1,22 @@
|
|||||||
|
You are Selora AI for Home Assistant. You are given a USER REQUEST, the AVAILABLE ENTITIES list, EXISTING AUTOMATIONS, and (when relevant) RELEVANT DOCS. Decide which ONE of five response types the request needs, then reply with ONLY that type's JSON object — no narration, no markdown fences, no chain-of-thought.
|
||||||
|
|
||||||
|
ROUTING — decide act-vs-ask first:
|
||||||
|
- PREFER TO ACT. If a target entity and an action are identifiable from AVAILABLE ENTITIES, emit a command — even when the request carries a number (brightness %, temperature, etc.). A numeric value is NOT a reason to clarify.
|
||||||
|
- Only CLARIFY when there is no actionable target or verb ("help", "do something", "set the temperature" with no device named) OR several entities are equally valid with no sensible default.
|
||||||
|
- UTILITIES is ONLY docs-grounded help about Home Assistant itself: a pending update, why a device is unavailable/offline, or how to set up/configure an integration. A bare "help" with no HA topic is CLARIFICATION, never utilities. Never invent state or entities.
|
||||||
|
|
||||||
|
THE FIVE TYPES:
|
||||||
|
1) command — control a device now (turn on/off, set, lock/unlock, open/close, play/pause, dim).
|
||||||
|
{"c":[{"s":"<domain.service>","e":"<entity_id>","d":{<params>}}],"r":"<one short past-tense confirmation>"}
|
||||||
|
One c entry per (service, entity_id); s is "domain.action" and its domain must match e's domain; omit d when there are no params.
|
||||||
|
2) automation — save a recurring rule, schedule, or multi-step sequence (cues: every, when, at <time>, if, whenever).
|
||||||
|
{"intent":"automation","response":"<short reply>","description":"<one line>","automation":{"alias":...,"triggers":[...],"conditions":[...],"actions":[...]}}
|
||||||
|
3) answer — a question about the home's current state.
|
||||||
|
{"r":"<answer, using {entity_id} placeholders wherever live state is needed>","q":["<entity_id>",...]}
|
||||||
|
4) clarification — the request is too vague to act on.
|
||||||
|
{"q":"<one short clarifying question>","o":["<option>",...]}
|
||||||
|
Here q is a QUESTION STRING (not an entity list); o offers optional quick-reply choices.
|
||||||
|
5) utilities — docs-grounded Home Assistant help (see ROUTING).
|
||||||
|
{"r":"<advice, with {entity_id} placeholders for live state>","q":["<entity_id>",...],"src":["<doc_chunk_id>",...]}
|
||||||
|
|
||||||
|
Use canonical entity_ids from AVAILABLE ENTITIES, never the human alias. Reference only entity_ids that appear there. Output exactly ONE JSON object of ONE type. JSON only.
|
||||||
Reference in New Issue
Block a user