Releases: Hal0ai/hal0
Release list
hal0 v1.0.0-alpha.1
Preview — v1.0.0-alpha.1
First official preview release. This is an alpha-quality build for testing the preview channel infrastructure.
Included
- Release-policy module (tag classification for stable/preview/nightly)
- Atomic version synchronization (
scripts/set-version.py) - Preview manifest fields (backward-compatible
hal0.releases.v1additions)
Known issues
- Channel-downstream preview routing not yet merged into release workflow
- This release was cut manually; automated preview channel publication is in progress
Upgrade from
- 0.9.8 or later (FHS installs)
--deveditable checkouts tracking this tag
Rollback
Safe — standard hal0 update --rollback works.
hal0 v0.9.8-nightly.20260720084155
Nightly v0.9.8-nightly.20260720084155 — changes since v0.9.8-nightly.20260713085026:
- feat: define 1.0 workload-oriented seeded profiles (#1324) (e316360)
- feat(brain): first-class hal0-brain module + /api/brain/chat primary route (#1327) (a314dcb)
- Merge remote-tracking branch 'origin/main' into rework/p3-brain (8a46421)
- docs(board): sync inc-3 + FLAGS-own rows to merged state (#1326) (3ddd9f5)
- feat(slots): §7 slot-purity — fold chat_template into model; sunset the slot tier (#1325) (9bde4b7)
- feat(mcp): autogenerate admin route-map from live routes (route-id keyed, deny-by-default) (#1323) (a894188)
- feat(api): typed request bodies for models/slots command routes (#1322) (28f497c)
- Merge pull request #1321 from Hal0ai/rework/descar (56e06ae)
- fix(dashboard): stop slot cards flickering bare→enriched (#1320) (cf55333)
- docs(board): P3-runtime-db inc0-3 — bilingual id/name slot layer + seed split-brain fix (83374a5)
- feat(slots): bilingual name-or-id on-disk slot layer + identity-keyed seed (P3-runtime-db inc0-3) (c8820cf)
- docs: install & migration guide (styled) + link from getting-started + v1.0.0 (R5) changelog (69dc413)
- fix(dashboard): stop slot cards flickering bare→enriched on poll (4f309b2)
- docs(board): SLOT-B id-flip root-caused — id-keyed read-path is unbuilt (= P3-runtime-db lane) (534b1e8)
- docs(deploy): SLOT-B id-flip rehearsed on 143 — BLOCKED on runtime-version sequencing (cafcf99)
- docs(board): Phase 4 deploy hardening — m3-renderer + migrator-ctx code fixes (9d23450)
- fix(migrate): drop launch-shadowed context_size from slot-flags fold + divergence key (6f8266d)
- fix(gpu): derive GPU group-add GIDs from device-node owner, not group name (0b4e678)
- docs(deploy): R5 deploy-window record — §2 flag regression fixed (3 slots), rocm/vulkan bench (da9f41c)
- chore(descar): sync with main f2db4d6 (post-R5-collapse) (2a067c2)
- chore(descar): unfreeze — descar→main R5 promotion complete (main f2db4d6) (66563c4)
- Collapse rework/descar → main: R5 (Phase 3 complete — memory/Hermes, SWEEP, FLAGS-own centerpiece, deploy-validation) (f2db4d6)
- chore(descar): FREEZE marker — descar→main promotion in progress, hold pushes (1490739)
- Merge remote-tracking branch 'origin/rework/descar' into claude/hal0-rework-handoff-o4fcws (ea938bd)
- style: ruff format agent_commands.py (CI format gate) (cbc0fb7)
- style(cli): ruff format agent_commands.py (fix-forward wave format-check) (d1b5af3)
- docs(board): Phase 3 wave-4 (deploy fix-forward) — quadlet-restart, installer-gates, cli-hermes (0d0af38)
- fix(cli): agent status --json output + hermes ownership-drift doctor check (0aa0ca0)
- fix(installer): accurate keyring-quota diagnostic + require render group in GPU preflight (6d154bf)
- fix(quadlet): emit StartLimit* in [Unit] so slot restart-limiting applies (205d526)
- docs(handoff): R5 Phase 4 handoff — deploy window → v1.0.0 release cut (12a8c89)
- docs(deploy): R5 install-validation record — 471c365 on halo150/143 (8d15b04)
- fix(tests): migrate unit-rerender test to FLAGS-own semantics (471c365)
- docs(board): Phase 3 wave-3 merged — FLAGS-own centerpiece + migrator + v1.0.0 re-target (8903a43)
- chore(sunset): re-target deprecated-surface stamps v0.10.0 → v1.0.0 (0a0bc6e)
- feat(flags): flags own by models — strip slot flag/device/template surface + copy-on-stamp profiles + migrator (d4253f8)
- docs(board): Phase 3 wave-2 merged — reconcile-wire, sunset-ratchet, persona-flag (b486d7e)
- feat(agent): add explicit persona-overwrite flag to install hermes (5bf5a74)
- chore(sunset): ratchet deprecated flags/aliases/fields for next-release removal (48a9b48)
- fix(ports): wire reconcile_listeners into /api/ports so the reconcile pass runs (cdb9ef8)
- docs(board): Phase 3 wave-1 merged — mem→hindsight, brain-lane→lifespan, drift-watch+hermes-bump, SWEEP dead-code (8ffc990)
- refactor(boot): relocate 5 brain-lane steps from install into api lifespan (577ee46)
- feat(memory): rename hal0_memory_* → hindsight_{recall,retain,reflect} + local_external config (4c2e5f1)
- test(drift): add drift-watch fixtures + docs(runbook): hermes-bump procedure (23c9242)
- refactor(sweep): delete 4 grep-confirmed-dead symbols (7e06260)
- docs(board): Phase 2 wave-2 merged — P3-routers inc3 pull-extraction (dc0f5cf)
- refactor(models): extract pull/update-pull orchestration to pull_jobs service (15b9d34)
- docs(board): Phase 2 wave-1 merged — lifespan-split + settings-seam (4e711d5)
- fix(settings): wire the 2 placeholder pages onto real typed hooks (0b2417e)
- refactor(api): phase-split boot lifespan into typed, observable stages (a9b1ea9)
- docs(board): Phase 1 COMPLETE — wave-3 merged (dup+mock, UI-API-2, CONTRACTS) (59f6983)
- refactor(ui): useAuthExposure uses ENDPOINTS.authExposure const (no literal dup) (67d1dde)
- feat(ui): wire DuplicateModelDialog to POST /api/models/{id}/duplicate + realign mock fixtures (c156f81)
- fix(ui): wire ExposureTable to live GET /api/auth/exposure (UI-API-2) (55d1da4)
- docs(ui): reconcile endpoints.ts/CONTRACTS.md self-contradictions (SWEEP §2) (7a81af9)
- docs(board): Phase 1 wave-2 merged — missing-routes, NAMES-stale, slot-logs-redact (3a852df)
- feat(api): implement 3 declared-but-missing GET routes (0941df5)
- fix(ui): NAMES-stale — role-based slot selection over name literals (8a95b8c)
- fix(slots): redact secrets in /api/slots/{name}/logs/stream (bc2954e)
- docs(board): Phase 1 wave-1 merged — MCP-sync, CLI verbs, UI endpoints, VERS-flash, api-logs-redact (43ea03d)
- fix(logs): correct false claim that slot logs share journalctl_sse() (d1c78d7)
- fix(logs): redact bearer/client_id/_KEY secrets in /api/logs stream (2b5662d)
- feat(mcp): hand-add §4.3 tier-a admin catalog tools + reclassify memory reads (645aa49)
- refactor(ui): consolidate dead endpoint constants (SWEEP §2) (024147f)
- feat(cli): wire the missing R3/R4 verbs — auth, board, ports, slot rename, model default/update/pull-cancel (§5.2) (21153e4)
- fix(ui): kill VERS-flash — reconcile UI version to backend, build-time inject (dbf9a63)
- docs(board): Phase 0 γ fallout fixed + descar reconciled with main (b5ac6b0)
- Merge remote-tracking branch 'origin/rework/descar' into rework/descar (6466994)
- Merge remote-tracking branch 'origin/main' into rework/descar (24cd248)
- merge: combine both activity-log de-flakes — their toHaveCSS + mock warming, my emitAndAwait for the emit-loss family (abb6db2)
- fix(e2e): de-flake activity-log — emits raced the pane's remount/reconnect gap (9d362d2)
- docs(rework): drive-2 fable handoff + SWEEP inventory (overlooked-surface backlog) (4ee5b07)
- fix(ui): warm the mockFixtures import + de-race ActivityLog bounded-scroll spec (979f58e)
- docs: R5 sync assessment + graphify knowledge graph + drive-2 handoff (#1317) (f2d82d3)
- fix(e2e): memory Recall/Reflect specs page.route their POSTs — mock substitution is GET-only now (68181bc)
- docs(board): Phase 0 (Wave 0) landed — 6 correctness/security lanes ✔ (278b32a)
- merge(phase0): thread service bearer into provisioned Hermes MCP + memory calls (1c01dba)
- fix(hermes): thread service bearer into _mcp_memory_call (auth-on 401) (f1af296)
- merge(phase0): ship hal0.target, model-store PermRow, uninstaller sync (1bb9bf7)
- fix(install): ship hal0.target, add model-store PermRow, sync uninstaller to R3/R4 (b9c0b0e)
- merge(phase0): gate prod mock fallback by method+env, lazy-import fixtures (54ccd58)
- fix(hermes): thread service bearer into provisioned MCP headers (auth-on 401 fix) (23cb35c)
- fix(ui): gate prod mock fallback by method + env, lazy-import fixtures (268f96a)
- merge(phase0): CLI auth-bypass fixes + deprecated-verb help texts (d55aa36)
- fix(cli): route streaming verbs through authed helper; fix deprecated-verb help texts (1dedecd)
- merge(phase0): MCP PATCH support + _REST_MAP route-sync test (b878ff9)
- fix(mcp): support PATCH in _call_rest + pin _REST_MAP to real routes (65fe49f)
- merge(phase0): MCP client_id leak fix — hashed label, not raw bearer (5d7cd28)
- fix(mcp): derive client_id label from hash, stop stamping raw bearer into journald (54a19c8)
- docs(rework): land R5 endgame handoff + sync assessment onto descar (9b05dee)
- docs(graphify): commit curated knowledge-graph artifacts (report + wiki + analysis) (2ab6ae6)
- docs(handoff): route 8 graphify structural findings into R5 doc surfaces (cb16294)
- docs(rework): handoff — route merge/meld/fix follow-ups into the R5 docs (4dbd1ed)
- docs(board): UI-API-1 → ✔ (PR #1316 merged to descar 4f828fe) (270a35a)
- docs(board): UI-API-1 row points at PR #1316 (built) + descar→main rebase note (bec70ce)
- docs(board): P4-docs/P4-rules/§21.11 golden-paths → ✔; UI-API-1 audit finding recorded (4a2ee61)
- test(golden-paths): complete 15-scenario coverage map + deploy runbook (6aadb8d)
- docs(P4-rules): add anti-scar rules 10-11 (bit-twice CI lessons) (a26fdef)
- docs(P4-docs): collapse CONTEXT/AGENTS stubs into ARCHITECTURE.md (6e3c168)
- feat(models): UI-API-1 — close the managed-arg gap (launch + create + /validate) (#1316) (4f828fe)
- feat(models): HF update-check basename fallback + a11y + e2e coverage (#1315) (462b28f)
- Collapse rework/descar → main: HP-realtime inc-1 + default consumers + docs cleanup (b70b981)
- docs(board): HP-realtime inc-1 ✔ merged — voice endpoint live on descar (1af40b0)
- merge rework/realtime-inc1: WS /v1/realtime — server VAD + both LLM legs + MCP recipe (HP-realtime inc-1) (65df5cd)
- docs(board): default-consumers + docs-stragglers ✔ merged (93be1ab)
- test(realtime): event-contract suite + voice guide (HP-realtime inc-1) (7a2a820)
- feat(realtime): WS /v1/realtime engine core (HP-realtime inc-1) (62d2358)
- merge rework/default-consumers: fallback prefers the type's default model; create modal preselects it (57fee5b)
- feat(ui): preselect a slot's type default model on create (9aaf9b6)
- feat(registry): fallback_local_model ...
hal0 v0.9.8
Turnstone lands as a second heavyweight bundled agent alongside Hermes and
becomes a first-class companion service; the unified hal0-rocmfpx runner
becomes the default image for AMD GPUs (with an automatic slot migration on
update); and a memory security fix stops one agent deleting another's private
memories.
Highlights
- Turnstone — a second heavyweight bundled agent joins Hermes: a native
turnstone-serveron loopback :9129, installed into its own managed PyPI venv, coexisting with Hermes via relaxed single-pick (#1299). - Turnstone is a first-class companion service — it shows in the Services pane and the Overview health card with start/stop/restart controls, next to Hermes/Hindsight/OpenWebUI.
- hal0-rocmfpx is now the universal default runner for AMD GPUs — one unified image (Vulkan/RADV + HIP) replaces the per-lane toolboxes; CUDA and CPU-only lanes keep their lean images (#1297).
- Memory: private-visibility is now enforced on delete — in unified-bank mode one agent could delete another agent's
visibility:privatememory by id;deletenow applies the same fail-closed ACL as read/search/list.
Migrations
hal0 updateautomatically re-pins existing AMD-GPU slots from the oldamd-strix-halo-toolboxesimages to the unifiedhal0-rocmfpxrunner (no-op on CUDA/CPU lanes) (#1297).
Added
- Turnstone bundled agent — provisioning pipeline (managed PyPI venv, JWT-secret generation, model automap) plus hal0 provider/memory wiring, coexisting with Hermes (#1299).
- Turnstone companion-service registration — a
ServiceDef+ systemd health probe, so turnstone appears in/api/services(Services pane, full lifecycle actions) and/api/services/health(Overview card + sidebar status).
Changed
- hal0-rocmfpx as the default AMD-GPU image — basic seed profiles defer to a manifest-driven resolver that returns the unified runner for AMD lanes, the CUDA image for NVIDIA, and the lean toolbox for CPU-only (#1297).
- PyPI distribution is published under the name
hal0ai(the import package andhal0/hal0-agentconsole scripts are unchanged) (#1298).
Fixed
- Memory delete ACL —
deleteenforces thevisibility:privateowner check in unified-bank mode, so an agent can no longer delete another agent's private memory by (guessable) document id; unresolved ids are withheld fail-closed (#1302).
hal0 v0.9.8-nightly.20260713085026
Nightly v0.9.8-nightly.20260713085026 — changes since v0.9.3-nightly.20260708080920:
- release: v0.9.8 (#1304) (402a772)
- feat(images): rocmfpx as universal default + hal0 update slot migration (#1297) (aff99e5)
- feat(agents): add turnstone as a heavyweight bundled agent (coexists with hermes) (#1299) (aac5514)
- build(pypi): publish under distribution name "hal0ai" (#1298) (40d11d3)
- release: v0.9.7.3 (#1296) (dd888d2)
- claude/hal0 hardening honcho2 (#1295) (4338739)
- fix(installer): harden Honcho memory standup (F24/F26/F27/F28) (#1294) (c7c82a2)
- fix(installer): init hc_cur_tag before use (unbound var aborted Honcho standup) (#1293) (979e29e)
- fix(preflight): container smoke runs image bare, not
... true(distroless hello has no /bin/true → false-fail aborted install) (#1292) (7f41684) - fix(installer): harden install.sh + preflight.sh for environment robustness (#1291) (26ef24e)
- fix(agents): hermes single-pick bypass + opencode installer uninstall/PATH gaps (#1290) (0c54ca6)
- fix(container): docker/podman runtime portability (--replace, runtime probe, OpenWebUI skip reason) (#1289) (e94011d)
- fix(pymisc): honcho remediation opt-in path + guarded first-run sentinel write (#1288) (06582c8)
- fix(comfyui): FHS-aware fetch script + workflow asset resolution for wheel installs (#1286) (e1da0f4)
- fix(uninstall): sudo env forwarding + Honcho stack teardown (#1287) (4f49f26)
- fix(agents): FHS-aware installer path + crash-safe agent switch (#1285) (e3852f4)
- fix(doctor): FHS-aware preflight.sh resolution + skip own-port false failure (#1284) (01a7cb9)
- fix(memory): honcho render-env tolerates missing hal0-honcho unit (opt-in) (#1283) (628e8f2)
- test(coverage): entry-point smoke tests for five 0%-coverage modules (#1282) (84a21e7)
- docs: sweep stale ADR/Cognee refs, deprecate gttsize, document fresh-box journey (#1280) (ba72db2)
- feat(memory): dry-run item-count preview on unconfirmed bank delete (#1273) (6ea9d2e)
- fix(comfyui): ESRGAN source, hf-CLI compat, curated model orchestration (#1277) (b590ff2)
- test(hermeticity): isolate 19 tests that leaked live-host state (#1281) (81e450c)
- docs(adr-0023): scrub stale hal0/chat default refs → hal0/agent (#1276) (0da9f6e)
- docs(installer): loud LAN-only warning on the 0.0.0.0:8080 bind (#1275) (9b0162d)
- upstream controls + registry/UI: close #1148 #1150 #1157 + #1256 audit (#1279) (8a054d6)
- fix(cli+update): editable-install refusal, footgun gates, command consolidation (#1274) (baf386d)
- fix(slots): errored-slot restart recovery + drift/image_status false positives (#1278) (ba5a4f5)
- ci(agents): add opencode lane to the nightly agent-shim smoke (#1272) (53aae72)
- fix(slots): auto-recover a slot wedged in WARMING (stale-anchor watchdog) (#1269) (e9639de)
- fix(honcho-migrate): retry transient embed 5xx and surface the cause (#1268) (0f46fcc)
- feat(agents): add opencode as a bundled agent (hal0 provider + hindsight memory) (#1271) (3ab82e8)
- fix(mcp): give each gated tool call a unique approval id (#1270) (f3e1f02)
- fix(profiles): re-tune seed flags, add Vulkan embed/rerank lanes, harden seed guards (#1266) (a2c3a85)
- feat(agents): provision hal0-brain as a first-class profile, not just a persona (#1258) (ca016b1)
- fix(board-chat): reliable steward tool-calling — text-parse fallback + tool_model routing (#1265) (ae1740c)
- feat(memory): re-land Honcho provider UI on current main (supersedes closed #1253) (#1267) (e36ed5e)
- refactor(mcp): share per-tool param schemas between board_chat and admin MCP (#1264) (1380e39)
- fix(bench): resume eval runs across defers instead of restarting + duplicating records (#1261) (3abdc71)
- fix(hermes): restore uv Python provisioning reverted by #1257 + make UV_PYTHON_INSTALL_DIR traversable (#1259) (0a7de78)
- fix(ui): don't render pi-mark
with empty src when agent art is missing (#1263) (e4fedc6)
- fix(memory-cli): tolerate operation_id key in ops retry + scope migrate-unify retag to transferred docs (#1262) (b5f184f)
- fix(memory): enforce visibility:private on read so private memories aren't cross-agent readable (#1260) (ad4c58d)
- feat(cli): hal0 memory bank/ops/mm/recall + migrate unify (rebased onto main) (#1257) (00bf2e1)
- fix(memory): consolidate drifted Hermes memory plugin + stray-bank identity fix (#1245) (a73a2c2)
- feat(memory): unified-bank model with server-side tagging (#1244) (b80911f)
- feat(memory): self-hosted Honcho v3 as a per-agent memory provider (unified memory) (#1243) (05efc59)
- feat(agents): pi-coder agent — provisioning, hal0 provider/memory plugins, live Pi dashboard card (#1254) (f2f7c13)
- feat(bench): queue dropdown with lane, tool-eval, and tune options (#1255) (ae7ebcf)
- release: v0.9.7.2 (fd9f2d5)
- feat(hermes): provision a uv-managed Python when no system 3.11-3.13 exists (#1252) (29a4887)
- fix(hermes): require a 3.11-3.13 venv interpreter + floor hermes-agent at 0.16.0 (#1251) (2c91d3d)
- release: v0.9.7.1 (b51062d)
- test(e2e): stub /api/upstreams so the γ connections-v3 suite stops blanking (#1242) (bda3e9a)
- feat(memory): retry-failed-extractions button, consolidate fix, mental-model delete (#1235) (4b92283)
- feat(memory): replace HAL0_MEMORY_ENABLED env var with [memory].enabled + CLI toggle (#1240) (2a0750e)
- fix(hermes): auto-provision on fresh installs (unclaimed home + self-foreign gateway) (#1239) (ff370a1)
- feat(setup): keyboard-driven guided-setup TUI (arrow keys + clean pickers) (#1237) (0ed4d8b)
- feat(upstreams): add MiniMax and DeepSeek providers + fix filter spacing (#1236) (ee87dbc)
- release: v0.9.7 (baef535)
- docs(agents): add the hal0-brain steward section — how it works, why it runs on a 1B model, MiniCPM5-1B recommendation (#1234) (26b8526)
- fix(release): sign with Sigstore bundle so client verify survives cert expiry (#1159) — HOLD, needs test-release validation (#1189) (c0d5538)
- docs(release): v0.9.7 changelog + README updates (#1233) (d78486e)
- fix(api): stop model pulls from blocking a graceful hal0-api restart (#1225) (#1232) (9a14f5e)
- fix(npu): finalize trio routing aliases in /v1 models route (#1231) (71723b6)
- feat(seed): agent + brain slot seeds (#1230) (825290b)
- feat(flm): NPU FLM naming migration (#1229) (2533bd1)
- feat(upstreams): upstream model controls — reactive CRUD, model filters, enabled kill-switch, CLI + MCP + dashboard (#1228) (9af8b23)
- feat(brain-chat): Agents/Brain settings section + slot override for the steward chat (#1223) (96d117b)
- feat(brain): tool-use hardening — pause-on-approval, tool-history replay, arg schemas, global port registry (#1222) (6ac481d)
- feat(brain-chat): [brain_chat] guardrails — kill switch, read-only mode, config-backed loop knobs (#1221) (3c55bfb)
- feat(agents): safe capture of existing hermes installs — --adopt, fatal claim abort, foreign-gateway preflight, ownership reconcile (#1220) (02e86b0)
- fix(test): stop static slot seeding from polluting every test's zero-slot baseline (#1219) (2995b72)
- feat(ui): slots-page polish batch — headings, NPU activity tint, logs drawer, image-gen header + hal0 update owui (#1216) (7e213d2)
- feat(ui): add a Documentation button to the topbar (#1213) (513c56c)
- feat(install): seed static slot TOMLs on hal0-api startup, not just fresh install (#1218) (7190919)
- feat(install): seed agent + brain slots; agent replaces chat as the LLM anchor (#1217) (cdffb22)
- feat(mcp): enforce per-persona tool policy on the Brain's admin surface (#1215) (37102c2)
- feat(ui): model updates in the notification bell + always-visible Check updates (#1190) (fb78a75)
- feat(doctor): FLM/migration/profile audits + --force delete for seeded slots (#1214) (160d71c)
- fix(slots): canon FLM-trio shadow slots to flm-stt/flm-embed + startup reconcile (#1210) (9efa6b1)
- chore: address #1182 review follow-ups — dev bench state wiring, docstring, README de-rot (#1184) (9db49af)
- feat(agents): ship hal0-brain persona on every install path; retire coder seed (#1204) (df53b65)
- fix(flm): land host pulls in the resolved store + robust 0-byte progress (#1211) (16a9ccc)
- test(e2e): follow container controls into the ComfyUI card header (#1212) (591ed7a)
- feat(mcp): expand hal0-admin to the full platform surface for the Brain agent (#1208) (423fbd6)
- feat(ui): polish image-gen header + size Local endpoints heading (#1209) (34c5b8d)
- feat(ui): unify slots-page headings and slot-status indicators (#1205) (5ca97c4)
- feat(ui): full-height slot logs drawer + surface raw logs from the Logs page dropdown (#1207) (998d62a)
- feat(dash): topbar tidy — move Kanban to Services + brain icon for Agent Chat (#1206) (fb56946)
- feat(dash): unify Needs Attention + bell into one notifications source (#1195) (f94e454)
- fix(dash): real update-banner actions + NPU grid colours by the primary FLM slot (#1194) (42ba168)
- feat(scripts): push-dev.sh — push-based inner-loop deploy to an editable box (#1193) (9134a51)
- fix(flm): FLM model pulls die instantly in prod — uvloop rejects user/group spawn kwargs (#1192) (7afc52f)
- Merge branch 'docs/sync-v0.9.5.x' (36dfb5b)
- docs: sync docs with v0.9.5–v0.9.5.2 + add docs agent team (3853429)
- release: v0.9.6.1 (#1191) (b4260d1)
- fix(setup): treat apply-selections 409 as a recoverable no-op (#1158) (#1187) (a0875c6)
- fix(ui): add no-undef lint guard for dash JSX (#1170) (#1188) (d3aef5c)
- fix(ui): consolidate --warn color token to single source (#1156) (#1186) (3dac085)
- fix(flm): probe the slot's assigned model in verify_inference, not models[0] (#1171) (#1185) (57babc4)
- test(npu): align stale NPU occupancy γ-specs with the reworked card (#1183) (879ee52)
- fix: install/setup/update audit follow-ups — manifest pins on prod, honest NPU setup, doc + harness catch-up (#1182) (9e27b91)
- feat(chat): retarget agent-chat slide-out to the hal0-brain profile (#1180) (fa082b7)
- feat(models): HF update check + per-row indicator ...
hal0 v0.9.7.3
A robustness release. Fresh installs now adapt to the host — provisioning their
own Python and Node, tolerating podman or docker, and surviving hardened
umasks and non-root operation — and the self-hosted Honcho memory stack stands
up cleanly alongside Hindsight. Plus the pi-coder and opencode bundled agents,
the unified per-agent memory model, and a large batch of install/setup/agent
fixes surfaced by end-to-end reinstall testing on a clean box.
Highlights
- Installs self-heal across environments — auto-provision Python 3.12 and Node 20 LTS when missing, tolerate podman or docker, and survive hardened/root umasks (#1291, #1289).
hal0 doctorand bundled-agent installs work on packaged installs — FHS-aware resolution fixes the "packaged without scripts" and "could not locate preflight.sh" failures on non-editable installs (#1284, #1285).- Crash-safe agent switching — a failed
agent install --switchno longer bricks the running agent; it verifies the target first and rolls back on failure (#1285). - Self-hosted Honcho as a per-agent memory provider — unified memory, swappable and migratable with Hindsight per agent, with a clean opt-in standup (#1243, #1294, #1295).
- pi-coder and opencode join Hermes as bundled, single-pick agents (#1254, #1271).
Added
- Node.js LTS auto-provisioning in the installer (the dashboard build and pi-coder/opencode all need npm) plus a curated
qwen3-embedding-0-6bmodel for the memory pipeline (#1291, #1294). - Unified per-agent memory — self-hosted Honcho v3 provider and a unified-bank model with server-side tagging, plus
hal0 memory bank/ops/mm/recallandmigrate unifyCLI (#1243, #1244, #1257). - pi-coder agent (provisioning, hal0 provider/memory plugins, live dashboard card) and the opencode bundled agent (#1254, #1271).
- hal0-brain as a first-class profile (#1258), upstream controls CLI and UI (#1279), and a bench queue dropdown with lane/tool-eval/tune options (#1255).
Fixed
- Installer/preflight robustness — Python floor raised to 3.12 with auto-install, hardened-umask permissions, a container-runtime smoke test that no longer false-fails, and Node/disk/graphroot preflight gaps (#1291, #1292).
- Honcho standup — unbound-var abort, pgvector embedding-dim reconcile, compose-provider and migration ordering, AppArmor-in-LXC, and full
--purgeteardown (#1293, #1294, #1295, #1287). - docker/podman portability — runtime-appropriate slot units and a PATH-based runtime probe (#1289).
- Slots — errored-slot restart recovery and the WARMING watchdog, plus drift/image_status false positives (#1278, #1269).
- Memory — private-visibility enforcement on read, migrate-unify retag scoping, and Hermes plugin/bank identity consolidation (#1260, #1262, #1245).
- CLI — editable-install
hal0 updaterefusal, footgun confirmation gates, and command consolidation (#1274).
Changed
hal0 v0.9.7.2
Hotfix for a fresh-install blocker on Python-3.14-only distros (Ubuntu 26.04,
the Strix Halo target): the Hermes venv was built on 3.14, where pip silently
resolves the broken hermes-agent 0.15.2 wheel, and the gateway crash-looped on
ModuleNotFoundError: No module named 'hermes_cli.dashboard_auth'.
Highlights
- Fresh installs now work on Ubuntu 26.04 / Python 3.14 — when the system has no Python 3.11-3.13, hal0 provisions a uv-managed 3.13 for the Hermes venv automatically; zero manual steps (#1250, #1252).
Fixed
- Hermes venv could be built on Python 3.14, where every supported hermes-agent wheel is filtered out (
requires-python <3.14) and pip silently installs the broken 0.15.2 build. The interpreter resolver now accepts only 3.11-3.13, probespython3.13/python3.12/python3.11newest-first (hosts with only 3.12/3.13 on PATH were previously ignored), and preflight fails early with an actionable message instead of a crash-loop three phases later (#1248, #1251). - hermes-agent version floor raised to 0.16.0 — the 0.15.2 wheel imports
hermes_cli.dashboard_authbut ships without it; it can no longer be selected on any interpreter (#1247, #1251). - Broken venvs self-heal: a venv previously built on an unsupported interpreter is detected from its on-disk layout and rebuilt during
hal0 agent bootstrap hermes --repair— no manual removal needed (#1251).
Added
- uv-managed Python fallback for the Hermes venv: system Python always wins; when nothing qualifies and uv is on PATH, the install phase runs
uv python install 3.13into/var/lib/hal0/python(kept out of root's home so thehal0service user can execute it) and builds the venv on the managed interpreter (#1250, #1252).
hal0 v0.9.7.1
Hotfix for a fresh-install blocker — Hermes never auto-provisioned on a clean
curl | bash install — bundled with the dashboard, provider, and memory
improvements merged since 0.9.7.
Highlights
- Hermes now auto-provisions on a fresh install — no more manual
--adopt/--repair; all bootstrap phases complete and the gateway comes up managed, idle until a bot token is added (#1239). - MiniMax and DeepSeek are now one-click options in the upstream provider catalog (#1236).
- Keyboard-driven
hal0 setup— the guided setup TUI gets real arrow-key navigation and a clearer model picker (#1237). - Memory subsystem toggle —
hal0 memory enable/disable/statusreplace the old install-time env flag (#1240).
Added
- MiniMax + DeepSeek upstream catalog entries (OpenAI-compatible, bearer auth) — selectable in Slots ▸ Endpoints ▸ Add upstream with prefilled URL and auth (#1236).
- Arrow-key navigation in the guided-setup TUI: ↑↓/jk move, space toggles apps/agents, enter selects; scaffold/skip are clean navigable rows; numbered entry still works over a pipe or in CI (#1237).
hal0 memory enable/disable/statuscommands, plus the[memory].enabledconfig flag (#1240).- Retry failed graph extractions — a Memory-tab button that re-runs Hindsight's failed extraction operations, plus mental-model delete (#1235).
Fixed
- Hermes fresh-install provisioning (#1239, closes #1238). Two chained bugs aborted provisioning on every clean install: the api-lifespan seed populated
HERMES_HOMEbefore the bootstrap could claim it ("unclaimed HERMES_HOME"), and the gateway installer started a unit that hal0 then flagged as its own "foreign" gateway. The lifespan now stamps.hal0-managedbefore seeding and pre-writes the gateway secrets drop-in before the unit starts. - Model-filters spacing in the Endpoints upstream pane — the filter inputs no longer sit flush against the panel border (#1236).
- Graph-extraction consolidation reliability in the Memory tab (#1235).
Changed
- The memory subsystem is now gated by
[memory].enabledinhal0.toml(default on) instead of the installer-writtenHAL0_MEMORY_ENABLEDenv var; toggle it withhal0 memory enable/disable(#1240).
Migrations
- If you had disabled memory with
HAL0_MEMORY_ENABLED=0, that env var is now ignored (memory defaults on) — runhal0 memory disableto keep it off (#1240).
hal0 v0.9.7
The steward release. The dashboard's agent chat graduates from a
side-panel into a real control surface for the whole platform, external
LLM providers get a first-class management surface, and the FLM/NPU
stack settles onto canonical names with an automatic migration.
Highlights
- hal0-brain steward. The top-bar agent chat now drives every platform surface: it runs the full 74-tool
hal0-adminMCP catalog under a per-persona tool policy, pauses turns on gated tools for inline approve/deny, replays tool history across turns, and renders reasoning + tool cards inline (#1208, #1215, #1221, #1222, #1223). - Upstream model controls. A full management surface for external providers (OpenRouter, Anthropic, OpenAI, Google AI Studio, Ollama, custom) — reactive CRUD, per-upstream model filters, an
enabledkill-switch, and CLI + MCP + dashboard parity (#1228). - Graceful restarts keep your downloads. Model pulls no longer block a clean
hal0-apishutdown, so restarting mid-download no longer trips the 90s SIGKILL that was killing in-flight pulls (#1225). - FLM / NPU canonicalization. The NPU trio's shadow slots settle onto
flm-stt/flm-embed,/v1/modelsalias routing is finalized, and a naming migration +hal0 doctoraudits move existing installs onto the new scheme (#1210, #1229, #1231, #1214). - Agent is the new anchor. Seeded
agentandbrainslots replacechatas the default LLM anchor, and are seeded on every startup — not just fresh installs (#1204, #1217, #1218, #1230). hal0 updateverification actually works again. Release signing now dual-emits a Sigstore bundle (with an embedded Rekor timestamp) so cosign verification survives the short-lived Fulcio cert's expiry —curl … | bashinstalls and in-app updates were failing verification 100% of the time on v0.9.2/0.9.3/0.9.5. Fully back-compatible with already-deployed clients (#1159).
Added
- hal0-brain steward / agent chat as a control surface. The dashboard's top-bar agent chat becomes a first-class operator for the whole platform:
- Full
hal0-adminplatform surface (#1208). The admin MCP catalog grew from 28 to 74 tools, replacing the slide-out's old hardcoded ~25-tool list, so the Brain can drive every surface end-to-end:- Models —
model_inspect(read an HF repo before pulling), register / add-from-path, metadata edit (PUT), in-place HF re-pull (model_update), pull status/cancel, scan (+preview), catalogue, update-check, and store show/set/migrate. Pulls always land in the operator's configured[models].store(re-read per call; the tool descriptions state the contract). - Slots — load / unload / edit-config (PUT) / metrics / capacity / logs, on top of create / delete / restart / swap.
- Stacks — create / update / export / snapshot, joining apply / import / delete.
- Profiles — create / update, joining import / export / delete (author a profile straight from a model card).
- Settings & platform —
settings_get/schema/apply_plan/reload,upstream_list, and benchmark runs/status/queue reads plus gated enqueue/control. - Two new guards make catalog drift impossible to reintroduce: import-time catalog validation (classification ↔ REST-map ↔ annotations ↔ path-args must cohere) and a
build_serverregistration-completeness check.
- Models —
- Per-persona tool policy (#1215). The persona TOML's
tools_allowed+[persona.approval]tables — previously decorative on the sidebar path — become an enforced server-side overlay (admin.ToolPolicy, fnmatch globs over tool names):tools_allowedhides tools from the surface entirely;require_approvaltightens an autonomous tool behind the approval queue;auto_approve/default_policy=auto-approvegrants standing approval to gated tools;default_policy=neverrefuses gated calls outright. Precedence is hide > tighten > loosen > server verdict. APOLICY_NO_LOOSENfloor meansmodel/slot/stack/profile_delete, bulkmemory_delete,config_write, andprovider_credential_writecan never be loosened by a persona edit; denials are typed (mcp.tool_not_allowed/mcp.gated_tool_refused) with audit rows. - Tool-use hardening (#1222) — fixes for four failure modes seen live:
- Runaway generation — every round now carries
max_tokens(default 4096, payload-overridable); an uncapped completion against a slow local slot used to burn ~25k tokens and the 300s transport window, killing the turn before the first tool call. - Invisible approvals — gated calls no longer park silently on the queue; the loop emits an
approval_requiredSSE frame and pauses (with keepalive pings) until the operator approves/denies, then streams a secondtool_resultso the same turn continues (timeout falls back to the pending result). - Guessed argument names — high-traffic tools ship real parameter schemas + explicit descriptions (no more
model_inspectwithouthf_repo→ 400, ormodel_pullwithmodel_id='org/repo'→ 405); path args containing/are rejected with an actionable hint. - Round budget —
_MAX_ROUNDSraised 8 → 90 (with per-round caps it's now a runaway backstop, not a working limit multi-step sessions kept hitting). - UI —
useBoardChat.send()now replays full tool history (calls + results) in the outgoing conversation; previously it rebuilt from user/assistant text only, so after one tool chain the model saw no tool calls in its own history and learned to skip tools and hallucinate results. Adds inline approval cards, a Stop button, an auto-approve toggle, and new-session. - A global port-claim registry — the single authority for which slot owns which port — lands alongside in the same PR.
- Runaway generation — every round now carries
[brain_chat]server-side guardrails (#1221), enforced independently of the persona TOML (a persona edit can loosen the persona, never these):enabledis a hard kill switch (the endpoint refuses every turn — no LLM call, no board request);read_onlylets reads through but refuses every mutating/admin-write tool at the single_dispatch_toolchokepoint (unknown tools fail closed);max_rounds/completion_timeout_smove the loop budget + per-round transport timeout out of module constants into config, with the schema as the single source of truth.- Agents / Brain settings + slot override (#1223). A dashboard Settings → "Agents / Brain" section surfaces
[brain_chat](enabled + read-only toggles,max_rounds/completion_timeout_sinputs) plus a slot-override picker built from the live slots list — point the steward at any slot (e.g.hal0/nputo run it on the NPU chat slot). Model precedence is explicit: per-request model >[brain_chat]override > persona/default; the rows carry a live badge (brain_chat.*is apply-planimmediate).
- Full
- Upstream model controls — full management surface for external LLM providers (OpenRouter, Anthropic, OpenAI, Google AI Studio, Ollama, custom):
- Reactive CRUD:
POST/PATCH/DELETE /api/upstreams(create prefills from the provider catalog viacatalog_id);upstreams.tomlstays canonical — every write rewrites it atomically before touching the running registry. hal0 upstreamCLI group (list/show/create/update/delete/test/set-credentials);create --catalog openrouter --api-keywires a provider end-to-end in one command.- MCP admin tools
upstream_create/upstream_update/upstream_delete(gated) +upstream_test. - Dashboard Upstream providers panel (Slots → Endpoints / Connections): add-from-catalog form, write-only key entry, test-connection with latency + model count, enable/advertise toggles, filter editor with live preview, delete-with-confirm.
- Per-upstream
model_filters(modelsallowlist +include/excludeglobs, exclude wins) curating/v1/modelsand/api/modelsadvertising — dispatch stays unfiltered so hidden models remain addressable by name (per the 2026-07-06 upstream-model-filters spec). enabledkill-switch on every upstream:falseremoves it from dispatch routing and the model catalog while retaining config + credentials.auth_key_presentin upstream serializations — whether the declared env-var actually holds a key (drives the dashboard auth badge), distinct fromauth_configured.
- Reactive CRUD:
- Install / lifecycle:
- Safe capture of an existing Hermes install —
--adopt, a fatal claim abort, foreign-gateway preflight, and ownership reconcile so hal0 can take over an already-running gateway without clobbering it (#1220). - Static slot TOMLs plus the
agent+brainslots are now seeded on everyhal0-apistartup, not only on fresh install (#1217, #1218, #1230).
- Safe capture of an existing Hermes install —
hal0 doctorgrows FLM / migration / profile audits and a--forcedelete for seeded slots (#1214).- Dashboard / UI — slots-page polish: unified headings + status indicators, an NPU activity tint, a full-height slot-logs drawer surfaced from the Logs page, an image-gen header pass, and a Documentation button in the topbar (#1205, #1207, #1209, #1213, #1216).
scripts/push-dev.sh— a push-based inner-loop deploy onto an editable box (#1193).
Changed
- The default LLM anchor is the seeded
agentslot; thecoderseed is retired andhal0/chatresolves tohal0/agent, on every install path (#1204, #1217). - The NPU trio's shadow slots are canonicalized to
flm-stt/flm-embed, with a startup reconcile that migrates legacy names (#1210). - FLM host pulls land in the operator's resolved model store (#1211).
Fixed
- Model pulls no longer block a graceful
hal0-apirestart — the shutdown path drains pulls instead of being SIGKILLed at the 90s deadline, which was killing in-flight downloads (#1225). - FLM model pulls that died instantly in production — uvloop rejects the user/group spawn kwargs the pull worker passed (#1192).
- FLM host pulls now land ...
hal0 v0.9.6.1
Added
hal0-brainagent profile: a third seeded persona (alongsidehermesandcoder) that stewards the platform from the dashboard's agent-chat slide-out — its own memory namespace (private:hal0-brain), a hal0-heavy system prompt (slot lifecycle, model setup, benchmarking), and the dedicatedbrainslot (hal0/brain) as the default model.- Board/agent chat streams a
{type:"thinking"}SSE frame: explicitreasoning_contentand inline<think>…</think>blocks are split out of the reply and rendered as a folded "thinking" section instead of raw tags.
Changed
- The top-bar agent chat now embodies the
hal0-brainprofile: it runs onhal0/brain(falls back to theagentslot via the resolver chain), and an operator-editedhal0-brainpersona TOML overrides its system prompt/model without a code change. - Agent-chat suggestion chips are now platform-steward starters ("Help me create a new slot", "Download and set up a model", "Benchmark the model on a slot", "How's the hardware doing?").
- Agent-chat replies render markdown (fences, lists, headings, bold/italic/inline code, links); tool calls render as structured cards with args, live status, and a folded result — replacing the raw
→ tool({json})text rows.
Fixed
- FLM NPU slots no longer wedge in
warmingforever: the warm→ready inference sentinel now probes the slot's assigned model instead ofmodels[0]from FLM's full catalogue. Probing an arbitrary other model forced FLM to reload the wrong weights onto its single NPU context mid-gate and deadlocked the load (#1171). hal0 setupno longer aborts with a rawHTTPStatusErrortraceback whenapply-selectionsreturns409 Conflict: the setup CLI now treats a 409 (install already applied / a concurrent apply in flight) as a recoverable no-op with a clean message, and still raises on genuine errors (#1158).- Unified the dashboard warming-state color to a single canonical
--warn: #f2792btoken indashboard.css, removing the divergent local redefinitions inengine-panes.css/overhaul.cssso every warming indicator renders the same orange-yellow (#1156, #1155). - Added an ESLint
no-undefguard (enforced in CI) over the dash.jsxprototype so an undefined identifier like the one behind the Create Slot crash can't ship again (#1170).
hal0 v0.9.6
Highlights
- FLM NPU trio: the edit-slot drawer's Chat/ASR/Embed toggles now drive the running
flm serveprocess, with a full-catalogue chat model picker that downloads on demand. - FLM can run embed- or STT-primary (chat disabled) — a modality-aware readiness gate promotes the slot instead of wedging on a chat probe.
- Real FastFlowLM v0.9.44 toolbox image (
ghcr.io/hal0ai/hal0-toolbox-flm:0.9.44), rebuilt from the actual v0.9.44 binary. - NPU occupancy cards now glow purple when a slot is running (they used to read flat green regardless of state).
- hal0-bench is now in-tree: the
hal0.benchengine,/api/benchmarks, and a Benchmarks dashboard page.
Added
- NPU trio drawer wiring: ASR/Embed on-off toggles + a Chat model picker that lists the full FLM catalogue and pulls a not-yet-downloaded model on select (auto-applies on completion).
FLMProvider.verify_embed— a one-shot/v1/embeddingsreadiness sentinel used when a slot serves embeddings without chat.- hal0-bench in-tree port:
hal0.benchengine,/api/benchmarksroutes, and the Benchmarks dashboard tab (roster / runs / evals / run-queue).
Changed
- FLM toolbox pinned to
0.9.44acrossmanifest.json, the flm seed profile, and the capabilities catalog (contains FastFlowLM binary v0.9.44). - The warm→ready gate for FLM slots picks its sentinel by served modality: chat →
/v1/chat/completions, chat-off+embed →/v1/embeddings, ASR-only →/v1/modelsliveness. /api/slots/flm/modelsreturns the full FLM catalogue (installed + downloadable) with accurateinstalledflags, container-exec first with a host-probe fallback.
Fixed
- Chat toggle now actually gates the container:
container_specno longer passes the positional chat tag when[npu].chat=false(it was cosmetic on the container path). - Editing an NPU modality no longer clobbers
[model].default/context_size— the drawer sends the model as a nested[model]table so the backend merge preserves sibling keys. /api/npu/occupancyno longer 500s while the FLM slot is offline (the degraded single-tenant fallback'szip(..., strict=True)mismatchedslots_outvsflm_slots).- NPU cards read a clear purple "running" glow for up/resident slots and dim for offline; the coresident STT/embed sub-cards reflect the anchor's
[npu]toggles instead of the legacy shadow-slotenabledflag.