Skip to content

v2 rewrite status ​

The v2 rewrite was built on the v2 branch in phases and merged into release as version 2.0.0. Since the cutover (phase 5) the v1 crates are gone; they remain in git history.

PhaseScopeStatus
0Branch, crate skeletons, protocol, llm core types, testkit, prompts skeletondone
1LLM adapters (OpenAI-compatible, Anthropic), retries, fallback, catalog, costdone
2Config v2 + v1 import, policy engine, host adaptersdone
3Toolsdone
4Engine core (single agent)done
5GUI cutover, v1 parity, delete v1 cratesdone
6Multi-agent orchestration, worktree isolationdone
7MCP, LSP, rules, repo map, cachingdone
8Verification (discovery, evidence, badges, modes)done
9Commands, MCP prompts, workspace trustdone
10Sandbox, fault injection, flakiness, performance, docs, 2.0.0done

GUI redesign (shipped as 2.1.0) ​

StepScopeStatus
1Dead CSS and components removed; tokens, base, motion and materials in ui/src/styles/; native window translucencydone
2Title bar, sidebar, content sheet, scroll-edge fades, one-home-per-fact cleanupdone
3Title bar island with the companion, the waiting and context satellites and the island sheet (replaces the status line and Now card); sidebar with Home, Inbox and Projects; ui.companion settingdone
4Transcript, tool grouping, turn receipt (ui.task_report_view), approval card, composerdone
5Splash, first-run setup, project home and the Inbox (background notifications)done
6Agents panel, Changes panel, prompt inspector, palette, every Settings tab, worktree dialog, shell drawerdone
7Errors after edits (engine event)pending
8Remove legacy stylesheets and token aliasesdone

After 2.1.0: the pet (shipped as 2.2.0) ​

ItemStatus
The pet replaces the title-bar companion orb (components/pet/, pure rules in lib/domain/pet/); liveliness ui.companion moved to Settings → Petdone
Roaming between perches (home spot, composer, sidebar footer, the side panel's tab band, empty Inbox), drag, boop, nap, carddone
Growth: XP, levels, stages, accessories, tricks, saved in <data dir>/pet.json (pet_load, pet_save)done
[ui.pet] settings (name, look, roam), first-run Meet step, palette actionsdone
One Send/Stop button, helper sprites, inspector ring chart, sliding sidebar selection, full-window page frame for Settingsdone
Level-up celebration: confetti once nothing needs you (Lively only)done
The pet as the logo: boot splash (the pet wakes up, then flies into its place), app icons, favicon and About logo replace the Z markdone

After 2.1.0: the v4 dark frosted redesign (shipped as 2.2.0) ​

ItemStatus
Always dark: four material layers, one glass recipe, springs (lib/ui/springs.ts), motion on position and opacity only; macOS dark HUD vibrancy, Windows 11 dark Mica, solid elsewheredone
Side panel with Changes, Plan, Agents and Context tabs, replacing the separate panels and the full-window prompt inspectordone
Approvals and questions docked in the composer; the Plan ready · Review row and the Plan tabdone
Island shapes and the island card; turn actions, the work summary line, the live compaction row, the long-chat turn windowdone
QA in a browser against a mocked backenddone
Performance profile in headless Chrome of a 1,000-turn chat (__zengine.longChat(1000), dev builds only)done
Native checks: macOS vibrancy under tauri dev, Windows 11 Mica, a full Windows MSVC build, title-bar double-click on a real windowpending

After 2.2.0: the 3D pet (unreleased) ​

ItemStatus
The pet drawn in 3D with Three.js (lib/pet3d/, the only three import): one shared WebGL renderer, one frame loop that runs only while something plays, loaded lazily; pure helpers keyframes, rig3d, faceShapes, faceLayout, portrait in lib/domain/pet/done
The flat SVG pet (PetFlat) while the 3D pet loads, without WebGL and after the context is lost; the boot splash and helper sprites stay flatdone
The island's portrait of the roaming pet (IslandPortrait), replacing the empty nestdone
Checks by hand in headless Chrome against the real app with a mocked backend: every perch (hero, composer, empty Inbox, sidebar, side panel, island), drag and drop, a hard throw (lands dizzy), the card, Settings → Pet, first run, every look, stage, accessory, prop and mood, the island portrait and its tone ring, the flat pet without WebGL and after a lost context; frame counts: none while idle or asleep (blinks only), about 30 fps for mood loops, one draw per change under Reduce Motiondone
Fixes from those checks: the frame loop skipped a motion's last frame after a slow frame (eyes left mid-blink); a Svelte warning on every roaming frame; closing the card with Customize… threw an error (petGaze focus handler); the portrait now follows a slumped, sunk or sleeping body and the seed's size, keeps both eyes in its window at the widest turn and shows no gap above the crown; pets redraw when the screen's scale changes, also under Reduce Motiondone
The same checks in the native Tauri window (tauri dev)pending

After 2.2.0: experimental features and the decision layer (unreleased) ​

ItemStatus
Phase 1: the feature registry and [experimental] (Off, Shadow, On), Settings → Experimental; [decisions] with the sidecar, Test connection, trust gating and the opt-in dataset; the z-engine-decisions crate (rules, SystemOne and hybrid providers, calibration, cache, trace); the engine's decision service and seams; the Context tab's Decisions sectiondone
Context and cost uses: relevance-aware compaction, task-scoped history (the TaskView log record, the divider, Include full history), first-request context, file prefetch ([decisions.prefetch]), output trimming and search ranking (the Relevance seam)done
Per-task routing: effort per task, the fast model with [decisions.routing] allow_model_switch, subagents that inherit their model, the route chipdone
Safety uses: risk review (gate and injection notes), secret screening (the request seam, placeholders), custom decision rules ([[decisions.rules]]), loop guard ([decisions.loop_guard])done
Verification uses: completion check (the Claimed, not checked badge), check selectiondone
Suggestions and guidance: plan, standing-rule and review cards (Suggested, ResolveSuggestion), correction signal, skill and agent hints, repeated-question checkdone
The app: pet mood from turn tone (TurnToneJudged), inbox priority (UrgencyScored, gating Notification hooks)done
Native runtime: decisions.runtime = "native" behind the onnx Cargo feature; pinned english, multilingual and typed-decisions checkpoints; resumable, SHA-256-checked download in Settings → Experimentaldone
SystemOne wire fixes: the request carries model and max_len; scale scores given as a choice index are normalized; choices keep their order; max_len is part of the cache keydone
End-to-end tests (crates/z-engine-engine/tests/decisions_*.rs, below)done
Release builds with the native runtime (--features onnx in the release workflow)pending
The live A/B test and the offline benchmark against a real modelpending: not run, no credentials

Graduation criteria ​

Each feature graduates when it meets its criteria in crates/z-engine-config/src/features/registry.rs; results are recorded here. Every feature is Experimental, Off by default and supports Shadow.

Feature[experimental] idGraduation criteriaResults
Relevance-aware compactiondecisions_compactionAt least 15% fewer context tokens at equal task success, zero loss of mandatory context, p95 at or under 250 ms, fallbacks at or under 5%.pending
Task-scoped historydecisions_task_viewAt least 30% fewer input tokens per turn on chats of 5 or more tasks, no drop in task success, readbacks at or under 5%, and cost per task no higher once prompt caching is counted.pending
First-request contextdecisions_session_contextFewer tool rounds before the first useful read at equal task success, with no rise in first-request tokens.pending
Per-task routingdecisions_routingLower cost per task at equal task success, with the share of input read from the prompt cache not lower than with the flag off.pending
Loop guarddecisions_loop_guardCatches at least 80% of the loop fixtures with at most one needless reminder per 50 turns.pending
Completion checkdecisions_completion_checkAt least 90% precision on claim fixtures; a turn is never marked Verified by this feature.pending
Risk reviewdecisions_riskNo decision ever loosened (enforced by structure) and at most one needless escalation per 50 calls.pending
Relevant output trimmingdecisions_output_trimAt least 20% fewer tool-output tokens with no rise in re-runs of the same command.pending
Search rankingdecisions_search_rankAt least 25% fewer follow-up searches after a capped result.pending
File prefetchdecisions_prefetchFewer tool rounds before the first edit, with at least 60% of prefetched files read or edited in the task.pending
Plan suggestiondecisions_plan_suggestAt least half of the suggestions accepted, and at most one suggestion per 20 small requests.pending
Correction signaldecisions_user_signalAt least 85% precision on labeled correction messages.pending
Skill and agent hintsdecisions_hintsHinted skills or agents used in at least 40% of hinted tasks.pending
Standing-rule suggestionsdecisions_memory_suggestAt least half of the offers accepted.pending
Repeated-question checkdecisions_question_checkAt least 90% precision on repeated-question fixtures.pending
Review suggestiondecisions_review_suggestAt least 30% of suggestions accepted, and at most one suggestion per 10 turns.pending
Pet mood from turn tonedecisions_pet_moodAt least 80% agreement with hand labels of turn tone.pending
Inbox prioritydecisions_inbox_priorityAt least 80% agreement with hand-ranked urgency, with no urgent item left without a notification.pending
Check selectiondecisions_check_selectCheck time cut by at least 30% with no failing check skipped.pending
Custom decision rulesdecisions_custom_rulesp95 at or under 250 ms, and no rule ever changes a permission (enforced by structure).pending
Secret screeningdecisions_secret_screenAt least 95% recall on credential fixtures with at most one false alarm per 100 tool results.pending

Decision layer tests and benchmarks ​

End-to-end tests in crates/z-engine-engine/tests/ (run by cargo test --workspace):

  • decisions_all_on.rs: every feature On against a decision model that always answers confidently with the answer that acts, over a scripted session (two tasks, approvals, a denied command, a stop with changes, a compaction, a question). It completes with no lost events, every call the policy asks about still asks, a denied call never runs or asks, no turn is cancelled and no turn is marked Verified.
  • decisions_fallback.rs: every feature On against a decision model that is down or never answers in time; the main model gets the same requests, the same calls ask and finish the same way, and verification ends as with every feature Off.
  • decisions_shadow.rs: every feature in Shadow against a confident model; the session runs exactly as with every feature Off while the trace records what each feature would have done, marked as shadow.
  • One file each for compaction, task view, first-request context, loop guard, secret screening and the repeated-question check.

The live A/B test (ignored) runs small fixture tasks against a real model twice, every feature Off and then the chosen features On, and prints tokens, cost and tasks solved per arm. Results: pending (not run; no credentials here).

bash
ZENGINE_MODEL=... ZENGINE_API_KEY=... ZENGINE_DECISIONS_URL=... \
  cargo test -p z-engine-engine --test decisions_ab_live -- --ignored --nocapture

Optional: ZENGINE_PROVIDER, ZENGINE_BASE_URL, ZENGINE_DECISIONS_KEY, ZENGINE_DECISIONS_ALLOW_REMOTE=1 and ZENGINE_AB_FEATURES (comma-separated ids; default every available feature).

The offline benchmark (crates/z-engine-decisions/tests/eval.rs) always validates the hand-labeled fixtures; with ZENGINE_DECISIONS_URL pointing at a laya-serve, its ignored run compares rules, the raw model and the hybrid gate per question kind, and prints a fitted [decisions.calibration.<question>] block per kind to paste into settings.toml. Results: pending.

bash
ZENGINE_DECISIONS_URL=http://127.0.0.1:8000 \
  cargo test -p z-engine-decisions --test eval -- --ignored --nocapture

Optional: ZENGINE_DECISIONS_KEY, ZENGINE_DECISIONS_CHECKPOINT, ZENGINE_DECISIONS_ALLOW_REMOTE=1, ZENGINE_DECISIONS_THRESHOLD (default 0.8) and ZENGINE_DECISIONS_CALIBRATION (question=temperature[:threshold],...).

Known gaps ​

  • Only the main conversation is secret-screened; titles, summaries and other side requests to the fast model are not.
  • The Claimed, not checked badge, route chips and suggestion cards are live only; they are not kept when a chat is reopened.
  • There are no operating-system notifications: inbox priority only sorts the Inbox and gates the user's own Notification hooks (they fire for high urgency or when unscored).
  • The native model pins are community ONNX exports (onnx-community/laya*-ONNX) at pinned revisions, not files published by Laya's authors; NOTICE covers attribution.
  • The native runtime is not the default, and release builds need --features onnx (the release workflow doesn't pass it yet).
  • Some parts act without the decision model by design: the loop guard's repeat detection, secret screening's pattern detectors and review suggestion's path rules. Everything that needs the model behaves as Off when it fails, times out or is unsure.

After 2.0.0: documentation website ​

Release checks (2.0.0) ​

  • cargo clippy --workspace --all-targets -- -D warnings: no warnings.
  • cargo test --workspace: 1268 passed, 0 failed, on two consecutive runs (the second alongside an extra engine suite for CPU load); the engine suite also passed five repeated runs.
  • Frontend: 303 vitest tests, svelte-check 0 errors, oxlint clean, production build succeeds.
  • Timing harnesses (release): a 5,000-message, 22 MB session opens in about 55 ms; token estimation over 2,000 messages takes about 4 ms.

After 2.0.0: documentation and review fixes ​

ItemStatus
User guide (docs/user-guide/)done
How Z Engine works (docs/how-it-works/)done
Documentation contract (docs/AGENTS.md), docs-maintainer subagent, CI link checkdone
Untrusted workspaces can only make settings stricter; sessions and the settings screen share one ruledone
AGENTS.md corrected: network exceptions, steps for adding a tool, updater listenerdone
Custom check ids in auto_checks, turning off built-in language servers, daily catalog refresh, fetch_extract wordingdone

Open findings from the user-guide review:

  • Settings, agents, commands and skills edited outside the app reach an open chat only when it is reopened or something is saved in Settings: there is no file watcher, and reloadExtensions has no caller.
  • Skill allowed-tools and command disable-model-invocation are parsed but have no effect.

Decisions and notes ​

  • Protocol and config TypeScript is generated by ts-rs into crates/z-engine-gui/ui/src/lib/protocol/ via .cargo/config.toml; regenerate with cargo test -p z-engine-protocol -p z-engine-config -p z-engine-llm.
  • z-engine-tools depends on z-engine-policy for the Action type used in permission decisions (documented in AGENTS.md).
  • The Tauri shell depends only on z-engine-engine, z-engine-protocol and z-engine-config; GUI-only queries live in z-engine-engine/src/engine/queries/.
  • v2 never modifies v1 config.toml files; it reads settings.toml and imports v1 settings. Spent cost stays in session totals after rewinds.
  • The v1 design documents have been removed; git history keeps them.