Changes merged to the main branch since the last tag. They are part of the source candidate, not of any published package.
Release record
What changed, and in which version.
Two facts sit at the top of this page: the newest published release, and the version the source tree currently declares. Everything below is the repository's own CHANGELOG.md, section by section.
- Latest published release
- v0.10.1 · published Oct 8, 2026Release page
- Source candidate
- v0.10.2 · unreleasedFull notes in CHANGELOG.md
v0.10.2
Added
- /pet on makes the animated GPUI whale the main terminal view, with the existing message box, queued messages and permission controls always available. F5 or /pet inspect opens streamed replies, errors and the current session's agents; Escape returns to the same pet view and draft. /pet off restores the ordinary shell. The selected view is remembered across launches, and the inspector can copy the last finished reply with c or its footer action. Motion preferences and unknown…
- A Terminal work dock (/workbar terminal) shows the model’s live PTY sessions in the current workspace, with ANSI-styled output, session selection, resize and direct keyboard input. Ctrl+N starts a separate shell; Alt+Down and Alt+Up switch sessions, and Esc returns input to the composer. Input remains bound to the selected shell’s identity across resets. Available on macOS and Linux; Windows names its current PTY limitation. The view shows a bounded output tail rather than…
- When a Plan turn completes with open To-do steps, the TUI asks "How do you want to continue?". Work (Ask) and Work (Auto-Review) switch to Work with that permission and send "Go ahead with the plan."; Keep planning, or Esc, stays in Plan and sends nothing; typed feedback is sent as the next message in Plan. The question appears only when the composer is empty, no message is queued, no other view is open and no agent is focused. If the chosen permission cannot be applied (for…
- /undo force undoes a request whose end was never recorded, including edits made to its files since it started. Any other /undo option is reported as unknown and nothing runs.
- /diff opens the changed lines in a pager titled "Changes since session start", rendered like an edit in the transcript, under the file list and stat it already printed. Patch text is cut at 256 KiB and left out past 200,000 changed lines; the message then says "The diff is too long to show in full. The file list above is complete."
- The model can stop a background shell command it owns: task_shell_wait takes cancel=true. That call needs approval, its card reads "Stop a running shell command", and it is refused when sent together with gate. The notices for a command moved to the background now name cancel=true.
- Each turn tells the model which configured MCP servers it may load, so a question can reach a server you never named. The note lists enabled, allowed server names only (at most 24), asks the model to tool_search a name before searching files or saying no tool exists, and is recorded again only when the list changes or a compaction dropped it. No server is started and no schema is loaded until the model searches.
- codewhale mcp --help lists the mcp subcommands with their descriptions and shows Usage: codewhale mcp [OPTIONS] <COMMAND>.
- A file-edit approval card in a folder with no Git repository says "No Git repository here, so edits ask first." on its question row. It appears only for write_file, edit_file and apply_patch in the posture where the same edit runs without a card inside a repository.
- Clients that start codewhale serve can declare a telemetry surface with CODEWHALE_TELEMETRY_SURFACE, including vscode-extension. This changes the session label while preserving collection settings and opt-outs (#6916).
Changed12 of 15 entries shown
- Goal mode has no step limit when [goal].max_steps is omitted or 0. Positive values still limit the goal, with values above 100000 clamped to 100000 (#6512).
- MCP and provider browser sign-in allow 15 minutes for the callback instead of five (#6865).
- /undo takes back your last request in one step, in every access mode: every file the request changed, plus the request and its reply. It no longer asks for /trust on. Its report names the request and lists each file as restored, brought back or removed, then says whether the request and its reply were removed from the conversation. It changes nothing if one of those files was edited after the request, and names the file. For a request whose end was never recorded it changes…
- /restore without trust mode or Full Access now says that it rolls every file in the folder back to the chosen point, including later edits, that nothing was changed, and that /undo takes back only your last request.
- /diff compares the workspace with this session's first restore point, so it works in a folder that is not a git repository and includes files created since. Before the session has a restore point it shows the workspace repository's uncommitted changes. With neither, it says "Nothing to compare yet: this session has not saved a restore point here, and this folder is not a git repository." instead of "No changes since session start".
- /trust on and /trust off leave a line in the transcript saying what changed; with --save the line also says what was saved for the folder.
- After Esc Esc the footer reads "Files not changed. /undo puts them back. Conversation rewound." in place of "Rewound to previous user message — edit and Enter to resend".
- /relay writes the session relay to .codewhale/handoff.md, the path a new session reads first, instead of .deepseek/handoff.md. A relay left at the legacy path is still read when the primary file is absent, and the prompt block names the file that was actually read.
- The highlighted option on an approval card is tagged (Enter), so the default [3 / d / n] Don't allow (Enter) row shows what Enter does.
- The edit approval card shows only the lines an edit changes, as - and + rows. Lines shared at the start and end of the old and new text are left out, the "replace this" and "with this" sub-labels are gone, and a single edit is no longer headed "edit 1". Lines left out on a side are counted (... (+2 more lines)); several edits are numbered and each shows its first changed line.
- The three-row preview of a successful command shows one opening row and two closing rows, preferring closing rows that say something passed or failed. Blank rows no longer take a slot, and lines hidden after the last row shown are counted in a "lines omitted" marker too.
- An MCP server the session has not needed yet shows as not connected in /mcp instead of "not started", with its transport and "Not connected in this session yet. Connects when a tool is needed, or connect it now." in place of zero tool counts.
Fixed12 of 19 entries shown
- ChatGPT sign-in omits the unsupported originator query parameter that caused an initial invalid_authorize_request rejection. PKCE, account binding and granted-scope validation remain in the same login flow (Refs #6925).
- The shared pet pauses its world and periodic checkpoint writes after its viewers, producer and audio leases expire. Authenticated clients wake it through the existing request queue; its active clock waits for the next frame instead of polling every two milliseconds (Refs #6728, #6155).
- Ctrl+B and steering input release foreground and background shell waits without stopping their commands. Requests apply only to active waits in the current session, including waits on several tasks; a later wait starts fresh (#6909).
- Concurrent starts and reconnects for one workspace converge on one Engine owner. A receipt retired during a reconnect is observed again before attaching; files with unsafe ownership, permissions or multiple links are still refused.
- Local plugin updates retain a private recovery backup while replacing the installed version. Plugin actions marked as requiring approval keep that requirement in Full Access, and internal installation paths are excluded from discovery.
- Concurrent Files API creates publish a workspace file exclusively. A competing creator gets HTTP 409 and cannot overwrite the winning file.
- Cache usage accounts for cache-write tokens as well as cache hits when deriving cache misses. Missing cache details remain unknown, and explicit zero counts stay zero (#6913).
- /profile usage, switching, success and failure replies follow the UI language across all 15 complete locale packs (#6919).
- On Windows npm installs, the model's environment identifies node.exe as the launcher and warns that killing it by name also stops npm-launched Codewhale sessions. It points to stopping servers by PID or port instead (#6906).
- Persistent terminal commands keep interactive stdin available while preserving current-shell variables, directories and aliases. Human answers no longer consume the command's completion marker. Oversized raw input batches are refused before writing in canonical line-input mode, across the Terminal dock, terminal/send and the Runtime API.
- A stream request that receives no response headers in time now adds that the provider may still be loading the model, and names codewhale config set stream.open_timeout_secs 180 (up to 300). codewhale doctor --probe-api reports a live check that ran out of time as "Not confirmed" with "API check got no answer in time", not as a failed credential, and no longer points at replacing the key (#6889, thanks @BX166).
- When the same model fails with the same upstream HTTP status on two or more requests in a row, the error adds how many times it has failed, that a provider can list a model that is not serving requests, and to choose another model with /model. The count starts over when a request opens its stream, fails another way, or uses a different model or status (#6889, thanks @BX166).
Maintenance
- CI and test harness only, no change to the shipped binaries. The persistence backlog budget takes five samples and judges enqueue time on the fastest one. Linux CI jobs install apt packages through scripts/ci-apt-install.sh, which bounds every wait and drops an unreachable Azure mirror.
Contributors
- @SparkofSpike — live network-policy updates and goal milestone hand-back (#6928, #6930).
- @dajiaohuang — cache-write token accounting (#6913).
- @gaord — embedder telemetry surface support (#6916).
- @Lstarsky0 — localized profile replies (#6919).
- @jayanthvee — Windows npm launcher guidance (#6906).
- @asto18089 — callback-window reference fix for longer browser sign-in waits (#6865).
- @BX166 — reported three provider failure modes measured on AICraft's own traffic: a cold model that reads as a dead connection, a listed model that fails on every call, and a response cut at the output ceiling that reports stop (#6889).
v0.10.1
Contributor integration and reliability12 of 21 entries shown
- OAuth retry diagnostics omit provider-controlled error fields; account replacement stops before retired Weixin receipts exceed their storage limit, preserving uncertain work for local review.
- Resuming a session through a symlink to the same workspace no longer shows a false workspace-change warning.
- Runtime clients can read one tool call's actual workspace changes and reviewed skill details (thanks @gaord, #6817 and #6869). In-flight snapshot pairs remain pending; missing objects and corrupt repository metadata are distinguished.
- Search accepts valid preferred locales, and image dimensions describe the same bytes sent to the model (thanks @asto18089, #6860 and #6858). Automation deletion keeps its definition until cleanup succeeds, and compaction preserves its original summary anchor (#6864 and #6857).
- Config/status/permission commands share portable contracts while the host retains mutation authority; queue workers acknowledge a scheduled retry for temporary first-claim contention and fail honestly on corruption (thanks @aboimpinto, #6832).
- Indefinite questions, approvals and elevation waits survive the TUI watchdog. Answers get time to resume the current turn; settled requests disappear by identity. Thanks @7jrxt42BxFZo4iAnN4CX for #6872.
- Configured approval expiry belongs to the held Engine request; hiding or covering its card cannot restart the deadline, and a late queued answer cannot approve an expired call.
- Tool discovery keeps the highest-ranked matches when a result batch exceeds the existing cache bounds, preserving search order and the 16 KiB limit (adapted from @AdityaVG13's #6393).
- Model-switch receipts now translate their session-only saving note in every complete locale pack; the three save commands remain directly usable (thanks @Lstarsky0, #6875).
- Route-save receipts and /workspace replies are translated in every complete locale pack, the Operate descriptions in fourteen packs match the current English, pt-BR, es-419 and ca regain accents six strings had lost, and long CJK text can be shortened between Han and kana characters (thanks @Lstarsky0, #6884, #6885, #6886, #6888, #6887 and #6882).
- Long sessions keep less in memory: superseded journal entries beyond a retained window move out of the live session into a per-session archive under sessions/.journal-archive/ (kept until the session is deleted; /branch can still restore an archived entry), ended shell operations in the work graph are capped, and session metadata larger than the first read no longer forces a full-file read. Not yet addressed: usage fingerprints in session metadata still grow without a bound,…
- Weixin bridge threads belong to the account and chat that created them: resuming or listing another chat's thread is refused. An explicit /new after an account replacement keeps the old account's private receipt and never replays its prompts automatically. Replies never reuse another bot account's Weixin context token, and /threads lists this chat's own older threads even when other chats have newer ones.
Added12 of 25 entries shown
- /plugin doctor reports stale built-in records and snapshots, and /plugin doctor --fix retires them. A user plugin, a snapshot a running process names, and a snapshot inside the grace window are kept. The previous state.json is kept as state.json.pre-gc.
- Reviewed plugins can declare named OpenAI-compatible OAuth routes. The host owns PKCE, refresh and credential storage, and checks the review at each request (docs/PLUGIN_PROVIDERS.md, #6805). The provider capability advances plugin review policy to v5 (v6 with the extension host): older receipts require explicit review again.
- OrcaRouter account sign-in uses PKCE and saves the same durable API key as manual setup; its live catalog keeps chat-capable rows and stated pricing and modality facts (#6867).
- Experimental TypeScript extension host. New in this release and off by default: turn it on with [features] extension_host = true. A plugin that declares a native TypeScript or JavaScript entry (the Cordis / DeepSeek Harness plugin model) can contribute tools, slash commands, pre-execution policy hooks, additive prompt sections, reviewed skill roots and plugin-local JSON state. Its code runs in a separate Node process, never inside Codewhale (Bun is an opt-in through…
- Extension slash commands: a plugin registers one with ctx.commands.register. It runs only when you type it, and it can show text, or submit a prompt as your next message that then goes through the ordinary turn and its approvals; it cannot call the model or a tool itself. A command cannot take a built-in command's name or another plugin's, and a markdown command with the same name wins. A command has 30 seconds, and commands are TUI only: the Runtime API does not list or run…
- Native extension authoring now includes policy listeners through ctx.on('tools/pre-execute', ...), prompt sections through ctx.prompt.registerSection, and persistent plugin-local state through ctx.storage. Listeners may deny, ask, revise input or annotate; Rust rechecks revised calls and owns every approval. Prompt sections cannot replace the system prompt, and state exposes no session history or secret API. These services share the experimental, off-by-default host's…
- Extension tool input is checked against the JSON Schema the tool registered, before an approval card appears and again before the call reaches the host. An invalid call comes back to the model as an error naming what to correct, and the host never sees it. A schema that cannot be compiled, or that refers outside itself with $ref, is refused when the tool registers.
- An extension tool can call eligible core tools with core/call while the ordinary turn is running that extension tool. Rust supplies a temporary invocation ticket and still owns planning, hooks, permissions, approval and execution. The ticket is bound to the plugin owner, host process and live invocation, and expires when that invocation ends or is revoked. Commands, timers and plugin startup cannot use this path. Approval grants are scoped to the reviewed extension build; a…
- Extension host identity is checked at startup: its reported trust tier and built-in module digests must match the process Rust launched. Built-in host code and third-party plugins have separate process and owner namespaces; the table pins the MCP protocol module and the execution harness. The optional Host MCP backend uses the official TypeScript SDK for framing, with Rust retaining credentials, network/process access, approval and exact operation tickets. All 30 existing…
- A plugin can declare several native entries (native.paths, up to 64). They activate in order under one owner, so one disable, review change or crash takes all of them down. If one entry fails to activate, only that entry is withdrawn: entries that already activated stay registered, and the plugin fails only when no entry activates.
- Plugin settings and context: [plugins."<name>".config] in your own config.toml is passed to the plugin's apply(ctx, config) and checked against its exported Config schema; a project's .codewhale/config.toml cannot set it, a change applies at /plugin reload, and /plugin show lists the keys, never the values. Tools and commands also receive the workspace of the session that called them and a private data directory for the plugin under ~/.codewhale/extension-host/data/plugins/;…
- codewhale --enable <feature> and --disable <feature> work from the codewhale command; before, they were rejected. Put the flag before any subcommand.
Changed
- The pinned prompt header follows the turn at the top of the viewport, handing over to the previous prompt as you scroll, and clicking it jumps back to the message it names (#6830, thanks @SparkofSpike).
- On Windows, every PowerShell command the shell tool starts now passes -ExecutionPolicy Bypass for that process only, so a local Restricted or AllSigned policy no longer blocks multi-line commands. (.ps1 script tools still start as powershell -File without the flag and remain subject to the local policy.) A policy set by Group Policy still wins and the command is refused with PowerShell's own message. Scripts that a command invokes run under the same process-scoped setting.…
- The website's not-found page now uses the Codwhale poster and typo joke, with English/Chinese recovery links to home and docs (#6419, #6420).
- Script tools can no longer approve themselves or replace built-in tools (founder decision D4). A # approval: auto line in a script under ~/.codewhale/tools (or [tools].plugin_dir) is ignored: the tool follows the session's approval setting like a script with no approval: line, and the runtime log and /plugin tools name each script that still declares it. A [tools.overrides] entry of type = "script" or type = "command" keyed by a built-in tool name is refused, and the…
- Complete the session slash-command group’s shared command boundary, including /structcopy, so all seventeen commands can compile independently of the TUI. Host operations remain behind capability interfaces; command behavior and upstream tool-execution identity safeguards are preserved (#6792, #6145).
- The declared minimum Rust version is now 1.89. It said 1.88, which could not build Codewhale: a locked dependency (serde-saphyr) needs 1.89, the version CI's minimum-version job builds. The install guides and the npm wrapper's build-from-source hints now say 1.89 too.
- A deferred tool's first call now runs when its arguments already carry every required field and no field the schema does not declare, instead of costing a retry turn. Malformed calls still get the schema, and approvals, deny lists and Plan mode apply first. Argument types are still checked by the tool itself. Sub-agents preload tools named in allowed_tools (#6494, #6437).
- Automatic Git status and review reads need Git 2.31 or newer, because they pin Git's runtime configuration overrides. With an older Git they refuse with "upgrade Git (2.31 or newer)". This covers the read-only Git tools (git_status, git_diff, git_log, git_show, git_blame, verify); Git writes you ask for keep their existing behavior (docs/dependency-maintenance.md).
Fixed12 of 96 entries shown
- A failure Codewhale can name is no longer labelled an internal fault. An HTTP 400/405/409/413/422 rejection, an out-of-credits 402, the context-budget stop and a turn's own step or wall-clock ceiling now carry an input or budget label, and a bare ERROR from a provider is reported as an unreadable error instead of a warning. Refs #6843.
- A transient upstream failure reported as an error frame inside a successful response is retried within the stream retry budget. When the budget is spent the turn fails once with an error card instead of an amber warning that promised a retry. Refs #6795.
- Diff lines and tool output wrap at grapheme boundaries, so emoji families, skin tones and variation selectors no longer split across lines (#6829, thanks @Lstarsky0).
- Twelve translated packs now translate the context inspector's making-room and anchors rows, the Ctrl+O hint and the /turn inspect and /advisor descriptions instead of showing English (#6831, thanks @Lstarsky0).
- Optional MCP servers can be found before they connect. tool_search matches the query against configured, enabled server names (or an mcp_<server>_ prefix), connects up to eight matches within the existing boot wait, and returns their real tool schemas. Before, a lazily started server exposed no tools, so the model could never trigger its connection (#6828). Known limit: a query that names neither the server nor mcp does not wake it. codewhale mcp connect, validate and tools…
- /undo and /restore <N> refuse while a turn is running in the workspace, instead of rewriting files under it.
- --resume <id> after a crash recovers that session's interrupted turn from its crash checkpoint, as --continue does.
- /resume <file> keeps an imported session on the current provider route instead of a default one.
- When the stall watchdog recovers a turn, the Engine's turn is ended too, so the next message is accepted (#6800).
- A turn you cancel, or one the stall watchdog recovers, releases input at once. A message still waiting to reach the engine is abandoned and its unsent text returns to the composer, instead of holding input for up to 60 seconds (#6800).
- A transient upstream failure that a gateway reports as an error frame inside a successful response ("Provider returned an empty response") is retried under the stream retry budget when nothing had streamed; authentication and invalid-model frames still fail at once (#6795).
- Long sessions no longer start every turn late. Once a workspace held more than 50 undo snapshots (around the sixteenth turn), each new snapshot rebuilt the whole snapshot history before the provider request, about 2.8 s per turn in a small workspace. Old snapshots are now dropped half a window at a time.
Security12 of 49 entries shown
- On Windows, a broad Node kill (Stop-Process -Name node, Get-Process node | Stop-Process, taskkill /IM node.exe, and wrapped, aliased or nested-shell forms) is held by a built-in safety floor: it is refused in Full Access, Auto-Review and Never, and asks in other modes. With the npm launcher such a command ended this and every other npm-launched Codewhale session without cleanup. Stopping a server by PID or port still works, and the Windows installer and archives have no Node…
- code_execution and js_execution now start their Python and Node interpreters through the same permission-aware launcher as other commands, so the session's execution policy applies to them. Before, an approved call ran the interpreter directly, and its working directory was the only boundary (#6820, thanks @Guan0923).
- A saved task without an auto_approve field is no longer treated as auto-approved. Updates accept HTTPS URLs only from the release-host allow-list, and the npm release-asset check bounds every request.
- Workspace profile and skill discovery, pasted-image and screenshot writes, configuration redaction receipts and legacy configuration migration reject linked paths where they could redirect access outside their intended roots. Additional .codewhale writers use the existing confined filesystem helpers. The VS Code file opener checks the file's real path, and PowerShell temporary scripts receive unpredictable names and owner-only permissions.
- Update development tooling to Undici 7.29.1 or newer for GHSA-w293-vg96-wgc3, brace-expansion 1.1.21/5.0.12 for GHSA-q2hr-2g5m-vwhr, and the VS Code extension's markdown-it to 14.3.2 for GHSA-253c-mchw-3w2r.
- Harden local runtime browser sessions, fleet SSH trust, agent continuation ownership, task gate approval, plugin tool registration, bridge action tokens, and release metadata credential forwarding. Browser sessions recover across reloads and new tabs, and stream tickets retry after transient failures. Fleet SSH known-host checks support OpenSSH 7.x and later. SSH host configs using host_key_fingerprint must migrate to known_hosts with verified host keys; the unsupported…
- Harden workspace instruction, note, and anchor file access with shared no-follow reads and writes. Compaction loads pinned anchors only from trusted workspaces. Validate registry skill names before selecting cache paths.
- Tighten approval, execution and endpoint boundaries. Computer control consent and script calls need an exact decision a person gave on that call's own approval card. Automatic Git status and review reads go through the sanitized review command and stay pinned to the checked Git executable. Config backups, config dumps, MCP listings and notification payloads share one sensitive-key vocabulary; structured exports now also redact keys ending in key. Python a child model writes…
- Chat bridges (Feishu, Telegram, WeCom, Weixin) accept an approval decision only from the person who started that turn, for the Runtime's current pending approval. After upgrading, approvals for turns that were already running are decided from the TUI, and WeCom /allow no longer takes remember.
- More file and network paths are confined. A task or automation created by a session without full shell access no longer inherits the host's shell default. The commit planner, oversized-paste backup and project harness notes do not read or write through links that leave the workspace. On Unix, audit and approval logs are created owner-only and are not opened through a link; Windows is unchanged. Lane ids must be plain names. A session file is refused when it records a…
- Nothing is written through a .codewhale that is a link. The workflow run journal no longer creates .codewhale/ or its ledger when it is opened, only when the first record is written, and reads and appends through the same confined helpers as other state. The sub-agent coordination lock refuses a linked .codewhale or state directory before it creates anything.
- On Unix, the runtime log is created owner-only and is not opened through a link; an existing log from an earlier version is tightened. The .reconcile.lock and current.json.lock lock files are created owner-only too.
Removed
- Flags, settings and tool parameters that did nothing are gone (#6516). --output-mode is hidden. It is still accepted, prints a warning, and is ignored.
- The dispatcher no longer exports DEEPSEEK_* copies of its CODEWHALE_* variables. A DEEPSEEK_* variable you set yourself is still read.
- lane start and workflow run --runtime vm|ci are rejected before a lane is created. Older lane records for those runtimes still load.
- The control socket's relaunch verb is removed; it always returned an error.
- The speech tool drops stream. stream=true used to fail; it is now ignored, a complete audio file is written, and the result no longer carries "stream": false. The finance tool drops market, and a call that still passes it has it ignored.
- [context].enabled, the seam-manager keys and tui.terminal_probe_timeout_ms no longer load; old configs that carry them still start. The [workshop] docs now describe bounded spillover instead of a synthesis sub-agent.
- About 2,650 lines of workflow code that nothing ran are deleted: the replay executor, the review-repair loop and experimental search. The replay_diverged status they produced goes with them. The isolated Runtime Chat prompt and the legacy YOLO alias list each have one owner now (#6517).
Experience
- Typing a first message with no model connected leaves a line in the transcript that says the message was not sent and opens the provider picker.
- First run picks a chat-capable Ollama model instead of the alphabetically first tag, and says plainly when no model is available yet.
- codewhale doctor leads and ends with one verdict and the next step, and gives the update command for how you actually installed Codewhale. Command-line usage and errors say codewhale.
- The approval card leads with a plain summary of the action, such as "Run cargo test", and shows workspace-relative paths. The footer labels its values.
- /status warns when the session's pinned model is no longer in its provider's live model list (#6035).
- Error messages give one true sentence and one next step. The TUI's English copy says agent, Fleet, Permissions and Work consistently, help lists one summary per row, provider rows without a key say "needs key", /setup says what it sets up, and the pet tank rests when it is offline.
- ACP clients can see the Permissions setting the server started with, including Full Access and how to turn it on, but cannot select it (#6310).
- GET /v1/commands tells clients each command's argument shape, so they do not re-derive composer behaviour from the usage string (#6230).
Fleet and agents
- codewhale fleet run <spec> --check runs every validation a real run would and stops there: nothing is created, launched or spent.
- A queued agent says why it is waiting, for example when launches are throttled after provider rate limits, and when it stops waiting (#6277).
- Read-only agents can run chained inspection commands (a leading cd, &&, ;, echo separators, 2>/dev/null), and a refused command now names the rule it broke and what to do instead. Durable Fleet workers accept the same read-only commands as in-session agents. An agent's time budget starts when it launches, and a queued agent that never gets a slot says it never started (#6015).
- Stopping an agent that writes files keeps and names the work it had changed, as a budget stop already did (#5529).
- workflow(fleet:) runs Fleets saved from the Fleet UI, and finds workspace Fleets under .codewhale/fleets.
- The runtime API can stop a delegated agent run from the desktop.
- A finished agent's answer is no longer cut off. Its row and its completion notification show the first sentence of its result instead of a ## Summary heading or its last tool, and opening the agent shows the whole result, or the full reason it stopped, even when no transcript was captured. Each agent also has one name: a workflow task's label or its dispatch name appears on the rows, the notification and the runtime API alike, never its internal id (#6565).
- The mobile page shows the thread's agents: a strip naming each one, its state, and what it is doing or what it found, rebuilt when the page reconnects. Sub-agent prompt caching now counts toward the session, including cache-write-only reports. PRICE and /cache show the parent, agents and combined hit rates, each labelled and weighted by all input tokens, including cache writes, and the footer cache N% still means this conversation's own requests (#6565).
- The dock's GIT, FILES and NOTES views are real. GIT shows the branch and where it stands against its upstream, the changes (with their paths one Enter away), linked worktrees and the last five commits, and it keeps updating during a turn while it is open. It says "not a git repository" only when that is true, reports failed status probes in the view and composer instead of claiming a clean tree, counts every unmerged path as a conflict, and only says there are more changed…
- Background work tells you when it ends, even between refreshes. Batched shell notices cover explicitly backgrounded or detached commands; foreground results and old completions from before this TUI session stay quiet. Switching sessions does not repeat a completion notice. Batched notices count completed, failed and stopped work separately; shell commands, task prompts and errors stay in the app, away from lock-screen notifications. A running dev server no longer holds the…
Plugins
- Codewhale no longer appends plugin recommendations to your messages to the model. Suggestions appear in one place, follow one switch and one budget, and never advertise built-in plugins, generic words or plugins for another operating system.
- /plugin dismissals lists the plugins suggestions skip, and /plugin dismissals reset [<name>] brings them back.
- Tools from reviewed plugins that declare themselves read-only no longer ask for approval on every call.
- The bundled Computer Use plugin is 0.12.0, the published upstream release 8435692 (#6303, #5856). On macOS the agent uses its own pointer and never drives your cursor. app_script refuses shell escapes. Clicks on irreversible actions such as pay, send or delete need confirmation. Consent decisions cannot ride inside run_actions or trajectory replay, and trajectories redact secure fields. Also new: a shared-computer control lease that pauses agent input while a person drives,…
- The bundled first-party catalog pins marketplace revision ae3dd2255a9a266365c6125a084f511eb26bc04d. It lists Computer Use 0.12.0 and adds Codewhale for Chrome (Chromewhale) 0.3.0 as a developer preview: you load its Chrome extension unpacked, and like every catalog plugin it installs disabled and untrusted until you review it.
Release reliability
- Upload the complete release into a private draft, verify every asset's size and SHA-256 digest, then publish. An interrupted retry cannot reuse stale same-size bytes. CNB and GHCR version tags follow canonical publication.
- Ubuntu Lighthouse bootstrap requires trusted SSH source CIDRs or an explicit public-SSH opt-in before changing the host; malformed IPv6 and broad default networks are refused.
CI
- State-touching command and session tests seal HOME, USERPROFILE and CODEWHALE_HOME onto temporary directories. Unsealed session, snapshot, artifact, composer-history, audit and built-in-plugin storage resolves into an isolated test root; unrelated worker threads cannot borrow another test's home seal. A child-process sentinel checks that known command-dispatch and session cases leave their ambient home unchanged.
- Fork pull requests stay under the macOS runner limit and the Actions cache stays under its cap.
- Release candidates and releases share one parity gate, and a release tag without a release-candidate receipt is refused.
- Budget ratchets block same-repository pull requests unless the pull request updates the budget with a receipt.
- A CodeQL advanced-setup workflow is ready for when the repository switches from default setup.
Contributors12 of 21 entries shown
- @AdityaVG13 — supplied the discovery-cache priority correction adapted from #6393, keeping highest-ranked tools through cache overflow. Its broader echo and fork-inheritance draft remains open.
- @7jrxt42BxFZo4iAnN4CX — reported indefinite questions cancelled by the TUI watchdog and supplied the timer evidence (#6872).
- @hodeswildsmith455-boop — added OrcaRouter account sign-in with PKCE and its live chat catalog (#6867).
- @LIghtJUNction — added reviewed plugin-provided AI routes with host-owned OAuth PKCE credentials and request-time authority checks (#6805).
- @Guan0923 — accepted case-insensitive HTTP(S) schemes in config doctor without rewriting the configured URL (#6819), and routed the Python and JavaScript execution tools through the session's execution policy (#6820).
- @harryvgiunta — added Yolo-Auto as a bundled OpenAI-compatible host, starting on the vendor's recommended qwen3.8-flash model (#6408).
- @asto18089 — contributed the integrated runtime liveness, context, search, JavaScript execution, stopship scout and pet repairs, preserving their original contributor commits (#6799); made context rule and chain-segment source labels repository-relative (#6739).
- @qiuYliangM — made provider-bound project instruction and constitution labels stable across directory moves and kept their absolute paths in operator reports (#6799).
- @zhuowp — supplied the process-scoped PowerShell execution-policy repair adapted for Codewhale, preserving machine and user Group Policy precedence (#6745).
- @Andrea-Bruno — designed the Superfast Decision Gate and contributed its off-by-default shadow classifier (#6604, #6603).
- @aiapienthusiast — added Cheaper Inference to the bundled provider catalog (#6761).
- @gaord — let a client fork a thread at a named turn (#6580), let undo roll back files for the turn it is undoing (#6483), stopped resume and fork from duplicating threads and sessions (#6406), exposed user-defined provider routes to native clients (#6404), and kept a fork going when a turn lost its tool call (#6664).
v0.10.0
Contributors12 of 19 entries shown
- @AdityaVG13 — fixed composer wrapping, tab/caret placement, pasted and editor-returned draft history, painted-column transcript copying, explicit terminal foregrounds, headless user-input tool availability, and engine synchronization after importing foreign sessions (#6369, #6363, #6365).
- @aboimpinto — moved the TUI session-export slice onto shared command contracts (FEAT-025): a session-export contract facet with one shared sanitizer, /export routed through the facet, pinned with baseline-captured goldens and gates (#6096).
- @BX166 — contributed the AICraft provider template and its documentation (#6171). It was closed unmerged, but it is what surfaced the decision to stop special-casing named OpenAI-compatible hosts (#6289).
- @7jrxt42BxFZo4iAnN4CX — reported the session-retention defects behind archive-past-the-cap and empty-session cap occupancy (#6136, #6137), the resume-failure design behind durable transcript errors (#6138), and the gaps behind the opt-in approval timeout (#6101), codewhale exec --hooks (#6099), Markdown drag-copy (#6156), and the browsable, current-aware session picker (#6014); the goal token-budget hard stop (#6013) and the fleet no-progress guard shared with child workers…
- @Lstarsky0 — reported TUI tests reading machine state instead of hermetic fixtures; the lock_test_env remedy from that report shaped two more hermetic fixes, for the shared UI fixtures and the compaction budget test (#5359).
- @Lujc0523 — reported /hooks edit splitting keystrokes between the editor and the composer, fixed by pausing the TUI input pump inside the editor handoff (#6165).
- @Statter — reported the Gemini /models failure that now surfaces the provider's reason instead of an empty error (#6173).
- @sequico — reported the ACP session/new ids that session/load could not resolve, fixed by minting resolvable session ids (#6174).
- @bevis-wong — reported the mid-run engine freeze behind the bounded turn-end foreground-child join, and the resume path that re-ran identical tool-call repair on every load instead of persisting it (#6184, #6185).
- @gaord — recorded the mode each turn ran in (#6321, harvested), stated the approval posture a task thread starts on (#6386), and rebuilt the runtime-API thread summary in one store pass (#6376).
- @zhuowp — preserved chat roles across compaction, protected user turns on recompaction, and kept the operate contract intact (#6286).
- @h3c-hexin — rate-limit-adaptive subagent launch scheduling: the DynamicGate that replaces fixed spawn pacing under provider throttling (#6055, harvested).
Security
- Children never inherit desktop or computer-control tools. Desktop control is the most machine-wide capability in the catalog, and a verifier child inherited it by default: on 2026-09-17 one opened the host Terminal and typed a blocked shell command into the user's live session. A family classifier now removes those tools when a child's registry is built, so they are neither eager nor searchable, and execute_full refuses the family at dispatch for every child role —…
- Runtimes can be held to an organization's plugin allowlist. A managed policy document (managed-policy.json beside state.json, or CODEWHALE_MANAGED_POLICY_PATH) lists the plugin ids a Runtime may run, with an allow_unlisted flag and a schema version checked exactly like PluginStateFile. Enforcement sits inside apply_state, so a plugin enabled before the policy arrived — or hand-edited to enabled: true — never comes back enabled, and enable() re-reads the document so a policy…
- Approving an apply_patch "for the session" is now scoped to the file you approved. The grouping key that scopes a session grant was built by a second, weaker patch parser that read only +++ b/ headers and the replace array: it saw no target at all for the documented apply_patch{path, patch} override, for --no-prefix diffs, or for delete-only diffs, and collapsed every one of them to a single shared key. One approval therefore pre-approved every later patch of that shape, to…
Added12 of 34 entries shown
- The statusline's performance readings survive compact mode and are separately configurable. Measured TTFT and average output rate are kept when space allows, shedding help text and secondary counts first; the existing metrics and statusline settings gain individual toggles with a migration that preserves a legacy single setting, and the picker takes mouse selection and scrolling.
- Tasks can be given their own run name. NewTaskRequest, TaskRecord and TaskSummary carry an optional name, stored as given and omitted when absent, so queues still fall back to prompt-derived titles; a run started by an automation inherits the automation's name. The tasks tool's schema extends backward-compatibly and legacy records decode through serde defaults (APPS-153).
- The offline catalog seeds Xiaomi's MiMo 2.6 family — xiaomi/mimo-v2.6-pro and xiaomi/mimo-v2.6-flash — so a first boot without network no longer shows the 2.5 generation. DEFAULT_XIAOMI_MIMO_MODEL stays on mimo-v2.5-pro, and the rows carry no price because the published rates cover only sk- keys.
- Native memory is a reviewed store, not a model-writable file. codewhale-memory backs the TUI with SQLite as the authority instead of Markdown, and the remember tool now only *proposes* candidates — a model can no longer write an active memory. /memory and the Runtime API commit through remember_reviewed, and a Context Lens surface (/v1/memory/lens, /lens/actions, /events) shows what was kept and why.
- Code mode (Experimental, default off): execute_tools runs a JavaScript program against a QuickJS host surface so a model can express several tool calls as one program. Nested calls must be read-only and auto-approved; the tool is hidden in Plan and refused under worker authority. Enable with [features] code_mode.
- Two new built-in providers: ZenMux (ZENMUX_API_KEY) and CSDN 星图 (CSDN_API_KEY, Coding Plan quota billing), each with its own key slot, bootstrap model and catalog rows.
- The Runtime API gained the surface a native client actually needs: jobs with stdin, kill and cursor reads; context, secrets, git, diagnostics, targets, LSP and voice routes; GET /v1/commands for the slash-command catalog; GET /v1/workspace/instructions; account-wide GET /v1/approvals; plan and to-do inventory; /v1/settings/schema, with POST /v1/config now persisting every declared settings.toml key rather than a curated allowlist; PTY byte replay, resize and exit; and tool…
- Codewhale holds the host's idle-sleep assertion while a turn is in flight (caffeinate -i on macOS, systemd-inhibit on Linux), so an unattended machine no longer sleeps mid-run. You will see one child process per turn.
- Skills are reachable in one call. The pinned ## Skills index told the model to call load_skill, but the tool was deferred behind tool_search, so it was never in the tool array that instruction was printed beside: using a skill cost a discovery hop, a name="list" round trip, and a change:tool_surface re-pin of the whole cached prefix. load_skill is now eager for the parent and for children — +244B of pinned catalog, measured, against a re-prefill avoided every time a skill is…
- A fast lane the router cannot serve now says so. provider_router_candidates answers cheap: None for any pair its tables do not know, and a Faster/Auto child on such a pair used to run at the parent's model and price with no receipt and no way to tell "single tier by design" from misconfiguration. The fallback stays — the router must not invent a model — but the spawn receipt now carries a fallback_note naming the unserved lane, the same field the pinned-provider fallback…
- The runtime API tells a replayed submission apart from a new admission: POST /v1/threads/{id}/turns answers 200 with idempotent_replay: true when the operation key was already used, and 201 for a fresh admission, so a client that retries after a dropped response can no longer create a second turn. The flag is absent on a fresh admission, so responses existing clients already parse are unchanged (#76).
- Event streams resume where they left off: journal frames carry their durable seq as the SSE event id:, Last-Event-ID is honoured as the cursor when no explicit since_seq is asked for, and RuntimeCapabilities advertises event_stream_resume so a client can gate its reconnect controls on the capability instead of discovering it from a missing id (#76).
Changed12 of 24 entries shown
- SearXNG results rank by the score the instance returns rather than by arrival order. Integers, floats and numeric strings are accepted; anything else, including NaN and infinity, becomes 0.0, and rows are stable-sorted descending before the result cap, so equal scores keep instance order. The docs now spell out the self-hosting requirement: the separate process must expose search.formats: [json] — an HTML-only instance answers 403 — and bind to loopback or a policy-allowed…
- The bundled skills move to a new generation. social-media and health leave the shipped pack (phone-export workflows rather than everyday skills), feedback moves to docs/skills/ beside contributor onboarding, and the exact earlier body is retained so only unmodified shipped skills upgrade — a skill you have edited is left in place. Google OAuth scopes and client setup, Photos exports, forgetting limits, Spotify playback and plugin reload guidance were corrected in the same…
- auto is a declared default, not a guess about your wording. Reasoning effort no longer maps request vocabulary to tiers (debug/error to Max, search to Low) and Auto routing no longer infers cheap-versus-big from phrasing: both resolve the configured default, with [auto] cost_saving as the explicit opt-in. The same prompt now costs the same thing twice.
- A workflow's shared token budget is opt-in. [workflow] default_token_budget applied a silent 120,000-token cap across a run and all of its children, and a fan-out died at the limit with no hand-back; the default is now 0, meaning no shared cap.
- The shipped deepseek-flash route speaks the Responses endpoint, and xhigh effort maps to high per the vendor's own table.
- Streamed text is paced at a steady rate rather than inheriting the provider's SSE chunking, so output reveals at a readable beat instead of in bursts.
- Broadening shell access for a conversation now requires an idle conversation and rejects a stale-workspace check before it commits. Engines advertise a thread_shell_consent capability, so an older Engine reads as unsupported to a native consent client instead of silently accepting.
- One base prompt now serves every host. HEADLESS_BASE_PROMPT was a second hand-maintained rendering of the same constitution — the drift pattern this repo forbids by convention — so headless runs compose the same BASE_PROMPT plus language and output layers that interactive runs do, skipping only host chrome (execution profile, authority recap). Headless and interactive can no longer disagree about what the agent is.
- BASE_PROMPT gains a Bearing article, which changes how the agent talks to you: the user is a peer who gets honesty rather than deference, a blocking gate is named plainly instead of dressed up as refusal, bad code is called bad, a crude request is carried out without a lecture, and an apology appears only when there is something to apologize for — not as punctuation. It also states that the request is the whole mandate, so the scope law opens with what is yours to do before…
- The ocean reads as animals rather than a mechanism. The school used to translate as one rigid body — bob phase and tail pose were staggered per fish, but horizontal position was locked to an exact wedge offset — so each fish now eases a dot fore and aft of its slot on its own slow period, and the formation breathes while it travels. Bubble emission and the per-animal periods are hash-jittered instead of sharing one clock. Amplitude stays an order of magnitude under the…
- Extensions keeps the exact-content plugin review on the panel: confirming a bundle's digest re-reads the inventory, so the row you just reviewed reports its new trust state and offers Enable instead of leaving you in the transcript with a stale "not reviewed" row.
- Underwater motion ticks at the cadence the frame limiter actually draws (the atmosphere interval while only the water moves, the authored 80 ms ocean cadence inside the interactive cap while a turn streams), and the event loop wakes exactly for the next tick instead of on the next idle poll. Idle water no longer requests frames it cannot draw or quantizes its cadence to the poll interval; reduced motion, Ghostty, tmux and the six-second idle settle are unchanged.
Removed
- The host no longer parses prose into goals. Ten phrasings and a clause allow-list ("make it your /goal to …") were turned into durable goals before the provider call; prose now reaches the model, which calls create_goal when a goal is useful. The deterministic path — a leading /goals <objective> — is unchanged.
Fixed12 of 46 entries shown
- A thinking fold is an absolute choice again. The stored bit was relative to the display preference (folded ^ !(verbose || thinking_default_expanded)), so every recorded choice flipped meaning the moment a preference changed: turning thinking_default_expanded on closed a block the reader had explicitly expanded (#5847). An explicit tri-state intent now records Expanded or Collapsed outright, and the absence of an entry means untouched, so the preference baseline decides that…
- A terminal byte-stream cursor past the head is clamped instead of echoed back. read_since returned a future cursor as next_cursor, so a client that continued from it skipped every byte the stream produced before reaching that position — permanently. The start position now clamps to total: a future cursor reads nothing, is not a gap, and hands back the head.
- Language-server startup is bounded and a failed transport now terminates. The client waits for successful initialization before sending notifications, drains stderr without buffering it, bounds request queueing and replies, caps protocol frames, and fails pending requests when the child transport dies.
- Input no longer freezes for the rest of a turn when the engine's 32-slot op mailbox is full. The remaining input-path sends no longer await: droppable ops whose rejection is reported and retryable use try_send (CancelSubAgent, PreviewOutboundRequest, bang shell input, PurgeContext, and the single-op settings updates), ops that must land once the UI changed reserve first, and ChangeMode publishes its live authority even on a full channel. Must-deliver ordered transitions…
- The pet's whale is one body again. The same authored point set is checked in three copies plus four fixtures, and the Rust and TypeScript cores disagreed on particle positions from frame 0 while agreeing on every channel and constant — so the v1 fixtures had never matched what Rust produced and the conformance job had never passed. The bodies and the fixtures are reconciled and that check now runs green.
- The context meter and the auto-compact gate share one honest estimator. The status bar inflated ctx % by about half and disagreed with the gate, so "ctx 82%" could sit beside a /compact that refused to run; displayed percentages now read materially lower because they are correct (#6297).
- A steer the engine never accepted is queued for the next turn instead of being shown as held and then silently dropped with no turn and no answer (#6297).
- Every reqwest client routes through codewhale_release::tls. A bare Client::builder() panics under rustls with no installed provider; 17 call sites were swept.
- The bundled OpenAI-compatible hosts have their /provider rows back. Retiring the ProviderSetupTemplate layer moved SenseNova, Baseten, Groq, Cerebras, DashScope and Command Code into provider_descriptors.json and then wired that file to nothing, so six vendors silently lost their picker rows; AICraft never had one. Each descriptor is a row again, built through the same named-custom-provider builder a configured host uses, so endpoint, bootstrap model and "missing <ENV>"…
- A StepFun Step Plan subscription reaches its own catalog. A subscriber's base_url is https://api.stepfun.ai/step_plan/v1, but only the pay-as-you-go /v1 host was recognised as official, so the Step Plan host read as a custom endpoint, the catalog was withheld, and the picker showed 0 bundled and a guessed context window for a route whose console advertises step-5-preview at 1M context. The model was in the seed the whole time. All four StepFun hosts — global (.ai) and China…
- The send cue stopped strobing while you type. [↵] flickered between dim [·] and bold blue once per character at an ordinary typing pace. The cue was not lying — Enter really does insert a newline during the ~120 ms paste-safety window, which every keystroke re-armed — but that window only exists for terminals that deliver a paste as a burst of ordinary keystrokes. Ghostty, iTerm2, WezTerm, Windows Terminal and Terminal.app now skip the heuristic from the first keystroke…
- Shift+Tab sets the permission posture in Plan. Tab cycles the mode and Shift+Tab cycles Ask/Auto-Review/Full Access, but Plan refused the second one outright, so the key silently did nothing there and the two axes read as welded together. Plan's read-only guarantee comes from the mode, not the posture — authority maps (Plan, _, Bypass) to SandboxPolicy::ReadOnly and tool_catalog gates every write tool on mode != Plan — so the cycle now moves the durable Act/Operate baseline…
v0.9.13
Added
- /pet turns the terminal over to the Codewhale pet. /pet on (or bare /pet) gives the habitat the whole content viewport now and on every accepted turn, reveals the actual answer or error when the turn completes, and Escape returns to the composer without cancelling anything. /pet off closes the view and stops automatic entry while the durable companion keeps the pet alive; /pet appearance|window|source|sound|export|status address the shared companion. The pet no longer lives…
- Codewhale Computer Use 0.3.1 ships as its own notarized Mac app. Download the disk image from codewhale.net/computer-use or the v0.3.1 release (Codewhale-Computer-Use-0.3.1-macos-universal.dmg, drag into Applications; the ZIP stays for the in-app updater). The bundled plugin and the first-party marketplace pin the same 0.3.1 sources, so the app, the computer-use plugin and /mcp see one implementation.
Fixed12 of 90 entries shown
- The website's Computer Use download page resolves its state without the GitHub API (using GITHUB_TOKEN only when bound), and every page regenerates on the Worker again: the Open Graph image route read brand SVGs at import time, which the Workers runtime cannot do, so codewhale.net had been serving its build-time snapshot.
- Operate can run structured workflows directly, with named phases, model assignments from Fleet, prerequisite results and shared budgets. Independent steps run together; dependent work waits for its required results and gates. Detached runs return their outcome to the owning conversation, and headless sessions stay alive between phases until the final handback is consumed.
- Computer Use 0.3.1: mouse actions no longer steal focus or reclaim the foreground when the user switches apps mid-action; background typing, scrolling and selection use semantic input, and screenshots stay scoped to the targeted app. The bundled plugin and the first-party marketplace pin carry the same 0.3.1 sources. A registered macOS helper stays in charge of input through its Pause and Stop controls; an unavailable registered helper produces an error instead of silently…
- The Fleet editor uses the standard model picker to manage sub-agent model and thinking assignments. Enter edits the selected row without changing the running session's model. Unconfigured providers are refused, failed saves retain the previous assignment, and a changed or removed team file must be reopened before a pick can overwrite it.
- The provider catalog includes Baseten and the other compatible-provider templates as selectable rows, opening their existing prefilled setup forms. DeepSeek routes with clock-based pricing show the current peak or off-peak tier beside session cost, with translated labels.
- Extensions, teams, workflows and automations support mouse-wheel scrolling. Plugin and MCP rows have keyboard enable/disable controls and two-step removal; MCP OAuth can retry with narrower scopes after a scope rejection.
- Healthy sub-agents continue after an ordinary parent reply. Headless runs keep the existing Engine alive for child results within the run deadline. Explicit cancellation remains authoritative when result queues are full or a completion starts a followup turn.
- Sub-agent followup supports multiple targets and all parked children, keeps old IDs connected to their current continuation, and saves continuation identity before starting work. Repeated followup does not fork duplicates.
- Sub-agents validate declared output files and distinguish real edit claims from file citations and unrelated workspace changes. Disjoint file claims can run together; overlapping writers receive the actual conflict and remedies. Explicit read-only shell analysis requires an enforcing native sandbox and refuses execution when that protection is unavailable.
- Delegation depth stays absolute through saved profiles, nested workers and continuations. Per-call token, step and time limits narrow inherited limits; continuation retains ancestor usage and deadlines. Workers reserve room for one tools-disabled partial report inside those limits, then run the declared- output checks. Missing usage or unavailable reporting room produces an explicit fallback; partial work is never marked complete.
- Agent rosters and detail pages have bounded output, visible continuation and descendant relationships, and usable handles for full diagnostic evidence. Completion receipts include measured worker and descendant token usage, count each continuation once, and distinguish unreported usage from zero.
- Localization and native helper builds resolve the active checkout when the build script runs, so a shared Cargo target keeps working after a worktree moves or is removed.
Changed12 of 13 entries shown
- How long Codewhale waits for a human is configurable. [tools] user_input_timeout_seconds governs the wait for an approval decision or a request_user_input answer; it was a hardcoded 300 seconds, which silently cancelled the work of anyone who stepped away mid-task. An explicit 0 waits indefinitely, the value is clamped to 24 hours, and omitting the key keeps the previous 300-second default. Documented in docs/CONFIGURATION.md (#6003).
- docs/PROVIDERS.md lists every beginner setup template, not the four it happened to mention when the page was written. Baseten, Groq, Cerebras and Command Code have shipped as supported OpenAI-compatible hosts for a while and appeared nowhere in the provider documentation, which reads from outside exactly like not supporting them. The page now carries the full table — host, default model and key env for each — and states the rule it follows: a plain Chat Completions backend…
- Reasoning capability for the Kimi coding routes and the qwen3.x Model Studio deep-thinking ids is catalog data now rather than hardcoded match arms, and model_reasoning_capability reports a model nothing knows about as unknown instead of silently not reasoning-capable. model_supports_reasoning keeps its bool shape for existing callers, where unknown still reads as false. The ids that have no cited source yet keep their literal arms (#6032).
- The website uses Shannon Sans with versioned local font assets and retained serif, monospace, and language fallbacks. Terminal fonts are unchanged.
- codewhale metrics reports recorded model requests and stream recovery separately from provider-reported token usage, with coverage for missing and duplicate receipts. Status messages and cumulative snapshots do not add requests or count tokens again.
- Runtime turn receipts retain the Engine's terminal model-request, stream-retry, and resume counters separately from displayed status and provider-reported usage. These counters do not count HTTP retries inside a provider client or establish provider billing.
- Initial tool definitions no longer repeat shell interpreter guidance and agent lifecycle/scope instructions in multiple description fields. Parameter schemas, approval rules and dispatch behavior are preserved. This reduces prompt schema size; it does not establish a provider billing regression.
- The built-in Computer Use plugin bundle is refreshed to the standalone plugin's 0.2.1 runtime (vendored from Hmbown/codewhale-cu-plugin at 724ad258): the native macOS accessibility backend with an a11y-first pointer strategy (covered points are refused, previews are drawn), the permission-owning desktop-app socket transport, remote computers over ssh and HarmonyOS HDC with contained temp handling, truthful win32 PowerShell failure reporting, and the shared allow-listed…
- /statusline drives the bottom chrome again. Since the 0.9.12 shell redesign the posture bar and the metrics line were built independently of tui.status_items, so every toggle in the picker except the balance fetch was decoration. Each remaining item now shows or hides exactly one thing: model, context_percent, cost, balance, cache, tokens and session_metrics are metrics-line segments, and mode is the posture bar's plan/act/operate chip. The status, agents, reasoning_replay,…
- The context reading is back on screen at every fullness. 0.9.12 painted ctx NN% only from 50% up, which left most of a session with no context signal at all; it now paints from 0% and keeps its warning colour from 80% up (#5950).
- A child agent parked because its parent's turn ended is shown as parked in the Agents panel, the sidebar and Agent Details, with resume_from / cancel as the recovery, instead of wearing the same "waiting for input" label as a child that asked a question. Parked work sorts below live and answerable work and no longer inflates the blocked chip; the receipts roster and the wire state gain parked (#5906, #5921).
- codewhale account keys set|remove|list no longer carry a hardcoded eight-provider list. Provider ids come from the control plane's public catalog (GET /api/model-providers), are validated locally against ^[a-z0-9][a-z0-9-]{0,63}$ before they reach a URL path, and list shows every catalog provider with its label and stored-key state. --from-local maps a catalog row onto the local runtime provider through the catalog's own runtimeProvider field, so a newly supported provider…
Fixed
- Five of the load-flaky tests tracked in #5929 no longer depend on shared state or live local daemons. Background-hook capture tests wait for the capture file to hold bytes instead of merely existing (the shell's > redirection creates the file empty before cat writes, which read as valid JSON: EOF under load); the session-picker acceptance test drives the real picker over a private store instead of a process-global CODEWHALE_HOME redirect that concurrent tests could observe…
- The posture bar states how long the session has been working and how long the current turn has run, distinguishing actively working from waiting on a tool, a sub-agent or the operator; the 0.9.12 shell had dropped the overall working-time indicator from the place a glancing user checks (#5914).
- A background runtime turn whose own store record could not be read, parsed or written (Failed to read turn …, Failed to read item …) was only a log line. The runtime now publishes a runtime.store_failure event naming the file, the root cause and the next action (move the file aside, or check free space and permissions); the TUI shows it as a warning toast and a transcript line, the task timeline records it, and the runtime API streams it. A turn whose own record is…
- An MCP token refresh that fails to parse the provider's answer keeps the endpoint's receipt — status line, content type, and a 200-byte excerpt with every credential-shaped value (access_token, refresh_token, client_secret, id_token, bearer schemes) masked before the cut — instead of rmcp's bare Failed to parse server response, so a provider outage answering an HTML 502 reads differently from a parser defect, and the login remedy stays named (#5926; remedy wording landed in…
Added12 of 22 entries shown
- POST /v1/threads/{id}/file-revert restores exactly one file from the exact tool:/pre-turn: snapshot the client selected, checking the reviewed file hash before and after the mandatory safety snapshot. Literal file names, regular files only, thread trust and active-turn admission are enforced, and patch-undo no longer forks a conversation whose file rollback failed (#6111, thanks @gaord; engine half of HengQuWorld/CodeWhale-VSCode#3).
- Authenticated Runtime API workspace file suggestions reuse TUI @file matching and discovery, with bounded queries/results and workspace-contained relative paths only (GET /v1/workspace/files/search, #6095, #6120, thanks @wuisabel-gif; reported by @LmeSzinc). Shared discovery now honors disabled symlink following for AI-tool directory scan roots too.
- Serply is available as an opt-in [search] provider for the Web tool (provider = "serply", key from [search] api_key or SERPLY_API_KEY). Preflight fails closed without a key; Firecrawl remains the default and existing configurations are unchanged (#6100, thanks @googio).
- Linux terminals: finishing a transcript or composer mouse selection copies the text to the PRIMARY selection without touching the regular clipboard, and middle-click inside the composer pastes PRIMARY at the pointer without submitting. Native X11 and Wayland data control are used through one bounded background worker; SSH sessions without a forwarded display keep their terminal's own selection behavior (#6116, thanks @dmt4).
- codewhale sessions export <id-or-unique-prefix> saves a .tar.xz archive with the durable record, portable session container, manifest and artifacts. Prefix exports preserve unfinished tool calls; confined reads reject linked artifact roots, and existing outputs require --force. Archives retain unredacted content; /load opens the extracted record without installing extracted artifacts (#6056, thanks @h3c-hexin and @asto18089).
- deepseek-flash (DeepSeek V4.1 Flash: 1M-token context, reasoning and tool calls) joins the catalog as DeepSeek's declared default, and the offline catalog seed matches it; the DeepSeek Pro listing no longer overstates the published price (#6025).
- DeepSeek's September 11 reversal is reflected in provider notices and cost estimates: V4 Pro remains available after September 14 at Pro rates. Explicit Pro selections remain unchanged; Flash remains the default (#6025, thanks @ronohara).
- Native plugin authoring guides now cover English and Chinese. The explicit offline converter supports selected portable Skills and static Streamable HTTP MCP declarations from OpenCode and DSH. Unsupported executable hooks, automatic OAuth and policy-bearing configurations are refused; generated bundles still require native installation, review and trust. Legacy SSE fallback is not reproduced (#5827, requested by @giancarlocp).
- Signed cloud model facts can refresh provider capabilities and prices while preserving verified cached data when a refresh fails. A dispatched request keeps its selected price snapshot so later catalog updates cannot change its recorded cost (#5752).
- Saved sessions preserve exact provider routes. Auxiliary model calls settle their usage once against the route and price snapshot that executed them, including recovery, rather than resolving a new price at completion (#5726, #5848).
- [tui].posture_bar and [tui].metrics_line accept full, compact, or hidden, also available through /config. Compact preserves the existing rows' essential fields; hidden returns their space to the transcript (#5973).
- Optional model-bound tool-output redaction opt-out, with two explicit startup confirmations and a receipt bound to the readable config contents and modification time. Unconfirmed requests keep masking enabled; routing and stored goal summaries remain redacted (#5982, thanks @SparkofSpike).
Contributors12 of 23 entries shown
- @LmeSzinc — requested Runtime API access to the TUI's fuzzy file search (#6095).
- @googio — added the Serply web-search provider (#6100).
- @dmt4 — requested Linux copy-on-select and middle-click paste (#6116).
- @Gabriel-Degret — reported that saved agent profiles were silently ignored when spawning sub-agents (#6117).
- @nightt5879 — Gemini signature recovery guidance and transport regressions (#6081).
- @c020627 — Chinese documentation link repairs (#6080).
- @h3c-hexin and @asto18089 — GLM-5.3 reasoning controls and tool-gating/documentation fixes (#6051, #6052).
- @Hmbown — dependency updates (#6057) and the Gemini signature recovery report (#6048).
- @gaord — contributed the file-scoped restore endpoint and the trust-gated whole-tree rollback (#6111), Fleet schema inspection, role precedence and worker deliverable receipts, and linked the community VS Code frontend (#5944, #5945, #5946, #5992).
- @goransh-walia — contributed the propose-only commit-planning rework (#5870).
- @7jrxt42BxFZo4iAnN4CX — documented turn budgets and goal configuration, and reported gaps in command discovery, Fleet navigation, human waits, state hooks, history and provider routing (#5996, #5952, #5954, #6003, #6004, #6006, #6007).
- @SparkofSpike — contributed two-stage consent for opting out of model-bound credential redaction (#5982).
Notes
- DeepSeek V4 Pro continues after September 14. The vendor reversed its earlier retirement notice. Codewhale preserves Pro selections and Pro pricing; deepseek-flash remains the default for new direct DeepSeek configurations.
- Upgrading from 0.9.12 with Computer Use trusted and enabled: the bundle's content hash changes with the 0.2.1 refresh, so the plugin deactivates and asks for a fresh review — that is the designed fail-closed path for a desktop-driving plugin. Re-trust it from the Plugins page.
- The multiline-paste fix restores v9.11 behavior on terminals that accept EnableBracketedPaste but deliver pastes as keystrokes (reported on Windows 11 / PowerShell). Verified at the input-contract level and in CI; a manual paste check on a real Windows terminal is still welcome — please comment on #5981 with your terminal if anything still misbehaves.
v0.9.12
Added
- Computer use ships with the binary. The computer-use plugin — 38 tools across macOS, Windows, Linux and HarmonyOS, accessibility-first observation with pixel fallback, screenshots, zoom, screen recording, and registered remote computers over ssh and hdc — is embedded in Codewhale and written to $CODEWHALE_HOME/builtin-plugins on first run, so every install channel carries it. It lists as builtin · not-reviewed and stays disabled until you review and enable it: shipping it is…
- Alibaba Model Studio joins the data-driven provider table as an openai-compatible descriptor: international compatible-mode endpoint, DASHSCOPE_API_KEY credential, live /v1/models discovery. Qwen 3.8 Flash and Qwen 3.8 Max arrive through the catalog authority — never a hard-coded id.
- Concentrate: first-class opt-in BYOK Responses gateway with live models discovery and typed SSE streaming (#5725).
- Cloud dispatch: remote runner offloads coding agent tasks to isolated cloud sandboxes with machine token auth and structured job tracking (#5701, #5712).
- Per-session control socket: config-gated [control_socket] table binds <sessions-dir>/<session-id>/control.sock per running session, exposing message, interrupt, relaunch, and status JSON-RPC verbs (#5533, #5831).
Changed12 of 26 entries shown
- Anonymous usage counting is on by default. The 0.9.11 release asked first; 0.9.12 counts the same aggregate version/platform, session, feature and error totals unless you turn it off, and says so once at first launch (policy notice version 5, schema 3, notice_version replacing consent_version). Every recorded opt-out stays off: a durable telemetry = false, a decline recorded under the old opt-in notice, unreadable privacy state, and the CODEWHALE_TELEMETRY=0 / --telemetry…
- The launch screen is our own card take: a thin top line ⑂ branch path; a centred bordered card with the whale mark, Codewhale + version, one announcement line only when it is true (the no-model warning, or MCP news), and the menu New worktree / Resume session / Changelog / Quit with their real chords right-aligned. Enter runs the highlighted entry, Up/Down move it, and typing goes straight to the composer. The card dissolves on the first keystroke or command (≤240 ms,…
- The work surface sits under the composer by default, keeping history readable and leaving the stage unencumbered (#5809).
- Skills command shapes: FEAT-022 command shapes and retained-host validation (#5825, #5829).
- After the card dissolves, the working screen shows ⑂ branch path with ⋮ MCP n/m on the right, the transcript starts with the ◆ session_start receipt (naming the configured session-start hooks), and the composer's bottom rule carries model (effort) · permission — the route's one launch reading. The posture bar and metrics line appear only once a session exists.
- Vocabulary: fleet is the public term and Pod is retired from copy — roster, setup, detail, worker-runtime and managed-API messages now say Fleet (/fleet canonical, /pod alias). The workflow wire accepts the canonical role spellings (general/explore/planner/reviewer/implement/ test/advisor) with the pre-rename ones kept as load-time aliases, and serializes canonical names.
- The Operate mode-picker hint is shortened to fit 80 columns.
- The footer always shows the permission posture; when only one chip fits, the permission chip outranks the mode word (#5796).
- Local Ollama: the header names a model only when the local catalog can serve it, and says unknown until it knows. The startup mark, web and app icon carry the new side-view prompt-eye whale (#5795).
- One focus owner: Tab and Shift+Tab work regardless of what is in the composer; Alt shortcuts survive mid-draft; Ctrl+Tab no longer cycles the mode by accident (#5798).
- Tool cells carry their own state: a running, failed or warned tool reads as such in the transcript itself, with per-entry rail dots and family-coloured glyphs (#5799).
- Web: docs hub with task search, shared empty/loading/error states, an offline-to-back-online banner, /changelog in every locale, and real 404s with correct metadata (#5743).
Contributors12 of 22 entries shown
- hexin (@h3c-hexin) — provider-native web search across four routes (#5682, #5683, #5685, #5687), authoritative edit-last-turn boundaries (#5621), Kimi Code k3-256k (#5622), post-compaction input-token reporting (#5623), and preserving the scheduled model selection in automations (#5650).
- 秋月凉梦 (@qiuYliangM) — co-authored the edit-last-turn boundary fix (#5621), Kimi Code k3-256k support (#5622), and post-compaction input-token reporting (#5623).
- Isabel Wu (@wuisabel-gif) — live session token totals (#5624), persisted context-pressure warnings (#5629), discoverable Fleet roster editing (#5604), the capability-gated cursor accent (#5599), and /copy for the latest completed response (#5692).
- Paulo Aboim Pinto (@aboimpinto) — Windows verbatim-path operands preserved through POSIX word splitting (#5610), the plugins group moved onto the command shapes (#5657), FEAT-022 skills command shapes with retained-host validation (#5825), and FEAT-020 plugin command shapes re-landed on main (#5865).
- Alex Musichen (@musichen) — a stable DeepSeek heading in the configured-view model picker, keeping every official catalog model for the active provider visible (#5689).
- @gaord — the GET /v1/fleet/profiles runtime API endpoint, reusing the FleetManager validation path (#5688).
- Sh1Zuku (@SparkofSpike) — corrected English documentation inaccuracies and the first zh_hans translations for the Tier-2 docs (#5613).
- @M-Maciej — goal continuation cadence (#5591) and the per-session control socket (#5533, #5831).
- Serephus (@serephus) — nixpkgs update (#5669).
- @whp233 — wire = responses|anthropic for openai-compatible custom routes and opencode-zen muse-spark (#5716, landed as #5719).
- Gabriel Degret (@Gabriel-Degret) — found the reasoning-only retry gap and built the first fix; landed as the [reasoning_only] retry ceiling with a request-scoped nudge (#5867).
- @huangxianzhan — the x-opencode-session header for OpenCode Go and Zen gateways (#5868).
Added12 of 38 entries shown
- Native ChatGPT sign-in for the openai-codex route: codewhale auth chatgpt opens a browser PKCE flow and stores refreshable tokens in Codewhale-owned credentials — no Codex CLI install required. /auth chatgpt-revoke clears them off the event loop (#5784, #5778).
- MCP servers and plugins can be connected self-serve from the session: a unified auth flow with rotation-safe token handling, a spoken authorization URL, and catalog refresh when stored credentials stop working (#5747).
- /operate reads match the landed CWC OperateRecord contract: fetching an absent record returns a truthful not-found view instead of a fabricated operation, and an empty evidence path is rejected rather than resolving to the workspace directory (#5703).
- TUI: scheduled automations project into the top strip (⏱ N scheduled · M running, compact ⏱ N·M) with typed HistoryCell::Automation receipts when a run this session watched settle. /automation acknowledges failures. The merged footer does not carry the work fact (#5748).
- The app-server can listen on a unix domain socket and advertise a daemon/attach handshake, so a local client can attach to an already-running engine instead of spawning its own. The socket is created with owner-only permissions and stale sockets are reclaimed on start. Non-unix hosts return a typed unsupported-platform refusal; the Windows named-pipe endpoint is named but not yet implemented (#5749).
- The engine's internal Op/Event types and the wire protocol's Op/ EventMsg now carry a compile-enforced twin for every variant: adding an engine variant without a protocol counterpart fails the build instead of drifting silently. Internal durability work — no user-visible surface change yet (#5751).
- Machine tokens: with CODEWHALE_API_KEY set, the CLI authenticates as the Codewhale account with no local session file and no browser — the CI authentication path, with a typed token shape and redaction (#5721).
- Compaction publishes a structured survival contract for session-tree journal entry types (crates/tui/src/compaction/SURVIVAL_CONTRACT.md) and fails closed when the last user round, tool results, /anchor text, or checkpoint receipt would vanish (#4394, #5782).
- Internal: codewhale-config gains RouteAuthoritySnapshot, one immutable authority that owns a compiled provider catalog together with the route resolver projected from it, so a picker, a readiness view, and an execution path can no longer resolve against different catalog snapshots without a type-level signal. Resolution still goes through the sole resolver; the returned receipt distinguishes an exact catalog row, a custom-endpoint route whose provider facts are deliberately…
- Computer session records now count only time a provider actually accepted the session as active, at per-second granularity. Idle, queued, stopped, and teardown time are excluded, and a session whose allocation does not match a standard profile is refused rather than recorded approximately. Covered by hermetic fixtures; no live provider call and no deploy (#5781).
- Website: the public site moves to the Tideline deep-ocean design language (dark by default with an opt-in light documentation sheet, palette grounded in the TUI's WHALE_* tokens) and the new whale brand mark across the favicon, app icons, web manifest, nav wordmark, and social card (#5573).
- Add codewhale dispatch / /dispatch so a local session can propose a Codewhale cloud agent against an explicit github, cnb, or gitee remote. Confirmation is required; missing credentials fail closed; cloud jobs share the existing /jobs surface as kind=cloud. See DAYTONA_CLOUD_DISPATCH.md.
Changed12 of 14 entries shown
- Provider-native web search now applies domain constraints before accepting an attempt, discards generated answers when returned citations violate those constraints, and preserves the caller's configured/local timeout as an independent fallback budget (#5681).
- Idle session metrics omit zero facts (0 turns, LLM 0s) until the runtime has evidence. Working chrome says in the current instead of a generic working.
- Deleted nine uncompiled runtime_contract/ staging files. Live contracts remain model.rs and termination.rs.
- The first #5587 dead-code sweep converts audited test-only helpers to #[cfg(test)], keeping production builds free of test-only APIs without changing runtime behavior.
- /plugin reload is now discoverable when on-disk plugin bundles change: the next send and /plugin list nudge once with Run /plugin reload to apply instead of silently keeping the stale catalog (#5579). Trust is unchanged; this does not auto-reload.
- Context-pressure warnings and critical alerts now remain visible in sticky UI status until compaction or explicit dismissal, instead of disappearing into scrolling turn metadata (#5620).
- The Runtime thread store defaults to a per-session root ($CODEWHALE_HOME/sessions/<id>/runtime) so multiple Codewhale processes on one machine no longer share one owner lock (#5630). The exclusive lock is unchanged; CODEWHALE_RUNTIME_DIR still selects a shared root when that is intended.
- Session token totals now include display-only per-model-call deltas while a turn is running, including input/output and cache-class counters; the authoritative TurnComplete totals still reconcile exactly once (#5581).
- Transcript focus now exposes per-block actions: y copies content, Y copies the rendered metadata view, Enter opens a fullscreen block pager, and r opens raw detail; the existing Tasks rail shortcuts remain unchanged (#5551).
- Provider neutrality (#5588): model resolution of omitted/aliased models is now provider-relative, OpenAI-native defaults no longer route through another provider's table, CLI credentials stay provider-scoped, and NVIDIA credentials no longer leak into the DeepSeek keychain. Neutrality test matrices exercise several providers instead of standing in with one.
- The workflow engine module was decomposed out of its 3k-line mega-file into journal, usage, and report modules plus a module directory (#5586 slices 1a/B/C/D), a semantic-free move verified by normalized-content hash.
- Compaction refusals are now always named (#5577): silent holds under context pressure are gone, the context meter honors the same provider-billed prompt the trigger uses (T1), and prune outcomes are projected without cloning the transcript.
Fixed12 of 25 entries shown
- Read-only Fleet workers no longer send "action": {"enum": null} in their projected bash schema. The read-only projection probed the action enum with a mutating index, which auto-vivified the key on schemas that have no action property, and strict OpenAI-compatible validators then rejected the whole request (null is not of type "array"). The probe is non-mutating now, in both the read-only projection and the Run arm next to it, and a regression test walks the whole projected…
- Fast typing no longer corrupts the composer. The paste-burst heuristic ran on every session until a real bracketed paste arrived, holding, buffering, retro-grabbing, and absorbing Enter on timing guesses; it is now fallback-only (gated off when the terminal provides bracketed paste), the retro-grab is deleted, and Enter on held command text flushes and submits.
- Ctrl+C works on the pre-session launch menu and speaks everywhere: the first press arms the two-second exit window with a visible localized "Press Ctrl+C again to quit" hint (previously silent), the second exits. The worktree name input keeps Ctrl+C as cancel-input.
- Fresh interactive sessions no longer leave a phantom one-message duplicate behind. The TUI claimed one session id (Runtime store lock, turn-start crash checkpoint) while the engine minted a second one; the first SessionUpdated re-keyed the App, the completion commit cleared only the engine id's checkpoint, and codewhale --continue later "recovered" the orphaned checkpoint as a duplicate session instead of the real one. The engine now adopts the host-owned id at spawn…
- Website: /signin, /signup, and /auth/callback are locale-aware public routes instead of localized 404s. Sign-in and create-account send the person to the CWC app; OAuth callbacks hop to app.codewhale.net with the query intact; /login and /register are aliases. Local CLI use is not presented as requiring an account (#5767).
- The sandbox read deny-list matches a rule's resolved path as well as its literal spelling. On macOS /etc and /var are symlinks into /private, so a read of /private/etc/sudoers walked around the /etc/sudoers rule, and a rule written against a symlinked directory never fired for the real path that canonicalize and the process cwd hand back.
- Background shells are first-class work-strip rows (▾ Shells N) you can open, watch, and cancel by the shell_* id on the row. /jobs cancel all cancels running shells; it no longer looks up a task named all. The composer hourglass crumb no longer stands in for a shell surface.
- codewhale logout and /logout now clear the Codewhale account session and the Daytona secret slot, not only provider API keys. The TUI crate's leftover login --api-key path no longer claims to save a key.
- Account sessions no longer read the macOS Keychain. Unsigned or rebuilt codewhale binaries were a new Keychain ACL principal every time, so codewhale web and the TUI popped a password dialog on start. Sessions now use ~/.codewhale/secrets/secrets.json (mode 0600), the same store as provider keys. Extracted from the Keychain-retirement half of #5632.
- Hardened the dispatcher-side config parse the same way: ConfigStore loads and project-config parsing now deserialize ConfigToml on a dedicated 16 MiB-stack thread (the guided-setup save path could overflow a 2 MiB worker stack the same way the TUI's ConfigFile parse did), and the #5585 setup-confirm toast test runs its runtime on an equally sized thread instead of overflowing the default libtest stack.
- Fixed detached interactive agents reporting worker usage after the parent turn ends with the usage missing from the session/live /cost total (#5597): interactive turns acquire an owner-scoped runtime usage lease, late usage enters the session cost pool without reopening the sealed mailbox, and worker/session/reload accounting share one hashed response identity so retried deliveries stay exactly-once.
- Fixed the sub-agent fiasco class: in-workspace absolute git -C no longer trips the read-only child shell gate with a coherent bounded gate (#5595), turn end parks turn-owned children resumably instead of silently cancelling them (#5596), stale write-claims are released by liveness with coordinate release (#5562), and the verifier role description matches its real surface (#5562).