Diagnosed from a live voice session (e1b9fccb90ef): three turns failed with
raw '**Error:** HTTP 502 ... hermes-{claude,codex}-broker' because the agent
pod hosting the model brokers rolled mid-conversation. The voice client
correctly refused to speak the error envelope, but it then dropped the user's
utterance and forced them to repeat it three times.
Voice: on a TRANSIENT provider error (5xx/502/'error sending request'/timeout),
conversation mode now re-runs the errored turn in place through the app's own
regenerate action (which truncates the errored turn — no duplicate user
message) and stays in Thinking so its cues cover the reconnect gap. Bounded to
MAX_TRANSIENT_RETRIES (2); a non-transient error or an exhausted budget still
drops cleanly to 'let's try that again — listening'. The raw error is never
spoken. New probe scenarios cover retry-then-recover and the bounded-then-drop
path; source contract updated.
STT: the same session mis-transcribed 'CUI' as 'cue'. Prime the default
initial_prompt with the domain acronyms the user uses (CUI, FOUO, DoD, NIST,
CMMC, FIPS, RMF, POA&M, ATO, SBU) so they bias to uppercase forms.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
The orb watermark was showing the character art as a low-opacity square
box (visible edges) — the earlier screen-blend pass washed it into a glow
that lost the artwork. Reprocess the source into a feathered circle so the
orb crops it (no box) while keeping the character's facial features, paint
it with a normal blend at 0.5 opacity, and size it to fill the orb (inset
7%). Verified by compositing over a simulated orb before shipping.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
- REGRESSION FIX: the full-screen conversation overlay stopped opening
(fell back to the inline bar) because building the language/output
selectors inside the overlay's single try/catch could throw on a
phone (navigator.mediaDevices/setSinkId). Overlay build is now two
phases: the essential orb+captions+controls attach first; the
selectors attach after, each guarded, so a selector failure omits only
that control and never the visualization. Control builders can no
longer throw (inert hidden fallback); device enumeration is async
after attach. Regression test covers mediaDevices-undefined and
enumerate-rejects.
- Orb watermark is the processed Hermes character glyph (feathered,
circle-cropped, screen-blended so the face glows on the dark orb), no
more white/grey box.
- Output selector defaults to the loudspeaker (excludes the
communications/earpiece endpoint that the open mic otherwise forces);
explicit choice sticks; hidden gracefully where setSinkId is
unsupported. 291 voice tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
Final conversation-mode polish from mobile testing:
- The Workspace toggle now sits with the chat/Telegram nav at every
width: nav.rail on desktop, the top app titlebar on mobile. The
floating pill that pushed the mobile composer's control row (and the
conversation-mode button) off screen is gone - a fallback exists only
for headless DOMs and is pinned to a top corner, never over the
composer.
- The conversation orb watermark is now the Hermes character avatar
(static/hermes-agent-192.png) instead of the caduceus staff.
- User-facing 'hands-free' copy renamed to 'Conversation mode'.
- Wrong-voice fix: strongReplyLanguage flagged Spanish on a single
accented char, so an English reply naming European cities (Zürich,
Málaga) overrode the correct English STT detection and was spoken by
the Spanish voice. Detection now requires density (Cyrillic >=4 at
>=50%, or inverted punctuation / >=2 accents corroborated by Spanish
stopwords); plain English always speaks English, forced language wins,
accent-free Spanish still routes via trusted STT. 294 voice tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
Three conversation-mode fixes in one pass:
- Interim-acknowledgement truncation: when an interim ack folds into the
hidden worklog segment mid-speech, the retained unspoken tail is now
flushed and spoken in full before the Thinking transition, and a
distinct follow-up message is chunked from its own start and queued
after the interim drains (no more 'stops after the first clause, rest
resurfaces with the next message').
- Natural thinking fillers: brief per-language interjections (Umm/Hmm/
One sec; Mmm/A ver; Хм/Секунду) on genuine >1.9s thinking gaps only,
non-repeating, answer-preempting, mute-aware.
- One unified audio sink for every spoken output (reply, cues, fillers,
WAV fallback) - fixes cues playing the loudspeaker while the reply
used a different output - plus a tidy corner output-device selector
(enumerateDevices + setSinkId, feature-detected, session-only) styled
like the language selector. Also realigns two STT-server decode-param
assertions to the dict form from the STT tuning commit. 284 voice tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
The 'Something went wrong - listening' state with no spoken answer was a
false positive: readAssistantTurn flagged the whole turn as an error if
ANY segment was error-stamped - including a recovered/transient tool
error or a cancellation notice from an earlier interim - and threw away
the real answer that the same turn produced. Error now surfaces only
when the turn yielded no spoken answer at all; a turn with real content
is spoken normally. Softened the genuine-error label to the friendlier
'Let's try that again - listening'. New probe scenarios lock both: an
error segment alongside an answer speaks the answer (error=false), and
an error-only turn still reports the error.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
Real root cause (confirmed against the live build-24 DOM): the caption
and TTS extraction fell back to turn.textContent whenever a settle-frame
race left no readable answer segment, scraping the avatar letter,
author name and 'Processed 13s' chip - and that truncated reply made
TTS speak only the first segment then drop to Listening even with the
mic muted (the muted-mic first-sentence-stop). Extraction now prefers
each answer segment's data-raw-text, else the answer .msg-body only
(excluding thinking/tool/worklog/role chrome), and the textContent
fallback is gone; a genuinely mid-flight reply retries briefly so the
whole thing is read before the overlay drains. Also: the app's own
caduceus mark embedded in the conversation orb as a subtle watermark; a
corner language selector (Auto + en/es/ru) that forces both the STT
hint and the reply voice; and a thinking affordance after 2.5s of dead
time. New response probe + extraction test prove caption==body-only and
that every sentence reaches the TTS queue. 278 voice-lane tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
- Captions read message bodies only (the scraper was concatenating
avatar, author and worklog chips); multi-segment interim turns are
now speakable and drive clean speak-to-thinking-to-speak cycles when
playback drains mid-turn.
- Dynamic endpointing: complete-looking partials (3+ words or terminal
punctuation) endpoint at the base window; the long hold remains only
for one-two-word fragments. A stale-busy 10s settle wait on every
post-error send is gone.
- Both overlay captions are bounded, touch-scrollable regions with
follow-tail; caps raised for long turns.
- Error envelopes are never spoken or captioned; errored turns run
resyncCapture (fresh STT session on the hot mic).
- Workspace toggle now lives in the sidebar rail (floating button only
below the rail breakpoint).
- False 'session unavailable' toast root-caused: the router continuity
guard shows it on a 409 that fired when a transient profile-listing
failure failed closed into a fake cross-profile mismatch; the patcher
now answers from the alias cache and never claims a default-vs-named
mismatch while aliases are unconfirmed.
- Barge-in sends carry a one-line cut-point marker with the last spoken
sentence; visible-history truncation judged infeasible client-side.
- Language switching works end-to-end: sticky per-session STT language
hint (restarting an unused next session on switch), reply voice from
script evidence, STT detection, then stopword heuristic; cues and WAV
fallback share the turn language.
- Legacy CI guards: node skip for the DOM probe, ffmpeg/codec skips for
the container-fallback test. 272 voice-lane tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf
Round 2 from live testing:
- Clipping root cause: the endpointer committed on any 1.1s pause, so a
thinking pause after a sentence opener sent one word; a speculative
Whisper pass over that fragment then stalled the real commit ~3.5s on
the Jetson. Young utterances now hold a 1.8s endpoint until 1.2s of
speech accrues, speculative decode waits for 700ms of speech, resume
is unconditional after any 650ms gap (server-frozen snapshots can
never reach commit), and the noise floor is capped so playback echo
cannot deafen onset. Deterministic capture-continuity probe added.
- Stitching now fires for any barge-cancelled send within 20s,
regardless of partial assistant output.
- The workspace-drawer dark rectangle was our own HUX chrome resolving
undefined theme tokens (--bg-primary) to a flat box; bootstrap.css
bridges the real theme tokens and drops a compositor-hazard
backdrop-filter.
- Conversation mode: full-viewport hands-free overlay with an energy-
driven orb (mic RMS in, TTS activity out), state caption, live
transcript/reply captions, mute and exit controls, focus trap,
Escape, reduced-motion support. Voice lane 253 tests green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvMSXH8VH2tMWXanb8SJdf