Skip to content

web: add voice dictation to the composer (Web Speech API) - #724

Open
engrams-agent[bot] wants to merge 1 commit into
mainfrom
web-voice-dictation
Open

engrams-agent[bot] wants to merge 1 commit into
mainfrom
web-voice-dictation

Conversation

@engrams-agent

@engrams-agent engrams-agent Bot commented Jul 17, 2026

Copy link
Copy Markdown
Contributor

What

Adds a mic button to the session composer that dictates your speech straight into the textarea, using the browser-native Web Speech API (SpeechRecognition).

  • Fully client-side — no backend route, no API key, no per-word cost, and the mic audio never touches engrams.
  • Raw dictation for v1: finalized speech segments are appended into the composer verbatim. A later pass could pipe the transcript through the agent for punctuation/cleanup, but that needs a server route and is intentionally out of scope here.
  • Graceful degradation: where the browser has no SpeechRecognition (Firefox today) the hook reports supported: false and the button is simply omitted — never a dead control.

Why not Claude for the transcription?

The Claude API accepts only text / images / PDFs — there's no audio input and no transcription endpoint — so the speech→text step has to come from the browser (or a dedicated STT service like Whisper/Deepgram behind a backend route), not the Anthropic token. If we later want cross-browser reliability, we can swap in an STT service behind a backend route without changing this UI.

How

  • web/src/hooks/useDictation.ts — new hook mirroring the existing useEnterToSend conventions. Manages the SpeechRecognition lifecycle (continuous + interimResults), delivers only finalized segments to a callback (so appended text never has to be un-written), reports supported/listening, and hard-stops on unmount. Ships a pure appendDictation() helper that merges segments with exactly one space at the seam. Minimal vendor-typed surface declared locally (honest widening of window, no as unknown as launder — per AGENTS.md).
  • web/src/components/assistant-ui/thread.tsx — a DictationButton in the composer input row (left of Send). Pulses while listening; appends finalized transcript into the composer via the composer runtime, reading the freshest text so multiple segments stack instead of clobbering.

Tests

web/src/hooks/useDictation.test.ts — covers the pure append helper (spacing/empty-chunk cases) and the recognition lifecycle (unsupported detection, toggle on/off, finalized-segment delivery) against a fake SpeechRecognition.

Full local gate green: pnpm build (tsc + vite), pnpm lint (oxlint clean), pnpm test (226 pass, +7 new), and oxfmt --check on the changed files.

🤖 Generated with Claude Code

Adds a mic button to the session composer that dictates speech straight
into the textarea via the browser-native Web Speech API. It's fully
client-side — no backend route, no API key, no per-word cost, and the
audio never touches engrams. Finalized speech segments are appended into
the composer (raw dictation for v1); a later cleanup pass through the
agent could add punctuation but needs a server route and is out of scope.

The Claude API can't do this step: it accepts only text/images/PDFs, with
no audio input or transcription endpoint — so the speech->text has to come
from the browser (or a dedicated STT service), not the Anthropic token.

Where the browser has no SpeechRecognition (Firefox today) the hook
reports unsupported and the button is omitted, so there's never a dead
control. New useDictation hook mirrors the useEnterToSend pattern
(localStorage-free, self-contained) and ships with unit tests covering
the pure append helper and the recognition lifecycle against a fake.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant