Repository navigation
web: add voice dictation to the composer (Web Speech API) - #724
Open
engrams-agent[bot] wants to merge 1 commit into
Open
engrams-agent[bot] wants to merge 1 commit into
engrams-agent[bot] wants to merge 1 commit into
Conversation
Adds a mic button to the session composer that dictates speech straight into the textarea via the browser-native Web Speech API. It's fully client-side — no backend route, no API key, no per-word cost, and the audio never touches engrams. Finalized speech segments are appended into the composer (raw dictation for v1); a later cleanup pass through the agent could add punctuation but needs a server route and is out of scope. The Claude API can't do this step: it accepts only text/images/PDFs, with no audio input or transcription endpoint — so the speech->text has to come from the browser (or a dedicated STT service), not the Anthropic token. Where the browser has no SpeechRecognition (Firefox today) the hook reports unsupported and the button is omitted, so there's never a dead control. New useDictation hook mirrors the useEnterToSend pattern (localStorage-free, self-contained) and ships with unit tests covering the pure append helper and the recognition lifecycle against a fake. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds a mic button to the session composer that dictates your speech straight into the textarea, using the browser-native Web Speech API (
SpeechRecognition).SpeechRecognition(Firefox today) the hook reportssupported: falseand the button is simply omitted — never a dead control.Why not Claude for the transcription?
The Claude API accepts only text / images / PDFs — there's no audio input and no transcription endpoint — so the speech→text step has to come from the browser (or a dedicated STT service like Whisper/Deepgram behind a backend route), not the Anthropic token. If we later want cross-browser reliability, we can swap in an STT service behind a backend route without changing this UI.
How
web/src/hooks/useDictation.ts— new hook mirroring the existinguseEnterToSendconventions. Manages theSpeechRecognitionlifecycle (continuous+interimResults), delivers only finalized segments to a callback (so appended text never has to be un-written), reportssupported/listening, and hard-stops on unmount. Ships a pureappendDictation()helper that merges segments with exactly one space at the seam. Minimal vendor-typed surface declared locally (honest widening ofwindow, noas unknown aslaunder — per AGENTS.md).web/src/components/assistant-ui/thread.tsx— aDictationButtonin the composer input row (left of Send). Pulses while listening; appends finalized transcript into the composer via the composer runtime, reading the freshest text so multiple segments stack instead of clobbering.Tests
web/src/hooks/useDictation.test.ts— covers the pure append helper (spacing/empty-chunk cases) and the recognition lifecycle (unsupported detection, toggle on/off, finalized-segment delivery) against a fakeSpeechRecognition.Full local gate green:
pnpm build(tsc + vite),pnpm lint(oxlint clean),pnpm test(226 pass, +7 new), andoxfmt --checkon the changed files.🤖 Generated with Claude Code