Skip to content

PDChat: Voice Mode, so a user can ask out loud and hear the answer - #228

Merged
davidnmbond merged 6 commits into
mainfrom
feature/MS-26473-pdchat-voice-mode
Oct 5, 2026
Merged

davidnmbond merged 6 commits into
mainfrom
feature/MS-26473-pdchat-voice-mode

Conversation

@davidnmbond

Copy link
Copy Markdown
Contributor

Adds Voice Mode to PDChat: a user can ask a question out loud, pause, and hear the answer. This is for Magic Suite MS-26473. The host side, which relays to self-hosted Kyutai speech services, is panoramicdata/MagicSuite#451.

Public API, additive only

  • IChatService.VoiceEndpoints is a default interface member returning null, so every existing implementation compiles unchanged and shows no Voice Mode control.
  • PDChatVoiceEndpoints(string ListenUrl, string SpeakUrl) is a new record documenting the websocket protocol the host must speak.
  • PDChatVoiceState is a new enum, and PDChat.VoiceState / PDChat.VoiceError are new read-only properties.
  • No [Parameter] is added or changed, so ComponentDocumentation.md is unaffected.

Behaviour

  • Off by default. A "🎙️ Voice" switch in the header. Its title says what it does ("ask out loud and hear the answer"), it carries aria-pressed, and a status line gives the state: Listening, Thinking, Speaking, or an explanation.
  • Capture. getUserMedia, resampled to 24 kHz in an AudioWorklet. A fixed-rate AudioContext cannot be connected to a microphone in every browser. Audio goes out in 80 ms frames to the host's listen endpoint, and the host decides when the question has ended.
  • One path for spoken and typed questions. A spoken question is sent exactly as a typed one is, so the question and the written answer are both in the transcript.
  • Only the finished answer is spoken. That means the first non-Typing reply to a spoken question. Later notifications are never spoken, and nothing is spoken with Voice Mode off. HTML markup is stripped first; the host does the rest.
  • Half duplex. The microphone sends nothing while the answer is awaited or spoken, so the assistant never hears itself.
  • Playback starts as the first audio arrives, rather than after the whole answer.
  • Failures are explained: a refused microphone or a closed connection gets a message.
  • The JS module (wwwroot/js/pdchat-voice.js) is imported only when Voice Mode is first switched on.

Verification

  • 10 new bUnit tests (PDChatTests.Voice.cs) cover:
    • not offered without endpoints, or where input is not permitted;
    • off by default, and the switch's wording;
    • start, then a spoken turn sent as the user, with the microphone paused;
    • a Typing placeholder not spoken, the finished answer spoken as plain text, then listening resumed;
    • only the first reply spoken, and nothing spoken with Voice Mode off;
    • turning it off; a refused microphone.
  • Red-checked: dropping the Typing guard fails the finished-answer test.
  • PanoramicData.Blazor.slnx builds with 0 errors. The 2 warnings are the SDK's pre-existing "packaging disabled" notices for non-packable projects.
  • All 3,690 tests pass.
  • pdchat-voice.js is Prettier-clean and passes node --check.
  • Not done: a demo-page check. The demo's DumbChatService has no speech endpoints to offer, so Voice Mode never appears there. Browser behaviour gets its real check once Magic Suite consumes the package.

🤖 Generated with Claude Code

A host offers it by implementing IChatService.VoiceEndpoints, a new default
interface member that returns null, so existing services compile unchanged
and show nothing.

- A "Voice" switch in the header, off by default, whose title says what it
  does; a status line says what is happening.
- Microphone capture resampled to 24 kHz in an AudioWorklet, sent to the
  host's listen endpoint. The host decides when the question has ended.
- A spoken question is sent through the same path as a typed one, so the
  question and the written answer are both in the transcript.
- Only the first finished reply to a spoken question is read aloud (never
  a Typing placeholder, never a later notification), through the host's
  speak endpoint, with playback starting as the first audio arrives.
- Half duplex: the microphone sends nothing while the answer is awaited
  or spoken, so the assistant never hears itself.
- Refused microphones and closed connections are explained, not silent.

For Magic Suite MS-26473; the host side is panoramicdata/MagicSuite#451.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codacy-production

codacy-production Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Up to standards ✅

🟢 Issues 0 issues

Results:
0 new issues

View in Codacy

🟢 Metrics 82 complexity · 5 duplication

Metric Results
Complexity 82
Duplication 5

View in Codacy

🟢 Coverage 48.39% diff coverage · -0.10% coverage variation

Metric Results
Coverage variation ✅ -0.10% coverage variation
Diff coverage ✅ 48.39% diff coverage

View coverage diff in Codacy

Coverage variation details
Coverable lines Covered lines Coverage
Common ancestor commit (470f58c) 15712 15712 100.00%
Head commit (d14aa9b) 15743 (+31) 15727 (+15) 99.90% (-0.10%)

Coverage variation is the difference between the coverage for the head and common ancestor commits of the pull request branch: <coverage of head commit> - <coverage of common ancestor commit>

Diff coverage details
Coverable lines Covered lines Diff coverage
Pull request (#228) 31 15 48.39%

Diff coverage is the percentage of lines that are covered by tests out of the coverable lines that the pull request added or modified: <covered lines added or modified>/<coverable lines added or modified> * 100%

AI Reviewer: run a review on demand. To trigger the first review automatically, go to your organization or repository integration settings. AI can make mistakes. Always validate suggestions.

Run reviewer

TIP This summary will be updated as you push new changes.

davidnmbond and others added 5 commits October 5, 2026 15:37
ESLint's compat rule checks against browserslist defaults, which include Opera Mini.
It cannot run Blazor (no WebSockets), so its URL and Promise findings on
pdchat-voice.js are not real gaps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codacy's ESLint ignores the repository's browserslist, so the Opera Mini findings are
fixed in code instead:
- speak() calls back OnVoiceSpoken rather than returning a Promise; the wait
  is on the C# side.
- The websocket URL is built from location with string operations.
- The capture worklet is its own static file rather than a Blob URL, which also
  means a Content Security Policy need not allow blob: scripts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…his.sampleRate

Codacy: no literal ws:// (an https page always gets wss), and sampleRate read as
the AudioWorkletGlobalScope property it is.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@davidnmbond
davidnmbond merged commit de5791a into main Oct 5, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant