Write summaries in the transcript's dominant language - #7960
Conversation
Co-Authored-By: John <john@fastrepl.com>
|
I'll fix CI failures and address comments from users with write access that start with 'Devin'.
|
There was a problem hiding this comment.
👀 1 finding needs your review
Devin reviewed the finding on 4140f5e and left it for you. Click a finding below to jump to its comment.
For your review (1)
| if let Some(detected) = detector.detect(text).map(|info| info.lang()) | ||
| && let Some(index) = languages.iter().position(|language| *language == detected) | ||
| { | ||
| weights[index] = weights[index].saturating_add(weight); | ||
| } |
There was a problem hiding this comment.
🟡 Mixed-language segments skew summary language
When a segment mixes languages, dominant_language credits every character to the detector's single label. Same-speaker words can share a segment, so a minority language can win the meeting's vote.
Learn more
The transcript renderer can merge adjacent words from the same speaker into one segment collect_segments. Detection here classifies that whole segment once and adds all its characters to the detected language, including characters spoken in other languages. A code-switching speaker can therefore outweigh a larger amount of text actually spoken in another candidate language, causing the summary prompt to request the wrong language.
Example: Suppose a merged segment contains 55 French and 45 English non-space characters and is classified French. A separate segment has 30 English characters. French gets 100 votes and English gets 30, even though English accounts for 75 characters overall.
Recommended fix: Split long or mixed segments into smaller detection windows before tallying, ideally at word or sentence boundaries; aggregate only the character counts of each window's detected language. Add a code-switching regression test that verifies the overall winner.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
This is a real limitation, but it only bites when a speaker switches languages a lot inside one segment. A segment gets whatlang's majority label, so it misreads overall dominance only when the switching is lopsided across segments. Splitting segments into sentence- or word-window chunks before voting would fix that, at the cost of more detection calls and noisier labels on short windows. I've left this to the PR author to decide whether it belongs in this PR or a follow-up.
Summary
Intent: Summaries were always written in the main language (
ai_language), even when most of the meeting was spoken in one of the user's additional languages. Now the summary is written in whichever of those languages dominates the transcript. Titles follow the same rule (detected from the enhanced note), so the title matches the summary. ANLG-447anlg_language::dominant_language(newdominantfeature,whatlang0.18.0) runswhatlangon each transcript segment, limited to the candidate languages, and weights each result by its non-whitespace character count. The candidate with the highest total wins, and ties go to the main language.None(keeps the main language) when there's no usable text, or when any candidate isn't supported bywhatlang(e.g.bs,ms). This avoids silently voting against a language we can't detect.en-UScomes back unchanged.ai_language; this change covers desktop only.Demo
N/A. The change is in how the language is chosen for the prompt (covered by unit tests). The UI is unchanged.
Verification
cargo test --locked -p language --features dominantcargo clippy --locked -p language --features dominant --all-targets --no-deps -- -D warnings, and the same fortauri-plugin-templateexport_typesto regenerate bindingspnpm -F desktop typecheck; ran vitest onenhance-transform.test.tsLink to Devin session: https://app.devin.ai/sessions/345f8b3f50be4a8eb25b5b97ff760f48
Open in Devin Desktop: https://app.devin.ai/desktop/session/345f8b3f50be4a8eb25b5b97ff760f48?variant=devin
Requested by: @ComputelessComputer
Summary by cubic
Summaries are now written in the dominant spoken language of the transcript instead of always the main
ai_language. Titles follow the same rule, so they stay consistent with summaries.ai_language.Written for commit 4140f5e. Summary will update on new commits.