What should the engine do?
Normalize BCP-47 region and script tags when resolving preferredAudioLanguages,
preferredSubtitleLanguages, and nativeSubtitlePreferredLanguages, while preserving the
distinctions that matter to each track kind.
This should stay entirely on the structured language field. The engine should not infer a
language from a track title or sidecar filename.
Current behaviour
The public external-subtitle descriptor documents ExternalSubtitleTrack.language as accepting
“BCP-47 / ISO 639”, but the preference path ultimately reaches languageMatches, which currently
does case-insensitive exact comparison plus fixed ISO/name synonym sets.
That leaves common structured tags unmatched:
preference track language current result
en en-US no match
en-US eng no match
ar-SA ara no match
zh-Hant zh-TW no match
zh-Hans zh-CN no match
The same matcher is shared by the first-frame audio pick and post-load subtitle pick. The native
subtitle DEFAULT ordinal also calls it directly, so the host-overlay and native/PiP paths can make
the same miss independently.
This is a follow-up to #73 rather than a request for host-side compensation. main at
b1b3c2add9d535739a1e450db05d2979906ff3c9 still has this behaviour; the downstream integration is
pinned to 7.10.0 (ca8512cd579b770d96dec8cea23507a5bb81128a).
Proposed matching semantics
Keep preference-array order as the outer priority, but rank matches within one preference:
Common normalization
- Treat ISO 639-1, ISO 639-2/B, ISO 639-2/T and a well-formed BCP-47 primary subtag as the same
language identity where they are equivalents.
- Exact normalized BCP-47 identity wins over a same-language fallback.
- Nil, empty,
und, free-form names and unrecognized labels remain non-matches.
- Do not parse
TrackInfo.name, ExternalSubtitleTrack.name, or filenames.
Audio
- Prefer an exact full tag, then fall back to the same canonical language when only the region
differs or one side carries no region (en-GB > en/eng > another en-* candidate for an
en-GB preference).
- Do not collapse distinct language identities merely because they are related: for example
yue
and nan must not become generic zh. A generic macrolanguage fallback, if supported, should be
weaker than an exact language match.
Subtitles
- Prefer an exact tag/script, then the same explicit script, then an unspecified-script track of
the same language.
- A track with an explicitly opposite script is not a fallback.
- For Chinese subtitle tags, explicit region can carry the script distinction used by media
servers and sidecar naming: zh-CN / zh-SG are Simplified and zh-TW / zh-HK / zh-MO are
Traditional, equivalent to zh-Hans and zh-Hant respectively.
- Bare
zh / chi / zho is script-unspecified. It may be a weaker fallback for either preference,
but must not be treated as explicitly Simplified.
- Language specificity should rank before the existing
subtitlePickRank; the existing
full > SDH > forced > commentary and text > bitmap ranking then resolves tracks at equal language
specificity.
One implementation trap: Locale.Language(identifier: "zh").script currently infers Hans on
Apple platforms. Using that inferred script directly would turn an unspecified zh track into an
explicit Simplified track. Script specificity therefore needs to come from an explicit script
subtag or an agreed explicit-region mapping, not from likely-subtags expansion of a bare language.
Reproduction / acceptance examples
Pure selection tests should be sufficient; no media fixture or decoder path is involved.
// Audio: exact region beats a generic/same-language fallback even when it is later in track order.
preferredAudioLanguages = ["en-GB"]
tracks = ["en-US", "eng", "en-GB"]
// expected: en-GB
// Subtitle: explicit matching script wins; generic is fallback; opposite script is rejected.
preferredSubtitleLanguages = ["zh-Hant"]
tracks = ["zh-CN", "zho", "zh-TW"]
// expected: zh-TW
preferredSubtitleLanguages = ["zh-Hant"]
tracks = ["zh-CN", "zho"]
// expected: zho
preferredSubtitleLanguages = ["zh-Hant"]
tracks = ["zh-CN"]
// expected: nil (subtitles stay off)
The same BCP-47 resolution should choose the DEFAULT rendition for
nativeSubtitlePreferredLanguages, so inline and PiP/AirPlay do not disagree.
Why this belongs in the engine
The engine already owns preference-order resolution, subtitle descriptor ranking, external-track
registration, and native rendition default selection. Reproducing ISO aliases and match ranking in
each host makes those paths drift, especially when external subtitles are added by media servers or
future download providers.
There is already robust BCP-47/ISO normalization work in AudioLanguageMap for emitted HLS/fMP4
metadata, but that helper intentionally collapses region/script to an ISO 639-2/T code. The selection
path needs to retain enough explicit-tag information to rank region/script before falling back.
Area
Public API behaviour / audio / subtitles
Host app / integration context
Custom iOS media client using structured Emby sidecar metadata today, with provider-neutral
downloaded subtitle tracks planned. The host wants to pass metadata into Aether and rely on the
engine's existing selection APIs rather than maintain a second matcher.
Would you be willing to open a PR?
Maybe, with guidance on the generic-language fallback and whether audio and subtitle matching should
share one ranked representation or two policies over a common parser.
What should the engine do?
Normalize BCP-47 region and script tags when resolving
preferredAudioLanguages,preferredSubtitleLanguages, andnativeSubtitlePreferredLanguages, while preserving thedistinctions that matter to each track kind.
This should stay entirely on the structured
languagefield. The engine should not infer alanguage from a track title or sidecar filename.
Current behaviour
The public external-subtitle descriptor documents
ExternalSubtitleTrack.languageas accepting“BCP-47 / ISO 639”, but the preference path ultimately reaches
languageMatches, which currentlydoes case-insensitive exact comparison plus fixed ISO/name synonym sets.
That leaves common structured tags unmatched:
The same matcher is shared by the first-frame audio pick and post-load subtitle pick. The native
subtitle DEFAULT ordinal also calls it directly, so the host-overlay and native/PiP paths can make
the same miss independently.
This is a follow-up to #73 rather than a request for host-side compensation.
mainatb1b3c2add9d535739a1e450db05d2979906ff3c9still has this behaviour; the downstream integration ispinned to 7.10.0 (
ca8512cd579b770d96dec8cea23507a5bb81128a).Proposed matching semantics
Keep preference-array order as the outer priority, but rank matches within one preference:
Common normalization
language identity where they are equivalents.
und, free-form names and unrecognized labels remain non-matches.TrackInfo.name,ExternalSubtitleTrack.name, or filenames.Audio
differs or one side carries no region (
en-GB>en/eng> anotheren-*candidate for anen-GBpreference).yueand
nanmust not become genericzh. A generic macrolanguage fallback, if supported, should beweaker than an exact language match.
Subtitles
the same language.
servers and sidecar naming:
zh-CN/zh-SGare Simplified andzh-TW/zh-HK/zh-MOareTraditional, equivalent to
zh-Hansandzh-Hantrespectively.zh/chi/zhois script-unspecified. It may be a weaker fallback for either preference,but must not be treated as explicitly Simplified.
subtitlePickRank; the existingfull > SDH > forced > commentary and text > bitmap ranking then resolves tracks at equal language
specificity.
One implementation trap:
Locale.Language(identifier: "zh").scriptcurrently infersHansonApple platforms. Using that inferred script directly would turn an unspecified
zhtrack into anexplicit Simplified track. Script specificity therefore needs to come from an explicit script
subtag or an agreed explicit-region mapping, not from likely-subtags expansion of a bare language.
Reproduction / acceptance examples
Pure selection tests should be sufficient; no media fixture or decoder path is involved.
The same BCP-47 resolution should choose the DEFAULT rendition for
nativeSubtitlePreferredLanguages, so inline and PiP/AirPlay do not disagree.Why this belongs in the engine
The engine already owns preference-order resolution, subtitle descriptor ranking, external-track
registration, and native rendition default selection. Reproducing ISO aliases and match ranking in
each host makes those paths drift, especially when external subtitles are added by media servers or
future download providers.
There is already robust BCP-47/ISO normalization work in
AudioLanguageMapfor emitted HLS/fMP4metadata, but that helper intentionally collapses region/script to an ISO 639-2/T code. The selection
path needs to retain enough explicit-tag information to rank region/script before falling back.
Area
Public API behaviour / audio / subtitles
Host app / integration context
Custom iOS media client using structured Emby sidecar metadata today, with provider-neutral
downloaded subtitle tracks planned. The host wants to pass metadata into Aether and rely on the
engine's existing selection APIs rather than maintain a second matcher.
Would you be willing to open a PR?
Maybe, with guidance on the generic-language fallback and whether audio and subtitle matching should
share one ranked representation or two policies over a common parser.