Skip to content

Add UDP/RTP multicast playback with FFmpeg software decoding - #7

Merged
yunnysunny merged 3 commits into
mainfrom
feat/udp-rtp-multicast
Oct 3, 2026
Merged

yunnysunny merged 3 commits into
mainfrom
feat/udp-rtp-multicast

Conversation

@yunnysunny

Copy link
Copy Markdown
Member

做了什么

新增 udp:// / rtp:// 组播直播支持,并引入 FFmpeg 软解补齐盒子硬解常缺的编码。
期间在模拟器上实测暴露出三个既有缺陷,一并修了。

详见 docs/multicast-udp-rtp.md(配置方式、地址写法、实现要点、排查手册)。

组播播放

  • 自建 MulticastDataSource。Media3 的 UdpDataSource 是 final 且缺三样直播必需的能力:可配 SO_RCVBUF、按活动网卡 joinGroup(盒子常同时有 eth0 和 wlan0,join 错网卡一个包都收不到)、以及 rtp:// 支持(Media3 的 RTP 耦合在 RTSP 会话里)。
  • 独立收包线程 + 环形缓冲(PacketRingBuffer)。UDP 是推模式,内核缓冲一满就永久丢包;原先由 ExoPlayer 的 Loader 线程直接 receive(),而那条线程还要做 TS 解复用和内存分配,GC 一停顿或解码器一挣扎就变成突发丢包。实测 rtpLost 由 71 降到 0。
  • RTP 剥头 + 乱序重排,收包路径零分配。源标注 rtp:// 但实际推裸 TS 时自动回退透传。
  • 地址规范化 + udpxy 改写。统一 VLC 风格 udp://@、SSM 风格 rtp://@src@group 等写法;配置代理后改走 HTTP 单播——这是绝大多数家宽环境下唯一可行的方案。TV 设置页和 Web 管理页都能配。
  • 断流自动重新入组。IGMP 成员关系被上游剪掉是「播几分钟后卡住、重进频道就好」的真正原因,重进之所以有效是因为它重新 join 补发了 Membership Report。现在应用自己做,且有次数上限,真死掉的流照常换源。
  • TS 解复用调参收窄到仅组播。MODE_SINGLE_PMT 的「只有一个 PMT」假定对来源不明的 HTTP TS 不成立,会导致 PID 映射走样。
  • 组播与 HTTP/HLS 两套缓冲档共用同一内存池,换台只切阈值、不重建播放器。

FFmpeg 软解

  • 引入 nextlib-media3ext,补上组播 TS 高频的 MPEG-2 视频和 MP2/AC3/DTS 音频。
    ⚠️ 其版本号格式为 <media3版本>-<nextlib版本>,升级 Media3 时必须同步升级它,否则运行时抛 NoSuchMethodError。
  • 新增解码方式设置:硬解优先(默认)/ 软解优先 / 仅硬解。切换后立即重建播放器并续播。
  • release 包限定 ARM ABI(debug 保留全部四个以便模拟器调试),APK 约 7.7MB → 18.9MB。

顺带修掉的既有缺陷

  • 频道列表惯性滚动时切分组会崩溃。回收带焦点的行触发 rootViewRequestFocus(),焦点落到分组列表项上并同步调用 setAdapter(),而那个列表还在回收中 → Cannot call removeView(At) within removeView(At)。
  • 起播成功后卡住再也不恢复。起播超时在 STATE_READY 时被取消后从未重新武装,之后的卡顿完全无人看管。新增看门狗,优先原地重开当前线路(单线路频道换源等于放弃)。
  • 诊断能力:吞吐心跳、断流告警、SO_RCVBUF 被内核压缩告警、重缓冲计数,以及 debug 构建挂载 ExoPlayer 官方 EventLogger。

测试

  • 新增 10 个测试类共 198 个单元测试,全部通过
  • assembleDebug / assembleRelease 均成功
  • 真机/模拟器实测:组播直收可用,RTP 解析 malformed=0,自动重新入组生效

已知限制 / 待确认

  • 不支持 SSM(指定源组播),m3u 里的源地址段会被忽略。
  • SwitchableLoadControl 的组播档开了 prioritizeTimeOverSizeThresholds,因为 DefaultLoadControl 的字节计数是 per-instance 的、组播委托看不到。副作用是时间戳错乱时内存上限只能靠重缓冲看门狗兜底。彻底解决需要自己实现 LoadControl,会动到目前工作正常的 HLS 主路径,本次未做。
  • 测试环境是 x86 模拟器镜像,1080p 持续掉帧 8~17 fps(数据链路已验证干净:rtpLost=0、overflow=0、不再重缓冲),瓶颈是该镜像的解码能力。建议合入前在真实电视盒子上复核一轮组播播放。

🤖 Generated with Claude Code

Multicast playback
- udp:// and rtp:// channels via a purpose-built MulticastDataSource.
  Media3's UdpDataSource is final and lacks what live IPTV needs: a
  configurable SO_RCVBUF, interface-aware joinGroup (boxes with both eth0
  and wlan0 receive nothing if the wrong one is joined), and rtp:// at all
  (Media3's RTP is coupled to RTSP sessions).
- Packets are received on a dedicated thread feeding a bounded ring buffer.
  UDP is push-based, so a full kernel buffer means permanent loss; letting
  ExoPlayer's Loader thread do receive() meant a GC pause or a struggling
  decoder turned straight into burst loss. Measured: rtpLost 71 -> 0.
- RTP header stripping (CSRC/extension/padding) plus a 32-packet reorder
  window, zero-allocation on the receive path. Streams that declare rtp://
  but carry raw TS fall back to passthrough automatically.
- Address normalisation for the VLC (udp://@) and SSM (rtp://@src@group)
  spellings, and optional udpxy rewriting to HTTP unicast, which is the
  only workable option on most home networks. Configurable on TV and in
  the web admin page.
- Stalls rejoin the multicast group in place before giving up. IGMP
  membership gets pruned upstream, which is why re-entering a channel used
  to fix it; the app now does that itself, bounded so a dead stream still
  switches source.
- Live TS extractor tuning (MODE_SINGLE_PMT) is scoped to multicast only;
  its single-PMT assumption is unsafe for TS of unknown origin over HTTP.
- Separate buffer profiles for multicast and HTTP/HLS sharing one memory
  pool, so switching does not rebuild the player.

FFmpeg software decoding
- nextlib-media3ext supplies decoders boxes commonly lack, notably MPEG-2
  video and MP2 audio. Its version encodes the Media3 version and must be
  upgraded in lockstep.
- Decoder mode setting: hardware first (default), software first, or
  hardware only. Changing it rebuilds the player and resumes the channel.
- Release builds keep ARM ABIs only; debug keeps all four for the emulator.

Robustness
- Fix a crash when switching groups during channel-list fling. Recycling a
  focused row triggers rootViewRequestFocus(), which lands on a group item
  and synchronously calls setAdapter() on the list still being recycled.
- Add a watchdog for stalling after playback has started. The start
  timeout was cancelled on READY and never rearmed, so a later stall was
  unsupervised. Recovery restarts the current source before switching,
  since switching is equivalent to giving up on a single-source channel.
- Diagnostics: throughput heartbeat, stall and SO_RCVBUF warnings, rebuffer
  counter, and ExoPlayer's EventLogger in debug builds.

198 unit tests pass; assembleDebug and assembleRelease both succeed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 37af29dd8f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

if (playbackState == Player.STATE_READY) {
cancelTimeout();
cancelRebufferWatchdog();
consecutiveStallRecoveries = 0;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve stall recovery count across transient READY

When a broken source briefly reaches STATE_READY after each restart and then falls back to STATE_BUFFERING, resetting consecutiveStallRecoveries here causes every watchdog invocation to be treated as the first attempt. This is exactly the short READY→BUFFERING cycle diagnosed below, and it means MAX_STALL_RECOVERIES is never reached, so the player can restart the same bad source forever instead of switching sources.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确认属实,已在 59e983e 修复。

这条特别值得记:我上一轮正是从日志里诊断出「重开 → READY 15ms → 又卡」这个形态才加的看门狗,却又在 STATE_READY 里无条件 consecutiveStallRecoveries = 0,等于让看门狗在它最该起作用的场景里失效。

改法:READY 不再重置预算,只有持续播放超过 STALL_RECOVERY_RESET_AFTER_MS(30 秒)才算真正恢复。换频道、手动切线路、自动换源仍然给满新预算。

新增两个回归用例:

  • transientReadyDoesNotRefillStallRecoveryBudget —— 短暂 READY 不续杯,预算耗尽后必须换源
  • sustainedPlaybackRefillsStallRecoveryBudget —— 稳定播放后应重新获得预算

if (cm == null) {
return null;
}
Network active = cm.getActiveNetwork();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Join the interface that carries the multicast group

On devices where the IPTV multicast network is not Android's default network—for example, IPTV on Ethernet while Wi-Fi remains the default Internet connection—getActiveNetwork() selects the wrong interface. Joining a multicast group on that interface normally succeeds, so the fallback is never attempted, but no stream packets arrive; the new direct-multicast feature therefore fails in the multi-interface scenario this selection logic is intended to support.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确认属实,已在 59e983e 修复。

这条是六条里最有价值的。讽刺的是 resolveActiveMulticastInterface() 这个方法本来就是为多网卡场景写的,却在最典型的多网卡场景里选错网卡:IPTV 常接在没有公网的以太网口上,Android 不会把这种网络当作活动网络;此时在 WiFi 上 joinGroup 会成功(不抛异常),但一个组播包都收不到,而我那个基于异常的回退永远不会触发。

改法:在所有可用组播网卡上一并 join(活动网络优先,其余随后),筛选条件为非 loopback、已 up、支持组播、且有 IPv4 地址。代价只是几个 IGMP 报文。一张都没成功时才回退到系统路由。leaveGroup 和重新入组同步改为遍历列表,列表在 open() 后不再修改,可跨线程安全读取。

日志里的 iface= 现在会列出全部已加入的网卡,例如 iface=eth0+wlan0。

throw wrap(new IOException("Data source closed"),
PlaybackException.ERROR_CODE_IO_NETWORK_CONNECTION_FAILED);
}
long waitMs = socketTimeoutMs * 2L;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Wait through the final multicast rejoin attempt

With the defaults, the reader times out at 3 seconds, rejoins twice, and only declares failure around 9 seconds, but the consumer gives up after socketTimeoutMs * 2 (6 seconds). For an initially silent or stalled stream, this races with the second rejoin and usually aborts playback without giving that final attempt any time to receive a packet, defeating the configured two-attempt recovery budget.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确认属实,已在 59e983e 修复。

算过了:收包线程首次超时 3s + 重入组 2 次各 3s ≈ 9s 才判定断流,而消费侧只等 socketTimeoutMs * 2 = 6s。最后一次重入组还没来得及收到包就被放弃,等于我白配了那个预算。

改法:抽出 consumerWaitMs(socketTimeoutMs, maxRejoinAttempts),按 (重入组次数 + 1) 个超时周期再加 2 秒余量计算。收包线程彻底失败时会写 readerError 并关闭环形缓冲,消费侧会被立刻唤醒,所以这个超时只是兜底。

新增 consumerOutwaitsTheWholeRejoinBudget,把「消费侧必须等得比收包线程久」这个不变量钉死,以后任一侧改参数都会红。

Comment on lines +610 to +613
if (player == null || player.getPlaybackState() != Player.STATE_BUFFERING) {
return;
}
onRebufferStall(timeoutMs);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Suppress stall recovery while playback is paused

If the activity is paused while the player is buffering after a previous successful start, the already-armed watchdog still passes this state-only check and calls playCurrentSource(), which unconditionally sets playWhenReady back to true. A stalled stream can therefore restart and resume audio/video in the background after PlayerActivity.onPause() explicitly paused it; the watchdog should also require playback to be requested or be canceled from pause().

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确认属实,已在 59e983e 修复。

后台偷偷续播是实打实的问题,不只是多耗电——用户按了暂停或退出,音频却自己回来了。

改法三处:

  • pause() 取消看门狗
  • resume() 若仍处于 STATE_BUFFERING 则重新武装(否则暂停期间的卡顿在恢复后就没人看管了)
  • 看门狗触发时额外校验 player.getPlayWhenReady(),作为第二道防线

新增 pauseCancelsRebufferWatchdog 和 resumeRearmsWatchdogWhileBuffering。

Comment on lines +82 to +85
if (diff >= SEQ_HALF) {
// 序号早于当前期望:迟到或重复,直接丢弃
latePackets++;
return;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reset reordering when the RTP sender changes sequence space

When an RTP sender restarts mid-session with a new, lower initial sequence number, every packet whose modular difference falls in the backward half of the sequence space is classified as merely late and discarded. Because the implementation does not track SSRC changes or repeated backward discontinuities, playback can drop up to 32,768 packets before the new sequence catches the old expectation instead of recovering immediately from the sender restart.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确认属实,已在 59e983e 修复。

把边界算清楚了:新起点落在当前期望值之前的半个序号空间时才会中招(落在之后的话 diff >= capacity 会判为跳变并自动重置)。最坏丢 32767 个包,按实测 760 pkt/s 约 43 秒黑屏。

改法:连续迟到达到 LATE_PACKETS_BEFORE_RESYNC(64)即认定序号空间已变,按新起点重新同步。没有按 SSRC 判断是因为 SSRC 在 RtpPacketUtil 里尚未解析,而连续迟到这个判据对「发送端重启」和「源切换」都成立,更通用。

阈值 64 在 760 pkt/s 下约 85ms,正常网络抖动不会连续迟到这么多次。配了反向用例 occasionalLatePacketsDoNotTriggerResync 确保零星迟到包不会误触发。

Comment on lines +454 to +456
boolean multicastOrigin = MulticastUrlUtil.isMulticastStreamUrl(rawUrl);
String url = MulticastUrlUtil.resolvePlaybackUrl(rawUrl, preferenceManager().getUdpxyProxyBase());
currentResolvedUrl = url;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Bypass the HLS disk cache for udpxy streams

When the optional live-TS disk cache is enabled, resolving a multicast URL to HTTP here sends the unbounded udpxy response through HlsSegmentPrefetcher's playback data source. Its LoggingPlaybackDataSource selects cacheDelegate even for URLs that are not .ts segments, so a continuous udpxy stream is written to the 96 MB cache and continuously evicted rather than merely streamed, causing sustained disk I/O and cache churn for as long as the channel plays.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确认属实,已在 59e983e 修复。

这条我特地跨文件核对了 LoggingPlaybackDataSource.open():isTsRequest 为 false 时 bypassCache 保持 false,直接落到 activeDelegate = cacheDelegate,确实会把 udpxy 的无界流写进 96MB 磁盘缓存并持续淘汰。盒子的闪存不该这么耗。

改法:条件从 if (bypassCache) 放宽为 if (bypassCache || !isTsRequest),非 .ts 请求一律走上游。这同时也符合该功能本来的定位——设置项就叫「使用硬盘缓存预取直播分片」,缓存对象本就该限定为 ts 分片。

副作用:fMP4 封装的 HLS 分片也不再走缓存。考虑到该设置默认关闭且文案明确限定 ts,这个取舍是合理的。

yunnysunny and others added 2 commits October 3, 2026 21:38
Six issues flagged on the PR, all confirmed against the code.

- Stall recovery budget was reset on every STATE_READY, but the failure
  mode it guards against is exactly "restart, READY for 15ms, stall
  again". The cap was therefore never reached and a bad source could be
  restarted forever instead of switched. Only sustained playback (30s)
  refills it now.
- The watchdog could fire while the activity was paused and set
  playWhenReady back to true, resuming audio in the background. pause()
  now cancels it, resume() rearms it if still buffering, and the
  watchdog checks playWhenReady before acting.
- Interface selection used getActiveNetwork() alone. IPTV is commonly on
  an Ethernet port with no internet, which Android does not make the
  active network; joining on Wi-Fi then succeeds without error and no
  packets arrive, so the exception-based fallback never triggered. The
  group is now joined on every usable multicast interface.
- A sender restarting with a lower RTP sequence number put every packet
  in the backward half of the sequence space, so all of them were
  discarded as late - up to 32767 packets, roughly 43 seconds of black
  screen. 64 consecutive late packets now resync to the new sequence.
- The consumer waited socketTimeoutMs * 2 (6s) while the reader needs
  about 9s to exhaust its rejoin budget, so playback was abandoned before
  the final attempt could receive anything. The wait is now derived from
  the rejoin budget.
- udpxy-proxied multicast is not a .ts request, so LoggingPlaybackDataSource
  still routed it through the cache delegate, writing an unbounded live
  stream into the 96MB disk cache and evicting it continuously. Non-TS
  requests now bypass the cache.

206 unit tests pass; assembleDebug and assembleRelease both succeed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Setup Android SDK has been failing with "Failed to find package 'tools'"
before any code is compiled. android-actions/setup-android defaults its
packages input to 'tools platform-tools', and Google has since removed the
obsolete 'tools' package from the SDK repository, so sdkmanager exits 1.

Unrelated to the branch contents: main's last green run predates the
removal, and the step runs before checkout output is ever built.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@yunnysunny

Copy link
Copy Markdown
Member Author

评审反馈处理完毕

6 条行内评审全部逐条核对过代码,均属实,已在 59e983e 修复,每条都有对应回归用例。各条的具体改法见我在对应行的回复。

# 等级 问题 影响
1 P1 恢复预算被瞬时 READY 重置 坏线路无限重开,永不换源
2 P1 只认 getActiveNetwork() 选网卡 IPTV 走以太网的盒子收不到流
3 P1 暂停后看门狗把播放拉起 后台偷偷续播
4 P1 RTP 发送端重启后序号判为迟到 最坏约 43 秒黑屏
5 P2 消费侧 6s 就放弃,收包线程要 9s 重入组预算白配
6 P2 udpxy 无界流被写入磁盘缓存 盒子闪存持续写入+淘汰

其中 1、2 两条尤其值得注意,都是「代码看起来正确、却在它本该覆盖的场景里恰好失效」:

  • 第 1 条让我上一轮刚加的卡顿看门狗,在实测日志中诊断出的那个确切故障形态下失灵;
  • 第 2 条让一个专为多网卡场景编写的方法,在最典型的多网卡场景(IPTV 接无公网以太网口)里选错网卡,且因为 joinGroup 不抛异常,基于异常的回退永远不会触发。

这两类问题靠读代码很难发现,模拟器上也测不出来(单网卡、源不会重启)。

另外:CI 已修复

本 PR 的 CI 之前连续失败,原因与本分支内容无关——android-actions/setup-android 的 packages 默认值含已被 Google 下架的 tools 包,Setup Android SDK 步骤在任何代码被编译之前就退出 1。main 最后一次绿的构建早于那次下架。

已在 2d12a20 把 packages 收窄为 platform-tools,build.yml 和 release.yml 同步修改(后者只在打 tag 时触发,本 PR 跑不到,但同样的雷埋着,发版时才炸就太晚了)。

现 CI 通过:build pass 2m38s。

如果你希望 CI 修复单独走一个 PR,我可以把那个提交摘出来。

当前状态

  • 206 个单元测试通过
  • assembleDebug / assembleRelease 均成功
  • CI 绿

仍建议合入前在真实电视盒子上过一轮组播播放:目前所有实测都在 x86 模拟器镜像上完成,该环境 1080p 持续掉帧 8~17 fps(数据链路已验证干净:rtpLost=0、overflow=0、不再重缓冲),瓶颈是镜像的解码能力,但真机的网络环境与解码能力都与之差距很大。

@yunnysunny
yunnysunny merged commit 2ebf609 into main Oct 3, 2026
1 check passed
@yunnysunny
yunnysunny deleted the feat/udp-rtp-multicast branch October 3, 2026 13:47
@yunnysunny
yunnysunny restored the feat/udp-rtp-multicast branch October 3, 2026 13:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant