Add UDP/RTP multicast playback with FFmpeg software decoding - #7
Conversation
Multicast playback - udp:// and rtp:// channels via a purpose-built MulticastDataSource. Media3's UdpDataSource is final and lacks what live IPTV needs: a configurable SO_RCVBUF, interface-aware joinGroup (boxes with both eth0 and wlan0 receive nothing if the wrong one is joined), and rtp:// at all (Media3's RTP is coupled to RTSP sessions). - Packets are received on a dedicated thread feeding a bounded ring buffer. UDP is push-based, so a full kernel buffer means permanent loss; letting ExoPlayer's Loader thread do receive() meant a GC pause or a struggling decoder turned straight into burst loss. Measured: rtpLost 71 -> 0. - RTP header stripping (CSRC/extension/padding) plus a 32-packet reorder window, zero-allocation on the receive path. Streams that declare rtp:// but carry raw TS fall back to passthrough automatically. - Address normalisation for the VLC (udp://@) and SSM (rtp://@src@group) spellings, and optional udpxy rewriting to HTTP unicast, which is the only workable option on most home networks. Configurable on TV and in the web admin page. - Stalls rejoin the multicast group in place before giving up. IGMP membership gets pruned upstream, which is why re-entering a channel used to fix it; the app now does that itself, bounded so a dead stream still switches source. - Live TS extractor tuning (MODE_SINGLE_PMT) is scoped to multicast only; its single-PMT assumption is unsafe for TS of unknown origin over HTTP. - Separate buffer profiles for multicast and HTTP/HLS sharing one memory pool, so switching does not rebuild the player. FFmpeg software decoding - nextlib-media3ext supplies decoders boxes commonly lack, notably MPEG-2 video and MP2 audio. Its version encodes the Media3 version and must be upgraded in lockstep. - Decoder mode setting: hardware first (default), software first, or hardware only. Changing it rebuilds the player and resumes the channel. - Release builds keep ARM ABIs only; debug keeps all four for the emulator. Robustness - Fix a crash when switching groups during channel-list fling. Recycling a focused row triggers rootViewRequestFocus(), which lands on a group item and synchronously calls setAdapter() on the list still being recycled. - Add a watchdog for stalling after playback has started. The start timeout was cancelled on READY and never rearmed, so a later stall was unsupervised. Recovery restarts the current source before switching, since switching is equivalent to giving up on a single-source channel. - Diagnostics: throughput heartbeat, stall and SO_RCVBUF warnings, rebuffer counter, and ExoPlayer's EventLogger in debug builds. 198 unit tests pass; assembleDebug and assembleRelease both succeed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 37af29dd8f
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if (playbackState == Player.STATE_READY) { | ||
| cancelTimeout(); | ||
| cancelRebufferWatchdog(); | ||
| consecutiveStallRecoveries = 0; |
There was a problem hiding this comment.
Preserve stall recovery count across transient READY
When a broken source briefly reaches STATE_READY after each restart and then falls back to STATE_BUFFERING, resetting consecutiveStallRecoveries here causes every watchdog invocation to be treated as the first attempt. This is exactly the short READY→BUFFERING cycle diagnosed below, and it means MAX_STALL_RECOVERIES is never reached, so the player can restart the same bad source forever instead of switching sources.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
确认属实,已在 59e983e 修复。
这条特别值得记:我上一轮正是从日志里诊断出「重开 → READY 15ms → 又卡」这个形态才加的看门狗,却又在 STATE_READY 里无条件 consecutiveStallRecoveries = 0,等于让看门狗在它最该起作用的场景里失效。
改法:READY 不再重置预算,只有持续播放超过 STALL_RECOVERY_RESET_AFTER_MS(30 秒)才算真正恢复。换频道、手动切线路、自动换源仍然给满新预算。
新增两个回归用例:
transientReadyDoesNotRefillStallRecoveryBudget—— 短暂 READY 不续杯,预算耗尽后必须换源sustainedPlaybackRefillsStallRecoveryBudget—— 稳定播放后应重新获得预算
| if (cm == null) { | ||
| return null; | ||
| } | ||
| Network active = cm.getActiveNetwork(); |
There was a problem hiding this comment.
Join the interface that carries the multicast group
On devices where the IPTV multicast network is not Android's default network—for example, IPTV on Ethernet while Wi-Fi remains the default Internet connection—getActiveNetwork() selects the wrong interface. Joining a multicast group on that interface normally succeeds, so the fallback is never attempted, but no stream packets arrive; the new direct-multicast feature therefore fails in the multi-interface scenario this selection logic is intended to support.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
确认属实,已在 59e983e 修复。
这条是六条里最有价值的。讽刺的是 resolveActiveMulticastInterface() 这个方法本来就是为多网卡场景写的,却在最典型的多网卡场景里选错网卡:IPTV 常接在没有公网的以太网口上,Android 不会把这种网络当作活动网络;此时在 WiFi 上 joinGroup 会成功(不抛异常),但一个组播包都收不到,而我那个基于异常的回退永远不会触发。
改法:在所有可用组播网卡上一并 join(活动网络优先,其余随后),筛选条件为非 loopback、已 up、支持组播、且有 IPv4 地址。代价只是几个 IGMP 报文。一张都没成功时才回退到系统路由。leaveGroup 和重新入组同步改为遍历列表,列表在 open() 后不再修改,可跨线程安全读取。
日志里的 iface= 现在会列出全部已加入的网卡,例如 iface=eth0+wlan0。
| throw wrap(new IOException("Data source closed"), | ||
| PlaybackException.ERROR_CODE_IO_NETWORK_CONNECTION_FAILED); | ||
| } | ||
| long waitMs = socketTimeoutMs * 2L; |
There was a problem hiding this comment.
Wait through the final multicast rejoin attempt
With the defaults, the reader times out at 3 seconds, rejoins twice, and only declares failure around 9 seconds, but the consumer gives up after socketTimeoutMs * 2 (6 seconds). For an initially silent or stalled stream, this races with the second rejoin and usually aborts playback without giving that final attempt any time to receive a packet, defeating the configured two-attempt recovery budget.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
确认属实,已在 59e983e 修复。
算过了:收包线程首次超时 3s + 重入组 2 次各 3s ≈ 9s 才判定断流,而消费侧只等 socketTimeoutMs * 2 = 6s。最后一次重入组还没来得及收到包就被放弃,等于我白配了那个预算。
改法:抽出 consumerWaitMs(socketTimeoutMs, maxRejoinAttempts),按 (重入组次数 + 1) 个超时周期再加 2 秒余量计算。收包线程彻底失败时会写 readerError 并关闭环形缓冲,消费侧会被立刻唤醒,所以这个超时只是兜底。
新增 consumerOutwaitsTheWholeRejoinBudget,把「消费侧必须等得比收包线程久」这个不变量钉死,以后任一侧改参数都会红。
| if (player == null || player.getPlaybackState() != Player.STATE_BUFFERING) { | ||
| return; | ||
| } | ||
| onRebufferStall(timeoutMs); |
There was a problem hiding this comment.
Suppress stall recovery while playback is paused
If the activity is paused while the player is buffering after a previous successful start, the already-armed watchdog still passes this state-only check and calls playCurrentSource(), which unconditionally sets playWhenReady back to true. A stalled stream can therefore restart and resume audio/video in the background after PlayerActivity.onPause() explicitly paused it; the watchdog should also require playback to be requested or be canceled from pause().
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
确认属实,已在 59e983e 修复。
后台偷偷续播是实打实的问题,不只是多耗电——用户按了暂停或退出,音频却自己回来了。
改法三处:
pause()取消看门狗resume()若仍处于STATE_BUFFERING则重新武装(否则暂停期间的卡顿在恢复后就没人看管了)- 看门狗触发时额外校验
player.getPlayWhenReady(),作为第二道防线
新增 pauseCancelsRebufferWatchdog 和 resumeRearmsWatchdogWhileBuffering。
| if (diff >= SEQ_HALF) { | ||
| // 序号早于当前期望:迟到或重复,直接丢弃 | ||
| latePackets++; | ||
| return; |
There was a problem hiding this comment.
Reset reordering when the RTP sender changes sequence space
When an RTP sender restarts mid-session with a new, lower initial sequence number, every packet whose modular difference falls in the backward half of the sequence space is classified as merely late and discarded. Because the implementation does not track SSRC changes or repeated backward discontinuities, playback can drop up to 32,768 packets before the new sequence catches the old expectation instead of recovering immediately from the sender restart.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
确认属实,已在 59e983e 修复。
把边界算清楚了:新起点落在当前期望值之前的半个序号空间时才会中招(落在之后的话 diff >= capacity 会判为跳变并自动重置)。最坏丢 32767 个包,按实测 760 pkt/s 约 43 秒黑屏。
改法:连续迟到达到 LATE_PACKETS_BEFORE_RESYNC(64)即认定序号空间已变,按新起点重新同步。没有按 SSRC 判断是因为 SSRC 在 RtpPacketUtil 里尚未解析,而连续迟到这个判据对「发送端重启」和「源切换」都成立,更通用。
阈值 64 在 760 pkt/s 下约 85ms,正常网络抖动不会连续迟到这么多次。配了反向用例 occasionalLatePacketsDoNotTriggerResync 确保零星迟到包不会误触发。
| boolean multicastOrigin = MulticastUrlUtil.isMulticastStreamUrl(rawUrl); | ||
| String url = MulticastUrlUtil.resolvePlaybackUrl(rawUrl, preferenceManager().getUdpxyProxyBase()); | ||
| currentResolvedUrl = url; |
There was a problem hiding this comment.
Bypass the HLS disk cache for udpxy streams
When the optional live-TS disk cache is enabled, resolving a multicast URL to HTTP here sends the unbounded udpxy response through HlsSegmentPrefetcher's playback data source. Its LoggingPlaybackDataSource selects cacheDelegate even for URLs that are not .ts segments, so a continuous udpxy stream is written to the 96 MB cache and continuously evicted rather than merely streamed, causing sustained disk I/O and cache churn for as long as the channel plays.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
确认属实,已在 59e983e 修复。
这条我特地跨文件核对了 LoggingPlaybackDataSource.open():isTsRequest 为 false 时 bypassCache 保持 false,直接落到 activeDelegate = cacheDelegate,确实会把 udpxy 的无界流写进 96MB 磁盘缓存并持续淘汰。盒子的闪存不该这么耗。
改法:条件从 if (bypassCache) 放宽为 if (bypassCache || !isTsRequest),非 .ts 请求一律走上游。这同时也符合该功能本来的定位——设置项就叫「使用硬盘缓存预取直播分片」,缓存对象本就该限定为 ts 分片。
副作用:fMP4 封装的 HLS 分片也不再走缓存。考虑到该设置默认关闭且文案明确限定 ts,这个取舍是合理的。
Six issues flagged on the PR, all confirmed against the code. - Stall recovery budget was reset on every STATE_READY, but the failure mode it guards against is exactly "restart, READY for 15ms, stall again". The cap was therefore never reached and a bad source could be restarted forever instead of switched. Only sustained playback (30s) refills it now. - The watchdog could fire while the activity was paused and set playWhenReady back to true, resuming audio in the background. pause() now cancels it, resume() rearms it if still buffering, and the watchdog checks playWhenReady before acting. - Interface selection used getActiveNetwork() alone. IPTV is commonly on an Ethernet port with no internet, which Android does not make the active network; joining on Wi-Fi then succeeds without error and no packets arrive, so the exception-based fallback never triggered. The group is now joined on every usable multicast interface. - A sender restarting with a lower RTP sequence number put every packet in the backward half of the sequence space, so all of them were discarded as late - up to 32767 packets, roughly 43 seconds of black screen. 64 consecutive late packets now resync to the new sequence. - The consumer waited socketTimeoutMs * 2 (6s) while the reader needs about 9s to exhaust its rejoin budget, so playback was abandoned before the final attempt could receive anything. The wait is now derived from the rejoin budget. - udpxy-proxied multicast is not a .ts request, so LoggingPlaybackDataSource still routed it through the cache delegate, writing an unbounded live stream into the 96MB disk cache and evicting it continuously. Non-TS requests now bypass the cache. 206 unit tests pass; assembleDebug and assembleRelease both succeed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Setup Android SDK has been failing with "Failed to find package 'tools'" before any code is compiled. android-actions/setup-android defaults its packages input to 'tools platform-tools', and Google has since removed the obsolete 'tools' package from the SDK repository, so sdkmanager exits 1. Unrelated to the branch contents: main's last green run predates the removal, and the step runs before checkout output is ever built. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
评审反馈处理完毕6 条行内评审全部逐条核对过代码,均属实,已在
其中 1、2 两条尤其值得注意,都是「代码看起来正确、却在它本该覆盖的场景里恰好失效」:
这两类问题靠读代码很难发现,模拟器上也测不出来(单网卡、源不会重启)。 另外:CI 已修复本 PR 的 CI 之前连续失败,原因与本分支内容无关—— 已在 现 CI 通过:build pass 2m38s。 如果你希望 CI 修复单独走一个 PR,我可以把那个提交摘出来。 当前状态
仍建议合入前在真实电视盒子上过一轮组播播放:目前所有实测都在 x86 模拟器镜像上完成,该环境 1080p 持续掉帧 8~17 fps(数据链路已验证干净: |
做了什么
新增
udp:///rtp://组播直播支持,并引入 FFmpeg 软解补齐盒子硬解常缺的编码。期间在模拟器上实测暴露出三个既有缺陷,一并修了。
详见
docs/multicast-udp-rtp.md(配置方式、地址写法、实现要点、排查手册)。组播播放
MulticastDataSource。Media3 的UdpDataSource是final且缺三样直播必需的能力:可配SO_RCVBUF、按活动网卡joinGroup(盒子常同时有 eth0 和 wlan0,join 错网卡一个包都收不到)、以及rtp://支持(Media3 的 RTP 耦合在 RTSP 会话里)。PacketRingBuffer)。UDP 是推模式,内核缓冲一满就永久丢包;原先由 ExoPlayer 的 Loader 线程直接receive(),而那条线程还要做 TS 解复用和内存分配,GC 一停顿或解码器一挣扎就变成突发丢包。实测rtpLost由 71 降到 0。rtp://但实际推裸 TS 时自动回退透传。udp://@、SSM 风格rtp://@src@group等写法;配置代理后改走 HTTP 单播——这是绝大多数家宽环境下唯一可行的方案。TV 设置页和 Web 管理页都能配。MODE_SINGLE_PMT的「只有一个 PMT」假定对来源不明的 HTTP TS 不成立,会导致 PID 映射走样。FFmpeg 软解
nextlib-media3ext,补上组播 TS 高频的 MPEG-2 视频和 MP2/AC3/DTS 音频。<media3版本>-<nextlib版本>,升级 Media3 时必须同步升级它,否则运行时抛NoSuchMethodError。顺带修掉的既有缺陷
rootViewRequestFocus(),焦点落到分组列表项上并同步调用setAdapter(),而那个列表还在回收中 →Cannot call removeView(At) within removeView(At)。STATE_READY时被取消后从未重新武装,之后的卡顿完全无人看管。新增看门狗,优先原地重开当前线路(单线路频道换源等于放弃)。SO_RCVBUF被内核压缩告警、重缓冲计数,以及 debug 构建挂载 ExoPlayer 官方EventLogger。测试
assembleDebug/assembleRelease均成功malformed=0,自动重新入组生效已知限制 / 待确认
SwitchableLoadControl的组播档开了prioritizeTimeOverSizeThresholds,因为DefaultLoadControl的字节计数是 per-instance 的、组播委托看不到。副作用是时间戳错乱时内存上限只能靠重缓冲看门狗兜底。彻底解决需要自己实现LoadControl,会动到目前工作正常的 HLS 主路径,本次未做。rtpLost=0、overflow=0、不再重缓冲),瓶颈是该镜像的解码能力。建议合入前在真实电视盒子上复核一轮组播播放。🤖 Generated with Claude Code