Skip to content

fix stream RX in RX threads on LAG member interfaces - #403

Merged
GIC-de merged 1 commit into
rtbrick:devfrom
vorco:fix-lag-member-rx-thread
Oct 3, 2026
Merged

GIC-de merged 1 commit into
rtbrick:devfrom
vorco:fix-lag-member-rx-thread

Conversation

@smpettit

Copy link
Copy Markdown

Stream packets received on a LAG member with RX threads enabled are dropped without being counted if the frame is larger than 4074 bytes. The stream is then reported as never received, although the frames arrive on the interface, so it looks like loss in the device under test. This affects, for example, downstream jumbo streams to access sessions on a LAG, which can be configured since jumbo-frames support was added (#385).

Cause

bbl_rx_thread() looks up the network and access interfaces on the receiving member, but these are bound to the LAG interface. The lookup fails, so every stream packet received on the member is redirected to the main thread through the TXQ. TXQ slots hold at most BBL_TXQ_BUFFER_LEN (4074) bytes; redirect() rejects larger frames with IO_ERROR, which the packet_mmap RX job does not handle, so they are dropped without a counter.

Fix

Resolve a LAG member to its LAG interface in bbl_rx_thread(), as bbl_rx_handler() already does on the main thread. Stream packets on LAG members are then handled in the RX threads instead of all going through the main thread.

Testing

  • Builds without warnings, ctest 4/4 passed (-DBNGBLASTER_TESTS=ON, RelWithDebInfo, gcc 13.3, Ubuntu 24.04).

  • A/B with the same config, only the binary changed: access session (static IPoE, QinQ) behind a LAG with one LACP member, rx-threads: 2, io-mode: packet_mmap_raw, jumbo-frames: true, one stream per frame size and direction at 1000 pps for 60 s. Received / sent:

    frame size downstream, before downstream, after upstream, before and after
    1500 63026 / 63026 60965 / 60965 all received
    2000 63026 / 63026 60965 / 60965 all received
    4000 63027 / 63027 60965 / 60965 all received
    9000 0 / 63027 60965 / 60965 all received

    In both runs the member NIC's port.rx_size_big counter increased by exactly the 2000, 4000 and 9000 byte downstream packets sent, so the 9000 byte frames did arrive before the fix.

🤖 Generated with Claude Code

Stream packets received on a LAG member with RX threads enabled were
dropped without being counted if the frame was larger than 4074 bytes,
so the stream was reported as never received even though the frames
arrived on the interface. This affects, for example, downstream jumbo
streams to access sessions on a LAG, which can be configured since
jumbo-frames support was added (rtbrick#385).

bbl_rx_thread() looked up the network and access interfaces on the
receiving member, but these are bound to the LAG interface. The lookup
failed, so every stream packet was redirected to the main thread
through the TXQ, whose slots hold at most BBL_TXQ_BUFFER_LEN (4074)
bytes; larger frames were rejected by redirect() without a counter.

Resolve a LAG member to its LAG interface, as bbl_rx_handler() already
does on the main thread. Stream packets on LAG members are then handled
in the RX threads instead of all going through the main thread.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants