Skip to content

Linux: support RESOLVE_NO_XDEV when openat2 is unavailable #423

Description

@amnabbouti

TL;DR: On Linux, cap-primitives falls back to its manual path resolver when
openat2() is unavailable, but the fallback cannot enforce
RESOLVE_NO_XDEV. Downstream users therefore cannot safely reject mount
crossings on systems where openat2() returns ENOSYS.

Context: This surfaced through bootc's open_dir_noxdev() usage under QEMU
linux-user, but the missing behavior is general and is not specific to bootc or
QEMU. The downstream report is
bootc-dev/bootc#1481.

Current behavior

cap-primitives/src/rustix/linux/fs/open_impl.rs uses openat2() for its fast
path and falls back to manually::open() when the syscall is unavailable. The
manual resolver preserves capability containment and resolves symlinks one
component at a time, but it has no NO_XDEV policy.

As a result, cap-std-ext implements open_dir_noxdev() with a direct
openat2() call. It handles EXDEV, but propagates ENOSYS instead of using a
capability-safe fallback.

Reproduction

Component Value
Host and VM x86_64, Fedora 44
Container aarch64 Fedora bootc
Affected emulator QEMU linux-user 8.2.9
Working emulator QEMU linux-user 10.2.2
cap-primitives 4.0.3
cap-std-ext 5.1.2

With QEMU 8.2.9 registered through binfmt_misc:

$ podman run --rm --arch arm64 \
    quay.io/fedora/fedora-bootc:latest \
    bootc container lint
error: Linting: Function not implemented (os error 38)

The trace reports unsupported AArch64 syscall 437 immediately before the
failure. That syscall is openat2, same image passes under QEMU 10.2.2,
which implements openat2.

Required behavior

The fallback must reject crossing any mount, including same filesystem bind
mounts and mounts reached through symlinks, while retaining the manual
resolver's containment and race-resistance properties.

A final st_dev comparison is insufficient because a bind mount may have the
same device and inode as its source. Checking only the final descriptor is also
insufficient: an intermediate component or symlink may cross a mount even when
the final component is not a mount root. Each opened component must therefore
be checked before symlink interpretation or descent.

If reliable mount information is unavailable, the operation should return an
explicit unsupported error rather than silently perform a normal open.

Proposed direction

  1. Keep openat2() as the fast path and add RESOLVE_NO_XDEV when requested.
  2. On ENOSYS, reuse the existing manual component resolver rather than add a
    separate path walker.
  3. Check every newly opened descriptor for a mount crossing before following a
    symlink or descending into that component.
  4. Return EXDEV for a detected crossing and fail closed when mount identity
    cannot be determined reliably.
  5. Expose this as an existing-directory operation through the public layer the
    maintainers consider appropriate, such as cap-fs-ext. Existing opens
    should remain unchanged when the policy is not requested.

STATX_ATTR_MOUNT_ROOT appears suitable for the fallback, under QEMU 8.2.9 it
correctly identified all 153 mount points visible in /proc/self/mountinfo, a
same filesystem bind mount, and descriptors held across mount and detach
operations. STATX_MNT_ID was not reliable in this environment because QEMU
reported the field as available but returned zero for tested paths.

The initial API should be limited to opening existing directories, applying a
general OpenOptions flag to effectful operations such as truncation would
require a broader design because validating the result after opening is too
late.

References

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions