Skip to content

feat: bucketized futexes - #2469

Open
zyuiop wants to merge 5 commits into
hermit-os:mainfrom
zyuiop:feat/bucketized-futexes
Open

feat: bucketized futexes#2469
zyuiop wants to merge 5 commits into
hermit-os:mainfrom
zyuiop:feat/bucketized-futexes

Conversation

@zyuiop

@zyuiop zyuiop commented Jun 6, 2026

Copy link
Copy Markdown
Contributor

While tracking #2468, I initially suspected the futex lock to be an issue, so I applied the "Todo" and made a bucket list instead of the single lock.

I have based this off on #2468 so that we get performance results that make sense.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmark Results

Details
Benchmark Current: 49d45ab Previous: 2e23902 Performance Ratio
startup_benchmark Build Time 86.58 s 80.34 s 1.08
startup_benchmark File Size 0.80 MB 0.80 MB 1.00
Startup Time - 1 core 0.72 s (±0.02 s) 0.75 s (±0.02 s) 0.97
Startup Time - 2 cores 0.76 s (±0.02 s) 0.74 s (±0.02 s) 1.04
Startup Time - 4 cores 0.74 s (±0.02 s) 0.74 s (±0.02 s) 1.00
multithreaded_benchmark Build Time 90.47 s 82.11 s 1.10
multithreaded_benchmark File Size 0.90 MB 0.86 MB 1.05
Multithreaded Pi Efficiency - 2 Threads 88.15 % (±10.21 %) 85.89 % (±6.61 %) 1.03
Multithreaded Pi Efficiency - 4 Threads 44.03 % (±3.55 %) 43.43 % (±2.56 %) 1.01
Multithreaded Pi Efficiency - 8 Threads 25.58 % (±1.84 %) 25.76 % (±1.53 %) 0.99
micro_benchmarks Build Time 84.33 s 80.40 s 1.05
micro_benchmarks File Size 0.91 MB 0.86 MB 1.05
Scheduling time - 1 thread 66.71 ticks (±2.07 ticks) 62.65 ticks (±4.06 ticks) 1.06
Scheduling time - 2 threads 37.29 ticks (±4.98 ticks) 34.08 ticks (±4.10 ticks) 1.09
Micro - Time for syscall (getpid) 3.46 ticks (±0.54 ticks) 3.45 ticks (±0.58 ticks) 1.00
Memcpy speed - (built_in) block size 4096 76134.60 MByte/s (±52957.15 MByte/s) 82448.38 MByte/s (±56997.13 MByte/s) 0.92
Memcpy speed - (built_in) block size 1048576 29919.11 MByte/s (±24149.80 MByte/s) 30585.98 MByte/s (±24707.84 MByte/s) 0.98
Memcpy speed - (built_in) block size 16777216 28420.96 MByte/s (±23417.75 MByte/s) 26340.06 MByte/s (±21720.96 MByte/s) 1.08
Memset speed - (built_in) block size 4096 76454.50 MByte/s (±53215.96 MByte/s) 82292.76 MByte/s (±56891.50 MByte/s) 0.93
Memset speed - (built_in) block size 1048576 30673.11 MByte/s (±24593.70 MByte/s) 31323.85 MByte/s (±25145.86 MByte/s) 0.98
Memset speed - (built_in) block size 16777216 29161.94 MByte/s (±23841.85 MByte/s) 27104.68 MByte/s (±22209.94 MByte/s) 1.08
Memcpy speed - (rust) block size 4096 72907.88 MByte/s (±51026.49 MByte/s) 74097.96 MByte/s (±51811.44 MByte/s) 0.98
Memcpy speed - (rust) block size 1048576 29768.67 MByte/s (±24146.58 MByte/s) 30361.60 MByte/s (±24602.37 MByte/s) 0.98
Memcpy speed - (rust) block size 16777216 28654.25 MByte/s (±23608.00 MByte/s) 27625.34 MByte/s (±22806.88 MByte/s) 1.04
Memset speed - (rust) block size 4096 73314.30 MByte/s (±51323.91 MByte/s) 74373.47 MByte/s (±51976.48 MByte/s) 0.99
Memset speed - (rust) block size 1048576 30517.99 MByte/s (±24580.34 MByte/s) 31110.89 MByte/s (±25033.24 MByte/s) 0.98
Memset speed - (rust) block size 16777216 29406.75 MByte/s (±24038.64 MByte/s) 28386.93 MByte/s (±23265.03 MByte/s) 1.04
alloc_benchmarks Build Time 84.10 s 74.76 s 1.12
alloc_benchmarks File Size 0.88 MB 0.87 MB 1.00
Allocations - Allocation success 91.38 % 91.31 % 1.00
Allocations - Deallocation success 100.00 % 100.00 % 1
Allocations - Pre-fail Allocations 61.60 % 61.44 % 1.00
Allocations - Average Allocation time 10577.88 Ticks (±143.96 Ticks) 5860.58 Ticks (±98.43 Ticks) 1.80
Allocations - Average Allocation time (no fail) 11212.09 Ticks (±175.80 Ticks) 6554.81 Ticks (±92.86 Ticks) 1.71
Allocations - Average Deallocation time 3582.78 Ticks (±489.29 Ticks) 1805.01 Ticks (±250.35 Ticks) 1.98
mutex_benchmark Build Time 85.09 s 79.82 s 1.07
mutex_benchmark File Size 0.91 MB 0.86 MB 1.05
Mutex Stress Test Average Time per Iteration - 1 Threads 12.64 ns (±0.62 ns) 12.10 ns (±0.41 ns) 1.04
Mutex Stress Test Average Time per Iteration - 2 Threads 99.22 ns (±5.86 ns) 40.26 ns (±1.68 ns) 2.46

This comment was automatically generated by workflow using github-action-benchmark.

@mkroening mkroening left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The code looks okay to me. Unless it is too much effort, could you make the bucket map implementation generic? Did you measure any speedup in any stress test with this?

CC: @joboet for the original implementation.

@mkroening mkroening self-assigned this Aug 22, 2026
@zyuiop

zyuiop commented Aug 24, 2026

Copy link
Copy Markdown
Contributor Author

I can indeed make it generic, it's just that AFAIK there is no other usecase in the code for now for such a map. So this may be a case of "premature abstraction" :D

But if you insist, I'll gladly do it :)

Comment thread src/synch/futex.rs
}

fn hash_key(v: usize) -> usize {
let v = (v >> 3).to_be_bytes();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you're trying to remove the zero bits resulting from AtomicU32's alignment, then you have to shift by 2, not 3. Since you're hashing anyway, I don't think this is necessary anyway.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I was trying to remove the last 3 bits from the address, which for some reason I believed to always be 0

But you are correct that this is in any case not needed since we hash. It may be the remain of an attempt to not use a hash function at all for better performance.

If it produced a more or less uniformly distributed distribution, (addr >> 3) % N would likely be better as it is less intensive than computing a hash.

Let me know what you think :)

Comment thread src/synch/futex.rs Outdated
type Bucket = InterruptSpinMutex<TaskListBucket>;

#[repr(transparent)]
struct TaskListBucket(LinkedList<BucketElem>);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using a linked-list here is a bit unfortunate – if there are a lot of tasks, the $O(n)$ lookup will start to matter. You could try using hashbrown::HashTable as an inner map to avoid having to recompute the hash while keeping hashmap-like lookup performance. Or alternatively, use a BTreeMap – that will also reduce memory usage as addresses become unused.

@zyuiop zyuiop Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think there is the same reasoning here... Wanting to avoid doing two hashes of the key. But I can switch to a BTreeMap.
And also it was likely easier to think in terms of iterators on a linked list than on a map

@zyuiop
zyuiop force-pushed the feat/bucketized-futexes branch from 2f2342e to d3446a9 Compare September 8, 2026 16:50
@zyuiop
zyuiop requested a review from joboet September 8, 2026 16:51
@zyuiop
zyuiop force-pushed the feat/bucketized-futexes branch from d3446a9 to 49d45ab Compare September 9, 2026 13:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants