feat: bucketized futexes - #2469
Conversation
There was a problem hiding this comment.
Benchmark Results
Details
| Benchmark | Current: 49d45ab | Previous: 2e23902 | Performance Ratio |
|---|---|---|---|
| startup_benchmark Build Time | 86.58 s |
80.34 s |
1.08 ❗ |
| startup_benchmark File Size | 0.80 MB |
0.80 MB |
1.00 ❗ |
| Startup Time - 1 core | 0.72 s (±0.02 s) |
0.75 s (±0.02 s) |
0.97 |
| Startup Time - 2 cores | 0.76 s (±0.02 s) |
0.74 s (±0.02 s) |
1.04 ❗ |
| Startup Time - 4 cores | 0.74 s (±0.02 s) |
0.74 s (±0.02 s) |
1.00 |
| multithreaded_benchmark Build Time | 90.47 s |
82.11 s |
1.10 ❗ |
| multithreaded_benchmark File Size | 0.90 MB |
0.86 MB |
1.05 ❗ |
| Multithreaded Pi Efficiency - 2 Threads | 88.15 % (±10.21 %) |
85.89 % (±6.61 %) |
1.03 |
| Multithreaded Pi Efficiency - 4 Threads | 44.03 % (±3.55 %) |
43.43 % (±2.56 %) |
1.01 |
| Multithreaded Pi Efficiency - 8 Threads | 25.58 % (±1.84 %) |
25.76 % (±1.53 %) |
0.99 |
| micro_benchmarks Build Time | 84.33 s |
80.40 s |
1.05 ❗ |
| micro_benchmarks File Size | 0.91 MB |
0.86 MB |
1.05 ❗ |
| Scheduling time - 1 thread | 66.71 ticks (±2.07 ticks) |
62.65 ticks (±4.06 ticks) |
1.06 |
| Scheduling time - 2 threads | 37.29 ticks (±4.98 ticks) |
34.08 ticks (±4.10 ticks) |
1.09 |
| Micro - Time for syscall (getpid) | 3.46 ticks (±0.54 ticks) |
3.45 ticks (±0.58 ticks) |
1.00 |
| Memcpy speed - (built_in) block size 4096 | 76134.60 MByte/s (±52957.15 MByte/s) |
82448.38 MByte/s (±56997.13 MByte/s) |
0.92 |
| Memcpy speed - (built_in) block size 1048576 | 29919.11 MByte/s (±24149.80 MByte/s) |
30585.98 MByte/s (±24707.84 MByte/s) |
0.98 |
| Memcpy speed - (built_in) block size 16777216 | 28420.96 MByte/s (±23417.75 MByte/s) |
26340.06 MByte/s (±21720.96 MByte/s) |
1.08 |
| Memset speed - (built_in) block size 4096 | 76454.50 MByte/s (±53215.96 MByte/s) |
82292.76 MByte/s (±56891.50 MByte/s) |
0.93 |
| Memset speed - (built_in) block size 1048576 | 30673.11 MByte/s (±24593.70 MByte/s) |
31323.85 MByte/s (±25145.86 MByte/s) |
0.98 |
| Memset speed - (built_in) block size 16777216 | 29161.94 MByte/s (±23841.85 MByte/s) |
27104.68 MByte/s (±22209.94 MByte/s) |
1.08 |
| Memcpy speed - (rust) block size 4096 | 72907.88 MByte/s (±51026.49 MByte/s) |
74097.96 MByte/s (±51811.44 MByte/s) |
0.98 |
| Memcpy speed - (rust) block size 1048576 | 29768.67 MByte/s (±24146.58 MByte/s) |
30361.60 MByte/s (±24602.37 MByte/s) |
0.98 |
| Memcpy speed - (rust) block size 16777216 | 28654.25 MByte/s (±23608.00 MByte/s) |
27625.34 MByte/s (±22806.88 MByte/s) |
1.04 |
| Memset speed - (rust) block size 4096 | 73314.30 MByte/s (±51323.91 MByte/s) |
74373.47 MByte/s (±51976.48 MByte/s) |
0.99 |
| Memset speed - (rust) block size 1048576 | 30517.99 MByte/s (±24580.34 MByte/s) |
31110.89 MByte/s (±25033.24 MByte/s) |
0.98 |
| Memset speed - (rust) block size 16777216 | 29406.75 MByte/s (±24038.64 MByte/s) |
28386.93 MByte/s (±23265.03 MByte/s) |
1.04 |
| alloc_benchmarks Build Time | 84.10 s |
74.76 s |
1.12 ❗ |
| alloc_benchmarks File Size | 0.88 MB |
0.87 MB |
1.00 ❗ |
| Allocations - Allocation success | 91.38 % |
91.31 % |
1.00 ❗ |
| Allocations - Deallocation success | 100.00 % |
100.00 % |
1 |
| Allocations - Pre-fail Allocations | 61.60 % |
61.44 % |
1.00 ❗ |
| Allocations - Average Allocation time | 10577.88 Ticks (±143.96 Ticks) |
5860.58 Ticks (±98.43 Ticks) |
1.80 ❗ |
| Allocations - Average Allocation time (no fail) | 11212.09 Ticks (±175.80 Ticks) |
6554.81 Ticks (±92.86 Ticks) |
1.71 ❗ |
| Allocations - Average Deallocation time | 3582.78 Ticks (±489.29 Ticks) |
1805.01 Ticks (±250.35 Ticks) |
1.98 ❗ |
| mutex_benchmark Build Time | 85.09 s |
79.82 s |
1.07 ❗ |
| mutex_benchmark File Size | 0.91 MB |
0.86 MB |
1.05 ❗ |
| Mutex Stress Test Average Time per Iteration - 1 Threads | 12.64 ns (±0.62 ns) |
12.10 ns (±0.41 ns) |
1.04 |
| Mutex Stress Test Average Time per Iteration - 2 Threads | 99.22 ns (±5.86 ns) |
40.26 ns (±1.68 ns) |
2.46 ❗ |
This comment was automatically generated by workflow using github-action-benchmark.
17f45b2 to
2f2342e
Compare
|
I can indeed make it generic, it's just that AFAIK there is no other usecase in the code for now for such a map. So this may be a case of "premature abstraction" :D But if you insist, I'll gladly do it :) |
| } | ||
|
|
||
| fn hash_key(v: usize) -> usize { | ||
| let v = (v >> 3).to_be_bytes(); |
There was a problem hiding this comment.
If you're trying to remove the zero bits resulting from AtomicU32's alignment, then you have to shift by 2, not 3. Since you're hashing anyway, I don't think this is necessary anyway.
There was a problem hiding this comment.
I was trying to remove the last 3 bits from the address, which for some reason I believed to always be 0
But you are correct that this is in any case not needed since we hash. It may be the remain of an attempt to not use a hash function at all for better performance.
If it produced a more or less uniformly distributed distribution, (addr >> 3) % N would likely be better as it is less intensive than computing a hash.
Let me know what you think :)
| type Bucket = InterruptSpinMutex<TaskListBucket>; | ||
|
|
||
| #[repr(transparent)] | ||
| struct TaskListBucket(LinkedList<BucketElem>); |
There was a problem hiding this comment.
Using a linked-list here is a bit unfortunate – if there are a lot of tasks, the hashbrown::HashTable as an inner map to avoid having to recompute the hash while keeping hashmap-like lookup performance. Or alternatively, use a BTreeMap – that will also reduce memory usage as addresses become unused.
There was a problem hiding this comment.
I think there is the same reasoning here... Wanting to avoid doing two hashes of the key. But I can switch to a BTreeMap.
And also it was likely easier to think in terms of iterators on a linked list than on a map
2f2342e to
d3446a9
Compare
d3446a9 to
49d45ab
Compare
While tracking #2468, I initially suspected the futex lock to be an issue, so I applied the "Todo" and made a bucket list instead of the single lock.
I have based this off on #2468 so that we get performance results that make sense.