Skip to content

stt: faster PD nearest neighbors with shared-coordinate handling - #11602

Merged
maliberty merged 15 commits into
The-OpenROAD-Project:masterfrom
The-OpenROAD-Project-staging:pdr-dbl-monotone
Oct 2, 2026
Merged

maliberty merged 15 commits into
The-OpenROAD-Project:masterfrom
The-OpenROAD-Project-staging:pdr-dbl-monotone

Conversation

@maliberty

Copy link
Copy Markdown
Member

Summary

Builds on #11351 by @gzz2000, which replaces the O(n²) Prim-Dijkstra nearest neighbor search with the double monotone chains algorithm from "Provably Optimal Planar Pareto Nearest Neighbor Search with Double Monotone Chains" (Guo et al, DATE'26). Their commits are kept on this branch.

The paper assumes distinct coordinates. Pins often share x or y (rows, sites, macro edges), and there the chain walk linked every pair of pins on a vertical line: 499,500 edges for 1000 pins instead of 999. This PR fixes that:

  • Neighbors are defined by an empty closed bounding box, so a pin on the boundary also excludes the pair. This is safe for PD: routing through such a pin w changes the cost by (α−1)·d(u,w) ≤ 0.
  • Coincident pins are merged and linked to the first pin at their location.
  • Pins are swept a column of equal x at a time. Column mates are linked to each other and bound the rows that can hold left neighbors.
  • Each staircase step is found with a max segment tree over rows instead of chain pointers. Chain pointers can need O(k) updates per inserted pin when a column has k pins.
  • Neighbor lists are sorted, so the PD search no longer depends on the order the sweep found them.

The search moves to src/stt/src/pdr/src/neighbors.{h,cpp}. A new gtest, TestNeighbors, compares it against a brute-force search on tie-heavy random inputs and covers collinear, grid, column-plus-staircase and duplicate-pin cases. It is registered in both CMake and Bazel.

Impact

With distinct coordinates the neighbor sets are identical to the old code. With shared coordinates the sets are smaller. On 10k pins:

input old entries new entries old time new time
random 311,093 310,692 271 ms 22 ms
rows/sites 391,053 159,238 127 ms 9 ms
100×100 grid 1,019,700 39,600 57 ms 1 ms
vertical line 19,998 19,998 53 ms 1 ms

The pd1 and pd_gcd goldens only change which of two coincident pins carries a connection. Wire length and path depth are unchanged for every net. ORFS QoR has not been run yet.

gzz2000 and others added 9 commits September 30, 2026 23:55
Signed-off-by: Zizheng Guo <gzz_2000@126.com>

Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Signed-off-by: Zizheng Guo <gzz_2000@126.com>

Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
The double monotone chains neighbor search assumes distinct
coordinates.  With shared x values it linked every pair of pins on a
vertical line (499,500 edges for 1000 pins instead of 999), which is
common for pins on rows, sites and macro edges.

Neighbors are now defined by an empty closed bounding box, so a pin
on the boundary also excludes the pair.  This is safe for PD since
routing through such a pin w changes the cost by
(alpha - 1) * d(u,w) <= 0.

The search keeps the paper's sweep in x with staircase walks but:
- merges coincident pins, linking them to the first at the location
- sweeps a column of equal x at a time, linking column mates and
  bounding the neighbor rows by them
- finds each staircase step with a max segment tree over rows rather
  than chain pointers, which can need O(k) updates per pin when a
  column has k pins

Each neighbor list is sorted so the PD search is independent of the
sweep order.  The search moves to neighbors.{h,cpp} and has a gtest
comparing it against brute force on tie-heavy inputs.

The pd1 and pd_gcd goldens only change which of two coincident pins
carries a connection; wire length and path depth are unchanged.

Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
@maliberty maliberty self-assigned this Oct 1, 2026
@github-actions github-actions Bot added the size/L label Oct 1, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new pdr_neighbors library implementing an efficient double monotone chains algorithm for nearest neighbor search, replacing the previous O(N^2) implementation in pd.cpp. It also adds comprehensive unit tests and updates build configurations and test expectations. The review feedback focuses on performance optimizations in neighbors.cpp, including adding fast-path checks in the segment tree queries, replacing std::make_tuple with direct coordinate comparisons in the sorting comparator to avoid overhead, and pre-reserving capacity for the column rows vector to prevent redundant reallocations.

Comment thread src/stt/src/pdr/src/neighbors.cpp
Comment thread src/stt/src/pdr/src/neighbors.cpp
Comment thread src/stt/src/pdr/src/neighbors.cpp
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
@maliberty
maliberty marked this pull request as ready for review October 1, 2026 05:45
@maliberty
maliberty requested a review from a team as a code owner October 1, 2026 05:45
@maliberty
maliberty requested a review from eder-matheus October 1, 2026 05:45
@maliberty

Copy link
Copy Markdown
Member Author

@gzz2000 I would be interested in your feedback on Claude's solution to the aligned points issue from your PR.

@oharboe

oharboe commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

This is a really neat optimization. Prim-Dijkstra brings back memories from university, and seeing it get a provably optimal neighbor search with proper handling of shared coordinates is very cool.

I was curious what the impact would be on XiangShan, so I looked at per-step perf profiles I had lying around from master on XSTile (asap7, about 4 M nets, 48 threads). This is how much of each step is in the Steiner tree code:

step samples in pdr:: in any Steiner code (stt, Flute, pdr)
3_4_place_resized 0.3 % 0.3 %
3_5_place_dp 1.0 % 1.2 %
4_1_cts 0.4 % 1.0 %
5_1_grt 0.1 % 0.1 %
estimate_parasitics -placement 0.2 % 0.3 %

Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
@maliberty

Copy link
Copy Markdown
Member Author

You need a very high fanout net to see a big benefit but some designs have those (eg global reset).

@maliberty maliberty mentioned this pull request Oct 1, 2026
4 of 5 tasks
@oharboe

oharboe commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

We tried this on XiangShan (KunMingHu, asap7, hardened per block) and saw little effect. Here's why, in case it helps anyone else measuring it.

XiangShan's reset is asserted asynchronously and released synchronously. In the generated Verilog, 994 always blocks are @(posedge clock or posedge reset) and none uses a synchronous if (reset). The release is pipelined by ResetGen (3 flops) and the AsyncResetSynchronizers, but those sit in XSTop and at the clock-crossing queues, outside the tile. Inside each hardened block, reset is therefore one very-high-fanout net with a pin on every flop: the case you describe.

It only stays that way until the first repair_design, though. We cap fanout at 32 with set_max_fanout, so the resizer breaks reset into a buffer tree of small nets. After that, CTS, global route and detailed route never see a high-fanout net. The faster search only gets used on the few Steiner trees built for the whole unbuffered net, during timing-driven global placement and that first repair_design. Going by the table above, that's seconds, against stages that take hours.

We also false-path reset (set_false_path -from [get_ports reset]). For now we treat reset as a solved problem that isn't interesting to study, so we keep its recovery/removal checks out of the timing picture. We may come back to it later. It isn't why we see little effect, though: the net is still estimated and buffered with its full fanout.

We'd expect a large benefit where a high-fanout net stays unbuffered: an ideal network, dont-touch, or a flow with no fanout limit.

@gzz2000

gzz2000 commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

I re-run part of the flow regression locally on the latest pdr-dbl-monotone commit (I had to ln -s ~/install/OpenROAD/bin/openroad ./build/bin/ before I run, which looks a bit hacky. Also I could not run it without the flow while it is documented to run all tests. Maybe I missed some configurations or the README documentation is stale).

There are cases that do not pass. Looks like one of them is a numerical difference:

zizhengg@5900ea8871e9:~/OpenROAD$ ./test/regression flow
------------------------------------------------------
Flow
gcd_nangate45 (tcl) pass
gcd_sky130hd (tcl) pass
gcd_sky130hd_fast_slow (tcl) pass
gcd_sky130hs (tcl) pass
gcd_asap7 (tcl) pass
ibex_sky130hd (tcl) *FAIL* DRT::max_slew_slack  -6% <  -6%
ibex_sky130hs (tcl) pass
aes_nangate45 (tcl) pass
aes_sky130hd (tcl) *FAIL* DRT::ANT::errors   1 >   0
aes_sky130hs (tcl) pass
aes_asap7 (tcl) pass
tinyRocket_nangate45 (tcl) pass
...

I think a segment tree to handle shared-coordinates handling (as done by Claude) is a bit overkill for this problem.

If our goal is a deterministic result and to avoid time complexity degradation on ties, an easier solution exists: we can run the algorithm on tie-breaker-embedded coordinates, such as letting x' := (x coordinate, y coordinate, pin id), y' := (y coordinate, x coordinate, pin id).

If we want to match the original brute force exactly: the original algorithm also has its quirks. Notably, its neighbor relations might be asymmetric (A->B does not imply B->A). Maybe we don't need to replicate that completely?

@maliberty

Copy link
Copy Markdown
Member Author

Identical replication of the brute force result is not essential.

@maliberty

Copy link
Copy Markdown
Member Author

From Claude on your suggestion:

Thanks for the suggestion. I tried it: your chain walk on rank coordinates x' = (x, y, id), y' = (y, x, id). It's correct (it's a superset of the closed-box neighbor set) and it fixes the vertical line. But the tie-break acts like a diagonal shear: x + εy, y + εx. Pins on the box edge move inside for lower-left/upper-right pairs and outside for upper-left/lower-right pairs. So the anti-diagonal pairs blow up. At 10k pins: a 100×100 grid gives 1.98M entries (this PR 39.6k, old O(n²) 1.02M) and rows/sites gives 1.20M (this PR 83k). Using one shear per direction would make each direction exact. But the upper-left and lower-left walks share chain pointers, so each pass would still step through the other direction's bad links, about O(n^1.5) on grids. The segment tree keeps both directions output-sensitive. I agree we don't need to match the brute force exactly; the closed-box relation here is two-way. I'll look into the ibex/aes flow differences.

@maliberty

Copy link
Copy Markdown
Member Author

The ibex delta is just noise in detailed routing.

The smaller neighbor sets change Prim-Dijkstra trees on nets with
shared coordinates, which global routing hits on gcell-snapped pins.
Placement and resizing results are unchanged; routing differs from
there on.

- grt: pd1, colocated_pins, overlapping_edges and top_level_term3
  logs and route guides.
- rsz: repair_fanout7_skip_pin_swap and
  repair_setup_split_load_hier_only logs (small slack/TNS changes).
- ibex_sky130hd and aes_sky130hd metrics and limits, regenerated from
  bazel runs.  ibex DRT::ANT::errors and aes DRT max slew/cap limits
  keep their previous values so both master and this branch pass.

Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
…onotone

Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
@openroad-ci
openroad-ci requested review from a team as code owners October 2, 2026 03:41
@openroad-ci
openroad-ci requested a review from a team as a code owner October 2, 2026 03:41
@github-actions github-actions Bot added size/XL and removed size/L labels Oct 2, 2026
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
@maliberty
maliberty merged commit 7331ca9 into The-OpenROAD-Project:master Oct 2, 2026
20 checks passed
@maliberty
maliberty deleted the pdr-dbl-monotone branch October 2, 2026 20:47
@maliberty

Copy link
Copy Markdown
Member Author

I've looked at the secure CI results and there are some variations but nothing noteworthy.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants