Conversation
Signed-off-by: Zizheng Guo <gzz_2000@126.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Signed-off-by: Zizheng Guo <19143357+gzz2000@users.noreply.github.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Signed-off-by: Zizheng Guo <gzz_2000@126.com> Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
The double monotone chains neighbor search assumes distinct
coordinates. With shared x values it linked every pair of pins on a
vertical line (499,500 edges for 1000 pins instead of 999), which is
common for pins on rows, sites and macro edges.
Neighbors are now defined by an empty closed bounding box, so a pin
on the boundary also excludes the pair. This is safe for PD since
routing through such a pin w changes the cost by
(alpha - 1) * d(u,w) <= 0.
The search keeps the paper's sweep in x with staircase walks but:
- merges coincident pins, linking them to the first at the location
- sweeps a column of equal x at a time, linking column mates and
bounding the neighbor rows by them
- finds each staircase step with a max segment tree over rows rather
than chain pointers, which can need O(k) updates per pin when a
column has k pins
Each neighbor list is sorted so the PD search is independent of the
sweep order. The search moves to neighbors.{h,cpp} and has a gtest
comparing it against brute force on tie-heavy inputs.
The pd1 and pd_gcd goldens only change which of two coincident pins
carries a connection; wire length and path depth are unchanged.
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
There was a problem hiding this comment.
Code Review
This pull request introduces a new pdr_neighbors library implementing an efficient double monotone chains algorithm for nearest neighbor search, replacing the previous O(N^2) implementation in pd.cpp. It also adds comprehensive unit tests and updates build configurations and test expectations. The review feedback focuses on performance optimizations in neighbors.cpp, including adding fast-path checks in the segment tree queries, replacing std::make_tuple with direct coordinate comparisons in the sorting comparator to avoid overhead, and pre-reserving capacity for the column rows vector to prevent redundant reallocations.
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
|
@gzz2000 I would be interested in your feedback on Claude's solution to the aligned points issue from your PR. |
|
This is a really neat optimization. Prim-Dijkstra brings back memories from university, and seeing it get a provably optimal neighbor search with proper handling of shared coordinates is very cool. I was curious what the impact would be on XiangShan, so I looked at per-step
|
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
|
You need a very high fanout net to see a big benefit but some designs have those (eg global reset). |
|
We tried this on XiangShan (KunMingHu, asap7, hardened per block) and saw little effect. Here's why, in case it helps anyone else measuring it. XiangShan's reset is asserted asynchronously and released synchronously. In the generated Verilog, 994 always blocks are It only stays that way until the first We also false-path reset ( We'd expect a large benefit where a high-fanout net stays unbuffered: an ideal network, dont-touch, or a flow with no fanout limit. |
|
I re-run part of the flow regression locally on the latest pdr-dbl-monotone commit (I had to There are cases that do not pass. Looks like one of them is a numerical difference: I think a segment tree to handle shared-coordinates handling (as done by Claude) is a bit overkill for this problem. If our goal is a deterministic result and to avoid time complexity degradation on ties, an easier solution exists: we can run the algorithm on tie-breaker-embedded coordinates, such as letting x' := (x coordinate, y coordinate, pin id), y' := (y coordinate, x coordinate, pin id). If we want to match the original brute force exactly: the original algorithm also has its quirks. Notably, its neighbor relations might be asymmetric (A->B does not imply B->A). Maybe we don't need to replicate that completely? |
|
Identical replication of the brute force result is not essential. |
|
From Claude on your suggestion: Thanks for the suggestion. I tried it: your chain walk on rank coordinates x' = (x, y, id), y' = (y, x, id). It's correct (it's a superset of the closed-box neighbor set) and it fixes the vertical line. But the tie-break acts like a diagonal shear: x + εy, y + εx. Pins on the box edge move inside for lower-left/upper-right pairs and outside for upper-left/lower-right pairs. So the anti-diagonal pairs blow up. At 10k pins: a 100×100 grid gives 1.98M entries (this PR 39.6k, old O(n²) 1.02M) and rows/sites gives 1.20M (this PR 83k). Using one shear per direction would make each direction exact. But the upper-left and lower-left walks share chain pointers, so each pass would still step through the other direction's bad links, about O(n^1.5) on grids. The segment tree keeps both directions output-sensitive. I agree we don't need to match the brute force exactly; the closed-box relation here is two-way. I'll look into the ibex/aes flow differences. |
|
The ibex delta is just noise in detailed routing. |
The smaller neighbor sets change Prim-Dijkstra trees on nets with shared coordinates, which global routing hits on gcell-snapped pins. Placement and resizing results are unchanged; routing differs from there on. - grt: pd1, colocated_pins, overlapping_edges and top_level_term3 logs and route guides. - rsz: repair_fanout7_skip_pin_swap and repair_setup_split_load_hier_only logs (small slack/TNS changes). - ibex_sky130hd and aes_sky130hd metrics and limits, regenerated from bazel runs. ibex DRT::ANT::errors and aes DRT max slew/cap limits keep their previous values so both master and this branch pass. Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
…onotone Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
Signed-off-by: Matt Liberty <mliberty@precisioninno.com>
|
I've looked at the secure CI results and there are some variations but nothing noteworthy. |
Summary
Builds on #11351 by @gzz2000, which replaces the O(n²) Prim-Dijkstra nearest neighbor search with the double monotone chains algorithm from "Provably Optimal Planar Pareto Nearest Neighbor Search with Double Monotone Chains" (Guo et al, DATE'26). Their commits are kept on this branch.
The paper assumes distinct coordinates. Pins often share x or y (rows, sites, macro edges), and there the chain walk linked every pair of pins on a vertical line: 499,500 edges for 1000 pins instead of 999. This PR fixes that:
The search moves to
src/stt/src/pdr/src/neighbors.{h,cpp}. A new gtest,TestNeighbors, compares it against a brute-force search on tie-heavy random inputs and covers collinear, grid, column-plus-staircase and duplicate-pin cases. It is registered in both CMake and Bazel.Impact
With distinct coordinates the neighbor sets are identical to the old code. With shared coordinates the sets are smaller. On 10k pins:
The
pd1andpd_gcdgoldens only change which of two coincident pins carries a connection. Wire length and path depth are unchanged for every net. ORFS QoR has not been run yet.