Interface sign-off / 06
UCIe 3.0 Signal Integrity: 64 GT/s Chiplet Links vs PCIe
UCIe is the open standard for wiring chiplets together inside one package. Its links are millimetres long, so loss, which dominates a board link, barely matters. What runs out first is room along the die edge to place enough bumps, followed by the energy each bit costs, and the crosstalk between hundreds of single-ended wires packed side by side.
UCIe at a glance
What is UCIe?
Universal Chiplet Interconnect Express: a die-to-die interface, published by the UCIe Consortium, that lets chiplets from different designers talk inside one package. It defines the physical layer, a die-to-die adapter, and the mapping of protocols such as PCIe and CXL onto it, plus a raw mode for anything else.
What is the difference between UCIe and PCIe?
The distance they cross, and so what they spend their effort on. PCIe crosses a board, tens of centimetres, with a few fast differential lanes, heavy equalisation and a clock recovered from the data. UCIe crosses millimetres inside a package, with many slower single-ended lanes, a clock sent alongside the data and little or no equalisation, and its scarce resources are die edge and energy per bit. They are not rivals: UCIe can carry the PCIe and CXL protocols over its own physical layer. More below.
How fast is UCIe?
4, 8, 12, 16, 24 and 32 GT/s per lane, and UCIe 3.0 adds 48 and 64 GT/s (C-64). A lane is one single-ended, one-way wire; a module groups 16 lanes on a standard package and 64 on an advanced one (C-65).
Standard, advanced and 3D: what is the difference?
The package. A standard package routes between dies through an organic substrate, with bumps 100 to 130 µm apart and reach up to 25 mm. An advanced package uses a silicon bridge or interposer, with bumps 25 to 55 µm apart and reach up to 2 mm (C-66, C-67). UCIe-3D stacks one die on another with hybrid bonding, at a pitch of 10 µm or less (C-66).
How much bandwidth does UCIe give?
The consortium quotes it per millimetre of die edge, counting both directions together: 28 to 224 GB/s/mm on a standard package and 165 to 1317 GB/s/mm on an advanced one across 4 to 32 GT/s, and 370 and 2634 GB/s/mm at 64 GT/s (C-68, C-76). The panel below turns that into the edge a given bandwidth needs.
Does UCIe use PAM4?
No. Even at 48 and 64 GT/s it stays NRZ, with a forwarded clock at a quarter of the data rate (C-70). A wire this short has little loss to save, so the eye margin that PAM4 would cost is not worth paying.
Best viewed on a laptop or desktop. These panels are built so you can move a slider and watch several charts answer at once. A phone has no room to put them side by side.
—
What UCIe is
A large processor used to be one die. Past a certain size that stops working: the lithography field limits how big one die can be, and a large die yields worse than several small ones. So the design is split into chiplets, and the question becomes how they talk. Inside a package they are millimetres apart, not the tens of centimetres of a board link, and that changes almost every signal integrity decision.
UCIe vs PCIe: why a die-to-die link is not a SerDes
A PCIe or Ethernet SerDes is built for a long, lossy channel: a few fast differential lanes, heavy equalisation, and a clock recovered from the data. Each lane is expensive in power and area, and that is acceptable because there are few of them and the channel leaves no choice.
Inside a package the trade turns over. The channel is short enough that loss hardly matters, and what is scarce is the edge of the die, because every wire needs a bump and bumps can only be packed so tightly. So UCIe goes wide and simple: many single-ended lanes, each at a modest rate, with the clock sent alongside the data on its own wires. A forwarded clock sees nearly the same delay and noise as the data beside it, so much of the jitter cancels, and no clock-recovery loop is needed. That saves power and latency, which is the point of the whole design.
The x16 module of a standard package at 32 GT/s carries 512 Gb/s in each direction: sixteen lanes at 32 Gb/s (C-65). A PCIe 6.0 x16 link carries twice that, but over a board, with sixteen differential pairs and the equalisation to cross it.
Shoreline: the number that sizes the design
If wires per millimetre of edge are fixed by the bump pitch, bandwidth per millimetre is that count times the data rate. The consortium quotes it directly: 224 GB/s/mm on a standard package and 1317 GB/s/mm on an advanced one at 32 GT/s, for bump pitches of 110 and 45 µm (C-68). The bumps are about 2.44 times finer, and the bandwidth per millimetre is about 5.9 times higher, because a finer pitch packs more rows of bumps into the same depth as well as more along the edge.
Those figures count both directions together. The consortium’s chair gives the size of a module at 32 GT/s (C-76): a 64-lane advanced-package module is 0.389 mm wide and sends 256 GB/s each way, and 512 GB/s over 0.389 mm is 1316 GB/s/mm, the table’s figure. A 32-lane standard-package module is 1.143 mm wide and sends 128 GB/s each way, and 256 GB/s over 1.143 mm is 224 GB/s/mm. Count one direction only and both come out at half the table’s figure. So for a link that has to carry a given bandwidth each way, the edge it needs is twice what the one-way figure divided by the shoreline suggests.
The table’s own figures confirm the linear scaling with rate: divide the 32 GT/s numbers by eight and you get 28.0 and 164.6 GB/s/mm, against the 28 and 165 it lists for 4 GT/s. Above 32 GT/s the advanced package keeps scaling, 2634 GB/s/mm at 64 GT/s, twice its 32 GT/s figure. The standard package does not: 370 GB/s/mm, only 1.65 times its 32 GT/s figure (C-68). The overview does not say why.
That is the number an architect sizes a chiplet with. Moving 4 TB/s in total across one edge, 2 TB/s each way, at 32 GT/s takes about 3.04 mm of edge on an advanced package and about 17.9 mm on a standard one. A large die has a few tens of millimetres per side, and memory, I/O and power all want some of it.
When does a 2 mm wire need termination?
The critical-length rule says that at these edge rates even 2 mm is a transmission line. An open receiver can still work, if every echo has died away before the receiver samples, in the middle of the bit. So the first test is the round trip against half a unit interval.
With a dielectric constant near 4, tpd is about 6.7 ps/mm, so a 2 mm advanced-package wire has a round trip of about 26.8 ps. At 16 GT/s half a unit interval is 31.25 ps, and the echoes settle before the sample. At 32 GT/s it is 15.6 ps: the first echo now arrives after the sample, and an open receiver samples the overshoot the driver’s mismatch leaves. For a 40 Ω driver on a 50 Ω line the first arrival overshoots by 2·50/90 − 1 = 11.1% of the swing, and in the panel below, where earlier bits leave their own residue, the worst mid-bit sample is off by about 12.5%. A driver matched to the line would leave none, but driver impedance moves with process, voltage and temperature. A receiver termination removes the echo whatever the driver does, which is the robust answer once rates rise, and UCIe 3.0 requires it on both package types at 48 and 64 GT/s (C-70).
Half a unit interval is the test for a driver close to the line’s impedance, where the first echo is nearly all there is. Each round trip multiplies the echo by ΓSΓL, so the further the driver is from 50 Ω, the more round trips the echo needs to die away, and the more of them have to fit before the sample. Set the panel to 1.5 mm at 16 GT/s with a 20 Ω driver: the round trip is only 0.32 of a unit interval, and the worst sample still misses by about 20%, because the second arrival undershoots just before the sample. A 25 mm standard-package wire has a round trip of about 335 ps: 1.34 unit intervals at 4 GT/s and 5.36 at 16 GT/s, so its reflections have to be controlled at every rate.
Drag the driver to 50 Ω with the receiver open and the echo vanishes at any rate: a matched driver absorbs it, which is series termination. Move it away and the open receiver samples the overshoot. Terminate the receiver and the driver stops mattering. The model is a lossless 50 Ω line with resistive ends and no pad or receiver capacitance, so it shows the mechanism, not a sign-off margin.
Inside a 48 and 64 GT/s link
UCIe 3.0 doubles the rate without changing what a lane is. What changes is the clock, the training and a little equalisation, all published in the consortium’s overview (C-72 to C-74).
The clock. At 64 GT/s the forwarded clock runs at 16 GHz and at 48 GT/s at 12 GHz, a quarter of the data rate, and quarter rate is the only option there. At 24 and 32 GT/s the table lists both a half-rate and a quarter-rate clock, and from 4 to 16 GT/s only half rate. Each is listed with two phases, 90° and 270° at half rate and 45° and 135° at quarter rate. Deskew, aligning each lane to the clock, is required at every rate from 12 GT/s up and optional only at 4 and 8 GT/s (C-72).
The training. During link training the receiver calibrates its clock phases, then tries the transmitter’s equaliser presets and keeps the one with the best receive eye margin (C-73). The six presets are three-tap FFE settings, one pre-cursor and one post-cursor tap around the main one. Their taps always add up to 1 in magnitude, so the high-frequency swing stays fixed while the low-frequency level drops; the drop is the de-emphasis:
| Preset | C(−1) | C(0) | C(+1) | low-frequency level | de-emphasis |
|---|---|---|---|---|---|
| P0 | 0 | 1 | 0 | 1.00 | 0 dB |
| P1 | −0.05 | 0.95 | 0 | 0.90 | 0.92 dB |
| P2 | 0 | 0.9 | −0.1 | 0.80 | 1.94 dB |
| P3 | −0.05 | 0.85 | −0.1 | 0.70 | 3.10 dB |
| P4 | 0 | 0.8 | −0.2 | 0.60 | 4.44 dB |
| P5 | −0.05 | 0.75 | −0.2 | 0.50 | 6.02 dB |
Six presets topping out at 6 dB is very little next to a board SerDes, which can equalise tens of decibels. That is the short channel showing again: there is little loss to undo.
The power. The target at these rates is 0.5 to 0.75 pJ/b, of which roughly 40% goes to the transmitter, 40% to the receiver and 20% to common circuits (C-74): at 0.5 pJ/b, about 0.2 pJ/b each end and 0.1 shared. To save some of it, UCIe 3.0 lets the transmitter, not only the receiver, adjust clock-to-data skew during runtime recalibration, using the wider range it already has from link start-up (C-75).
Where a 15.6 ps unit interval goes
With loss out of the way, what closes the eye of a forwarded-clock link is timing: every lane in a 64-lane module has to land its bit inside the window of the clock that samples it. At 64 GT/s that window is 15.6 ps, and length alone eats it fast. One millimetre of mismatch on a 6.7 ps/mm interposer is 6.7 ps, 43% of the unit interval. Training removes that static part: UCIe requires deskew at every rate from 12 GT/s up (C-72).
What deskew cannot remove is what changes from bit to bit:
- Crosstalk-induced jitter. A victim’s delay depends on what its neighbours are doing, because coupled lines switching together and apart travel at different speeds. With hundreds of lanes side by side that is a pattern-dependent timing error on every lane.
- Supply noise that is not common. The forwarded clock cancels jitter that hits clock and data alike. Noise that reaches a data driver but not the clock driver, or arrives at a different time, does not cancel; see supply-induced jitter.
- Residual reflections, from the panel above, as voltage noise that the slope of the edge turns into timing error.
- The receiver’s own window, the setup and hold time its sampler needs.
A full timing budget lays these out against the unit interval as a DDR timing budget does. This page does not build one, because every term depends on the package and the circuits, and the consortium publishes none of them. What carries over is the structure: because deskew is trained, the static terms are removed before the eye is measured, and the dynamic ones have to fit in what is left at the target error rate.
Crosstalk, and the bump field
Hundreds of single-ended wires, millimetres long and packed as tightly as the bump field and the routing layers allow, couple. Each victim sees every aggressor around it, and with no differential pair to cancel common noise, the crosstalk budget is spent on spacing, shielding wires and the return path. The return current flows through the ground bumps nearest the signal, so the ratio of signal bumps to ground bumps, and where the grounds sit, is as much an SI decision as trace spacing. Switching many lanes together also pulls current from the same local supply, which is simultaneous switching noise on a supply shared by many drivers.
Energy per bit is the budget
Die-to-die power is quoted in picojoules per bit, and multiplying by the bit rate gives watts. The consortium’s targets are 0.5 pJ/b on a standard package up to 16 GT/s and 0.75 at 32 GT/s and above, and 0.25 to 0.5 pJ/b on an advanced package, rising with rate (C-69). Moving 1 TB/s at 0.3 pJ/b costs 2.4 W; 4 TB/s costs 9.6 W. The higher rates buy edge with energy: 64 GT/s halves the edge an advanced package needs, at 0.5 pJ/b against 0.3, 1.67 times the energy per bit.
UCIe-3D: from an edge to an area
Stack one die on another and the interface is no longer limited to an edge; it can use the whole overlap. The number of wires then grows as the inverse square of the bump pitch. The consortium gives about 4 TB/s/mm² at a 9 µm pitch and about 300 TB/s/mm² at 1 µm (C-71). The square law checks it: (9/1)² = 81, and 81 times 4 TB/s/mm² is 324 TB/s/mm². With that many wires none of them has to be fast: UCIe-3D runs at up to 4 GT/s (C-64), and the consortium keeps its circuits simple and its frequency low, at the dies’ own logic clock, because power is what matters; the near-zero distance between the dies removes most of the wire’s parasitics (C-77). The target is below 0.05 pJ/b at a 9 µm pitch (C-69). The SI problem becomes a power-delivery and thermal one.
What it takes to sign off
- Lane-to-lane skew against the forwarded clock, across the whole module, because the module works only if every lane lands inside the window of the clock it shares.
- Crosstalk from the worst-case neighbourhood, with every adjacent lane switching, not a single aggressor.
- Supply noise during simultaneous switching, as timing jitter: supply-induced jitter on a forwarded clock is only partly common to the data.
- Reflections at the chosen rate: whether the round trip still fits inside half a unit interval, or termination and equalisation are needed.
- Training margins: the link trains its clock-to-data alignment and, at 48 and 64 GT/s, its equaliser presets, so margin is reported after training, as it is for DDR5.
Go deeper: the arithmetic, and what the figures do and do not say
Edge for a bandwidth. Edge (mm) = bandwidth (GB/s) / shoreline (GB/s/mm). At or below 32 GT/s the panel scales the consortium’s 32 GT/s shoreline linearly with rate, which fixed geometry implies and the table’s 4 GT/s figures confirm to within rounding; at 48 and 64 GT/s it uses the table’s own figures.
Which direction? The table does not say, but the module sizes do (C-76): only the total of both directions reproduces the table, 224 and 1316 GB/s/mm. The panel’s bandwidth is that total, so a link that must carry 2 TB/s each way is a 4 TB/s link on the panel. One loose end: the 0.389 mm advanced-package module is quoted at a 55 µm pitch and the table at 45 µm, yet the width reproduces the table; neither source explains the difference.
Energy figures are targets, not measurements of any implementation. The table gives none for a standard package at 24 GT/s, so the panel shows the range between its neighbours.
Error rates. At 48 GT/s the bit error rate target is 10−15 and at 64 GT/s 10−12, with CRC and replay to recover errors (C-70). A raw error rate that high is acceptable only because the link layer catches and resends what the physical layer gets wrong, the same bargain as PCIe 6.0 makes with its FLIT and FEC.
Equalisation, at 48 and 64 GT/s only: a three-tap transmit FFE (one pre-cursor, one post-cursor), a first-order passive receive CTLE, and an optional one-tap receive DFE (C-70). Compare a board SerDes, which relies on equalisation at every rate.
The round-trip estimate uses 6.7 ps/mm, the site’s figure for a dielectric constant near 4; an interposer oxide is close to that, and an organic build-up film somewhat lower. It sets a scale, not a sign-off number.
In the real world
The decision that matters most is made before any SI analysis: which package. At 32 GT/s an advanced package buys about 5.9 times the bandwidth per millimetre of edge at 40% of the energy per bit, and pays for it in package cost and assembly. Once that is chosen, the data rate trades edge against energy, and the SI work is to prove that a wide, dense, single-ended bus keeps its skew, crosstalk and supply noise inside a unit interval that shrinks as the rate rises. Read package families for the packaging side, and package interconnects for when a package wire stops being a lumped element.
Where the numbers come from
The UCIe Consortium’s overview deck and its UCIe 3.0 white paper, and module sizes from a 2025 talk by the consortium’s chair (ledger C-64 to C-77). The UCIe 3.0 specification is available from the consortium on request and has not been read for this page. Everything else is arithmetic on those figures, checked by the site’s number gate.
Related
Sources
- UCIe Consortium — overview deck, 2025, with the key metrics for UCIe 3.0 C-64
- UCIe Consortium — UCIe 3.0 white paper C-70
- Das Sharma, UCIe for on-package memory, Hot Interconnects (HOTI) 2025 slides C-76
Rows marked with a claim id are tracked in the claim ledger, which records what each source can and cannot establish.