SIPI

Interface sign-off / 07

Chiplet Interconnects: UCIe vs BoW vs AIB Compared

UCIe, BoW and AIB solve the same problem, wide and simple parallel links between dies a few millimetres apart, and they physically look alike: single-ended wires, a clock sent with the data, and little or no equalisation. They differ in how fast each wire runs, when a wire is terminated, and how much of the stack above the physical layer they define. At a given wire rate, how much bandwidth fits along a millimetre of die edge depends mostly on the bump pitch and how deep the layout goes: the published layouts differ more in depth than in how densely their bumps carry data.

At a glance

What are the chiplet interconnect standards?

Three open die-to-die standards with public documents are UCIe, from the UCIe Consortium; BoW (Bunch of Wires), from the Open Compute Project’s ODSA workstream; and AIB (Advanced Interface Bus), from the CHIPS Alliance. All three connect dies inside one package.

UCIe vs BoW: what is the difference?

Speed and scope. BoW runs 2 to 32 Gb/s per wire (C-79) and specifies the physical layer only, leaving protocols to other documents (C-99). UCIe runs 4 to 64 GT/s (C-64) and also defines a die-to-die adapter and the mapping of protocols such as PCIe and CXL. Electrically they are close relatives: a BoW slice is 16 single-ended data wires and a differential clock (C-78).

UCIe vs AIB: what is the difference?

AIB is slower and simpler: up to 6.4 Gb/s per wire in its second generation, with no maximum specified (C-90), and no line termination at all (C-98). It is a physical layer only (C-89). An interoperability guide describes what a PHY on an advanced package must do to serve both UCIe and AIB chiplets (C-97).

Why is HBM not in this comparison?

HBM is a memory interface between a processor and a DRAM stack, defined by JEDEC, not a general die-to-die link between any two chiplets. It solves a related problem with a different contract, so it is left out of a comparison of general die-to-die links.

Bandwidth per millimetre of edge, and per square millimetre of bumps exact UCIe-S — UCIe-A — BoW laminate — BoW interposer — AIB — —
Four published layouts, both directions counted
interposer or bridgeorganic substrateBoW, dashedbump view: a bar spans each standard’s published readings

—

Edge is what a floorplan spends; area is what the bumps allow
16 Gb/s
Every rate either standard defines; used by the edge view
40 µm
Used by the area view: every layout moved to this pitch
—

Three standards, one problem

Inside a package two dies sit millimetres apart, so a wire loses little and the expensive machinery of a board SerDes buys little. What is scarce is the edge of the die, because every wire needs a bump. All three standards therefore make the same trade: many single-ended wires at a modest rate, a clock sent alongside the data so no clock-and-data-recovery loop is needed, and as little equalisation as the channel allows. The UCIe page walks through why; this page compares the three on what their published documents say.

Side by side

UCIe 3.0 from the consortium’s overview and slides; BoW 2.0 and AIB 2.0 from their specifications (ledger rows in brackets)
UCIeBoWAIB
Published byUCIe ConsortiumOCP, ODSA workstreamCHIPS Alliance
Rate per wire4 to 32 GT/s; 48 and 64 GT/s in 3.0 (C-64)2 to 32 Gb/s in six modes; 24 and 32 new in 2.0 (C-79, C-80)up to 2 Gb/s Gen1, 6.4 Gb/s Gen2; no maximum specified (C-90)
Unit of widthmodule: 16 lanes standard, 64 advanced (C-65)slice: 16 data wires and a clock pair, or a half slice of 8 (C-78)channel: 20 to 80 data signals at pitches up to 55 µm, up to 320 at 10 µm (C-93)
Clockforwarded; half or quarter rate (C-72)forwarded, differential, half the wire rate (C-78, C-79)forwarded per channel, quasi-differential, DDR (C-92)
Terminationrequired at the receiver at 48 and 64 GT/s (C-70)none, source or double, by package and rate (C-81, C-83)none defined (C-98)
Reachup to 25 mm standard, 2 mm advanced (C-67)2 mm advanced; up to 25+ mm on laminate, doubly terminated (C-81)10 mm or less, no hard maximum (C-89)
Bump pitch100–130 µm and 25–55 µm (C-66)not specified; examples at 130 and 40 µm (C-87)classes up to 55, 52 and 20 µm (C-94)
Equalisation3-tap TX FFE and CTLE at 48 and 64 GT/s (C-70)RX CTLE required at 24 and 32 Gb/s (C-85)none specified (C-101)
Error or margin targetlink bit error rate 10−15 at 48, 10−12 at 64 GT/s (C-70)transmitter timing budget at 10−15 (C-86)an eye mask, ±50 mV at the far end in Gen2 (C-95)
Energy figuretargets, 0.25 to 0.75 pJ/b (C-69)targets, under 0.25 to 1 pJ/b (C-88)none in the specification (C-96)
Layers definedPHY, die-to-die adapter, protocol mappingPHY (C-99)PHY (C-89)

Read the table as three answers to one question, how fast can a simple wire go before it needs help. AIB’s headline rate is 6.4 Gb/s, with no maximum set, and it defines no termination. BoW goes to 32 Gb/s and adds termination and a receive CTLE as the rate and the wire length grow. UCIe goes to 64 GT/s and, at its top two rates, requires both termination and equalisation. One row needs care: the error and margin targets are three different kinds of thing, a link error rate, a probability for a timing budget and an eye mask, so their numbers do not compare directly.

Bandwidth per millimetre is mostly depth

The figure that sizes a chiplet is bandwidth per millimetre of die edge, and it is tempting to read it as a property of the standard. The published layouts say otherwise. BoW’s specification includes a worked example: eight transmit and eight receive slices at 16 Gb/s per wire carry 2.0 Tb/s each way, and on an interposer with a 40 µm bump pitch they take 1.60 mm of edge and 0.42 mm of depth (C-87). Counting both directions, that is 320 GB/s/mm. A UCIe-A module at the same 16 GT/s carries about 658 GB/s/mm (C-68, C-76), about 2.06 times as much, although BoW’s bumps are finer.

Edge density is the product of two things: how much bandwidth each square millimetre of bumps carries, and how deep the layout goes. The UCIe-A module is 1.585 mm deep (C-76), about 3.8 times the BoW example’s depth. Per square millimetre of its own bumps, at the pitches as built, it carries about 0.55 times as much as the BoW example, because its bumps are coarser. 3.8 times 0.55 is the 2.06: the gap along the edge is depth, partly given back by bump density.

Is one standard’s bump field simply denser? To ask that fairly, move each layout to a common pitch. The number of bumps in an area goes as the inverse square of the pitch, so a layout’s density moves by (its pitch / common pitch)², assuming the same bump arrangement and that its circuits still fit. For UCIe the answer is a range, not a number, because its published figures disagree: one slide labels the UCIe-A module 55 µm while the consortium’s table says its figures are for 45 µm, and the table’s own areal row (C-100) is about 1.3 times denser than the module rectangles. At a common pitch, UCIe-A’s readings run from about 31% below BoW’s example to 37% above it, and UCIe-S’s from about 29% to 6% below. BoW’s one layout sits inside UCIe-A’s range. So, at the same pitch, the published layouts are within about a third of each other per square millimetre, while along the edge they differ twofold. Try it in the panel: choose the bump view and drag the pitch.

Termination: three answers to the same wire

An unterminated receiver is the cheapest and lowest-power choice, and it works while every echo dies before the sample. BoW states the trade most clearly. On an interposer it expects no termination at any of its rates, recommending wires up to 2 mm (4 mm at 2 Gb/s), because interposer traces are short and much more resistive than laminate ones, so termination is not expected to help (C-81, C-82). The physics behind that, in this site’s reading rather than the specification’s words, is that a resistive wire shrinks each echo on its way back. On a laminate, whose traces are less resistive and at least about 3 mm long between neighbouring chips, an unterminated wire is recommended only at 2 Gb/s; a source termination reaches 20, 10 and 5 mm at 2, 4 and 8 Gb/s, and a double termination reaches 25 mm or more at every rate (C-81).

AIB, at a few Gb/s per wire over 10 mm or less, defines no termination at all (C-98). UCIe 3.0 lists a receiver termination as required on both package types at 48 and 64 GT/s (C-70). The termination panel on the UCIe page shows the mechanism on a lossless line.

What it takes to sign off

Go deeper: the arithmetic behind the panel

Both directions. All four layouts are counted the way the UCIe table counts shoreline: the total of both directions (C-76). BoW’s example is 256 data wires, 128 each way, at 16 Gb/s: 4096 Gb/s, 512 GB/s in all. Over 1.60 mm that is 320 GB/s/mm; over the 1.28 mm it needs without the optional AUX and FEC wires, 400. The panel uses the full slice, because a UCIe module’s width also includes its clock and other lanes.

Layouts used. UCIe-S: an x32, two x16 modules, 1.143 by 1.54 mm at 110 µm. UCIe-A: an x64 module, 0.389 by 1.585 mm, labelled 55 µm on the slide that gives it (C-76). BoW: the example link, 5.2 by 1.35 mm on laminate at 130 µm and 1.60 by 0.42 mm on an interposer at 40 µm (C-87), which is one layout at two pitches, not two layouts. The AIB specification gives no comparable layout (C-96), so it has no curve.

Scaling with rate. With the geometry fixed, bandwidth is proportional to the wire rate. UCIe-S at 48 and 64 GT/s is the exception: the consortium’s table gives 278 and 370 GB/s/mm (C-68), less than its 32 GT/s module scaled up. The deck does not say how its layout changes there, so the panel uses the table.

Bump-field density. Bandwidth over edge times depth, divided by the wire rate, gives 47.6 GB/s per mm² per Gb/s for BoW’s interposer example at 40 µm and 25.95 for the UCIe-A module rectangle. This counts the whole rectangle each layout occupies, power and ground bumps included.

Where the UCIe figures disagree. Moved to 40 µm, the UCIe-A rectangle is 25.95 × (55/40)² = 49.1 if its slide’s 55 µm label is right, and 32.8 if the table’s 45 µm applies to it. The table’s areal row gives 1646 GB/s/mm² at 32 GT/s and 45 µm (C-100), 51.4 per Gb/s, which is 65.1 at 40 µm. For UCIe-S the x32 rectangle gives 4.54 and the areal row 6.0, both at 110 µm. The areal row is about 1.32 times the rectangle for UCIe-S, and for UCIe-A at its 55 µm label, so the two sources likely draw the rectangle differently; neither says how. The panel shows every reading and shades the range.

BoW’s own density figure. BoW quotes a target of up to 2 to 12+ Tbps per mm of chip edge (C-88) without saying whether it counts one direction or both, so the panel does not use it.

In the real world

The standard is usually chosen for you, by the chiplets you can buy and the ecosystem you are joining, and the three are closer electrically than their names suggest. What remains an engineering decision is the layout: how fine a bump pitch the package allows, how deep the interface may reach into the die, and whether the wire is short and resistive enough to run open. Those set the bandwidth per millimetre, the power and the margin. Read UCIe for the termination and timing detail, and package families for the packaging side.

Where the numbers come from

BoW and AIB from their specifications: the BoW PHY Specification 2.0 (OCP, March 2023) and the AIB Specification 2.0.3 (CHIPS Alliance, June 2022), with an Intel guide to UCIe and AIB interoperability. UCIe from the consortium’s overview deck and its chair’s slides, not from the UCIe 3.0 specification, which is available on request and has not been read for this page. Every figure is quoted with its section in the claim ledger (C-78 to C-101); the comparisons are arithmetic on those figures, checked by the site’s number gate.

Related

Sources

Rows marked with a claim id are tracked in the claim ledger, which records what each source can and cannot establish.