Interface sign-off / 02
DDR5 Signal Integrity: Fly-by, ODT, DFE and Training
DDR5 runs a single-ended data bus at up to 8800 MT/s through wiring that a serial link would never accept: a connector, a shared wire with a second DIMM hanging off it, a clock that reaches each DRAM at a different time. Its signal integrity problem is reflections and timing, not loss, and its answer is to move termination, equalisation and calibration into the DRAM itself and to train the whole bus at power-up.
The panel above is one data bit on a two-slot server bus. The left chart is a map of the wiring against time: distance runs across, from the controller to the DRAM being written, with the branch to the other slot as its own strip, and time runs down. A travelling bit is a diagonal stripe. Where it meets a change of impedance, part of it turns round, and a reflection is a stripe running the other way. The column on the right is the same bit as the receiver sees it, on the same time axis. Every bump in that column is a stripe arriving. Most of this page is about where those stripes come from and what DDR5 does about each of them.
What DDR5 is
DDR5 SDRAM is JEDEC’s fifth double-data-rate memory generation, the main memory of servers, desktops and laptops. The standard is JESD79-5, first published in 2020 and revised since: JESD79-5C (April 2024) extended the timing definitions to 8800 MT/s, and the current edition, JESD79-5D, appeared in November 2025. Parts and platforms are qualified at particular speed bins, so the bin a given system runs is a property of the parts and the board, not of the generation.
At 3200 MT/s a bit lasts 312.5 ps; at 4800, 208 ps; at 6400, 156 ps; at 8800, 114 ps. The clock runs at half the transfer rate, so at 6400 MT/s one clock period, tCK, is 312.5 ps. Keep those numbers in mind: the wiring on this page is measured in hundreds of picoseconds, so every reflection lands one or more whole bits late.
| DDR4 | DDR5 | Why it matters here | |
|---|---|---|---|
| Transfer rates | 1600–3200 MT/s | 3200 MT/s upward, timing defined to 8800 | bits as short as 114 ps |
| VDD / VDDQ / VPP | 1.2 / 1.2 / 2.5 V | 1.1 / 1.1 / 1.8 V | less swing to spend |
| Data bus per DIMM | one 64-bit channel (72 with ECC) | two independent 32-bit subchannels (40 with ECC) | each subchannel trains and fails on its own |
| Burst length | 8 | 16 | 64 bytes from one subchannel |
| Receiver equalisation in the DRAM | none | 4-tap DFE on DQ | the DRAM cancels part of the ISI itself |
| CA, CS and CK termination | resistors on the module | on the DRAM die, programmable | termination is a setting, per DRAM |
| DRAM power regulation | on the mainboard | a PMIC on the DIMM | the DRAM rails’ noise is now the module’s |
| Training | write leveling, read training | adds CS and CA training, internal write leveling, LFSR read patterns, duty-cycle adjust | more of the bus is calibrated, not designed |
Two things have not changed, and they set the whole problem. The data lines are still single-ended, one wire per bit, because a 72- or 80-bit bus cannot afford a second wire each. And they still pass through a connector, onto a module the board designer did not lay out, often sharing their wire with a second module. If you are weighing DDR5 against LPDDR5X, the comparison is further down.
The module, from above
Two wiring styles on one module, for two different reasons. The data bus carries the fast signals, so it gets the cleanest possible path: point to point, short, one DRAM per wire. The command bus is slower but goes to every DRAM, and a star that split it ten ways would be a stub forest; a fly-by, one line passing each DRAM in turn, has no stubs longer than a via. The price is that the clock reaches each DRAM at a different time, which is what write leveling, further down, exists to fix.
The RCD also changes the command bus’s shape. The controller sends it 7 command and address pins per subchannel at double data rate; it drives the DRAMs with 14 pins at single data rate. Half the pins on the mainboard, at twice the rate, and the freed pins became grounds: the DDR5 RDIMM edge has noticeably more VSS pins than DDR4’s, because every signal on it is referenced to ground.
The signal: pseudo-open drain
DDR5’s data lines use pseudo-open-drain (POD) signalling, at 1.1 V. Read the figure as a voltage divider: every termination on the wire, in parallel, against the driver. That one fact explains a lot of what the panel does. Terminate more strongly, or terminate in two places at once, and the swing shrinks. Terminate less, and the swing grows but reflections are absorbed less well. Every termination choice on a DDR5 bus trades those two against each other.
Because the reference is internal and adjustable, the eye’s vertical centre is not a design value but a trained one. A VrefDQ step is 0.5% of VDDQ, 5.5 mV, fine against a few-hundred millivolt swing.
Two DIMMs on one wire
A server board usually has two slots per channel, and the data line runs from the controller past the first slot to the second: a daisy chain. With both slots filled, a write to the near DIMM still sends part of every bit down the line to the far one. At slot 1 the bit meets a junction. Some of it goes up into the near DIMM, where it was meant to go; some of it carries on to slot 2, and whatever is not absorbed there comes back, a few hundred picoseconds later, and arrives at the near DRAM on top of a later bit.
The arithmetic is worth doing once. In the panel’s model the far branch, from slot 1 to the far DRAM, is 10 mm of trace between slots (67 ps), the connector (25 ps), 20 mm on the module (134 ps) and the DRAM package (30 ps): 256 ps one way and 512 ps there and back. At 4800 MT/s that echo lands 2.5 bits after the bit that launched it; at 6400, 3.3 bits; at 8800, 4.5. The echo itself does not change. What changes is which later bit it lands on.
There is a second echo, and in the panel’s first scenario it is the larger one. The near DIMM’s own DRAM is lightly terminated there (RTT_WR 240 Ω), so it sends much of each bit back down its own branch. At slot 1 that reflection meets the mainboard and the far branch in parallel, a lower impedance than the line it arrives on, so it turns round inverted and reaches the DRAM again one round trip of the near branch later: connector, module trace and package, 189 ps each way and 378 ps there and back, about 1.8 bits at 4800 MT/s. That is the deep dip just after the cursor in the panel’s receiver column. Move the far slot and it stays where it is; terminate the target more strongly and it nearly vanishes, at the price of swing.
Choose “Other DIMM unterminated” in the panel and watch the far strip: the share of the bit that enters it bounces between the open end and the junction, a zig-zag that keeps leaking back into the path. The eye shuts. Turn the other DIMM’s termination back on and the zig-zag fades on its first trip, because the far DRAM now absorbs most of what reaches it.
Board makers’ manuals name the slot to fill first when a channel has one DIMM; ASUS’s, for example, says DIMM_A2 and gives no reason. The panel gives a likely one for a daisy-chained layout: the far slot leaves the shorter stub. Which slot is “far” depends on the board’s routing, so follow the manual rather than the slot number.
Termination is a schedule, not a resistor
On a DDR5 bus every DRAM carries its own on-die termination (ODT), programmable in steps of a 240 Ω reference: 240, 120, 80, 60, 48, 40 and 34 Ω, or off. What makes it powerful is that a DRAM applies a different value depending on what the bus is doing at that moment:
| Transfer | Controller | Near DIMM’s DRAM (the target) | Far DIMM’s DRAM |
|---|---|---|---|
| Write | drives, through Ron | RTT_WR | RTT_NOM_WR, if told of the write; otherwise RTT_PARK |
| Read | terminates, with its own ODT | drives, through its Ron: 34, 40 or 48 Ω | RTT_NOM_RD, if told; otherwise RTT_PARK |
| Idle | — | RTT_PARK | RTT_PARK |
A non-target DRAM learns that a transfer is under way from a command addressed to it, so the controller can switch it between its nominal and park values burst by burst. The board designer’s job is no longer to pick a resistor; it is to find the combination that leaves the best eye at every receiver, for every rank, in both directions.
That is what the panel’s Termination sweep does: it runs the bus with every pair of target and other-DIMM values, 64 simulations, and colours each by its worst-case eye height. Look for the shape of the answer rather than one number. Off on the other DIMM is always bad, because it leaves an open stub. Very strong termination everywhere is never best, because it divides the swing away. The best settings lie between those two, and they move when the data rate, the slot spacing or the target changes. That is why ODT values are swept, not calculated.
The DRAM equalises its own input
DDR4’s DRAM had no equaliser at all. DDR5’s has a decision feedback equaliser (DFE) on every data input, because at these rates the reflections described above no longer fit inside one bit. A DFE works on what the receiver has already decided: if the previous bit was a 1, and a 1 is known to leave a certain residue on the bit after it, subtract that residue before deciding the next one. Because it subtracts a known quantity rather than boosting high frequencies, it does not amplify noise, which is why it suits a reflective channel; a public DDR5 data sheet makes the same point.
Two limits follow, and the panel shows both. A DFE has four taps, so it reaches the four bits after the cursor and no further. The near DIMM’s own 378 ps echo lands on the second or third bit at these rates, inside its reach; the far branch’s 512 ps echo, if the far DIMM returns much of it, lands four and a half bits late at 8800 MT/s, between taps 4 and 5, and only part of it can be cancelled. And each tap has a range: a reflection larger than the tap’s limit is cancelled only up to it. In the “Two DIMMs at 6400” scenario the near DIMM’s echo is the third post-cursor and far larger than tap 3’s 60 mV, and taps 2 to 4 together may not exceed 60 mV either; the DFE taps readout shows what training was allowed to set, and the eye shows what was left.
The map shows where the DFE’s reach ends. The shaded band in the receiver column covers the four post-cursors it corrects; anything arriving below the band is past it. Making the stub shorter moves its echo up into the band. Making the bus faster pushes it down and out.
With the DFE on, the panel’s eye is what the slicer sees: the input after the summer has subtracted each correction. That eye has seams half a bit either side of the sampling instant, where the correction for one bit gives way to the next, and it is only meaningful between them. It also sits lower than the eye at the pin. With its post-cursors removed, a 0 settles at the same level whatever came before it and a 1 at that level plus the cursor, so training finds the reference near the middle of the cursor rather than the middle of the full swing.
The fly-by and write leveling
Back to the command bus. The RCD drives the clock along the fly-by, so DRAM 0 sees each clock edge first and DRAM 4 last. Each DRAM’s input capacitance also slows the line it hangs on. A line loaded with a capacitance Cload every pitch behaves like a slower line of lower impedance:
With a 50 Ω line at 6.7 ps/mm and DRAMs 12 mm apart, each pitch of line holds 1.61 pF. Add 1 pF of DRAM input and the line slows by a factor of 1.27: each gap now takes 102 ps instead of 80, the line looks like 39.3 Ω, and the clock reaches the fifth DRAM of the subchannel 410 ps after the first. At 6400 MT/s that is 1.31 clock periods; at 8800, 1.80. A designer who wants the loaded line to land at a chosen impedance routes it above that impedance to begin with.
The data strobes, meanwhile, come straight up from the edge connector and arrive at every DRAM at about the same time. If the controller launched every byte lane’s strobe the same way, it would be on time for DRAM 0 and hundreds of picoseconds early for DRAM 4. Write leveling fixes that per lane. Untick it in the panel to see the error it removes.
DDR5 adds a second step. Inside the DRAM, the path from the strobe pin to where data is captured is not matched to the clock’s path, so lining the strobe up with the clock at the pins is not quite enough. An internal cycle alignment step, also driven from the controller, adjusts the DRAM’s own write timing to absorb that difference. Between them, every lane’s strobe ends up aligned in phase and in cycle with its own DRAM’s clock, however far along the fly-by that DRAM sits.
The same chain explains DDR5’s command-bus termination. With the termination on the DRAM die, every DRAM on the chain has one, and a strap pin decides which group’s setting it uses: the DRAMs along the chain take a weak setting or none, and the last one, at the end of the fly-by, a strong one. Choose “No end termination” above to see what the last DRAM would otherwise do to everyone else’s clock.
Training: the controller never sees the eye
Everything so far assumes someone knows where the eye is. Nobody does. The controller has no oscilloscope at the DRAM’s pins; it has a delay line on each strobe, a register for each DRAM’s VrefDQ, and the ability to write a pattern and read it back. Training is the process of finding the eye with only those, at every power-up and again when the temperature moves.
The panel’s What training sees view is that process in one picture, a shmoo. For every strobe delay step and every VrefDQ step, send a PRBS burst and record pass or fail. The pass region is the eye as the controller knows it: blockier than the true eye, because Vref moves in 5.5 mV steps, and ragged at its edges, because real noise and jitter make marginal settings pass sometimes and fail sometimes. A trainer then picks a point inside it. The one in the panel takes the Vref with the widest passing run of delays and the middle of that run, one common rule among several.
DDR5 trains in stages, each one making the next possible: chip select and command/address first (otherwise no command can be trusted), then write leveling, then read training with patterns the DRAM itself generates, including an LFSR pseudo-random one, then write training, VrefDQ per bit and the DFE taps. A duty-cycle adjuster in the DRAM corrects the read strobe’s duty cycle, and a strobe interval oscillator lets the controller measure how far the DRAM’s internal strobe delay has drifted, so it can decide when to retrain.
What a pass means
The target is a raw bit error ratio of 10−16 per lane, before any correction. That is hard to picture until it is turned into time: at 6400 MT/s one lane at 10−16 errs about once every 18 days, and the 80 data lanes of one DIMM together about once every 5.4 hours. That is why server memory carries ECC on top, and why DDR5 adds CRC on reads as well as writes and single-error correction inside the DRAM.
Nobody measures 10−16 directly. Receiver and transmitter limits are validated at 10−9, which at 99.5% confidence needs 5.3 × 109 bits without an error (−ln(0.005)/10−9), 0.83 s of one lane at 6400 MT/s, and the eye is extrapolated from there. The DRAM’s receiver is tested with a stressed eye: a reference channel plus injected jitter and noise, adjusted until the eye at the slicer, after the DFE, is just the specified height and width. In the data sheet read for this page that is 95 mV by 0.25 UI at 3200 MT/s, narrowing to 57.5 mV by 0.23 UI at 6400, with the higher bins still to be defined in that revision. The panel draws that diamond on the eye. It is a statement about the receiver, not a mask for the channel: a channel whose eye at the slicer is bigger than the diamond, with everything else that closes it accounted for, is one the receiver is specified to handle.
The panel’s two eye readouts differ for a reason. “Eye, PRBS” is the opening of a 2047-bit pattern; “worst case” is the peak-distortion bound, the eye of the worst possible bit sequence for this pulse response. The worst sequence may never occur in a training pattern, which is one reason a bus can train well and still fail later on real data.
LPDDR5X vs DDR5
LPDDR5X and DDR5 are JEDEC memories of the same era, and from far away they look alike: single-ended data, a strobe travelling with it, everything trained at power-up. The difference is where each one lives, and nearly every electrical choice follows from that. LPDDR5X is soldered next to the processor, or stacked on its package, so its channel is a few centimetres of board with nothing in the way. DDR5 sits on a module in a socket, so its channel crosses a mainboard and a connector and often shares its wire with a second module.
The swing follows the channel. LPDDR5X signals at a VDDQ of 0.5 V (0.45 V on some parts) with low-voltage swing-terminated logic (LVSTL), and can drop to 0.3 V with termination off. That is affordable only because its channel is short and private: there is little loss, no connector and no stub to eat the margin. DDR5 keeps a 1.1 V pseudo-open-drain bus, because its signal has to survive the connector, the shared wire and the terminations along it.
The clocking differs. LPDDR5X separates a slow command clock from a fast write data clock that runs at two or four times its rate and only when it is needed; reads come back with their own strobe. DDR5 uses one differential strobe per byte, driven by whichever end is sending. Both make the data source-synchronous, so delay the data and its strobe share cancels, and both train the strobe into the eye at power-up.
What limits each is different. On LPDDR5X the enemies are crosstalk and simultaneous switching noise in a dense escape from a fine-pitch package, and skew between bits; the LPDDR5X page works through that budget. On DDR5 they are reflections from the connector and from the other module, the choice of termination for each DRAM, and the fly-by skew on the module: this page.
| LPDDR5X | DDR5 | |
|---|---|---|
| Where it sits | soldered beside the processor, on its package, or on an LPDDR5X CAMM2 | a module in a socket: RDIMM, UDIMM, SODIMM, clocked UDIMM, or a DDR5 CAMM2 |
| Fastest bin in the editions read | 8533 MT/s | timing defined to 8800 MT/s |
| Data width | 16 bits a channel; an x32 package carries two channels | two 32-bit subchannels a DIMM, 40 bits each with ECC |
| Signalling and VDDQ | LVSTL, 0.5 V (0.3 V with ODT off) | POD, 1.1 V |
| Data timing | WCK from the controller for writes, at 2:1 or 4:1 to CK; RDQS for reads | DQS, one per byte, driven by the sender |
| The data channel | point to point, about 10–30 mm, no connector on a soldered board | mainboard, connector and module, often past a second DIMM |
| Command and clock | straight from the controller | fly-by along the DRAMs, from the RCD on a registered DIMM |
| Equalisation in the DRAM | optional, advertised per device: a DFE and transmit pre-emphasis | a 4-tap DFE on every data input |
| Power regulation | the platform’s regulators, on a soldered board | a PMIC on the module |
| The SI problem | crosstalk, SSN and skew in a dense, short escape | reflections, termination and fly-by skew |
| What you trade | speed and power, with the capacity fixed when the board is built (unless it is on a CAMM2) | capacity, upgrades and ECC, paid for in signal integrity |
The line between them is blurring. JEDEC’s CAMM2 standard puts either memory on a compression-attached module with a common connector and different pinouts, and JEDEC aims LPDDR5/5X CAMM2s at notebooks and some server segments. A CAMM2 is still a connector, but a short, flat one, so it buys LPDDR5X upgradability while keeping most of its channel. The choice that remains is the old one: LPDDR5X where power and speed per watt rule and the capacity is fixed at build, DDR5 where capacity, upgrades and ECC matter and the board can pay for a longer channel.
Go deeper — how the panel solves the bus, the peak-distortion eye, and what it leaves out
The method of characteristics. A lossless line of impedance Z and delay T, seen from one end, looks like a voltage source of twice whatever wave is arriving, behind a resistor Z. That turns a network of lines into a sequence of small resistive problems. At each node and each time step, the node voltage is
where the capacitor enters as a conductance 2C/Δt with a history current JC (the trapezoidal rule). The panel steps this at 1 ps over the whole tree. It is exact for lossless lines whose delays are whole picoseconds, and the site’s model gate checks it against a frequency-domain solution of the same network, built independently from each line’s two-port admittance. At a three-way junction of equal lines it gives the textbook split: a third reflected, two thirds transmitted down each branch.
The peak-distortion eye. Sample the single-bit response at the cursor and at every whole bit before and after it: h0 and the hk. The worst 1 is the cursor plus every negative hk; the worst 0 is every positive one. The worst-case height is h0 − Σ|hk|, and the DFE replaces h1 to h4 with what is left after its taps. It is a bound, not a measurement, and on a long ringing tail it is pessimistic, because it adds every small echo with its worst sign at once.
Why reflections are the ISI. A data sheet for DDR5 describes the memory channel as reflection dominated, and the panel agrees: its lines are lossless, and still the eye closes. Loss would add to that on a long mainboard route, and crosstalk and simultaneous switching noise, both covered on their own pages, would close it further. The panel leaves out all three.
What else the panel leaves out. The mainboard and module lines are 40 Ω and the connector a 60 Ω, 25 ps section; the packages are 45 Ω sections; the DRAM die is 0.8 pF and the controller’s 1.0 pF. Each DIMM is single-rank, so there is one DRAM per wire; a dual-rank DIMM puts two on it, and its non-target rank is one more termination to schedule. The driver’s edge is a raised cosine of 30% of a bit, held between 30 and 70 ps. Noise and jitter appear only in the shmoo, as model choices: 5 mV and 2 ps rms. There are no ESD clamps and no package or line loss, so when a stub is left open the pin rings several hundred millivolts past both rails, further than real silicon would let it; read those eyes for their shape and timing rather than their peaks.
In the real world
Two DIMMs per channel costs speed. Server platforms commonly qualify a lower data rate with two DIMMs per channel than with one, and the panel shows why: the second DIMM’s stub, terminated or not, takes a share of every bit and returns some of it late. Check the platform’s population rules before assuming the headline speed.
Client modules above 6400 MT/s add a clock driver. On an unbuffered DIMM the controller drives the DRAMs’ clock directly, through the connector and along the fly-by. JEDEC’s clocked UDIMM and SODIMM (JESD323 and JESD324) put a clock driver on the module that regenerates the clock locally, which JEDEC presents as the step from 6400 toward 7200 MT/s on client modules. Command and address still come from the controller, so the rest of this page still applies.
Power is now on the module. The PMIC regulates VDD, VDDQ and VPP next to the DRAMs, from 12 V on server modules and 5 V on client ones. That shortens the DRAM’s power delivery path, and it means the module’s own regulator and decoupling are part of every data eye’s noise. A DDR5 signal-integrity simulation that ignores the module’s power network ignores a source of jitter that DDR4 put on the mainboard.
Debug starts from the training results. The controller knows, for every DRAM and every lane, where training landed and how wide the passing region was. A lane whose window is narrower than its neighbours’ points at its own routing; a whole DIMM that is narrow points at the slot, the population or the termination; a window that shrinks as the system warms points at timing drift. All of it is available without probing a pin.
Where the numbers come from
- JEDEC JESD79-5D, DDR5 SDRAM (Nov 2025), and JEDEC’s JESD79-5C announcement (Apr 2024) — editions and the 8800 MT/s timing extension; ledger C-35 and C-36. The standard itself is sold and has not been read.
- Micron, DDR5 SDRAM core data sheet (Rev. D, Oct 2022, stated JESD79-5B compliant) — supply voltages, ODT and driver values, VrefDQ range and step, DFE tap ranges, stressed-eye limits, BER requirements, write leveling; ledger C-37 to C-43.
- Micron white papers and technical briefs: “Introducing Micron DDR5 SDRAM: More Than a Generational Update”, “DDR5 SDRAM: New Features”, “DDR5: Key Module Features” and “DDR5: Client Module Features” — the DDR4 comparison, subchannels, PMIC input voltages, the RCD’s command bus and CA termination groups; ledger C-44 to C-49.
- JEDEC clocked UDIMM and SODIMM (JESD323, JESD324), as announced in October 2024 — ledger C-50.
- Micron, x32 automotive LPDDR5X data sheet and JEDEC’s CAMM2 announcement (Dec 2023) — the LPDDR5X side of the comparison; ledger C-53 and C-54, with C-2a, C-3 and C-4a from the LPDDR5X page.
- ASUS support FAQ, “How to install DRAM on motherboard” — the slot to fill first; ledger C-51.
- Platform memory population guides — that two DIMMs per channel usually run slower; ledger C-52, awaiting a source.
Related
- LPDDR5/5X Signal Integrity: WCK, ODT and Training — the mobile cousin: point to point, no DIMM
- Transmission-Line Reflections and Ringing — what one boundary does to a step
- TDR (Time-Domain Reflectometry): Reading PCB Impedance — reading a stub and a connector on a trace
- Termination Schemes: Series, Parallel, Thevenin and AC — why a terminated line stops ringing
- SerDes Equalization: CTLE, FFE, DFE and CDR — the DFE in a serial receiver
- Simultaneous Switching Noise and Ground Bounce — what an 80-bit bus switching together does to its reference
Sources
- JEDEC — LPDDR5 standard update announcement C-2a
- Synopsys — Key features designers should know about LPDDR5 C-3
- JEDEC — JESD79-5C DDR5 SDRAM standard update announcement C-35
- JEDEC — JESD79-5D DDR5 SDRAM document page C-36
- Micron — DDR5 SDRAM core data sheet (Rev. D, JESD79-5B compliant) C-37
- Micron — DDR5: Key Module Features (technical brief) C-44
- Micron — Introducing Micron DDR5 SDRAM: More Than a Generational Update C-45
- Micron — DDR5: Client Module Features C-47
- EDN — JEDEC unveils memory designs with DDR5 clock drivers C-50
- ASUS — How to install DRAM (memory) on motherboard C-51
- Micron — x32 Automotive LPDDR5X SDRAM data sheet (315b) C-53
- JEDEC — CAMM2 memory module standard announcement C-54
Rows marked with a claim id are tracked in the claim ledger, which records what each source can and cannot establish.