Shengyi Wei

AI DATA MOVEMENT · MEMORY SYSTEMS · INTERCONNECT · RTL

From the hardware's side, an inference run is a traffic pattern. I work on reading that pattern off real workloads, and on what follows from it: memory systems, on-chip and chip-to-chip networks, and the RTL underneath.

ONE TRANSFORMER LAYER DURING ONE DECODE STEP · B = 1

EVENT ORDER ONLY · NOT A LATENCY SCALE · SERIAL CACHE-FIRST DEPICTION

  1. RMSNORM
  2. Q/K/V PROJ
  3. ROPE (Q,K)
  4. KV APPEND
  5. ATTENTION
  6. O PROJ
  7. + LAYER INPUT
  8. RMSNORM
  9. MLP
  10. + POST-ATTN STATE
  11. UNSHOWN WORK

STAGE HIGHLIGHT MEANS "BEING EXPLAINED, POSSIBLY WAITING". COMPUTE-ACTIVE IS THE SEPARATE MARK ON THE COMPUTE BAND. Q/K/V AND MLP EXPAND INTO SEVERAL OPERATIONS; NORM SCALE VECTORS AND ON-CHIP ACTIVATION MOVEMENT ARE NOT DRAWN HERE.

VALID KV POSITIONS · THIS LAYER2048 → 2049PREFILLED BASELINE · NOT CAPACITY, NOT PAGES

T
5.700 s OF 16.200 s · ORDER, NOT LATENCY
STAGE
ATTENTION
IN FLIGHT
READ RETURN K[0:T] · ROUTE HBM->MC->BUF · NOW HBM->MC · AFTER MEMORY SERVICE K[0:T] · FOR SCORES = Q · Kt
BUFFERS
NO OPERAND HELD
COMPUTE
NOT READY · SCORES = Q · Kt STILL NEEDS K[0:T]
KV
VALID POSITIONS THIS LAYER 2048 -> 2049
DEMAND
KV READ, T VISIBLE POSITIONS (QUALITATIVE, NOT MEASURED)

ONE BLOCK MATRIX–VECTOR PRODUCT: y = W x

ILLUSTRATIVE MAPPING · NOT A GPU FLOORPLAN · DEPENDENCY ORDER, NOT A NOC SIMULATION

Projection mapping. One block matrix-vector product y = W x on a two-row by three-column array of tiles. A weight buffer runs across the top and feeds one trunk per column. Tile (r,c) owns the weight block W[J_c, I_r] and no other tile consumes it, even when the block passes through one on its way down the column. Each row has its own source carrying the activation slice x[I_r], delivered to the first tile of the row and forwarded along it, so one slice is reused by every tile in that row. A tile forms its local product z[r,c] only once both its own weight block and its row activation have arrived, and an outlined tile means it is computing. Partial sums travel down their column and are added into the tile below. Each column ends in its own accumulator holding its own distinct output slice, y[J_0], y[J_1] and y[J_2]; the array never collapses to a single scalar. This is a dependency order, not a cycle-accurate network simulation. The readout below states what is on a link at the current instant and which rows each column has so far.Projection mapping. One block matrix-vector product y = W x on a two-row by three-column array of tiles. A weight buffer runs across the top and feeds one trunk per column. Tile (r,c) owns the weight block W[J_c, I_r] and no other tile consumes it, even when the block passes through one on its way down the column. Each row has its own source carrying the activation slice x[I_r], delivered to the first tile of the row and forwarded along it, so one slice is reused by every tile in that row. A tile forms its local product z[r,c] only once both its own weight block and its row activation have arrived, and an outlined tile means it is computing. Partial sums travel down their column and are added into the tile below. Each column ends in its own accumulator holding its own distinct output slice, y[J_0], y[J_1] and y[J_2]; the array never collapses to a single scalar. This is a dependency order, not a cycle-accurate network simulation. The readout below states what is on a link at the current instant and which rows each column has so far.
T
3.600 s OF 5.750 s · DEPENDENCY ORDER
IN FLIGHT
PARTIAL SUM FROM ROW 0 DOWN COLUMN 1 · T01->T11 · AFTER T01 COMPUTES z[0,1] · FOR ACCUMULATE ROWS 0,1 IN COLUMN 1
OUTPUT
y[J0] PENDING · COLUMN SUM HOLDS ROWS 0 · y[J1] PENDING · COLUMN SUM HOLDS ROWS 0 · y[J2] PENDING · COLUMN SUM HOLDS ROWS -
RULE
A TILE COMPUTES ONLY WITH ITS OWN W BLOCK AND ITS ROW ACTIVATION IN HAND

ONE TRANSFORMER LAYER DURING ONE DECODE STEP, BATCH SIZE ONE, DRAWN FROM A VALIDATED EVENT LEDGER. EVENT ORDER ONLY: EQUAL-WIDTH STAGES ARE NOT EQUAL LATENCIES, AND NOTHING HERE IS A MEASUREMENT. THE SCHEDULE IS A CHOSEN SERIAL, CACHE-FIRST ILLUSTRATION — AN IMPLEMENTATION MAY PREFETCH, OVERLAP, FUSE, OR KEEP THE CURRENT K AND V ON CHIP INSTEAD OF WRITING THEM OUT AND READING THEM BACK. OPERAND DEMAND IS QUALITATIVE; NO ELEMENT HERE IS MEASURED HBM, NOC OR PE UTILIZATION.ONE BLOCK MATRIX–VECTOR PRODUCT y = W x, AS AN ILLUSTRATIVE MAPPING, NOT A GPU FLOORPLAN. TILE (r,c) OWNS W[J_c, I_r]; THE ROW ACTIVATION SLICE x[I_r] IS REUSED ALONG THE ROW; PARTIAL SUMS REDUCE DOWN THE COLUMN INTO THE OUTPUT SLICE y[J_c]. DEPENDENCY-VALID ORDER, NOT A CYCLE-ACCURATE NOC SIMULATION: NO ARBITRATION, BANDWIDTH OR CONTENTION IS MODELLED.

Life is a long trade, not a greedy algorithm.

READING NOTES9 ENTRIES

  1. Building a GPU by hand (2)
  2. Building a GPU by hand (1)
  3. Industrial-grade RTL style
  4. Encoding for high-speed serial links
  5. PCIe, part 4: the physical layer (1)
  6. PCIe, part 3: the data link layer
  7. PCIe, part 2: the transaction layer
  8. PCIe, part 1: overall architecture
  9. Skimming ASIC synthesis (1)

ALL READINGNOTES