Lecture 2 (9/2/2008) — Locality and Wires


Lecture 2 Scribe Notes (Bob_Ascott@mail.utexas.edu):

The scribe notes are complementary to the slides.

Class Announcements - (included on Blackboard elsewhere)

  • Next 3 lectures focus on Locality Mechanisms
Locality within a CPU
Another view of locality
Exploiting locality

Components of a CPU:

Basic in-order machine
ALU(L),
Registers(L),
CACHE(L),
Controller(L),
I/O-Memory,
Bus/Interconnect,
Memory Ctlr/LoadSt Q(L)
Added for OOO (Out of Order)/ Pipelined machine
Reservation stations(L),
Reorder buffer (ROB),
Renaming tables,
Pipeline Bypass Network

(L) indicates elements where locality is important

Registers (both architecture and micro-architecture)

CPU (Programming)
First from of storage
Compact coding in ISA
Lowest level cache
ASIC (Control)
Buffering
Latency
Power
Bandwidth

Processor Memory Performance Gap:

(From Hennessy-Patterson, “Computer Architecture”)

(summary of graph)

  • CPU performance (as measured by instructions per second) increasing @ 50% per year
  • Memory performance (as measured by reciprocal of read latency) increasing @ 7% per year
  • “Gap” is the increasing spread between these two projections
  • (Memory bandwidth graph would be similar to CPU performance - implying less of a gap than with latency)
  • Parallelism and locality are exploited to allow CPU performance to increase in light of the lag in Memory Performance.
  • Potential peak performance due to VLSI improvements historically grows at 78 % per year due to smaller, faster devices.

Chart of Microprocessor data from Intel (Pentium, P2, P3, P4 )

  • Locality improves power EXCEPT in transfering 64 bits of data across chip
  • Chart showing power required by
    64b FP operation
    Read 64 bits from 16K Cache
    xfer 64 bits 10mm across chip
    xfer 64 bits off chip
  • For technologies:
    65nm, 32nm, 16nm
    (Assumes scaling both size and frequency)
  • Locality improves bandwidth
    Wire Density
    Vias + repeaters restrict routing
  • Rule of thumb:
    Latency directly proportional to distance
    (not distance**2 due to repeaters)
    BW inversely proportional to distance
    Power directly proportional to distance
  • Latches vs Memory
    Use SRAM for > 2 or 4 elements
    larger + slower
    Latches for pipelines
    small + fast

ITRS Delay Estimation

(Graph from referenced paper)

  • shows normalized delay vs technology geometry
  • Gate delay FO4 (fan out of 4) - improving
  • Local wire delay (within 50K ckts) - improving almost as much
  • Global wire delay with repeaters - getting worse
  • Global wire delay without repeaters - much worse
  • Bandwidth chart
    shows Global across chip has > bw than
    local wires within 50K ckts
    However; Many more local wires available; hence overall bandwidth improves

“The Future of Wires” paper

  • Latency, Bandwidth, Power(not discussed in this paper)
  • Elements:
    Resistance ~ length ~ 1/Area
    Capacitance ~ length
    Inductance ~ complex
    Noise = sub-property
  • (Wires made tall and narrow to increase area AND increase horizontal density of wires.)
  • Let a=scaling factor (pitch of new technology)/(pitch of old technology)
  • Resistance/length ~ a**2
  • Capacitance/length decreases a little (both size of wire AND size of dielectric decrease)
  • Inductance - hard to measure AND small with respect to RC.

Rest of equations to be reviewed on next lecture

September 4, 2008