Lecture 2 (9/2/2008) — Locality and Wires
- Lecture 2 Scribe Notes (Bob_Ascott@mail.utexas.edu):
- Components of a CPU:
- Registers (both architecture and micro-architecture)
- Processor Memory Performance Gap:
- (From Hennessy-Patterson, “Computer Architecture”)
- Chart of Microprocessor data from Intel (Pentium, P2, P3, P4 )
- ITRS Delay Estimation
- (Graph from referenced paper)
- “The Future of Wires” paper
- Rest of equations to be reviewed on next lecture
- September 4, 2008
Lecture 2 Scribe Notes (Bob_Ascott@mail.utexas.edu):
The scribe notes are complementary to the slides.
Class Announcements - (included on Blackboard elsewhere)
- Next 3 lectures focus on Locality Mechanisms
Locality within a CPU
Another view of locality
Exploiting locality
Components of a CPU:
Basic in-order machine
ALU(L),
Registers(L),
CACHE(L),
Controller(L),
I/O-Memory,
Bus/Interconnect,
Memory Ctlr/LoadSt Q(L)
Added for OOO (Out of Order)/ Pipelined machine
Reservation stations(L),
Reorder buffer (ROB),
Renaming tables,
Pipeline Bypass Network
(L) indicates elements where locality is important
Registers (both architecture and micro-architecture)
CPU (Programming)
First from of storage
Compact coding in ISA
Lowest level cache
ASIC (Control)
Buffering
Latency
Power
Bandwidth
Processor Memory Performance Gap:
(From Hennessy-Patterson, “Computer Architecture”)
(summary of graph)
- CPU performance (as measured by instructions per second) increasing @ 50% per year
- Memory performance (as measured by reciprocal of read latency) increasing @ 7% per year
- “Gap” is the increasing spread between these two projections
- (Memory bandwidth graph would be similar to CPU performance - implying less of a gap than with latency)
- Parallelism and locality are exploited to allow CPU performance to increase in light of the lag in Memory Performance.
- Potential peak performance due to VLSI improvements historically grows at 78 % per year due to smaller, faster devices.
Chart of Microprocessor data from Intel (Pentium, P2, P3, P4 )
- Locality improves power EXCEPT in transfering 64 bits of data across chip
- Chart showing power required by
64b FP operationRead 64 bits from 16K Cachexfer 64 bits 10mm across chipxfer 64 bits off chip
- For technologies:
65nm, 32nm, 16nm(Assumes scaling both size and frequency)
- Locality improves bandwidth
Wire DensityVias + repeaters restrict routing
- Rule of thumb:
Latency directly proportional to distance(not distance**2 due to repeaters)BW inversely proportional to distancePower directly proportional to distance
- Latches vs Memory
Use SRAM for > 2 or 4 elementslarger + slowerLatches for pipelinessmall + fast
ITRS Delay Estimation
(Graph from referenced paper)
- shows normalized delay vs technology geometry
- Gate delay FO4 (fan out of 4) - improving
- Local wire delay (within 50K ckts) - improving almost as much
- Global wire delay with repeaters - getting worse
- Global wire delay without repeaters - much worse
- Bandwidth chart
shows Global across chip has > bw thanlocal wires within 50K cktsHowever; Many more local wires available; hence overall bandwidth improves
“The Future of Wires” paper
- Latency, Bandwidth, Power(not discussed in this paper)
- Elements:
Resistance ~ length ~ 1/AreaCapacitance ~ lengthInductance ~ complexNoise = sub-property
- (Wires made tall and narrow to increase area AND increase horizontal density of wires.)
- Let a=scaling factor (pitch of new technology)/(pitch of old technology)
- Resistance/length ~ a**2
- Capacitance/length decreases a little (both size of wire AND size of dielectric decrease)
- Inductance - hard to measure AND small with respect to RC.
