Problem:
Side note: Defining the problem is important. That way, we can see if their solution addresses the problem and helps inspire new solutions.
1. Bandwidth scales slower than arithmetic
a. Bandwidth: off-chip memory
i. Slow because pin count increases slowly
ii. You can signal faster, but energy per bit seems to have a lower bound.
1. You can get 80Gbits/sec with current technology, but the energy
is very high
2. Range: 0.1pJ/b is the minimum: 2Gbit/sec to 1pJ/b : 20csps
3. pJ/b – good metric because it’s the same as mW/Gbps
4. Typically used today is around 20pJ/b
iii. Today 200 Gbytes/s – high performance GPU
1. This is about 1Tb/s which is 1000 Gb/s. So at 20pJ/b its around
20W just to move data
iv. Optical interconnect is another way to improve bandwidth. Nanostores are
another way as well. Proximity communication is another interconnection
method. It uses inductive coupling.
2. Power efficiency overall
a. BW is where the power goes. More energy to move data than to compute
3. Current solution (caching) not sufficient
Users:
1. Supercomputer vendors, programmers
Solution:
1. Locality
2. Stream model
a. Infinite stream of data coming in, process it, infinite stream of data going out
i. Lots of input data
ii. Limited Window to process the data(frontier)
3. Ideally you get lots of compute within a window
4. Parallelism: if “stream” implies no dependencies between “chucks”
a. Caching doesn’t work well because there is limited temporal locality
b. “Nest” streams to get finer grained parallelism
5. Side notes:
a. Sync. Dataflow
i. Nodes wait for all the producers arrive before firing.
Thus there is order.
ii. You can better statically analyze the
system with sync. Dataflow
b. Other examples: Ptolemy, petrinets, khan process networks
i. Streaming model is similar to sync. Dataflow
c. Two important metrics:
i. Bisection and diameter bandwidth
1. Problem: large diameter slows down network because of congestion
ii. Alternative: all to all connect
1. Problem: too many wires, very costly
iii. Alternative: high radix
d. What an interconnection network is for: connect points, basically a switch
with input and output ports
i. Used to be crossbars. Telephone used to work. Doesn’t scale well: N^2
ii. Klauss network: build a large switch with smaller switches .
Creates a perfect switch with extra layers of switches
1. This is essentially what Merrimac is describing
iii. Reduces diameter, reduces cost
iv. Good switch: anywhere to anywhere with predictable BW and latencies
v. Ideal switch: interference free latency: cross bar is such a switch
6. You need to optimize the following: How many FPUs you can put: how much BW you can give them
Evaluation:
1. Simulator, single node
2. Custom benchmarks (small benchmarks)
3. Locality, hierarchy, use: how BW effects
4. No real demo of SRF
a. Really not showing caching doesn’t work. Evaluation does not prove SRF
is better than caching
5. KernalC/StreamC
6. Other side notes:
a. 3 stage software pipeline: read inputs, process, write outputs.
i. Reduce the amount of data you have to move and don’t waste time using
pipelining – balanced design
b. Vector processors deal with fine granularity words: strip mining and chaining
is similar, but it is at a different granularity. Chaining has limited forking.
i. Complexity of the hardware is what separates vector from streaming.
There is a hierarchy of control in streaming; not for vector.
ii. Streaming architecture is taking the advantages of vector and moving
the complexity to software
iii. Can buffer things in streaming, vector uses renaming and more complex
hardware to do this. Streaming is easier with more control
Comments:
Cell processor is similar to Merrimac. Since cell processor was build, Merrimac was not.
