Originally Posted by agenda2005
You are right! However, that type of design will require a sophisticated memory controller which need to schedule data transfer between each CPU L2 and the correct bank, therefore a 1 or more cycle delay. They will also require a crossbar switch to access data from each others L2.
Whereas a glueless design with shared L2 makes the work of the memory controller easier and no need for a crossbar. This is easier and more efficient, since A64 CPUs have an exclsive L2 cache.