so your saying the bottleneck is the thread dispatch and it's only being used at about 31.25% if it where redesigned to use all 1024 dispatches threads at once and not have those 5 alu's grouped. 5 alu is one SIMD. changing this to be all seprate alu shouldn't be too hard, the alu's them self are quite small already.
it's seem to me it's more like what ever is easy the programs will go for shorter times to code things.
easy isn't the best possible way to do things.





Reply With Quote
Bookmarks