MMM
Results 1 to 25 of 4519

Thread: AMD Zambezi news, info, fans !

Threaded View

  1. #11
    Xtreme Cruncher
    Join Date
    Jun 2006
    Posts
    6,215
    Quote Originally Posted by mAJORD View Post
    For the record, and another reference. going back to my Bobcat clock/clock comparisons
    Per core K10 is:

    Super pi 5% faster
    Fritz chess: K10 20% faster
    Cinebench 11.5: 49% faster

    at the same clk speed.

    which means, somehow Bulldozer is about the same performance :S , even on some of these SSE/FPU heavy benches.

    Has anyone determined the pipeline length details for Bulldozer? I believe Bobcat is 15 stages vs 12 for K10

    It is likely a bit longer to enable these high clockspeeds, but I'm still finding some of the results out of line. I know you didn't look up superpi, but it for one is, according to these results slower than Bobcat (as i've mentioned before), given the architectures seem to be similar, I find this quite odd.

    Even if the pipeline stages are longer than Bobcat's the massive amounts of Cache, much larger buffers, much wider more capable performance orientated FPU (Bobcat has a very trimmed down FPU due to its target market) , assumed more aggressive prefteching It certainly doesn't make much sense at this stage.

    I can understand similar IPC to Thurban given the higher frequency headroom, and trade-off's to achieve high performance / watt (all valid design decsisions), but these outlier results like Cinebench, Wprime, Fritz, are quite baffling
    I agree with you. I personally think something went horribly wrong between the Q4 of 2010 and Q2 of 2011 and now instead of a chip which has a thread count of 8 ,whose each "thread" is at least as strong as Thuban's core in MT and at least 10% stronger in ST workloads(both int and fp),we end up with a Bobcat level of performance per core,roughly of course, which can clock really high though ,but which scales poorly (1.7x or 1.8x over single thread). Even little bobcat ens up faster in select few benchmarks (per clock). To release this thing as a successor to relatively successful Phenom II line is very much crazy.
    I'm especially baffled by really slow FP unit. Bobcat's FPU looks like to be on a similar perf. level (per "thread") as Bulldozer's. Sounds crazy but it's true. Just forget K10.

    Quote Originally Posted by Apokalipse View Post
    That's why I think this is not final performance. Because why would AMD release something that's worse than what they had before? They'd have to be bat sheet insane to do that.
    I have no idea to be honest. Maybe they didn't expect to see this from their simulations? This thing was supposed to be a Thuban crushing machine,excelling at MT workloads and being 30-50% faster at the same TDP as per John Taylor. This is the post I wrote on AT forum few days ago:

    This video is from March 8 this year. It was shot at Cebit 2011. It features Macci and John Taylor @ AMD. Mr Taylor said @ 2:20 mark that bulldzoer was designed to deliver 30 to 50% more performance within the same TDP envelope and roughly the same die are versus the cpu it replaces. The video is here.

    What I want to know now is in which real world or synthetic benchmarks/workloads is Zambezi going to deliver 30-50% more performance than Thuban and with which magic dust is this going to happen? With the latest numbers it barely beats Thuban in highly MT workloads like cinebench(both old 10 and new 11.5) and handbrake. The difference ranges from tiny to 10 or so %. This is 8T vs 6C case,so best case scenario for Bulldozer. Bulldozer even runs at much higher clocks (both stock and Turbo). Where is 30% difference (just forget 50%)? Oh yes,AES and such are just outliers so those are corner cases.

    If it was supposedly designed to deliver 30-50% more performance and if Mr Taylor stated this in the context of very parallel workloads (which is legitimate ) then we can say Bulldozer failed since it can't overall outperform Thuban by more than 20%,let alone 30% or now astronomical 50%. In order to achieve this ,the Bulldozer that John Taylor talked about must be the same one from this slide (and no,this slide was not fake). What happened in the meantime ? How from this 30-50% throughput machine we ended up with barely faster than Thuban? Unless the Bulldozer John Taylor spoke about in the video was expected to run at 4.5-5Ghz stock clock and Turbo to 5.5Ghz ,all within 125W TDP envelope,then something else went wrong and now we get "this" (slower than Thuban at the same clock by 5-15%,depending on the app and barely beating it since it has weaker core scaling and only 1.33x more cores to make up the difference).
    Last edited by informal; 10-08-2011 at 01:03 PM.

Bookmarks

Bookmarks

Posting Permissions

  • You may not post new threads
  • You may not post replies
  • You may not post attachments
  • You may not edit your posts
  •