MMM
Results 1 to 25 of 4519

Thread: AMD Zambezi news, info, fans !

Hybrid View

  1. #1
    Xtreme Cruncher
    Join Date
    Jun 2006
    Posts
    6,215
    Quote Originally Posted by Olivon View Post
    The only thing I see is ppls who don't want to understand ...
    xsecret brings real infos and told us before that the way how AMD communicates around Bulldozer : Big frequencies and that's all ...
    That's really nice, especially for overclockerz on LN˛ and He, and non coldbug feature is really great, for sure. Extreme overclocking seems really fun on it.
    But for an H24 rig, I won't bet BD will be the SB killer many has wanted here.
    Well the only issue with what xsecret said ,that I see, is that he confirms that Chinese(VRzone) leaked numbers are true or close to what he has seen under NDA. Mind you that those numbers show 8120 @ 3.1Ghz scoring lower than Deneb X4 in 3dmark06 CPU subtest ,for instance. IMO this is hard to believe since even on paper,those 4 FPUs inside Zambezi are much more powerfully than 4 Deneb FPUs. Similar goes for Cinebench,wprime etc. The design has to be seriously borked in order to perform worse than quad core Deneb at similar clock(8C Bulldozer vs QC Deneb).

  2. #2
    Registered User
    Join Date
    Feb 2005
    Posts
    39
    Quote Originally Posted by informal View Post
    Well the only issue with what xsecret said ,that I see, is that he confirms that Chinese(VRzone) leaked numbers are true or close to what he has seen under NDA. Mind you that those numbers show 8120 @ 3.1Ghz scoring lower than Deneb X4 in 3dmark06 CPU subtest ,for instance. IMO this is hard to believe since even on paper,those 4 FPUs inside Zambezi are much more powerfully than 4 Deneb FPUs. Similar goes for Cinebench,wprime etc. The design has to be seriously borked in order to perform worse than quad core Deneb at similar clock(8C Bulldozer vs QC Deneb).
    High raw throughput for an FP unit is nice. But in order to use this power in real-world application, you need a frontend able to feed it correctly. That means massive code optimization and a good compiler, in best case. And keep in mind the horribly slow L1 Write-Through, probably added in order to remove a bottleneck in frequency scaling. Write-Through means your writing from the frontend to the L2 "through" the L1. So, seen from the frontend, the L1 write bandwidth is as "slow" as the L2 write bandwidth. The last ľarch to use that horrible trick was Netburst, with high frequencies in mind. Bulldozer comes with a L1 WT too and that point only could explain many disappointments from a performances point of view.

    Anyway, Macci speaks about a 5 GHz overclocking w/ air cooler. According to the roadmap, a 4.5 GHz CPU (turbo, but turbo is there in order to stay within TDP and we don't care about that in OC) planned in early 2012, so a 500 MHz overclocking doesn't seems so extreme. Is there any leaks related to power dissipation ?
    Doc_TB @ CanardPC.Com (FR)

  3. #3
    I am Xtreme
    Join Date
    Dec 2007
    Posts
    7,750
    Quote Originally Posted by xsecret View Post
    Anyway, Macci speaks about a 5 GHz overclocking w/ air cooler. According to the roadmap, a 4.5 GHz CPU (turbo, but turbo is there in order to stay within TDP and we don't care about that in OC) planned in early 2012, so a 500 MHz overclocking doesn't seems so extreme. Is there any leaks related to power dissipation ?
    thuban gets 1ghz overclock if you look at base clocks, or 500mhz OC if you look at turbo. so why compare what you can do with 1 core stock to 8 cores overclocked. also why compare the clocks of a future cpu with current cpu overclocking. the PII 940 was pretty crappy at overclocking compared to future models
    http://www.overclockersclub.com/revi...nomii940/4.htm
    940 got 3.755ghz at 1.55v
    http://www.overclockersclub.com/revi...omii_965/3.htm
    965 got 3.915ghz at 1.52v

    its very expected to see an extra 200-300mhz out of those models coming early 2012
    2500k @ 4900mhz - Asus Maxiums IV Gene Z - Swiftech Apogee LP
    GTX 680 @ +170 (1267mhz) / +300 (3305mhz) - EK 680 FC EN/Acteal
    Swiftech MCR320 Drive @ 1300rpms - 3x GT 1850s @ 1150rpms
    XS Build Log for: My Latest Custom Case

  4. #4
    maltrabob
    Guest
    From Muropaketti's article: "According to AMD, record processors were used in the test just before the mass production of manufactured B2 stepping test pieces".
    I am a bit lost in this translation, especially with what it says at the end of that sentence. Does it mean that B2 stepping was just a testing one and there is going to be a different one for the final production (B2.G rumored plus there is no revision shown on that CPU-Z shot)? Or is it just an imperfection in Google Translator? Anybody from Finland to shed some light on it, please?

    From Overclockers.com: "In passing conversation, Brian and Sami mentioned doing 5Ghz on air running fully multi-threaded benchmarks."
    So far so good ...
    Last edited by maltrabob; 09-13-2011 at 07:57 AM.

  5. #5
    Xtreme Addict Evantaur's Avatar
    Join Date
    Jul 2011
    Location
    Finland
    Posts
    1,043
    Quote Originally Posted by maltrabob View Post
    I am a bit lost in this translation, especially with what it says at the end of that sentence. Does it mean that B2 stepping was just a testing one and there is going to be a different one for the final production (B2.G rumored plus there is no revision shown on that CPU-Z shot)? Or is it just an imperfection in Google Translator? Anybody from Finland to shed some light on it, please?
    no, these were B2 ES

    I like large posteriors and I cannot prevaricate

  6. #6
    Xtreme Cruncher
    Join Date
    Nov 2005
    Location
    Netherlands
    Posts
    1,778
    Quote Originally Posted by xsecret View Post
    High raw throughput for an FP unit is nice. But in order to use this power in real-world application, you need a frontend able to feed it correctly. That means massive code optimization and a good compiler, in best case. And keep in mind the horribly slow L1 Write-Through, probably added in order to remove a bottleneck in frequency scaling. Write-Through means your writing from the frontend to the L2 "through" the L1. So, seen from the frontend, the L1 write bandwidth is as "slow" as the L2 write bandwidth. The last ľarch to use that horrible trick was Netburst, with high frequencies in mind. Bulldozer comes with a L1 WT too and that point only could explain many disappointments from a performances point of view.
    I can only understand the basics and have no way of verifying anything related to performance based on your input. So bear with me. But logic dictates that AMD wouldn't do this if avoiding the WT would net in higher performance overall. So what is there to compain about if the increased clocks make up for it apparently?

    It's another thing if they should have avoided this route from a design point of view and create a better cpu. Better as in able to keep up with thuban in single and multi thread performance. People are implying that BD will fail on both fronts, no? I just can't believe that, at all, personally.

  7. #7
    Xtreme Cruncher
    Join Date
    Jun 2006
    Posts
    6,215
    Quote Originally Posted by xsecret View Post
    High raw throughput for an FP unit is nice. But in order to use this power in real-world application, you need a frontend able to feed it correctly. That means massive code optimization and a good compiler, in best case. And keep in mind the horribly slow L1 Write-Through, probably added in order to remove a bottleneck in frequency scaling. Write-Through means your writing from the frontend to the L2 "through" the L1. So, seen from the frontend, the L1 write bandwidth is as "slow" as the L2 write bandwidth. The last ľarch to use that horrible trick was Netburst, with high frequencies in mind. Bulldozer comes with a L1 WT too and that point only could explain many disappointments from a performances point of view.
    So you re pretty sure Bulldozer will be slower than Thuban per core? And you are pretty sure you have final platform in your hands? If this is true then the design is truly broken is some way. Still doesn't make any sense to me. AMD knew the perf. level of Nehalem by middle of 2008 probably. They knew intel will just go up from there(Westmere,SB,SB-E,IB). And you are telling me that with all this foreknowledge they opted for Netburst-like design that is actually less competitive Vs Core generation 1 (Merom) while having only 15%-20% higher frequency potential than Family 10h ? This is ridiculous.

  8. #8
    Registered User
    Join Date
    Feb 2005
    Posts
    39
    Quote Originally Posted by informal View Post
    So you re pretty sure Bulldozer will be slower than Thuban per core? And you are pretty sure you have final platform in your hands? If this is true then the design is truly broken is some way. Still doesn't make any sense to me. AMD knew the perf. level of Nehalem by middle of 2008 probably. They knew intel will just go up from there(Westmere,SB,SB-E,IB). And you are telling me that with all this foreknowledge they opted for Netburst-like design that is actually less competitive Vs Core generation 1 (Merom) while having only 15%-20% higher frequency potential than Family 10h ? This is ridiculous.
    Was a P4 slower than a P3 ? Sometimes no, sometimes yes, depending on the software. The absolute performance is not something so important for AMD. The most important thing is money. And just money. Spending gazillions dollars in R&D to reach the performance of a CPU sold in very low quantities (and generating very low incomes) like the 990X is ridiculous. Bulldozer must solve two problems : 1/ Be able to gain performances (with frequency increases) at mid-term without spending more gazillions in another ľarch 2/ Compete with Intel *mainstream* CPUs (and not Extreme CPU) with a similar price/performance ratio.
    Doc_TB @ CanardPC.Com (FR)

  9. #9
    Xtreme Member
    Join Date
    Apr 2007
    Location
    Serbia
    Posts
    102
    Quote Originally Posted by xsecret View Post
    Was a P4 slower than a P3 ? Sometimes no, sometimes yes, depending on the software. The absolute performance is not something so important for AMD. The most important thing is money. And just money. Spending gazillions dollars in R&D to reach the performance of a CPU sold in very low quantities (and generating very low incomes) like the 990X is ridiculous. Bulldozer must solve two problems : 1/ Be able to gain performances (with frequency increases) at mid-term without spending more gazillions in another ľarch 2/ Compete with Intel *mainstream* CPUs (and not Extreme CPU) with a similar price/performance ratio.
    I agree. But, in BD architecture are less tradeoffs than Netburst.

    Quote Originally Posted by xsecret View Post
    High raw throughput for an FP unit is nice. But in order to use this power in real-world application, you need a frontend able to feed it correctly.
    Average IPC of most workloads isn't much more than 1 IPC on Thuban core. 4-way front end is more than enough to feed two threads.

    And keep in mind the horribly slow L1 Write-Through, probably added in order to remove a bottleneck in frequency scaling. Write-Through means your writing from the frontend to the L2 "through" the L1.
    No, that doesn't mean WT.
    Write Trough means that every write to the cache causes a synchronous write to the backing store. Because L2 is slower than L1, L1 must wait for L2 to write out data. But there is WCC (Write Coalescing Cache) to hold on data for later writing out. I can't see why the WT policy cache is so much issue with BD core. Ratio between loads and stores is arround 2:1. For every two loads, we have one store.

    So, seen from the frontend, the L1 write bandwidth is as "slow" as the L2 write bandwidth.
    Not quite, because of WCC.

    The last ľarch to use that horrible trick was Netburst, with high frequencies in mind. Bulldozer comes with a L1 WT too and that point only could explain many disappointments from a performances point of view.
    Again, I don't think so there is the problem with WT. L1D is WT, WCC is write buffer, and L2 is probably WB. Because of WCC there can be some issues with multiple write out streams. Also, we don't know what is behaviour of WCC when two integer cores writing data. There is probably WCC cache trashing.
    "That which does not kill you only makes you stronger." ---Friedrich Nietzsche
    PCAXE

  10. #10
    Xtreme Member
    Join Date
    Apr 2007
    Location
    Serbia
    Posts
    102
    Quote Originally Posted by xsecret View Post
    Was a P4 slower than a P3 ? Sometimes no, sometimes yes, depending on the software. The absolute performance is not something so important for AMD. The most important thing is money. And just money. Spending gazillions dollars in R&D to reach the performance of a CPU sold in very low quantities (and generating very low incomes) like the 990X is ridiculous. Bulldozer must solve two problems : 1/ Be able to gain performances (with frequency increases) at mid-term without spending more gazillions in another ľarch 2/ Compete with Intel *mainstream* CPUs (and not Extreme CPU) with a similar price/performance ratio.
    I agree. But, in BD architecture has less tradeoffs than Netburst. I expect per module min. same level of performance of K10, not 40-50% lower. Look at horrific chineese results of wprime. It is 65% slower than Thuban core per core and per clock. Something is wrong here. I still can't believe that it is true.

    Quote Originally Posted by xsecret View Post
    High raw throughput for an FP unit is nice. But in order to use this power in real-world application, you need a frontend able to feed it correctly.
    Average IPC of most workloads isn't much more than 1 IPC on Thuban core. 4-way front end is more than enough to feed two threads.

    And keep in mind the horribly slow L1 Write-Through, probably added in order to remove a bottleneck in frequency scaling. Write-Through means your writing from the frontend to the L2 "through" the L1.
    No, that doesn't mean WT.
    Write Trough means that every write to the cache causes a synchronous write to the backing store. Because L2 is slower than L1, L1 must wait for L2 to write out data. But there is WCC (Write Coalescing Cache) to hold on data for later writing out. I can't see why the WT policy cache is so much issue with BD core. Ratio between loads and stores is arround 2:1. For every two loads, we have one store.

    So, seen from the frontend, the L1 write bandwidth is as "slow" as the L2 write bandwidth.
    Not quite, because of WCC.

    The last ľarch to use that horrible trick was Netburst, with high frequencies in mind. Bulldozer comes with a L1 WT too and that point only could explain many disappointments from a performances point of view.
    Again, I don't think so there is the problem with WT. L1D is WT, WCC is write buffer, and L2 is probably WB. Because of WCC there can be some issues with multiple write out streams. Also, we don't know what is behaviour of WCC when two integer cores writing data. There is probably WCC cache trashing.
    Last edited by drfedja; 09-13-2011 at 08:45 AM.
    "That which does not kill you only makes you stronger." ---Friedrich Nietzsche
    PCAXE

  11. #11
    Registered User
    Join Date
    Feb 2005
    Posts
    39
    Quote Originally Posted by drfedja View Post
    Again, I don't think so there is the problem with WT. L1D is WT, WCC is write buffer, and L2 is probably WB. Because of WCC there can be some issues with multiple write out streams. Also, we don't know what is behaviour of WCC when two integer cores writing data. There is probably WCC cache trashing.
    WCC is a joke in the current BD implementation and is not able to catch up with the massive loss that comes from the L1D. The entire caching-system is lowering the performance of the ľarch. The L3 is a non-inclusive victim cache (L2 data are evicted to the L3) with data transfered from L3 to the L1D of the expected core without being copied to the L2. That mean high snoop traffic in order to keep the coherency correct. And snoop traffic is something really unwanted from a bandwidth/performance pov. There is a pardox here : The L1 is in Write-through, but you're not sure a data not in L2 is not the L1D of another core.

    Quote Originally Posted by AKM View Post
    Bullsh1t.
    Last edited by xsecret; 09-13-2011 at 09:07 AM.
    Doc_TB @ CanardPC.Com (FR)

  12. #12
    Xtreme Addict
    Join Date
    Feb 2008
    Posts
    1,209
    Quote Originally Posted by xsecret View Post
    WCC is a joke in the current BD implementation and is not able to catch up with the massive loss that comes from the L1D. The entire caching-system is lowering the performance of the ľarch. The L3 is a non-inclusive victim cache (L2 data are evicted to the L3) with data transfered from L3 to the L1D of the expected core without being copied to the L2. That mean high snoop traffic in order to keep the coherency correct. And snoop traffic is something really unwanted from a bandwidth/performance pov. There is a pardox here : The L1 is in Write-through, but you're not sure a data not in L2 is not the L1D of another core.
    Interesting reads here, but 3 questions:

    1.) why would AMD do such a thing (you wrote for better clock-scaling, that means higher clocks or better scaling?)

    2.) can it be optimized further

    3.)
    Write Trough means that every write to the cache causes a synchronous write to the backing store. Because L2 is slower than L1, L1 must wait for L2 to write out data.
    this contradicts somewhat
    The L3 is a non-inclusive victim cache (L2 data are evicted to the L3) with data transfered from L3 to the L1D of the expected core without being copied to the L2.
    or you mean this can be due to WCC? I am not a professional, but I would guess WCC doesnt take that long to preserve coherence... If anything at all, I'd guess the waiting of L1D for L2 can be a problem, but if stuff is written to L1D fast, and then some cycles later WCC writes to L2, I guess there will only be rare cases where
    The L1 is in Write-through, but you're not sure a data not in L2 is not the L1D of another core.
    really matters...

    I guess
    data transfered from L3 to the L1D of the expected core without being copied to the L2.
    will first speed up things, then
    snoop traffic in order to keep the coherency correct
    will need some time. All, to my eyes, depends on how this stuff is used, it can be faster in one case, and slow in the other... Mhm somewhere we came to that conclusion before

    Maybe the problem is because of
    The L1 is in Write-through, but you're not sure a data not in L2 is not the L1D of another core.
    , one core might need to wait for WCC to complete to ensure coherency?

    Ahhhh btw... CONGRATZ FOR NEW WR
    1. ASUS Sabertooth 990fx | FX 8320 || 2. DFI DK 790FXB-M3H5 | X4 810
    8GB Samsung 30nm DDR3-2000 9-10-10-28 || 4GB PSC DDR3-1333 6-7-6-21
    Corsair TX750W | Sapphire 6970 2GB || BeQuiet PurePower 450w | HD 4850
    EK Supreme | AC aquagratix | Laing Pro | MoRa 2 || Aircooled

  13. #13
    Xtreme Member
    Join Date
    Apr 2007
    Location
    Serbia
    Posts
    102
    Quote Originally Posted by xsecret View Post
    WCC is a joke in the current BD implementation and is not able to catch up with the massive loss that comes from the L1D. The entire caching-system is lowering the performance of the ľarch.
    How you can claim such things ? Are you chip architect of BD, or some experienced with low level code optimisation programmer to claim that the WCC is a joke or good enough patch for WT policy?

    The L3 is a non-inclusive victim cache (L2 data are evicted to the L3)
    So, on 10h is also non-inclusive, or exclusive, and K10 with L3 cache is faster than K10 without cache. Gain from such large L3 may be bigger if cache is faster, but K10 core run very well withot L3. Non-inclusive L3 isn't reason for low performance, especially in FP intensive, or integer intensive, number crunching apps like wprime or CB. With BD module there is 2MB of 16-way L2 cache which is enough for achieving acceptable performance levels. I don't think so that the SB will run so great without L3 cache. I also think that cache hieararchy of AMD K10 and BD is some kind of tradeoff for having good enough performance of low end, and APU products, like Llano and Trinity.
    with data transfered from L3 to the L1D of the expected core without being copied to the L2. That mean high snoop traffic in order to keep the coherency correct. And snoop traffic is something really unwanted from a bandwidth/performance pov. There is a pardox here : The L1 is in Write-through, but you're not sure a data not in L2 is not the L1D of another core.
    L2 is mostly inclusive of L1D caches, so the data from L1D caches in most cases is in L2, because of WT and mostly inclusive cache policy. But to have complete picture of the puzzle, we need to know how they solved problems with snoop traffic and how much from for eg. core 0 must snoop core 1 L1D through the L3.
    The question for AMD is what is the average hit rate of L1 copy in L2.
    "That which does not kill you only makes you stronger." ---Friedrich Nietzsche
    PCAXE

Bookmarks

Bookmarks

Posting Permissions

  • You may not post new threads
  • You may not post replies
  • You may not post attachments
  • You may not edit your posts
  •