I'm not even arguing about that so I don't see what your point is.
Printable View
I'm not even arguing about that so I don't see what your point is.
Funny, Hammer IMC can be "ganged" albeit the term "128-bit ganged operation" means two independent 64bit ctrls co-operate effectively acting as a 128bit ctrl.Quote:
Originally Posted by AMD
But, this not worth arguing, it's just semantics.
It's Japanese, not Chinese.Quote:
Originally Posted by savantu
Intel doing better than AMD is not a surprise, since Intel is quite a few times larger than AMD and has more than twice its R&D budget. AMD doing better than Intel certainly is. If AMD does better than Intel, then that just goes to show you how ineffective Intel is, right?
Real World Quad FX Speed & AgilityQuote:
Originally Posted by savantu
Our real world testing today disproves my preconceptions entirely, and shows that in quite a few cases, the FX-74 is as fast, or even faster than the QX6700. We ran the numbers over and over again, and the FX-74 simply has the horsepower to compete with the QX6700 in the most CPU intensive applications. When the systems are completely maxed out and you’re running rendering, encoding, or multi-media applications, the difference between the two CPU’s is quite minimal, usually between 1 and 3 percent.
http://enthusiast.hardocp.com/articl...VudGh1c2lhc3Q=
Thats not even K8L thats just what a K8 quad core would be like. Now if K8L is 40% faster the K8's how do you think york needs to run to beat K8L on just 2.9ghz. York will need 3.73ghz...oooops... thats what intel will do.
http://badhardware.blogspot.com/2006...y-rainier.html
Then 3.5ghz K8L will counter york with. Also to run HT3 at full speed 2.6ghz on the link is needed requireing 3.5ghz from the cpu cores. The math looks right according to this. Right now conroe is just 10 to 20% faster then K8's. In the real world its almost none at all. What do you thinks going to happen when K8L starts running at 3.5ghz stock? Its still 20% to 30% faster then core2 arcs. York can't do it then. They need core3 to win. However a new amd ark was announched for 2008, when that comes will core3 be enough to counter AMD's new ark?
AMD is only 1 to 3% slower then C2D now because of 4x4. What happens when you put 2 K8L's in that 4x4 it makes it a 8x8. Intel is going to have a hard time being better now at this rate. They need to lower QX6700 price if they want to stay ahead. AMD's are cheaper then ever. AMD is doing just fine infact they are not loseing at all. The benchmarks are in the link... 1 to 3% difference faster or slower then QFX can not be even close to a difference. Stop over hypeing C2D, QFX is a match for it from what we see now. Stop lieing don't lie to yourself its just childish. We see the truth. QFX is no better or worse then QX6700. They all have there advantages and disadvantages.
It is clear they are only 1 to 3% different from eachother. That is not stoming all over AMD that is barly keeping the lead agenst a much cheaper solution. Its a struggle. Stop kidding yourselfs. And when K8L makes 4x4 into 8x8 then we will see whos laughing. But all I need to do is drop a K8L into my AM2 without any mobo changing. If you want Quad core you need a new mobo not me. Intel needs CSI and Core 3 to compete with a 3.5ghz K8L. Cus york isn't enough. They wouldn't raise the core 800mhz for no reason X6800's will be killed by a K8L at 2.7ghz with only 2 cores. Thats Intel being scared, kind of extreme to raise the speed to such crazy speeds. If you think a X6800 will compare to K8L your wrong.
Thats why they are making york to stay compeditive until Core3 and csi. Then AMD will be in trouble again. However AMD isn't sitting around after K8L they have another ark that will be much better then K8L and its been in devenopment since 2003. K10 was canceled partually in favor for K8L a mix of K10 and some new ideas, replaced with K11 the new ark will be out in the end of 2008.
Somewhat true. But the QFX uses 50% more power! AMD should've just waited and released it's native quadcore cpu. No need to keep up with Intel on the quadcore front using 90nm technology.Quote:
We see the truth. QFX is no better or worse then QX6700
http://www.amd.com/quadcoredemo
K8L performance comparison
http://img86.imageshack.us/img86/5162/amd1fe5.jpg
http://img86.imageshack.us/img86/1534/amd2tt9.jpg
as i promised )) suppose it's not all for december
http://img146.imageshack.us/img146/282/amd3jj1.jpg
Yeah, that's pretty lame on AMD's part. Any demo would've been better than task manager.Quote:
Too bad AMD didn't run ONE SINGLE D$#@ BENCHMARK today. Just TASK MANAGER.
ROFLMAO!!! I've never before seen one person be so completely wrong on so many different points in one post before! Congrats! :toast:Quote:
Originally Posted by Serge84
You're not mentally challenged are you? I only ask because I wouldn't feel right about laughing so hard at you if you were.
In AMD's defense, it was clearly the best they could manage this early in K8L's development.Quote:
Originally Posted by brentpresley
Gotta love the estimated performance. No test is even run, just some guesses and ofcourse something that had to be better. Looks pretty sad if it ships around or after 45nm Intel products and it can barely if even beat todays Intel offerings.
2 x 3ghz 90nm K8s in dual socket beat Intel's QuadCore offerings today in 3d rendering and other
strongly multithreaded apps, the Altairs will simply crush them, be it on 45nm :P
Yeah, and they (90nm K8's) overclock to 3.5GHz with ease as well....:rofl: :rolleyes:Quote:
Originally Posted by alayashu
Oh my, yes. Take the one review that supports your hopes and dreams, dismiss all the others, horrendous power consumption and jet aircraft landing noise levels and dream the fanboy dream... I hope you buy a 4x4 rig, I sincerely do... :DQuote:
Originally Posted by alayashu
Fred, the 1500$ QX6700 its barely faster than than the 599$ FX70 in Cinebench, under Vista and it gets its ass kicked by the FX72 and FX74.
http://www.hwupgrade.it/articoli/cpu...ench_vista.png
The problem with XP is that it handles NUMA like :banana:, unlike Vista, therefore on Vista all the AMD dual socket systems (including QuadFX) get a nice performance boost.
-e
it was vista64 allready lol :>
That's not NUMA, that's from going into 64-bit mode. Cinebench is almost completely non-depedent on bandwidth, memory latency or cache size. How come there isn't a NUMA boost for the big megatasking test? In fact, it slows down.Quote:
Originally Posted by alayashu
http://www.hwupgrade.it/articoli/cpu...megatask_1.png
http://www.hwupgrade.it/articoli/cpu...sk_1_vista.png
And I bet NUMA won't stop the FX-74 from getting beat by the FX-62 in the vast majority of game benches.
QX6700 is faster in anything else...so unless you render in cinebench day and night. And if you do, who pays your power consumption? Also QX6700 aint 1500$. But its 999$, else FX70 aint 599$ and you still need a 480$ board and a super PSU vs normal cheap board and a standard PSU.Quote:
Originally Posted by alayashu
Maybe check techreport..they used XP64, or 2003 NUMA so to say:
http://techreport.com/reviews/2006q4...x/index.x?pg=1
You are grasping for straws :rolleyes:
And we all buy 4x4 for workstation tasks right? We wouldnt buy a dual opteron system or dual xeon for that. What an odd thought! And AMD markets 4x4 as a GAMING PLATFORM.
NUMA and the chained HT structure is the reason FX62 beats it silly in games. I think alot of people missunderstand what NUMA is about. Its alot faster for the CPU to fetch data via its own IMC than it is to travel over HT and through another IMC. Core<->IMC got alot more BW than HT can deliver and therefor gets a performance loss whiles its located remote. So want NUMA performance? Run some independent threads (Not game) that can use local memory of its host CPU. NUMA is alot nicer when its not the same workingset ;)Quote:
Originally Posted by accord99
Also there aint much HT bandwidth left in heavy multitask either, when CPUO needs to travel over HT (That it shares memory with CPU1 with) to reach chipset A and B that it again shares HT links with and chipsetB yet again get shared. CPU0->CPU1->ChipsetA->ChipsetB. Oh ye, it reminds me of old BNC networks! Clumsy design.
put QuadSLI on it and it's the ultimate gaming platform at ultimate price, hehe :DQuote:
And we all buy 4x4 for workstation tasks right? We wouldnt buy a dual opteron system or dual xeon for that. What an odd thought! And AMD markets 4x4 as a GAMING PLATFORM.
Though it would be nice to see some next gen multithreaded games on Vista64 /w QuadFX vs Kentsfield.
I bet the difference will be much lower than the most optimistic voices have said.
And the platform itself is nice for rendering guys, it performs impressively well
for its price, but mainly the upgradability to Octo Barcelona/Altair is what
probably will get AMD more sales than expected towards this category.
QuadSLi is useful for what? Even 8800GTX SLI on anything but 30"+ is a waste. Also the quadsli might be impacted by the 2 extra slots being 8x and the starved HT links.Quote:
Originally Posted by alayashu
Multithreaded games is still the same workingset, means NUMA is a problem rather than a benefit.
Also working in the animation business and making international movies. 4x4 is not nice for rendering guys, its too expensive to buy and run, too crappy, to
much of a spaceheater. Try use Maya or something and you see 4x4 get beaten badly. Even 3DsMax is not good. But for some reason people focus on cinebench...
And the upgrade path? You mean the AM2+ socket quadcores where you need a new 4x4 board not to get hit by yet another performance penalty stick? :fact:
in maya/mentalray the difference in the favour of QuadFX is the biggest among all the rendering softwares.Quote:
Also working in the animation business and making international movies. 4x4 is not nice for rendering guys, its too expensive to buy and run, too crappy, to much of a spaceheater. Try use Maya or something and you see 4x4 get beaten badly. Even 3DsMax is not good. But for some reason people focus on cinebench...
About the maya software/max scanline renderers, they are old poorly multithreaded softwares, as you probably know.
Cinebench9.5 reflects pretty good the performance CPUs in modern renderers as mentalray, vray, prman, etc.
true, not many need QuadSLI, not even DualSLI, not even QuadCPU, not even DualCPU. But the options exists for those who know they need them.Quote:
Originally Posted by Shintai
remains to be seen ;)Quote:
Multithreaded games is still the same workingset, means NUMA is a problem rather than a benefit.
watch the official AMD presentation, they say it clear all you'll have to doQuote:
And the upgrade path? You mean the AM2+ socket quadcores where you need a new 4x4 board not to get hit by yet another performance penalty stick? :fact:
is to change the 2x2C K8s with the new 2x 4C Barcelonas and update the
mainboard Bios. And they say it clear that there wont be any performance
penalty nor power consumption or whatever problems with it.
C2D/Woodcrest/Kentsfield/Clovertown here beats K8 silly in Maya. But maybe its the 8way dualcore opteron systems that sucks :(Quote:
Originally Posted by alayashu
Just remind me again WHY you would even run quadsli. And then remind me again what GFX cards CAN run quadsli and with what drivers. (GX2 cards dont count, it needs to be 4 single cards).Quote:
Originally Posted by alayashu
Nope, the FX62 making a joke out of even FX74 in all cases where its 2 threads or less is what is to be seen.Quote:
Originally Posted by alayashu
I never said you couldn´t do it. But using old HT instead of HT3 in dual quadcore..lol! Its like comparing Merom and Conroe and say FSB speed doesnt matter.Quote:
Originally Posted by alayashu
show me one single bench of maya/mentalray with an 3ghz Woodcrest vs 3ghz QuadFX (even if the second setupQuote:
C2D/Woodcrest/Kentsfield/Clovertown here beats K8 silly in Maya. But maybe its the 8way dualcore opteron systems that sucks
costs near twice less money) under Vista64 in wich the Woodcrest wins hands down.
Oh, and using all the cores, 4 threads please, and a full raytracing 10mil displaced polys scene, not a sphere or such.
Quad Monitor Setup it's a pleasure to work with in heavy scenes while animating, modelingQuote:
Just remind me again WHY you would even run quadsli.
and rendering in Maya and working with Photoshop on the textures :rolleyes:
/* the end with this postcount++ OT */
You really hang on to Vista64 dont you? You think its the saviour that will rescue 4x4 from looking crappy? Vista64 dont offer superduperultradeluxe new NUMA performance boost. We talk <1%.Quote:
Originally Posted by alayashu
How is it SLI and just dual monitors work out?Quote:
Originally Posted by alayashu
In this context they're referring to having 2 DIMMs working together to get 128 bit mode.Quote:
Originally Posted by largon
Details can be important too.Quote:
Originally Posted by largon
Quote:
Originally Posted by freeloader
Quote:
Originally Posted by shintai
To go from tape out to booting windows in 5-6 months is impressive, and not being able to look inside the machine is pretty common with early parts and standard practice in general for pre-release stuff, so I don't see why you're all getting worked up about this. Still ~6 months away from scheduled release.Quote:
Originally Posted by brentpresley
What kind of job does your friend have? (if we could know)Quote:
Originally Posted by brentpresley
there's around 14weeks gap between tap out and first silicon. assuming they are at a2 right now, it would still take about 3 respins before it hits the market. Not able to run benchs right now implies that they have a slightly worse chance for a May/June release.Quote:
Originally Posted by mesyn191
But K8L is starting to look great regardless. But the mainstream parts of K8L is yet to be seem.
So much for the green company and its performance/watt claims.Quote:
Originally Posted by alayashu
No one would argue about how HTT bus is superior to FSB, and even Intel knows it needs some sort of serial interconnects.
And what you have showed is not really relevent either. Comparing a 2P unit to an UP unit?
here is the translation of that siteQuote:
Originally Posted by alayashu
http://translate.google.com/translat...5findex%2ehtml
why translation, here is original english version:Quote:
Originally Posted by The Ghost
http://www.hwupgrade.com/articles/cp...orm_index.html
how about this? ;)Quote:
Originally Posted by Shintai
http://www.hwupgrade.com/articles/cp...as_1_vista.png
http://www.hwupgrade.com/articles/cp...as_1_vista.png
Serge84, Core2 is far faster than K8 in "real world" things, almost 40% clock-per-clock.
Quad FX can keep up Kentsfield in strong multitask due the huge FSB limitations.
I think K8L will be about 10% faster than Core2 clock-per-clock.
And 8x8 systems will carry the performance crown in desktop til the next generation.
But K8L Quad-Core won't cross the 3GHz wall in 65nm process due to the 125w AMD's top TDP.
Dual-Core K8L CPUs may get to about 3.4GHz.
you call gaming at 640x480 "real world"? ;) :DQuote:
Originally Posted by doompc
K8L is 65nm, they shoudl have no problem scaling it over 3GHz with Quad CoreQuote:
Originally Posted by doompc
Core2 is 65nm and Intel can't scale Kentsfield over 2.66Hz cause it eats up 130w.
Rev H quad-core is the biggest AMD die ever made (or at least the biggest since the Clawhammer).
65nm Quad-Core K8L may be just a little less power hungry than a 90nm Dual-Core K8 at the same frequency.
And because it's wider execution engine it may be dificult to scale frequency unless AMD go deeper pipeline...
Even Intel says its closer to 20% improvement in IPC over K8, they only claim a 40% IPC advantage over the P4.Quote:
Originally Posted by doompc
K8L comes with SiGe SOI which yields upto ~40% increase in transistor perf compared to vanilla 65nm silicon. It's up to AMD how will they make use of that advantage.Quote:
Originally Posted by doompc
brentpresley
It's funny you picked the ONLY one it won to put up here.
Let me correct you - i picked up ANOTHER one in addition to one from previous page.
If that 40% better transistor perf with the same power consumption is true, Rev G Brisbane may work at 3.6GHz with stock vcore.
But K8L is a different core.
Intel uses the SiGe layer since the 90nm process, they call it Strained Silicon.
AMD came with a better tecnology in rev E, the Dual Stress Liner, it's about the same, but a little better and without the expensive SiGe layer.
http://www.xbitlabs.com/articles/cpu...-venice_2.html
In AMD's 65nm process the SiGe layer is applied and then removed:
http://www.theregister.co.uk/2005/12...essed_silicon/
More on Dual Stress Liner:
http://www.lostcircuits.com/cpu/amd_venice/3.shtml
And on Strainer Silicon:
http://www.lostcircuits.com/cpu/prescott/2.shtml
First the QX6700 doesn't cost $1500, more like $1000-1100 (search Froogle). Second, $850 QX6600s will be here soon enough and generally outperform the more expensive $1000 FX74 even before overclocking them to levels that QFX simply can't touch. Third, QFX has a lot of additional platform costs that C2Q doesn't have. 1000W PSUs aren't cheap, neither are the $400 mobos or the water cooling required to make QFX noise levels tolerable. IMO it's not wise to try and argue QFX as a better value than C2Q because you will always lose that argument. QFX is anything but economical in both initial cost and operating cost.Quote:
Originally Posted by alayashu
QFX beating C2Q in one out of ~50 benchmarks under Vista64 is not terribly impressive. If all you do is run Cinema 4D and power useage, heat and noise aren't concerns, QFX may be a good fit for you. Otherwise C2Q offers far superior performance overall for lower initial cost and lower long term energy costs.
2.67Ghz Clovertown with 1333FSB is "only" 120W TDP. Yes they could scale it higher.Quote:
Originally Posted by doompc
65nm K8L quadcore will be like Intel i thin. Some 100-130W for the top bins.
Just a shame you cant directly add these things into CPU speed. Look at history for a fact. Its like the 250Ghz transistors at room temp. Its just different when you need a quarter of a billion of them.Quote:
Originally Posted by largon
yes, but even a 20% boost as a whole would be nice. brings our 3ghz oc'ed chips up to 3.6ghz ;)Quote:
Originally Posted by Shintai
You are basing this statement on the assumption that the 3.73GHz Yorkfield will be either as fast or just slightly faster than the 2.9GHz K8L. But you have not provided any evidence, nor have you given us any reason to believe that you would be in possession of such knowledge.Quote:
Originally Posted by Serge84
Looking at real world benchmarks over at Toms Hardware comparing the E6600 and the X2 4600+, both running at the same clock frequency, I'm getting a difference in between 15% to 35%. The biggest difference is in gaming and encoding. No one here believes that the core 2 duo and the K8 are close to tied in performance per clock. Neither do you.Quote:
Right now conroe is just 10 to 20% faster then K8's. In the real world its almost none at all.
The 40% increase has not been proven yet, but if we assumed it was:Quote:
What do you thinks going to happen when K8L starts running at 3.5ghz stock? Its still 20% to 30% faster then core2 arcs.
K8 = 100%
Core 2 = 100% + 20% = 120%
K8L = 100% + 40% = 140%
Difference between Core 2 and K8L = 140%/120%=117%
Thus... the K8L would be 17% faster than the Core 2 per clock.
Not all that impressive when you think about it, since it would take a 3.38GHz Core 2 (65nm) processor to beat a 2.9GHz K8L processor.
This is of course assuming that the percentages are reliable, which I'm not all that sure of. We have not yet heard of any K8L processors running at 3.5GHz stock.
let me see if I get your math correctly. you say that Conroe is just 20% faster than K8 and that K8L is just 17% faster than conroe.
And all the Intel Fanboys are screaming about Conroe being Jesus reborn...
and K8L is not a big deal how?
It was hypotheroetical.Quote:
Originally Posted by nn_step
Also how much faster is Core 2 over Core 1? Its just that people seems to expect K8L will be the jeesus reborn and do something super amazing.
K8L gonna be 40-50% faster than K8 per clock? And clock to 3.5Ghz? I smell jeesus..
That is about what Conroe does right now, K8L (dual core) might do the same and sport the benefits of the IMC.Quote:
Originally Posted by Shintai
There certainly are a lot of bold assumptions being made about K8L considering the known facts.
I keep hearing mid-2007 when AMD claims Q3.
I keep hearing 40-50% faster than C2D when AMD claims 40% faster than K8.
I keep hearing 3.5GHz when AMD claims 2.7-2.9GHz.
It's a little early to make such assumptions, isn't it? But since we're speculating... IMO the likely scenerio is K8L releasing in Q3-07 at 2.7-2.9GHz and up to 40% faster than K8 (same clock), 0-20% faster than C2Q (same clock and app dependent) but with C2Q clocked at 3.33GHz or more. Shortly afterward Intel launches 45nm Penryns at 3.6-4GHz and we're back to where we are now... watching AMD struggle to keep up with Intel while being one process shrink behind.
I know that AMD is once again claiming that they will narrow the gap from 12 to 6 months at 45nm but they've been claiming that for so long I'll have to see it to believe it. I hope they do though. It's their achilles heel.
Core 2 is about 0-20% faster than Core 1. Not 40% faster.Quote:
Originally Posted by doompc
AMD didnt claim 40%Quote:
Originally Posted by Fred_Pohl
Core2 is 15% to 35% faster than K8.
Assuming K8L core is as fast as Core2's, the IMC would make it perform 10% faster.
About 20% to 40% faster than K8.
And being made in 65nm process it may overclock "easealy" to 3.6GHz.
But when K8L appears Intel will be almost releasing Yorkfield and Wolfdale, they will be 45nm, cheaper and better overclockers.Quote:
Originally Posted by doompc
The thing is that ... AMD has to work on its transitors to clock higher ... the core 2 duo basically relies on prescott transistors, which were made to clock mad high. that's probably why core 2 duo clocks that high. just a lil speculation though.
a set of interesting facts for you.Quote:
Originally Posted by Saiyan[CHW]
1) more stages means more clock speed
2) Smaller process can mean more clock speed
3) K8 has less stages (12int, 17 FP) and is on a larger process(90nm) vs Conroe (14int, 22FP, 65nm)
4) K8 Max Clock speed is (on average) 3Ghz
5) Conroe's Max Clock speed is (on average) 3.3-3.5Ghz
Just thought I should put those things out
Saiyan, yes.
Intel is already sampling 45nm Core2 cpus (Penryn, for mobile).
So even if K8L is clock-to-clock faster than Core2, Wolfdale will clock far higher than K8L.
Let's see what the IMC can do to help and how good is AMD's 65nm process.
k8 max clock speed is from 2.7 to 3ghz on air, allendale is 3ghz to 3.4, and conroe is from 3.3 to 3.7ghz, thats true, but again consider that k8l is going to be clocked lower than 3ghz initially, and that penryn will be clocked higher than 3ghz, assumed considerably, and it still doesnt look good, except for 8x4. k8l will not be considerably faster than core 2, since it doesnt introduce any revolutionary changes in architecture, and without them, its impossible to see sudden drastic changes in performance. core 2 wasnt that drastic of a change from core 1, and k8 to k8l is a very similar switch.
IIRC AMD has put out some marketing propaganda claiming that K8L will be 40% faster than K8 in certain server BMs. Which is where I think K8L will really shine.Quote:
Originally Posted by Shintai
Quote:
Originally Posted by nn_step
Although I pretty much agree with most of your selected 'facts', IMO you've skewed the avg max clock speeds a bit. Getting 3GHz from an X2 seems to be every bit as difficult as getting 3.6GHz from a C2D. I'd say that on avg C2D OC's ~500MHz higher than X2, not 300-500MHz. Not meaning to split hairs here but 200MHz is 200MHz or to use AMD marketing speak, on avg C2D OC's to '7000+' vs '6000+' for X2.
My hope is that they are able to produce these kind of results in the dual core processors in future revisions of the 65 nm series...
Intel is AMD's competitor... K8L will be competing against 45nm Core2. People are speculating/predicting a lot of things about K8L and 45nm Core2... what's the big fuss?Quote:
Originally Posted by LOE
Are you alergic to the mention of Intel?
all the fanboys are didn't you knowQuote:
Originally Posted by Epsilon84
No, that was a simple K8 quadcore against dualcores.Quote:
Originally Posted by Fred_Pohl
max average whatQuote:
Originally Posted by brentpresley
screenshot
or 1M
or 32M
or 3D benches
or prime
It's been december for so long!!!! When can I hold a brisbane!?!?!
I want one now!
Tomorrow, by the looks of it. I'm probably going to get one pretty much ASAP, unless they're crap...Quote:
Originally Posted by Silvermirage
I wonder if just the NDA lifts tomorrow, or if E-tailers will have them for sale tomorrow (meaning they already have them in the channel.)
I think that tomorrow is the release.Quote:
Originally Posted by Silvermirage
I wanna see Nehalem, an Intel IMC will be interesting. :)
I stand corrected. My bad.Quote:
Originally Posted by Shintai
Probably not, they're starting to put so much cache on their chips an IMC won't give them very much in the way of performance gains. IMO thats an IMC's real benefit, it allows you to significantly reduce system RAM latency without requiring you to dedicate more and more CPU die space to cache, which is very important to AMD seeing as how they're almost always fab limited.Quote:
Originally Posted by Saiyan[CHW]
I think they'll reduce the cache size. (L2) Don't know anything about L1. Probably they'll gain performance.Quote:
Originally Posted by mesyn191
Oh they'll certainly gain performance, wether its significant and/or worth while vs. Intel's current practice of stupid size L2 cache or not is what is I'd question, for AMD it makes alot of sense though. I doubt they'll increase the L1 cache size at all, Intel doesn't seem to like doing this much, probably because it makes it more difficult to get a higher clock speed.
there is no reason to increase L1 Cache Size for AMD, increasing bandwidth on the other hand is a different matterQuote:
Originally Posted by mesyn191
Didn't suggest AMD would, I was talking about Intel.Quote:
Originally Posted by nn_step
AMDs and Intels L1 cache is very different. AMDs is twice as big, but basicly twice as slow. Overall the got about the same level of efficiency. And I doubt the benefit is worth it, if either AMD or Intel have increased it any further. I doubt cache is any problem for both in terms of clockspeed.Quote:
Originally Posted by mesyn191
Yes, Intel has always made processors with a small but vey fast L1.Quote:
Originally Posted by Shintai
I think that an Intel IMC will bring a lot of benefits... for example, more cache means more heat, if they include an IMC in Nehalem the processors won't like big L2, I think that in some years 2MB-4MB L2 will be samething standard for Intel and AMD.
Cache dont make much heat. On 16MB Tulsa its 0.75W per MB cache. Logic however does. Even with IMC Intels cache will reach new heights. Main memory is just way way too slow. Its something like 100GB/sec vs 10GB/sec not even to mention latency.Quote:
Originally Posted by Saiyan[CHW]
Amen. A lot of people seem to think that a IMC works miracles but it doesn't. IIRC AMD's IMC has only 10-15% lower latency than i975. That's nice but not even close to L2 or even L3 cache.
10-15 percent??Quote:
Originally Posted by Fred_Pohl
More like 20ns But only if you compare AMD IMC with the slowest DDR2 to 975X with high-end >1GHz and low latency settings.
*Ahem* umm Latency for Intel from CPU to Northbridge and to Memory and Back is about 200 - 300ns (depending on memory and quality of Northbridge)Quote:
Originally Posted by Fred_Pohl
Latency for a Socket AM2 Ranges from 40ns To 80ns (depending on the quality of the memory)
The reason that it doesn't seem to have much effect on Intel is because of the Prefetch design of the Processor. IMC can get away with a Simple prefetch unit, Intel's FSB method however can not.
The L1 cache is on the critical pathways of the CPU so it has a direct impact on how high you can get the clockspeed.Quote:
Originally Posted by Shintai
IIRC AMD's cache is "slower" than Intel's but not flat out half as fast, I thought it was Intel's L2 cache that really blew away AMD's in the K8? K8L's L3 should level the playing field on this IMO...
Like nn said cache doesn't generate a lot of heat, logic is the main culprit there. The only thing bad about having lots of cache is that it eats up your silicon budget for core logic and/or makes your die size massive thus lowering yields and increasing costs, otherwise I could care less if AMD/Intel/etc. were throwing 1GB L2/L3's on their cores, in fact I'd be all for it as then you could get rid of low bandwidth and high latency main system RAM.Quote:
Originally Posted by Saiyan[CHW]
IMO the only reason Intel would go the IMC route is because they want to dedicate more of their available silicon towards logic instead of cache (which makes sense given that were heading in the direction of more cores rather than more Mhz to increase performance, think 8+ Conroe cores on a single die, something has to give and I guess they've decided cache is it). Anyways thats _HIGHLY_ speculative, so feel free to poke holes in it.
Supposedly Intel is putting some pretty incredible amounts of cache on their recent chipsets as well...Quote:
Originally Posted by nn_step
Intel is 256bit, 8way. Amd is 128bit, 4way. Intel is more than twice as fast. However AMDs doublesize makes up for it so they are about equal in the hit/miss ratio.Quote:
Originally Posted by mesyn191
not really. Their amounts for their standard retail chips aren't huge by any stretch.Quote:
Originally Posted by mesyn191
You should try Power5's 256-512MB of L3 ;)
Doesn't the K8's dual ported L1 deliver 2x 128 bit reads or 1x 128 bit read and write per clock make that bus width distinction meaningless though?
By huge I was of course comparing to what you might find in the consumer class products, if you're going to make comparisons vs. enterprise class products (didn't think POWER5 was retail?) then yes it does seem pretty weak but that is to be expected.Quote:
Originally Posted by nn_step
right now BOTH AMD and Intel Fetch 16 bytes per cycle Per L1 Cache.Quote:
Originally Posted by mesyn191
regardless of all other bandwidth associated by the L2/L3. That is what is actually being sent into the Processing core.
Thus if you actually do the math, that means a max of 1.5 SSE4 calculations per clock issued. Conroe/K8/Whatever
K8 is an enterprise class product, Opterons remember?Quote:
Originally Posted by mesyn191
And it was kind of retail, you just couldn't buy it at newegg though..
Intel Core 2
AMD K8
L1 data cache
32 KB
64 KB
L1 instructions cache
32 KB
64 KB
L1 latency
3 clock cycles
3 clock cycles
L1 associativity
8-way
2-way
L1 TLB size
Instructions: 128 entries
Data: 256 entries
Instructions: 32 entries
Data: 32 entries
L2 latency
14 clock cycles
12 clock cycles
L2 associativity
16-way
16-way
L2 cache bus width
256 bit
128 bit
L2 TLB size
?
512 entries
Pipeline
14 stages
12 stages
x86 decoders
1 complex and 3 simple
3 complex
Integer execution units
3 ALU + 2 AGU
3 ALU + 3 AGU
Load/Store units
2 (1 Load + 1 Store)
1
FP execution units
FADD + FMUL + FLOAD + FSTORE
FADD + FMUL + FSTORE
SSE execution units
3 (128-bit)
2 (64-bit)
This table explains a lot of things right away. And the most important thing is that the processors with Core microarchitecture have “wider” architecture that allows processing more instructions per clock cycle than CPUs with K8 microarchitecture. Although the execution units of both competing processor architectures can process up to three x86 and x87 instructions per clock cycle, Core Microarchitecture should prove more efficient with SSE instructions. While K8 processors can perform only one 128bit command per clock, Core can process up to three commands like that.
Moreover, Core Microarchitecture boasts another great advantage: more advanced decoding system. Together with the four decoders, macrofusion technology allows decoding up to five instructions per clock (in an ideal case). The competitor processors can only decode three instructions simultaneously. All this indicates that the decoders of Core Microarchitecture based CPUs will be able to better load the processor execution units by performing up to four instructions per clock in the most optimal conditions. In this case the overall commands execution will go 33% faster than by K8 AMD processors.
Here I would like to also mention more efficient data processing algorithms of the CPUs on Core Microarchitecture. The advantages of this microarchitecture show themselves best in the data caching system. Although, the L1 cache of the Core based processors is smaller, it is more associative. And as for L2 cache, it is not only bigger but also has higher bandwidth. Moreover, the shared structure of the L2 cache memory is beneficial for multi-threaded workload.
An important addition to the data prefetch algorithms of the new Core based processors is the unique memory disambiguation technology that has no analogues in the competitor solutions. It makes the upcoming Intel processor more out-of-order (from the code prospective).
In fact, the only indisputable advantage of the AMD K8 microarchitecture that will survive the arrival of Core will remain the integrated memory controller that can definitely ensure lower latency during data processing. However, it is a very tough question if integrated memory controller will be enough for AMD to worthily oppose the Conroe processors, and we still have to answer it later. However, AMD engineers are not keeping their hands in pockets. The future Athlon 64 cores scheduled to come out in early 2008 will be free from some architectural bottlenecks. But, it is a different story and a different article.
http://www.xbitlabs.com/articles/cpu...preview_9.html
if do that, atleast link the source
http://www.xbitlabs.com/articles/cpu...o-preview.html
I did link the source at the bottom. Is this better?
http://www.xbitlabs.com/articles/cpu/print/amd-k8l.html
Instruction Fetching:
Each clock cycle the K8 processor is fetching instructions in aligned 16-byte blocks from the L1I instruction cache into a buffer where the instructions are extracted from the block and are then sent to the decoder’s channels. The 16 bytes per cycle fetch rate allows sending 3 instructions with an average length of up to 5 bytes for decoding on each cycle. In certain algorithms, however, the average instruction length in a chain may be bigger than 5 bytes.
Particularly, the length of a simple SSE2 instruction with register-register operands (for example, MOVAPD XMM0, XMM1) is 4 bytes. If an instruction uses indirect addressing (with a base register and an offset like in MOVAPD XMM0, [EAX+16]), its length increases to 6-8 bytes. In 64-bit mode a one-byte REX prefix is added to the instruction code when the additional registers are employed. Thus, SSE2 instructions may be as long as 7-9 bytes in 64-bit mode. An SSE1 instruction may be 1 byte shorter if it is a vector one (that is, operates on four 32-bit values), but is 7-9 bytes, too, under the same conditions if it is a scalar one (with one operand).
In this situation, the fetch rate of 16 bytes per cycle doesn’t seem fast enough to keep up the decoding speed at a rate of 3 instructions per cycle. This limitation is not important for the K8 processor because vector SSE and SSE2 instructions are decoded at a rate of 3 instructions per 2 clock cycles (or 1.5 instructions per cycle), which is enough to load the two 64-bit FPUs. In the future processor, a rate of at least 3 instructions per cycle must be maintained. Considering this, the fetching in 32-byte blocks, announced in the presentation of architectural innovations of the K8L, doesn’t seem excessive. If the succession of these long commands takes a few neighboring 16-byte blocks, then the average fetching tempo with 16-byte data blocks of 3 commands per clock cannot be achieved.
By the way, Conroe processors fetch instructions in 16-byte blocks, just like K8 processors do, so they can decode the instruction stream at a rate of 4 instructions per clock only when the average instruction length is no longer than 4 bytes. Otherwise the decoder cannot process not only 4 but even 3 instructions per clock. To fight this in short loops, the Conroe has a special 64-byte internal buffer that caches loops up to 64 bytes long (four 16-byte blocks) and allows fetching data in such loops at a rate of 32 bytes per cycle. If a loop is longer than 4 blocks, it cannot be cached in this buffer.
The fetching of the next block of instructions is done using the branch prediction mechanism if there are any branch instructions present. Branches are predicted in the K8 processor by means of simpler algorithms than those employed in the Conroe. For example, the K8 cannot predict alternating indirect branches (this may have a negative effect on the execution of object-oriented polymorph code) and is also doesn’t always predict correctly regular patterns. The branch prediction mechanism will be improved in the K8L, but there’s no detailed info about that yet. The branch tables and counters will probably be made larger, and the algorithm of predicting branches alternating in regular patterns may be improved.
AMD sure marketed it as such but that was largely considered a joke against something like POWER5, or PA-RISC, heck even Itanium.Quote:
Originally Posted by nn_step
Kind've retail? Never seen a "kind've" retail store before, must be local only...Quote:
Originally Posted by nn_step
It is a great list but it is missing one important point:Quote:
Originally Posted by Fred_Pohl
Core2 also has speculative prefetch for the L2 data cache.
That means Core2 can fill cache lines in that cache even before it is sure they will be needed. If the code execution later turns out to need other data, then it can discard it. The K8 cannot do this, the K8 can only prefetch when it is sure which cache line will be needed.
That is a huge deal when running code that is not super-optimized. Because in the traditional setup your L2 cache is pretty much never capable of having data that you didn't use before ready, and you will stall until the memory fetch finishes. With speculative prefetch you have a much higher chance of finding the data you want even if it wasn't in the cache from previous operations.
I think that without that feature any CPU cache above 1 MB is probably useless for general-purpose, not superoptimized code.
I haven't had time to confirm, but I think this plays a major role for one performance oddity I see with Core2: while it is normally 20% faster than AMD64 at the same clock, it sometimes goes to 50% and I have seen 90%. And that is code that is known to be totally free of SSE. I think prefetch is the most likely cause of this.
that is a very valid point, however in recent years the quality of compilers are making well Optimized code more and more common.Quote:
Originally Posted by uOpt
Now the direct implications of that prefetch would most prevalently appear in the compiling of software. However after the said compilation, there isn't going to be much of a noticeable difference. At the same time it is rare for the standard user to build applications from source, even though it is easy as hell with FreeBSD
I don't see where's the joke...Opteron are used in Cray supercomputers, for instance.Quote:
Originally Posted by mesyn191
No.Quote:
Originally Posted by nn_step
And I work on a compiler :)
Of course there's progress. But nobody comes close to compile optimally (not even for C/C++/Fortran, not to mention other languages) for an optimization nightmare like Netburst.
All that code that runs well on Netburst is either
- very simple so that the compiler can figure it out
- it is hand-optimized with SSE code that actually makes use of the parallelism in SSE. A dirty secret is that although most compilers make use of SSE by now and get a speedup out of it compared to that rotten x86 instruction set - but they were using it as a one-data piece at a time machine.
- Both (almost always), such as encoding/decoding.
Anyway, the true advantage of AMD64 over Netburst was that code that did not go through a lot of hand-optimizations performed well enough. Games, scripting languages etc.
And now Intel leapfrogged and Core2 actually does it better than AMD64. Run normal code fast. And that is what the user wants. Before whole classes of software have no time for hand optimizations, namely games.
That is true, actually. The 90% speedup case that I mentioned is actually for compilation of a large software system, when turning on all the compiler bells and whistles.Quote:
Now the direct implications of that prefetch would most prevalently appear in the compiling of software. However after the said compilation, there isn't going to be much of a noticeable difference. At the same time it is rare for the standard user to build applications from source, even though it is easy as hell with FreeBSD
%%
And I didn't realize until now that the Core2 has such a much larger TLB covered area in L1 compared to the AMD64. That might as well have to do with it, too.
So what is the number of L2 TLB entries?
Intel refuses to publish or reference that number anywhere.Quote:
Originally Posted by uOpt
We can make estimates but we can't know for sure.
Conroe's L1 cache does 1 256 bit load per cicle.
K8L's L1 cache allows 2 128bit loads and is twice as big.
So cache-wise K8L is a lot better than Core2.
Xbit doesn't say if K8L's decoders are any better than K8's, aside of decoding SSE instructions in less macro-ops (due to the widder 128 bit SSE units).
Can anyone confirm if it has 4 decoders?
If that isn't true, what is the point in fetching 32 bytes of code per cicle?
K8's 3 decoders decode 3 simple instructions per cicle, 1 simple and 1 double or 1.5 double instructions per cicle, that means it outputs 3 macro-ops per cicle.
Complex instructions thru the vectorpath also are limited to 3 macro-ops per cicle.
With average 4 Byte-long x86 instructions, a 32 Byte chunk contains 8 instructions, but K8L can only decode 3 per cicle (assuming it has 3 decoders) or less, so the bufers will be full in just a few cicles...
And when it comes to SSE execution, Conroe still leads, do 3x 128 bit SSE per cicle, K8L will do 2.
This should be very easy to see in the OS. The OS must know the number of TLB entries to initialize the page tables. I'll have a look tomorrow.Quote:
Originally Posted by nn_step
Now thats something interesting:
http://img5.pcpop.com/ArticleImages/.../000377896.jpg
128 bit L2 to NB bus, witch makes sense, since the IMC is 128 bit (dual channel).
Why isn't K8 like this?
In tandem with lots and lots of custom Cray logic soldered (something similar to Horus IIRC) on the motherboard so that they can scale past 8 sockets well (well is a relative term BTW, compared to Itanium or POWER5 the performance scaling is still pretty lousy).Quote:
Originally Posted by Piotrsama