GPU HighLevel ViewnVIDIAPCleGlobalMemoryCPUChipsetStreamingMultiprocessor(SM)A set of CUDA coresGlobal memory
© NVIDIA Corporation 2013 GPU High Level View Streaming Multiprocessor (SM) A set of CUDA cores Global memory
SMXGK110SMle(65,536x32-bit)Control unit4WarpScheduler8instructiondispatcherExecutionunit4192single-precisionCUDACoresBE64double-precisionCUDACoresGF32SFU,32LD/ST5SMemoryRegisters:64K32-bitoBFCacheBFL1+sharedmemory(64KB)BETextureAKBLtCachConstantknBnnd-OntTexTexTexTexTexTexTexTexTexTexTexTexTexTexTexTexSMAinn2013
© NVIDIA Corporation 2013 GK110 SM Control unit 4 Warp Scheduler 8 instruction dispatcher Execution unit 192 single-precision CUDA Cores 64 double-precision CUDA Cores 32 SFU, 32 LD/ST Memory Registers: 64K 32-bit Cache L1+shared memory (64 KB) Texture Constant
Kepler/Fermi Memory HierarchyDVIDIA3levels,very similartoCPURegisterSpillstolocalmemoryCachesSharedmemoryL1cacheL2cacheConstantcacheTexturecacheGlobalmemory2053
© NVIDIA Corporation 2013 Kepler/Fermi Memory Hierarchy 3 levels, very similar to CPU Register Spills to local memory Caches Shared memory L1 cache L2 cache Constant cache Texture cache Global memory
Kepler/Fermi Memory HierarchyDVIDIASM-1SM-0SM-NRegistersRegistersRegisters全业T金TTTT金L1&L1&L1&ccCTEXTEXTEXSMEMSMEMSMEM不不7L2Global Memory
© NVIDIA Corporation 2013 Kepler/Fermi Memory Hierarchy L2 Global Memory Registers C SM-0 L1& SMEM TEX Registers C SM-1 L1& SMEM TEX Registers C SM-N L1& SMEM TEX
BasicConceptsDVIDIATransferdataCPU MemoryGPU MemoryPCIBusGPUCPUOffloadcomputationGPU computing is all about 2 things:TransferdatabetweenCPU-GPU:Do parallel computing on GPU
© NVIDIA Corporation 2013 Basic Concepts PCI Bus Transfer data Offload computation GPU GPU Memory CPU CPU Memory GPU computing is all about 2 things: • Transfer data between CPU-GPU • Do parallel computing on GPU