GK110.SMSMXControlunitDisnateReaisto(65.536x32-bl4WarpScheduler8instructiondispatcherRExecution.unit乐乐192single-precisionCUDACores64double-precisionCUDACores小32SFU,32LD/STMemory乐Registers:64K32-bit乐5CacheL1+sharedmemory(64KB)TextureConstantTexToxToxTexTexTexTexTeTaTHTexTetTerT
GK110 SM Control unit 4 Warp Scheduler 8 instruction dispatcher Execution unit 192 single-precision CUDA Cores 64 double-precision CUDA Cores 32 SFU, 32 LD/ST Memory Registers: 64K 32-bit Cache L1+shared memory (64 KB) Texture Constant
GPU and Programming ModelGPUSoftwareThreadsareexecuted byCUDAcoresCUDACoreThreadThreadblocksareexecutedonmultiprocessorsThreadblocks donotmigrateSeveral concurrent thread blocks canreside ononeThreadmultiprocessor -limited bymultiprocessorresourcesMultiprocessorBlockAkernelislaunchedasagridofthreadblocksUpto16kernelscanexecuteonadeviceatonetimeGridDevice
GPU and Programming Model Software GPU Threads are executed by CUDA cores Thread CUDA Core Thread Block Multiprocessor Thread blocks are executed on multiprocessors Thread blocks do not migrate Several concurrent thread blocks can reside on one multiprocessor - limited by multiprocessor resources . Grid Device A kernel is launched as a grid of thread blocks Up to 16 kernels can execute on a device at one time
WarpWarpissuccessive32threadsinablockBlock.oE.g.blockDim=160Automaticallydividedto5warpsbyGPUE.g.blockDim=161IftheblockDimisnottheMultipleof32TherestofthreadwilloccupyonemorewarpBlocko32Threads32Threads-32ThreadsBlockWarps
Warp Warp is successive 32 threads in a block E.g. blockDim = 160 Automatically divided to 5 warps by GPU E.g. blockDim = 161 If the blockDim is not the Multiple of 32 The rest of thread will occupy one more warp Block 0 Warp 3 (96~127) Warp 4 (128~159) Warp 0 (0~31) Warp1 (32~63) Warp2 (64~95) Block 0 Warp 3 (96~127) Warp 4 (128~159) Warp 5 (160) Warp 0 (0~31) Warp1 (32~63) Warp2 (64~95) Block 32 Threads 32 Threads 32 Threads . Warps =
WarpSiMD:SameInstructionMultiDataThethreadsinthesameWarpScheduler0Warp Scheduler1warpalwaysexecutingthesameinstructionwarInstructionswill beissuedtowaroperationunitsbywarp
Warp SIMD: Same Instruction Multi Data The threads in the same warp always executing the same instruction Instructions will be issued to operation units by warp warp 8 instruction 11 Warp Scheduler 0 warp 2 instruction 42 warp 14 instruction 95 warp 8 instruction 12 . . . warp 14 instruction 96 warp 2 instruction 43 warp 9 instruction 11 Warp Scheduler 1 warp 3 instruction 33 warp 15 instruction 95 warp 9 instruction 12 . . . warp 3 instruction 34 warp 15 instruction 96
WarpLatencyiscaused bythedependencybetweentheneighborinstructionsWarpScheduler0Warp Scheduler1inthesamewarpInthewaitingtime,otherinstructiohswarpfromotherwarpscanbeexecutedwarp2warContextswitchingisfreeAlotofwarpscanhidememorylatency
Warp Latency is caused by the dependency between the neighbor instructions in the same warp In the waiting time, other instructions from other warps can be executed Context switching is free A lot of warps can hide memory latency warp 8 instruction 11 Warp Scheduler 0 warp 2 instruction 42 warp 14 instruction 95 warp 8 instruction 12 . . . warp 14 instruction 96 warp 2 instruction 43 warp 9 instruction 11 Warp Scheduler 1 warp 3 instruction 33 warp 15 instruction 95 warp 9 instruction 12 . . . warp 3 instruction 34 warp 15 instruction 96