Evolutionof CUDA-Enabled GPUsCompute 1.0: basic CUDA compatibilityG80Compute1.1:asynchronousmemorycopiesandatomicglobaloperationsG84.G86.G92.G94.G96.andG98Compute1.2:dramatically improvedmemorycoalescingrules,doubletheregistercountintra-warp voting primitives,atomic sharedmemoryoperationsGT21XCompute1.3:doubleprecisionGT20011
11 11 Evolution of CUDA-Enabled GPUs ◼ Compute 1.0: basic CUDA compatibility ⚫ G80 ◼ Compute 1.1: asynchronous memory copies and atomic global operations ⚫ G84, G86, G92, G94, G96, and G98 ◼ Compute 1.2: dramatically improved memory coalescing rules, double the register count, intra-warp voting primitives, atomic shared memory operations ⚫ GT21X ◼ Compute 1.3: double precision ⚫ GT200
CUDA成功案例236X17X146X18X100XInteractivelonicplacementTranscoding HDSimulation inAstrophysicsNMatlabusingvisualizationofformolecularvideostreamtobody simulationmex file CUDAvolumetric whitedynamicsH.264simulationonmatterfunctionGPUconnectivity149X47X20X24X30XFinancialGLAME@lab:AnUltrasoundHighly optimizedCmatchexactsimulation ofM-scriptAPIformedical imagingobject orientedstring matchingLIBORmodellinearAlgebraforcancermolecularto find similarwithswaptionsoperationsondiagnosticsdynamicsproteinsandGPUgene sequencet12
12 12 CUDA成功案例
提纲从GPGPU到CUDA并行程序组织并行执行模型CUDA基础存储器CUDA程序设计工具新一代FermiGPU13
13 13 提纲◼ 从GPGPU 到CUDA ◼ 并行程序组织 ◼ 并行执行模型 ◼ CUDA基础 ◼ 存储器 ◼ CUDA程序设计工具 ◼ 新一代Fermi GPU
并行性的维度a[1]aJ0]a[n]1维Xb[0]b[1]b[n]y=a+b/ly,a,bvectorsy[0]y[1]y[n]2维P=M×N.//P,M,Nmatrices3维CT or MRI imagingvexMnsaray14
14 14 并行性的维度 ◼ 1维 ⚫ y = a + b //y, a, b vectors ◼ 2维 ⚫ P = M N //P, M, N matrices ◼ 3维 ⚫ CT or MRI imaging a[0] a[1] . a[n] b[0] b[1] . b[n] y[0] y[1] . y[n] + + + = = = =
并行线程组织结构HostDeviceThread:并行的基本单位Grid 1Threadblock:互相合作的线程组KernBlockBlockBlockel1CooperativeThreadArray(CTA)(0, 0)(1, 0)(2, 0)允许彼此同步BlockBlockBlock(0,1)(1,1)(2, 1)通过快速共享内存交换数据以1维、2维或3维组织Grid'2最多包含512个线程Kernel2Grid:一组threadblock以1维或2维组织Block (1,1)共享全局内存Threanrhrea(1.0(2.0)3.0(0,0Kernel:在GPU上执行的核心程序ThrealThreacThreadThrea(1.1(2, 110:.OnekernelonegridThreadThreatThreadrea(1.2)(2, 2)(3,2)(0.24215
15 15 并行线程组织结构 ◼ Thread: 并行的基本单位 ◼ Thread block: 互相合作的线程组 ⚫ Cooperative Thread Array (CTA) ⚫ 允许彼此同步 ⚫ 通过快速共享内存交换数据 ⚫ 以1维、2维或3维组织 ⚫ 最多包含512个线程 ◼ Grid: 一组thread block ⚫ 以1维或2维组织 ⚫ 共享全局内存 ◼ Kernel: 在GPU上执行的核心程序 ⚫ One kernel one grid Host Kern el 1 Kern el 2 Device Grid 1 Block (0, 0) Block (1, 0) Block (2, 0) Block (0, 1) Block (1, 1) Block (2, 1) Grid 2 Block (1, 1) Thread (0, 1) Thread (1, 1) Thread (2, 1) Thread (3, 1) Thread (4, 1) Thread (0, 2) Thread (1, 2) Thread (2, 2) Thread (3, 2) Thread (4, 2) Thread (0, 0) Thread (1, 0) Thread (2, 0) Thread (3, 0) Thread (4, 0)