GPU ProgrammingBasicsDVIDIA
© NVIDIA Corporation 2013 GPU Programming Basics
How To Get StartnVIDIACUDA C/C++:download CUDA drivers&compilers&samples(AlllnOnePackage)freefrom:http:lldeveloper.nvidia.com/cuda/cuda-downloadsCUDAFortran:PGlOpenACC: PGl, CAPS, Cray
© NVIDIA Corporation 2013 How To Get Start CUDA C/C++: download CUDA drivers & compilers & samples (All In One Package ) free from: http://developer.nvidia.com/cuda/cuda-downloads CUDA Fortran: PGI OpenACC: PGI, CAPS, Cray
CUDA Programming BasicsnVIDIAHelloWorldBasicsyntax,compile&runGPU memory managementMalloc/freememcpyWritingparallelkernelsThreads&blockMemory hierarchy
© NVIDIA Corporation 2013 CUDA Programming Basics Hello World Basic syntax, compile & run GPU memory management Malloc/free memcpy Writing parallel kernels Threads & block Memory hierarchy
Heterogeneous ComputingNVIDIACProgramSequential ExecutionSerial codeHostWExecutesonbothCPU&GPUDeviceSimilartoOpenMP'sParallel kernelfork-joinpatternKernel0<<<>>>()AcceleratedkernelsCUDA:simpleextensionsSerial codetoC/C++DeviceParallel kernelKernel1<<<>>>()
© NVIDIA Corporation 2013 Heterogeneous Computing Executes on both CPU & GPU Similar to OpenMP’s fork-join pattern Accelerated kernels CUDA: simple extensions to C/C++ Device Grid 0 Block (0, 1) Block (1, 1) Block (2, 1) Block (0, 0) Block (1, 0) Block (2, 0) Host C Program Sequential Execution Serial code Parallel kernel Kernel0<<<>>>() Serial code Parallel kernel Kernel1<<<>>>() Host Device Grid 1 Block (1, 1) Block (1, 0) Block (1, 2) Block (0, 1) Block (0, 0) Block (0, 2)