GPU Architecture, CUDA & Hybrid Program
GPU Architecture, CUDA & Hybrid Programming
M. Baqir Kazmi 24642
Farhan Ahmad 26357
Imran Ali 30587
Ahmad Raza 26343
Dawood 26852
BSCS 7th Semester
Parallel & Distributed Computing
GPU Architecture
• GPU executes thousands of threads in parallel
• Designed for high performance computing
• Optimized for high throughput instead of low latency
Main Components of GPU
• Streaming Multiprocessors (SMs)
• CUDA Cores
• Memory Hierarchy (Registers, Shared, Global Memory)
GPU vs CPU
• CPU has few powerful cores
• GPU has thousands of small cores
• GPU is best for data-parallel tasks
CUDA Programming
• CUDA allows programmers to use GPU
• Enables massive parallel execution
• Reduces execution time
CUDA Programming Model
• Kernel: Function executed on GPU
• Threads: Smallest execution units
• Blocks: Group of threads
• Grid: Collection of blocks
Hybrid Programming
• Combination of MPI and OpenMP
• MPI for communication between nodes
• OpenMP for parallelism inside a node
GPU + Hybrid Programming
• MPI handles data distribution
• OpenMP manages CPU threads
• CUDA accelerates computation on GPU
Advantages
• High performance
• Better scalability
• Efficient utilization of hardware
Conclusion
GPU architecture, CUDA programming, and hybrid models
are essential for modern parallel computing.