Panduan Pemrograman CUDA
Sumber resmi dan komprehensif bagi para pengembang untuk belajar model pemrograman CUDA dan bagaimana menulis kode dengan kinerja tinggi yang berjalan di GPU NVIDIA. Panduan ini mencakup arsitektur platform, antarmuka pemrograman, fitur perangkat keras lanjutan, serta spesifikasi teknis.
Gambaran Umum Kursus
📚 Ringkasan Konten
Sumber resmi dan komprehensif bagi para pengembang untuk belajar model pemrograman CUDA dan cara menulis kode berkinerja tinggi yang dijalankan pada GPU NVIDIA. Panduan ini mencakup arsitektur platform, antarmuka pemrograman, fitur perangkat keras lanjutan, serta spesifikasi teknis.
Kuasai seni komputasi paralel dengan panduan standar industri untuk NVIDIA CUDA.
Penulis: NVIDIA Corporation
Ucapan Terima Kasih: Hak Cipta © 2007–2024 NVIDIA Corporation & afiliasi. Seluruh hak dilindungi.
🎯 Tujuan Pembelajaran
- Mendefinisikan peran host (CPU) dan perangkat (GPU) dalam sistem heterogen.
- Menjelaskan model pemrograman SIMT dan organisasi hierarkis dari thread, blok, dan grid.
- Membedakan antara PTX (Parallel Thread Execution) dan kode biner (cubins), serta menjelaskan bagaimana kompilasi Just-in-Time (JIT) memfasilitasi kompatibilitas.
- Mengembangkan dan Mengompilasi Kernel CUDA: Tulis fungsi global, konfigurasikan eksekusi dengan notasi tiga tanda panah, dan kelola alur kerja kompilasi NVCC.
- Mengoptimalkan Memori dan Perpindahan Data: Bedakan antara model memori Unified, Eksplisit, dan Mapped, serta implementasikan memori host yang terkunci halaman untuk transfer yang efisien.
- Mengelola Eksekusi Paralel: Gunakan CUDA Streams, Events, dan Cooperative Groups untuk mengelola tugas asinkron dan sinkronisasi operasi CPU-GPU.
- Melakukan aritmetika pointer kompleks dan mengidentifikasi bottleneck arsitektur (von Neumann vs. Harvard).
- Menerapkan pola eksekusi CUDA lanjutan, termasuk Peluncuran Kernel Bergantung Secara Programatik dan Transfer Memori Batch Heterogen.
- Menggunakan fitur khusus perangkat keras seperti Thread Scopes, Proxy Asinkron, dan Pipa untuk memaksimalkan konkurensi.
- Mengonfigurasi dan menyesuaikan performa Unified Memory menggunakan prefetching, petunjuk penggunaan, serta manajemen ukuran halaman.
Pelajaran 共 5 课时 · 预计 30.0h
Pelajaran
Lesson
This lesson introduces the fundamental shift from latency-optimized CPU architectures to throughput-oriented GPU computing. Students will learn to distinguish between these processing models and understand how the CUDA programming platform enables massive parallel execution for data-intensive tasks.
This lesson introduces the fundamentals of CUDA kernel development, focusing on the SIMT execution model and the use of the __global__ specifier to launch parallel functions on GPU Streaming Multiprocessors. Students will learn how to manage asynchronous kernel execution, handle device memory, and structure code to ensure effective hardware utilization.
This lesson explores the fundamental differences between von Neumann and Harvard architectures, focusing on how memory access pathways impact computational performance. Students will learn to identify the von Neumann bottleneck, understand the benefits of split-cache Harvard designs, and analyze how modern systems utilize a Modified Harvard Architecture to balance throughput with programming flexibility.
AI021: Optimization, Graphs, and Hardware Accelerators (Lesson 4) explores the shift from CPU-bottlenecked stream execution to GPU-autonomous workflows. Students will learn to utilize modern primitives like CUDA Graphs, lazy loading, and asynchronous memory prefetching to minimize host-side overhead and maximize hardware efficiency.
This lesson explores the technical reference and language extensions in CUDA, focusing on the relationship between virtual architectures (PTX) and real hardware (SASS). Students will learn to manage compute capabilities, utilize architecture-specific macros, and navigate language constraints to ensure code portability and performance.