คู่มือการเขียนโปรแกรม CUDA
แหล่งข้อมูลทางการและครอบคลุมสำหรับนักพัฒนาที่ต้องการเรียนรู้โมเดลการเขียนโปรแกรม CUDA และวิธีเขียนโค้ดประสิทธิภาพสูงที่ทำงานบนหน่วยประมวลผลกราฟิกของ NVIDIA คู่มือนี้ครอบคลุมสถาปัตยกรรมแพลตฟอร์ม ชุดคำสั่งการเขียนโปรแกรม คุณสมบัติฮาร์ดแวร์ขั้นสูง และข้อมูลเทคนิคต่างๆ
ภาพรวมคอร์สเรียน
📚 สรุปเนื้อหา
แหล่งข้อมูลทางการและครบถ้วนสำหรับนักพัฒนาที่ต้องการเรียนรู้โมเดลโปรแกรม CUDA และวิธีเขียนโค้ดประสิทธิภาพสูงที่ทำงานบนหน่วยประมวลผลกราฟิกของ NVIDIA คู่มือนี้ครอบคลุมสถาปัตยกรรมแพลตฟอร์ม หน้าที่การใช้งานของโปรแกรม การใช้คุณสมบัติฮาร์ดแวร์ขั้นสูง และรายละเอียดทางเทคนิค
เชี่ยวชาญศิลปะของการคำนวณแบบขนานด้วยคู่มือมาตรฐานในอุตสาหกรรมสำหรับ NVIDIA CUDA
ผู้เขียน: บริษัท NVIDIA Corporation
คำขอบคุณ: สิทธิ์การใช้งาน © 2007-2024 บริษัท NVIDIA Corporation และบริษัทในเครือ สงวนลิขสิทธิ์ทั้งหมด
🎯 เป้าหมายการเรียนรู้
- กำหนดบทบาทของโฮสต์ (โปรเซสเซอร์) และอุปกรณ์ (กราฟิกไพเพอร์) ภายในระบบไฮบริด
- อธิบายโมเดลโปรแกรมแบบ SIMT และโครงสร้างระดับชั้นของเธรด บล็อก และกริด
- แยกความแตกต่างระหว่างรหัส PTX (Parallel Thread Execution) และรหัสไบนารี (cubins) และอธิบายว่าการคอมไพล์แบบจุดเวลา (JIT) ช่วยให้เกิดความเข้ากันได้ได้อย่างไร
- พัฒนาและคอมไพล์เคอร์เนล CUDA: เขียนฟังก์ชัน global กำหนดการดำเนินการด้วยสัญลักษณ์สามชั้น และจัดการกระบวนการคอมไพล์โดยใช้ NVCC
- ปรับปรุงการจัดการหน่วยความจำและการเคลื่อนย้ายข้อมูล: แยกแยะโมเดลหน่วยความจำแบบรวม (Unified), แบบระบุชัด (Explicit), และแบบแมป (Mapped) และนำหน่วยความจำโฮสต์ที่ตรึงหน้า (page-locked) มาใช้เพื่อการส่งข้อมูลอย่างมีประสิทธิภาพ
- จัดการการดำเนินการแบบขนาน: ใช้สตรีม, อีเวนต์ และกลุ่มการทำงานร่วมกัน (Cooperative Groups) เพื่อจัดการงานแบบไม่ซิงโครนัส และทำให้การดำเนินการระหว่าง CPU กับ GPU ซิงโครนัสกันได้
- ดำเนินการคำนวณตัวชี้ที่ซับซ้อนและระบุจุดที่เป็นปัญหาในสถาปัตยกรรม (แบบ von Neumann ตรงกับแบบ Harvard)
- ใช้รูปแบบการดำเนินงานขั้นสูงของ CUDA เช่น การเรียกเคอร์เนลตามเงื่อนไขที่ควบคุมด้วยโปรแกรม และการถ่ายโอนข้อมูลแบบบัทช์ที่หลากหลาย
- ใช้คุณสมบัติเฉพาะฮาร์ดแวร์ เช่น ขอบเขตของเธรด (Thread Scopes), ตัวแทนแบบไม่ซิงโครนัส (Asynchronous Proxies), และพายเปล (Pipelines) เพื่อเพิ่มความสามารถในการทำงานพร้อมกันสูงสุด
- ตั้งค่าและปรับแต่งประสิทธิภาพของหน่วยความจำแบบรวม (Unified Memory) โดยใช้การดึงข้อมูลล่วงหน้า (prefetching), คำแนะนำการใช้งาน (usage hints), และการจัดการขนาดหน้า (page size management)
บทเรียน 共 5 课时 · 预计 30.0h
บทเรียน
Lesson
This lesson introduces the fundamental shift from latency-optimized CPU architectures to throughput-oriented GPU computing. Students will learn to distinguish between these processing models and understand how the CUDA programming platform enables massive parallel execution for data-intensive tasks.
This lesson introduces the fundamentals of CUDA kernel development, focusing on the SIMT execution model and the use of the __global__ specifier to launch parallel functions on GPU Streaming Multiprocessors. Students will learn how to manage asynchronous kernel execution, handle device memory, and structure code to ensure effective hardware utilization.
This lesson explores the fundamental differences between von Neumann and Harvard architectures, focusing on how memory access pathways impact computational performance. Students will learn to identify the von Neumann bottleneck, understand the benefits of split-cache Harvard designs, and analyze how modern systems utilize a Modified Harvard Architecture to balance throughput with programming flexibility.
AI021: Optimization, Graphs, and Hardware Accelerators (Lesson 4) explores the shift from CPU-bottlenecked stream execution to GPU-autonomous workflows. Students will learn to utilize modern primitives like CUDA Graphs, lazy loading, and asynchronous memory prefetching to minimize host-side overhead and maximize hardware efficiency.
This lesson explores the technical reference and language extensions in CUDA, focusing on the relationship between virtual architectures (PTX) and real hardware (SASS). Students will learn to manage compute capabilities, utilize architecture-specific macros, and navigate language constraints to ensure code portability and performance.