บทนำสู่การเขียนโปรแกรมด้วย ROCm และ HIP: บทเรียนปฏิบัติจริง
คู่มือปฏิบัติที่ทันสมัยเกี่ยวกับการเขียนโปรแกรมสำหรับ GPU ของ AMD โดยใช้ ROCm และ HIP ครอบคลุมชุดซอฟต์แวร์ทั้งหมด ตั้งแต่การติดตั้ง กระบวนการสร้าง การเขียนโค้ดแกน (kernel programming) การจัดการหน่วยความจำ การวิศวกรรมประสิทธิภาพ การใช้งานไลบรารี การย้ายโค้ดจาก CUDA และแนวทางการตรวจสอบข้อผิดพลาดในสภาพแวดล้อมผลิตภัณฑ์
ภาพรวมคอร์สเรียน
📚 สรุปเนื้อหา
คู่มือปฏิบัติที่ทันสมัยเกี่ยวกับการเขียนโปรแกรม GPU ของ AMD โดยใช้ ROCm และ HIP ครอบคลุมชุดซอฟต์แวร์ทั้งหมด ตั้งแต่การติดตั้ง การจัดทำโครงการ (build workflows) การเขียนโค้ดเคอร์เนล การจัดการหน่วยความจำ การวิศวกรรมประสิทธิภาพ การใช้งานไลบรารี การย้ายจาก CUDA ไปยัง HIP และแนวทางการตรวจสอบข้อผิดพลาดในสภาพแวดล้อมผลิตภัณฑ์
เรียนรู้การเขียนโปรแกรม GPU ของ AMD และการย้ายจาก CUDA ไปยัง HIP อย่างมืออาชีพด้วยบทวิเคราะห์เชิงเทคนิค
ผู้เขียน: EvoClass
คำขอบคุณ: เอกสารทางเทคนิคของ AMD เกี่ยวกับ ROCm และ HIP ซึ่งรวมถึงโครงการต่าง ๆ เช่น ROCm, HIP และ ROCm LLVM
🎯 เป้าหมายการเรียนรู้
- นิยาม HIP และบทบาทของมันในระบบนิเวศของ ROCm ด้วยประโยคเดียวที่กระชับ
- แยกแยะความแตกต่างระหว่าง ROCm (แพลตฟอร์ม), HIP (อินเทอร์เฟซ), และไลบรารีของ ROCm (ส่วนประกอบหลัก)
- ระบุระดับโครงสร้างของสถาปัตยกรรม ROCm จากฮาร์ดแวร์ถึงกรอบงานแอปพลิเคชัน
- อธิบายความสัมพันธ์ระหว่าง SDK ของ HIP กับแพลตฟอร์ม ROCm ในระบบปฏิบัติการต่าง ๆ
- ดำเนินการตามขั้นตอนการติดตั้งอย่างเป็นระบบ รวมถึงการตรวจสอบแมทริกซ์รองรับและตั้งค่าเส้นทางหลังติดตั้ง
- คอมไพล์และรันโปรแกรมตรวจสอบขั้นพื้นฐานเพื่อแก้ไขปัญหาทั่วไปเกี่ยวกับไดรเวอร์และสิทธิ์การเข้าถึงสภาพแวดล้อม
- เข้าใจว่ากลยุทธ์การจัดทำโครงการที่แข็งแรงมีความสำคัญอย่างไรในการประสานความเข้ากันได้ของรหัสต้นฉบับกับประสิทธิภาพเฉพาะสถาปัตยกรรม
- ประยุกต์ใช้
hipLaunchKernelGGLเพื่อเรียกใช้เคอร์เนลแบบสามารถย้ายได้แทนการใช้รูปแบบสามเหลี่ยมสามชั้นของ CUDA - ตั้งค่าโปรเจกต์ CMake ระดับองค์กรที่กำหนดเป้าหมายไปยังสถาปัตยกรรม ROCm ที่เฉพาะเจาะจง และจัดการความพึ่งพาไลบรารีภายนอก
- วิเคราะห์โครงสร้างของเคอร์เนล HIP และนำไปใช้สูตรพื้นฐานในการจัดลำดับเธรด
บทเรียน 共 10 课时 · 预计 30.0h
บทเรียน
Lesson
This lesson introduces the ROCm platform and the HIP programming model as a bridge for porting CUDA applications to AMD hardware. Students will learn how to use automated tools like hipify to migrate code while understanding the importance of architecture-aware tuning to achieve optimal performance.
This lesson covers the essential steps for installing and configuring the ROCm software stack, including dependency management, environment variable setup, and user permission requirements. Students will learn how to verify their system environment and ensure successful hardware-software communication through diagnostic tools and proper configuration.
This lesson explores the distinction between source portability and binary performance in the ROCm ecosystem, emphasizing that while HIP code is functionally portable, achieving peak throughput requires architecture-specific compilation. Students will learn to utilize the hipcc toolchain and CMake to manage build configurations that optimize code for specific hardware instruction sets.
This lesson introduces the HIP programming model, focusing on the transition from sequential CPU iteration to spatial GPU parallelism using the Parallel Pivot approach. Students will learn to map independent data tasks to thread grids, manage memory, and implement kernel execution with proper boundary checks and error handling.
AI024: Memory Management and Data Patterns (Lesson 5) explores the memory-centric nature of GPU performance, focusing on the Roofline Model and the critical importance of minimizing data movement between host and device. Students will learn to distinguish between memory-bound and compute-bound kernels while mastering strategies to optimize data residence and bandwidth utilization.
This lesson explores the transition from synchronous to asynchronous GPU execution, focusing on how to use HIP streams to decouple CPU and GPU tasks. Students will learn to optimize performance by implementing non-blocking memory transfers and kernel launches to maximize hardware utilization and eliminate execution bottlenecks.
This lesson introduces a systematic, data-driven approach to performance engineering on AMD GPUs, emphasizing the use of tools like rocprofv3 to identify bottlenecks rather than relying on intuition. Students will learn to follow a six-step scientific workflow to optimize memory access, instruction throughput, and hardware utilization while avoiding common performance "superstitions."
This lesson introduces the Library-First Engineering Principle, which emphasizes using optimized ROCm libraries like rocBLAS and rocFFT to reduce technical debt and ensure hardware portability. Students will learn to prioritize these vendor-tuned solutions over custom kernel development to achieve better performance and easier maintenance across evolving GPU architectures.
AI024: Porting CUDA Applications to HIP (Lesson 9) covers the systematic, incremental migration of CUDA code to the HIP platform using tools like HIPIFY-Clang and HIPIFY-Perl. Students will learn to distinguish between mechanical API translations and architectural optimizations, such as adjusting for warp-size differences, to ensure functional and performance parity on AMD ROCm hardware.
This lesson explores the GPU Developer’s Creed, which prioritizes functional correctness and architectural isolation over raw performance when working with ROCm and HIP. Students will learn to implement systematic debugging, testing, and CI/CD practices to ensure stable, reproducible, and accurate GPU kernel deployments.