บทนำสู่การเขียนโปรแกรมด้วยทริตอน: คู่มือปฏิบัติจริง
คู่มือการเรียนรู้ทางวิทยาศาสตร์ที่ครอบคลุม ออกแบบมาเพื่อให้เส้นทางการเรียนรู้แบบเต็มรูปแบบสำหรับ Triton ซึ่งเป็นภาษาและเครื่องคอมพิวเตอร์ที่ใช้ภาษา Python ในการเขียนเคอร์เนล GPU ที่กำหนดเอง หลักสูตรนี้ครอบคลุมโมเดลการเขียนโปรแกรม ไวยากรณ์ของภาษา ลักษณะทางตัวเลข และการปรับประสิทธิภาพ ตั้งแต่การบวกเวกเตอร์พื้นฐาน ไปจนถึงการดำเนินการรวมและแบ่งส่วนที่ใช้ในระบบเรียนรู้เชิงลึกสมัยใหม่
ภาพรวมคอร์สเรียน
📚 สรุปเนื้อหา
บทเรียนทางวิทยาศาสตร์ที่ครอบคลุม ออกแบบมาเพื่อให้เส้นทางการเรียนรู้อย่างครบถ้วนสำหรับ Triton ซึ่งเป็นภาษาและเครื่องมือแปลภาษาที่ใช้พื้นฐานจาก Python สำหรับเขียนเคอร์เนล GPU แบบเฉพาะเจาะจง หลักสูตรนี้ครอบคลุมโมเดลการเขียนโปรแกรม เอกลักษณ์ของภาษา ลักษณะเชิงตัวเลข และการปรับแต่งประสิทธิภาพ โดยย้ายจากแนวคิดพื้นฐานของการบวกเวกเตอร์ไปจนถึงการดำเนินการที่รวมกัน (fused) และแบ่งเป็นช่อง (tiled) ที่ใช้ในระบบเรียนรู้เชิงลึกสมัยใหม่
เชี่ยวชาญศิลปะในการออกแบบเคอร์เนล GPU ประสิทธิภาพสูง จากหลักการพื้นฐาน
ผู้เขียน: EvoClass
คำขอบคุณ: สารบัญของ Triton และ โครงการทริตอนบน GitHub
🎯 วัตถุประสงค์การเรียนรู้
- นิยามความหมายของ Triton และบทบาทของมันในระบบร่วมซอฟต์แวร์การเรียนรู้เชิงลึก
- แยกแยะความแตกต่างระหว่าง Triton กับ CUDA, โค้ดแบบเร่งด่วนของ PyTorch และการเขียนโปรแกรมระดับต่ำสำหรับ GPU
- ระบุงานที่เหมาะสมสำหรับการใช้งานกับ Triton และเข้าใจความสำคัญของการรวมเคอร์เนลและการเกิดจุดขัดข้อง (bottlenecks)
- ติดตั้งสภาพแวดล้อมของ Triton อย่างสะอาดและตรวจสอบชุดซอฟต์แวร์ให้เรียบร้อย
- ประยุกต์ใช้เคอร์เนลสำหรับการคัดลอกเวกเตอร์พื้นฐาน เพื่อยืนยันความถูกต้องของตรรกะสภาพแวดล้อมเทียบกับตรรกะเคอร์เนล
- ระบุและจำแนกประเภทจุดขัดข้องของ GPU เพื่อสนับสนุนการรวมเมธอดใน PyTorch
- นิยามอินสแตนซ์ของโปรแกรม และคำนวณมิติของกริดการเริ่มต้น 1 มิติโดยใช้
cdiv - ดำเนินการคำนวณตำแหน่งหน่วยความจำด้วยการคำนวณชี้ตำแหน่ง (pointer arithmetic) เพื่อจับคู่รหัสเฉพาะของโปรแกรม (
pid) กับตำแหน่งหน่วยความจำ - แยกแยะระหว่างเทนเซอร์ของ PyTorch (ข้อมูลเมตาข้อมูลฝั่งโฮสต์) กับเทนเซอร์ของ Triton (บล็อกระดับคอมไพเลอร์)
- คำนวณการจับคู่ระหว่างรหัสโปรแกรม (
pid) กับตำแหน่งหน่วยความจำเฉพาะโดยใช้tl.arange
บทเรียน 共 10 课时 · 预计 30.0h
บทเรียน
Lesson
This lesson introduces Triton as a bridge between high-level Python productivity and low-level CUDA performance, focusing on its tile-centric programming model. Students will learn how Triton automates complex hardware tasks like memory management and synchronization to enable the development of high-performance custom kernels.
This lesson introduces Triton as a high-performance, block-based programming model that bridges the gap between high-level PyTorch and low-level CUDA. Students will learn to configure a GPU development environment and utilize Triton to overcome memory bottlenecks by fusing operations and optimizing data movement within the GPU's SRAM.
AI023: The Triton Programming Model: Grids and Pointers (Lesson 3) introduces the block-based parallel paradigm, contrasting Triton’s efficient tile-level processing with the overhead of PyTorch’s eager execution. Students will learn to manage memory through pointer arithmetic and coordinate systems, enabling them to optimize GPU performance by minimizing global memory round-trips.
This lesson introduces the Triton programming model, focusing on the transition from scalar CUDA threads to vectorized program instances that operate on data blocks. Students will learn how to utilize program IDs (pid) for SPMD execution, manage memory offsets, and implement masking to handle data boundaries effectively.
This lesson introduces the parallel execution model for GPU programming, focusing on how to implement a vector addition kernel using block-based execution. Students will learn to identify performance bottlenecks—specifically memory-bound versus compute-bound operations—and optimize hardware utilization by managing occupancy and block size.
This lesson explores the Performance Paradox in Triton programming, explaining how fixed GPU launch overheads can make functionally correct code inefficient for small workloads. Students will learn to distinguish between latency-bound and throughput-bound operations, identify the importance of asynchronous execution in benchmarking, and apply strategies like workload batching to minimize the impact of the launch tax.
AI023: Introduction to Triton Programming — Beyond 1D: Why 2D Layout Awareness Matters (Lesson 7) This lesson explores how transitioning from 1D elementwise processing to 2D tiled grids allows Triton kernels to maximize spatial locality and hardware efficiency. Students learn to implement layout-aware kernels by utilizing strides and broadcasting to process data blocks, which is essential for high-performance operations like matrix multiplication.
This lesson explores reduction operations in Triton, focusing on how to collapse multi-dimensional tensors while managing memory layouts and hardware-level data dependencies. Students will also learn to implement numerically stable Softmax functions by addressing common floating-point issues like overflow and underflow.
This lesson explores the transition from memory-bound elementwise operations to compute-bound tiled matrix multiplication (GEMM) in Triton. Students will learn to optimize LLM performance by implementing 2D tiling, managing tensor strides to avoid memory access errors, and applying operator fusion to reduce global memory overhead.
This lesson explores the systematic optimization lifecycle for Triton kernels, focusing on the transition from functional correctness to hardware-aware performance. Students will learn to utilize debugging tools like the Triton interpreter, establish strong performance baselines, and implement autotuning strategies to maximize hardware utilization.