Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Chinese AI GPU Ecosystem Overview
- Comparative analysis of Huawei Ascend, Biren, and Cambricon MLU
- Contrasting CUDA with CANN, Biren SDK, and BANGPy paradigms
- Current industry trends and vendor ecosystem dynamics
Pre-Migration Preparation
- Auditing your existing CUDA codebase
- Defining target platforms and required SDK versions
- Installing toolchains and configuring the development environment
Code Translation Methodologies
- Translating CUDA memory access patterns and kernel logic
- Aligning compute grid and thread models
- Evaluating automated versus manual translation approaches
Implementation Specifics by Platform
- Leveraging Huawei CANN operators and custom kernels
- Utilizing the Biren SDK conversion pipeline
- Reconstructing models using BANGPy (Cambricon)
Cross-Platform Testing and Tuning
- Profiling execution metrics on each target platform
- Optimizing memory usage and comparing parallel execution strategies
- Tracking performance indicators and iterating on improvements
Oversight of Mixed GPU Environments
- Managing hybrid deployments across multiple architectures
- Implementing fallback mechanisms and device detection logic
- Introducing abstraction layers to enhance code maintainability
Case Studies and Industry Best Practices
- Transferring vision and NLP models to Ascend or Cambricon
- Adapting inference pipelines for Biren clusters
- Resolving version inconsistencies and API discrepancies
Conclusion and Forward-Looking Steps
Requirements
- Practical experience in programming with CUDA or GPU-based applications
- Comprehensive understanding of GPU memory models and compute kernels
- Proficiency in AI model deployment or acceleration workflows
Target Audience
- GPU Programmers
- System Architects
- Porting Specialists
21 Hours