Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Core Performance Concepts and Key Metrics
- Analyzing latency, throughput, power consumption, and resource utilization.
- Distinguishing between system-level and model-level bottlenecks.
- Profiling techniques for inference versus training scenarios.
Profiling Strategies on Huawei Ascend
- Leveraging CANN Profiler and MindInsight.
- Diagnosing kernel and operator performance issues.
- Understanding offload patterns and memory mapping.
Performance Profiling on Biren GPUs
- Utilizing Biren SDK monitoring features for performance insights.
- Exploring kernel fusion, memory alignment, and execution queue management.
- Conducting power and temperature-aware profiling.
Profiling on Cambricon MLUs
- Employing BANGPy and Neuware performance tools.
- Gaining kernel-level visibility and interpreting logs effectively.
- Integrating the MLU profiler with various deployment frameworks.
Graph and Model-Level Optimization Techniques
- Applying graph pruning and quantization strategies.
- Restructuring computational graphs via operator fusion.
- Standardizing input sizes and tuning batch parameters.
Memory and Kernel Optimization
- Improving memory layout efficiency and data reuse.
- Managing buffers efficiently across different chipsets.
- Applying platform-specific kernel-level tuning methods.
Best Practices for Cross-Platform Deployment
- Achieving performance portability through abstraction strategies.
- Developing shared tuning pipelines for multi-chip environments.
- Case study: Tuning an object detection model across Ascend, Biren, and MLU platforms.
Summary and Future Recommendations
Requirements
- Practical experience with AI model training or deployment pipelines.
- A solid grasp of GPU/MLU computing principles and model optimization techniques.
- Familiarity with fundamental performance profiling tools and key metrics.
Target Audience
- Performance Engineers.
- Machine Learning Infrastructure Teams.
- AI System Architects.
21 Hours