Get in Touch

Course Outline

Overview of Biren GPU Architecture

  • Biren platform introduction and key use cases.
  • Hardware composition: cores, memory structures, and compute clusters.
  • Comparative analysis with NVIDIA and AMD GPUs.

Configuring the Biren Programming Environment

  • Installation of the Biren SDK and runtime components.
  • Insights into the toolchain and compiler architecture.
  • Standard project structures and build workflows.

GPU Programming via the Biren Stack

  • Conceptualization of thread and block models.
  • Memory management strategies and data transfer mechanisms.
  • Kernel design and execution patterns.

Migration from CUDA to Biren

  • Methodologies for translating CUDA code.
  • Mapping common APIs and necessary adaptations.
  • Hands-on labs for code conversion and practice.

Debugging and Profiling Processes

  • Leveraging Biren’s debugger and profiler tools.
  • Techniques for identifying system bottlenecks.
  • Analyzing memory access patterns and applying optimizations.

Advanced Optimization Strategies

  • Thread scheduling and instruction pipelining techniques.
  • Utilizing loop unrolling and shared memory for efficiency.
  • Advanced kernel tuning to maximize throughput.

Case Studies and Practical Applications

  • Training models using Biren accelerators.
  • Porting and profiling vision or NLP models.
  • Performance benchmarking against CUDA/NVIDIA solutions.

Conclusion and Future Directions

Requirements

  • Solid understanding of GPU architecture and parallel processing principles.
  • Prior experience with CUDA, OpenCL, or comparable GPU programming environments.
  • Proficiency with deep learning frameworks such as PyTorch or TensorFlow.

Target Audience

  • HPC developers.
  • AI infrastructure engineers.
  • Performance optimization specialists.
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories