Get in Touch

Course Outline

Introduction to Multimodal LLMs in Vertex AI

  • Overview of multimodal capabilities available in Vertex AI
  • Gemini models and the modalities they support
  • Enterprise and research application scenarios

Setting Up the Development Environment

  • Configuring Vertex AI specifically for multimodal workflows
  • Managing datasets across different modalities
  • Hands-on lab: setting up the environment and preparing datasets

Long Context Windows and Advanced Reasoning

  • Understanding the mechanics of long-context workflows
  • Applications in strategic planning and decision-making
  • Hands-on lab: executing long-context analysis

Cross-Modal Workflow Design

  • Integrating text, audio, and image analysis
  • Chaining multimodal steps within pipelines
  • Hands-on lab: architecting a multimodal pipeline

Working with Gemini API Parameters

  • Configuring inputs and outputs for multimodal tasks
  • Optimizing inference speed and resource efficiency
  • Hands-on lab: adjusting Gemini API parameters

Advanced Applications and Integrations

  • Developing interactive multimodal agents and assistants
  • Connecting with external APIs and third-party tools
  • Hands-on lab: building a complete multimodal application

Evaluation and Iteration

  • Testing the performance of multimodal models
  • Key metrics for accuracy, alignment, and drift detection
  • Hands-on lab: assessing multimodal workflow effectiveness

Summary and Next Steps

Requirements

  • Strong proficiency in Python programming
  • Practical experience in developing machine learning models
  • Working knowledge of multimodal data types, including text, audio, and images

Target Audience

  • AI researchers
  • Senior developers
  • Machine Learning Scientists
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories