Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open Source Tools
- Key concepts and advantages of AIOps.
- The role of Prometheus and Grafana within the observability stack.
- The place of ML in AIOps: balancing predictive and reactive analytics.
Setting Up Prometheus and Grafana
- Installation and configuration of Prometheus for time series data collection.
- Building Grafana dashboards leveraging real-time metrics.
- Deep dive into exporters, relabeling, and service discovery mechanisms.
Data Preprocessing for ML
- Techniques for extracting and transforming Prometheus metrics.
- Preparing datasets suitable for anomaly detection and forecasting tasks.
- Utilizing Grafana’s transformation capabilities or Python-based data pipelines.
Applying Machine Learning for Anomaly Detection
- Introduction to basic ML models for outlier detection (e.g., Isolation Forest, One-Class SVM).
- Training and evaluating models on time series data.
- Visualizing detected anomalies within Grafana dashboards.
Forecasting Metrics with ML
- Developing simple forecasting models (introduction to ARIMA, Prophet, and LSTM).
- Predicting system load and resource consumption trends.
- Leveraging predictions to enable early alerting and informed scaling decisions.
Integrating ML with Alerting and Automation
- Defining alert rules based on ML outputs or dynamic thresholds.
- Managing Alertmanager and notification routing strategies.
- Automating scripts and workflows triggered by anomaly detection.
Scaling and Operationalizing AIOps
- Integrating with external observability tools (e.g., ELK stack, Moogsoft, Dynatrace).
- Operationalizing ML models within observability pipelines.
- Best practices for scaling AIOps initiatives.
Summary and Next Steps
Requirements
- A solid understanding of system monitoring and observability concepts.
- Practical experience working with either Grafana or Prometheus.
- Familiarity with Python and fundamental machine learning principles.
Target Audience
- Observability engineers.
- Infrastructure and DevOps teams.
- Monitoring platform architects and Site Reliability Engineers (SREs).