Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for Operations
- The evolution of IT automation: from static runbooks to reasoning agents
- Agent structure: reasoning loops, tool utilization, memory, and planning
- Determining when to automate versus retaining human oversight
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Creating your first operational agent: monitoring queries, diagnostics, and proposals
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log querying: integration with Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated incident triage: severity classification and routing
- Generating root cause hypotheses and gathering evidence
- Automated remediation actions: restarts, scaling, rollbacks, and failovers
- Developing incident runbook agents with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Action classification: read-only, low-risk, high-risk, and destructive
- Approval gates and escalation policies for critical operations
- Guardrail patterns: action allowlists, blast radius limitations, and rollback assurances
- Audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialized agents: triage, diagnosis, and remediation agents
- Inter-agent communication and shared context management
- Resolving conflicts when agents suggest contradictory actions
- End-to-end major incident simulation using multi-agent response
Observability and Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Evaluating decision quality: precision, recall, and time-to-resolution
- Feedback loops: learning from operator overrides and final outcomes
- Cost tracking and token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: APIs, webhooks, and scheduled tasks
- Phased autonomy rollout: transitioning from shadow mode to full auto-remediation
- Agent failure management: protocols for when the agent itself encounters issues
- Building the business case and measuring ROI for autonomous operations
Requirements
- Background in IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST APIs.
- Foundational knowledge of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers developing self-healing infrastructure.
- IT operations leaders assessing agentic AI for incident management.
14 Hours