Requirements
-
Experience: 8+ years in Platform Engineering, DevOps, or Site Reliability Engineering (SRE).
-
Cloud Expertise: Deep proficiency in AWS (IAM, CloudWatch, Bedrock, Lambda).
-
Observability Tools: Proven experience with Dynatrace, Jaeger, or Honeycomb, and distributed tracing standards.
-
AI/LLM Interest: Familiarity with the LLM lifecycle, including prompt execution, token usage, and frameworks like LangChain or AgentCore.
-
Automation: Advanced experience with Terraform and CI/CD pipeline design.
-
Collaboration: Experience working in an Agile environment with integrated tools like Microsoft Teams and Confluence.
User this when submitting candidates:
Also please check if the next candidate has some experience with at least 50% of below items:
• Implementation of Agents on Agentcore runtime
• Implementation of Agentic SDLC in Agentcore
• Understanding of Strands or any other Agentic AI framework like Langraph, Langchain or Crew AI
• Implementation of Bedrock Knowledge Base
• Implementation of Knowledge Graph
• Implementation of MCP servers in Agentcore
• Implementation of Agentcore Gateway
• Implementation of Agentcore Identity
• Implementation of Agentic AI Observability
• Implementation of Agentcore Evaluations
• Implementation of AWS Bedrock
• Implementation of AWS Bedrock Inference Profile
• Implementation of AWS Sagemaker
• AWS Services (Cloud) in General
• Terraform
Deiliverables:
Observability
-
Assess CloudWatch, X-Ray, Bedrock logging, AgentCore traces vs. agentic workflow requirements; produce gap analysis, Setup observability in Dynatrace
-
Design post-deployment validation pipeline for agents & MCP servers (deployment health + tool registration checks)
-
Implement distributed tracing & structured logging: LLM decisions, tool selections, sub-agent calls, MCP interactions
-
Evaluate LangFuse / LiteLLM proxy vs. AWS-native; deliver target-state observability architecture recommendation
Cost Tracking & TCO
-
Extend tagging taxonomy to cover agent runtimes, MCP servers, vector DBs, Bedrock token consumption per namespace
-
Design cost visibility model: aggregate agent, MCP, vector DB, and Bedrock token costs per team/department
-
Build CloudWatch (or equivalent) dashboards for per-team spend; configure AWS Budgets with alerting thresholds
-
Automate cost reports delivered via email / Microsoft Teams; implement anomaly detection rules
Monitoring & Alerting
-
Define P1–P4 alerting rules: deployment failures, runtime errors, tool invocation failures, MCP connectivity issues
-
Integrate alert notifications to Microsoft Teams channels and email; route by resource ownership tags
-
Author runbooks linked to every alert; publish in Confluence for developer self-service resolution
-
Evaluate AWS-native vs. third-party monitoring stack; deliver recommendation aligned to observability architecture
Security & Access Control
-
Assess current IAM + tagging approach for multi-team isolation; identify scalability gaps and risks
-
Evaluate Cedar policy engine (AgentCore) for fine-grained tool access control; document enterprise-scale gaps
-
Design scalable ABAC-based identity model for multi-team isolation without IAM policy sprawl; deliver Terraform modules
|
S.No
|
Qualifying Question
|
Mandatory
|
|
1
|
Is your candidate willing to relocate for this position?
|
Yes
|
|
2
|
Has your candidate answered the questions at the bottom for submission to CWC?
|
Yes
|