Cloud infrastructure (ideally AWS) with AI


Keylent
Dice Job Match Score™
👤 Reviewing your profile...
Job Details
Skills
- API
- Amazon DynamoDB
- Management
- Microsoft Certified Professional
- Incident Management
- Java
- Machine Learning (ML)
- Caching
- Collaboration
- Evaluation
- FOCUS
- Amazon Web Services
- Artificial Intelligence
- C++
- Routing
- Python
- Reasoning
- Regulatory Compliance
- Roadmaps
- IaaS
- Open Source
- Orchestration
- Prompt Engineering
- Systems Design
- Amazon S3
- Amazon SageMaker
Summary
Cloud infrastructure (ideally AWS) with AI -
REMOTE - Dallas Preferred - 15+ years of experience
Requirement
Proficiency building and operating on cloud infrastructure (ideally AWS): containerized services (ECS/EKS), serverless (Lambda), data services (S3, DynamoDB, Redshift), orchestration (Step Functions), model serving (SageMaker), and infra-as-code (Terraform/CloudFormation).
Significant software development experience in one or more languages (Python, C/C++, Go, Java); strong hands-on large-scale Python experience preferred.
Three or more years of designing, architecting, testing, and launching production ML systems (deployment/serving, evaluation, monitoring, data pipelines, fine-tuning workflows).
Practical LLM experience: API integration, prompt engineering, fine-tuning/adaptation, RAG, and tool-using agents (vector retrieval, function calling, secure tool execution).
Understanding of commercial and open-source LLMs and their capabilities (e.g., OpenAI, Gemini, Llama, Qwen, Claude).
Focus Area
Description
Build agentic AI systems
Design and implement tool-calling agents combining retrieval, structured reasoning, and secure action execution (function calling, change orchestration, policy enforcement) following the MCP protocol; engineer guardrails for safety, compliance, and least-privilege access.
Productionize LLMs
Build evaluation frameworks for open-source and foundational LLMs; implement retrieval pipelines, prompt synthesis, response validation, and self-correction loops for production operations.
Integrate with runtime ecosystems
Connect agents to observability, incident management, and deployment systems for automated diagnostics, runbook execution, remediation, and post-incident summarization with full traceability.
Collaborate with users
Partner with production and application teams to translate pain points into agentic AI roadmaps; define objective functions linked to reliability, risk reduction, and cost.
Enhance governance
Build validator models, adversarial prompts, and policy checks; enforce deterministic fallbacks, circuit breakers, and rollback strategies; instrument continuous evaluations.
Improve scale & performance
Optimize cost and latency via prompt engineering, context management, caching, model routing, and distillation; leverage batching, streaming, and parallel tool-calls to meet SLOs.
- Dice Id: 10423210A
- Position Id: 9056683
- Posted 2 days ago
Company Info
About Keylent
We have been involved with the industry for over 2 decades and have seen the up's and down's. We have weathered bad times and enjoyed good times by putting our client needs ahead of ours. We continue to do the same thing.
We take great care of our Talent Acquisition and Administrative staff who in turn put in their best work to fulfill our Consultant and Client needs.
Our Clients and our Consultants have a variety of choices and we are thankful that they have chosen Keylent.
Careers


Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs