LLM DevOps/Inference Engineer-
Client: Protégé
Location: REMOTE
Duration 12-18 months
Must Have:
Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD.
Strong AI & LLM
Recent healthcare industry exp (HIPAA, hl7, etc)
LinkedIn Page
Strong communication
Context:
We are looking for a DevOps/Inference Engineer for one of our clients building a healthcare-focused AI benchmark and evaluation suite. The initial target is for clinical prediction tasks including sepsis onset, days-to-death, and lab value trend forecasting, evaluated across multiple frontier and vertical-specific models. Role Summary You will be responsible for the infrastructure the benchmark harness runs on, the selfhosted model serving stack, and the reliability of the platform. This is a role for someone who can provision a GPU cluster in the morning and tune the inferencing model engine in the afternoon.
What You will Own
• Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD.
• Stand up the self-hosted inference track for the long tail of vertical healthcare models. This involves provisioning infrastructure for models serving on GPU compute with sensible batching, quantization where appropriate, autoscaling, and a standard onboarding path so adding new models takes hours, not weeks.
• Build the provider abstraction layer alongside the AI engineers so that APIbased models (OpenAI, Anthropic, Gemini, and the growing list beyond) and self-hosted models present a uniform interface to the harness. Rate limiting, retry and backoff, quota management, request/response logging, and cost attribution per run are your responsibility.
• Make benchmark runs reproducible and cost-optimized with pinned model and container versions, captured configuration, spot and reserved capacity strategy, and idle GPU elimination.
• Build the observability story with throughput, latency, token and GPU-hour cost, failure taxonomy, and per-model dashboards for monitoring.
• Support the surge model that the platform must let a burst of AI engineers land, run experiments, and leave without breaking anything or leaving orphaned resources behind.
• Contribute to Trusted Execution Environment (TEE) architecture. Evaluate AWS Nitro Enclaves and comparable confidential computing approaches for the bring-your-own-data / bring-your-own-model scenario, including attestation- gated key release and the practical constraints of running model inference inside an enclave.
Required Skills
• AWS infrastructure at production scale: EKS or ECS, EC2 GPU instance families (G5/G6, P4d/P5) and their capacity realities, VPC design, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service Quotas.
• Infrastructure as code: Terraform. No console-clicked production resources.
• Model serving and inference optimization: Hands-on experience working with LLMs. Practical command of batching strategy, KV cache behavior, quantization tradeoffs, and multi-GPU sharding.
• Container orchestration and GPU scheduling: ECS/EKS with GPU workloads, node autoscaling, and image build pipelines for CUDA-dependent stacks.
• Reliability and cost engineering. SLOs, alerting, and a demonstrated track record of optimizing cloud spend without cutting capability.
Desirable Skills
• AWS SageMaker endpoints and Bedrock.
• Hands-on experience with Python to contribute directly to the harness and the provider adapter layer.
• Confidential computing fundamentals: Enclaves, remote attestation, sealed key release, and the security boundaries of TEEs.
• Healthcare compliance posture: HIPAA-eligible service selection, BAA scope, audit logging, and the access-control mechanisms for PHI data.
• Security hardening, including image scanning and network egress control for a closed-loop environment.
Nice to Have
• Prior experience hosting medical imaging or multimodal models.
• Nitro Enclaves in production, or comparable TEE work.
• Experience supporting self-service environments, clean tenancy boundaries, and fast credential provisioning.