Job Title: Google Cloud Platform DevOps Engineer, AI Platform Delivery
Location: 100% Remote
Work Authorization: USC
Duration: 12+ months
Position Summary
This role sits within a regulated pharmaceutical technology environment and supports applications built on Google Cloud, with Anthropic Claude and Google Gemini models served through Vertex AI.
The ideal candidate has strong Google Cloud infrastructure experience, practical DevOps and DevSecOps depth, and the ability to support validated systems where deployment, configuration, monitoring, access control, and evidence capture must be inspection ready.
This is NOT a generic infrastructure support role. You will help ensure that cloud environments, CI/CD pipelines, deployment patterns, observability, security controls, and operational processes are reliable, reproducible, auditable, and aligned with validation expectations.
Key Responsibilities
Google Cloud Infrastructure and Platform Support
- Build, configure, and maintain Google Cloud environments supporting AI application delivery across development, test, staging, and production.
- Support Google Cloud Platform services relevant to the platform, including Vertex AI, BigQuery, Cloud Run or GKE, Cloud Storage, Dataflow or Composer, Document AI, Cloud Logging, Cloud Monitoring, Secret Manager, and related services.
- Partner with AI architects, engineers, data specialists, security teams, and delivery leadership to ensure platform architecture is scalable, secure, cost-conscious, and operationally supportable.
- Implement environment strategies that support regulated delivery, including clear separation between development, test, and production environments.
- Support infrastructure patterns for AI and data-intensive applications, including batch processing, event-driven workflows, APIs, retrieval pipelines, and model-backed services.
CI/CD, Release, and Environment Management
- Design, implement, and maintain CI/CD pipelines for application, infrastructure, and AI platform components.
- Support controlled deployment processes across validated and non-validated environments.
- Ensure infrastructure-as-code deployments are reproducible and generate evidence suitable for validation review where required.
- Coordinate release execution with engineering, QA, Technical Project Management, and Business Analyst teams.
- Support rollback procedures, release checklists, runbooks, environment readiness reviews, and deployment documentation.
- Partner with QA to ensure automated test evidence generated through CI/CD pipelines is complete, reliable, attributable, and suitable for validation use where applicable.
DevSecOps, Security, and Access Control
- Implement and maintain cloud security controls across IAM, service accounts, secrets management, network architecture, encryption, and least-privilege access.
- Partner with Security and Cloud Engineering on VPC Service Controls, CMEK, network segmentation, identity federation, and environment access strategy.
- Support secure handling of sensitive data, including PHI, PII, confidential business information, and regulated system data.
- Ensure deployment and operational processes align with company policy regarding source code handling, third-party data, confidential information, and AI-assisted development tooling.
- Support audit readiness by maintaining clear records of access controls, configuration changes, approvals, and operational procedures.
Observability, Reliability, and Incident Response
- Implement monitoring, logging, alerting, and dashboarding for AI-driven applications and supporting cloud infrastructure.
- Monitor system performance, latency, service availability, quota usage, rate limits, cloud consumption, and model dependency behavior.
- Support incident response for platform, infrastructure, deployment, and model-backed application issues.
- Help define and operationalize fallback behavior when upstream data sources, model services, retrieval systems, or cloud dependencies degrade.
- Maintain runbooks and operational documentation that enable repeatable support and clear escalation.
AI Platform and Model Operations Support
- Support operational deployment patterns for model-backed applications using Vertex AI, Gemini, Claude, and related AI services.
- Partner with AI engineers and architects on model routing, prompt/configuration deployment, evaluation release gates, version tracking, and promotion workflows.
- Support re-execution of evaluations when model versions, prompts, retrieval logic, data snapshots, or infrastructure configurations change.
- Help ensure AI outputs can be traced to relevant source records, model versions, prompt or configuration versions, and data snapshots where required.
- Support observability for model-backed systems, including latency, error rates, cost, usage patterns, evaluation results, and production performance drift.
Compliance, Validation, and Audit Support
- Work within computer system validation and computer software assurance practices for GxP-relevant systems.
- Support Installation Qualification activities by documenting environment configuration, infrastructure provisioning, component versions, dependencies, and deployment evidence.
- Provide objective evidence for validation activities, including screenshots, logs, deployment records, configuration outputs, pipeline results, database queries, and system-generated reports.
- Support change control by documenting infrastructure changes, deployment activities, environment updates, and operational impacts.
- Maintain documentation and evidence in a manner that supports internal audits, quality review, and health authority inspection readiness.
- Partner with QA to ensure validated environments remain representative, controlled, and reproducible.
Collaboration and Delivery
- Work closely with Technical Project Managers, Business Analysts, QA Leads, AI Architects, data engineers, application engineers, security teams, and client stakeholders.
- Participate in Agile delivery ceremonies and provide realistic estimates for infrastructure, deployment, DevOps, and operational work.
- Maintain Jira and Confluence documentation related to deployment tasks, runbooks, environment configuration, operational support, and release readiness.
- Surface risks early, especially those related to deployment readiness, environment drift, security gaps, validation evidence, cloud cost, or production supportability.
- Follow issues to closure across engineering, infrastructure, security, QA, and vendor teams.
Required Qualifications
- Bachelor’s degree in computer science, Information Systems, Engineering, or a related field, or equivalent practical experience.
- 5+ years of DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering experience.
- Hands-on experience supporting production systems on Google Cloud Platform.
- Experience with Google Cloud Platform services such as Vertex AI, BigQuery, Cloud Run, GKE, Cloud Storage, Dataflow, Composer, Document AI, Cloud Logging, Cloud Monitoring, IAM, Secret Manager, or related services.
- Strong experience building and maintaining CI/CD pipelines for application and infrastructure deployment.
- Practical experience with infrastructure as code, preferably Terraform.
- Experience with containerized application deployment using Kubernetes, GKE, Docker, or Cloud Run.
- Working knowledge of cloud security principles, including IAM, least privilege, secrets management, network controls, encryption, and environment segregation.
- Experience implementing observability using logs, metrics, traces, alerts, dashboards, and incident response procedures.
- Ability to document deployment steps, configuration evidence, operational procedures, and change details with precision.
- Strong communication skills and the ability to work across engineering, QA, security, delivery, and business stakeholders.
- Authorized to work in the United States without current or future sponsorship.
Preferred Qualifications
- Experience supporting AI, machine learning, LLM, or data-intensive applications in production.
- Familiarity with Vertex AI, Gemini, Claude, Model Garden, model deployment workflows, evaluation release gates, prompt/configuration versioning, or AI observability.
- Experience in pharmaceutical, biotech, medical device, healthcare, or another FDA-regulated environment.
- Familiarity with GxP, CSV, CSA, GAMP 5, 21 CFR Part 11, ALCOA+ principles, or validated system delivery.
- Experience supporting environments where audit trails, objective evidence, change control, and inspection readiness are required.
- Experience with VPC Service Controls, CMEK, private networking, identity federation, and secure cloud architecture.
- Experience with cloud and inference cost monitoring, FinOps practices, quota management, and usage forecasting.
- Experience supporting document processing, retrieval-augmented generation, structured and unstructured data pipelines, or enterprise content systems.
- Google Cloud certification such as Professional Cloud Architect, Professional Cloud DevOps Engineer, Professional Cloud Security Engineer, or Professional Machine Learning Engineer.
Core Competencies
- Operational discipline: builds systems and processes that are repeatable, observable, and supportable.
- Security mindset: treats access control, data protection, and environment isolation as design requirements.
- Validation awareness: understands that deployment evidence, configuration control, and auditability are part of the product.
- Technical pragmatism: balances speed, reliability, cost, and compliance without over-engineering.
- Ownership: follows infrastructure, deployment, security, and production issues through to closure.
- Clear communication: explains technical risks and operational constraints in language appropriate for engineers, QA, delivery leaders, and business stakeholders.