Hello
Position: Infrastructure Engineer
Location: Cambride, MA (Hybrid )
Job Introduction: We are seeking a hands-on Systems Admin who will be responsible for the secure, reliable, and compliant operation of cloud-hosted research applications, agentic AI systems, and scientific SaaS platforms. This role supports systems that integrate cloud infrastructure, containers, AI models, Model Context Protocol (MCP) servers, scientific databases, APIs, and third-party services. Partnering with scientists, developers, IT platform owners, cybersecurity, architecture, networking, data teams, and vendors, the administrator keeps pipeline and production services running reliably. The role does not own architecture, but must understand it well enough to operate, secure, troubleshoot, and maintain these applications effectively.
Job Responsibilities:
Administer research applications and infrastructure across Google Cloud Platform, AWS, and scientific SaaS environments, including cloud projects, IAM, networking, storage, quotas, and compute services, while maintaining separation among development, test, validation, and production environments
- Operate Cloud Run services, containers, registries, domains, DNS, TLS certificates, load balancing, and Identity-Aware Proxy, supporting secure connectivity among cloud environments, internal systems, scientific databases, SaaS platforms, model APIs, and laboratory networks
- Administer SaaS tenant settings, access, integrations, APIs, audit controls, and licenses, monitoring vendor releases, security advisories, and end-of-life notices
- Maintain inventories of approved agents, MCP servers, tools, data sources, integrations, permissions, credentials, and system owners; maintain and update tooling supporting agentic AI frameworks and orchestration libraries, monitoring performance and resolving system errors
- Administer user, group, service account, and machine-to-machine access using role-based access controls and least-privilege principles; rotate credentials and store secrets in approved services, eliminating hard-coded credentials
- Maintain container images, Dockerfiles, registries, and pinned dependencies, remediating vulnerabilities and deprecations while maintaining rollback procedures and deployment traceability
- Maintain logs, metrics, traces, alerts, dashboards, and automated health checks covering availability, performance, errors, model/tool failures, and cloud cost and usage
- Create runbooks and support incident triage, root-cause analysis, corrective actions, backup, restoration, and disaster recovery; record and manage incidents, requests, and changes through ServiceNow and Jira with clear ownership and traceability
- Maintain CI/CD pipelines and infrastructure-as-code configurations, applying source control, peer review, security scanning, and rollback controls, and help transition scientist-managed prototypes into documented, secure, production-ready deployments
- Apply enterprise security and risk standards, supporting architecture reviews, threat modeling, audits, and investigations; maintain audit logs, identify unmanaged integrations and shadow AI services, and partner with scientists, developers, security, cloud engineering, and vendors on operational readiness and lifecycle ownership
Job Requirements:
Bachelor's degree in computer science, information systems, engineering, technology, or a related field
- 8+ years of experience administering production systems in Google Cloud Platform, AWS, Azure, or a comparable cloud environment
- 5+ years of experience with Linux, containers, Docker, Python environments, APIs, networking, DNS, TLS, identity, secrets management, and serverless or container orchestration services such as Cloud Run or Kubernetes
- Knowledge of IAM, SSO, OAuth, service accounts, machine-to-machine authentication, CI/CD, infrastructure as code, monitoring, incident management, backup, recovery, patching, and vulnerability remediation
- Familiarity with ServiceNow for IT service-management processes and Jira for work intake, backlog management, issue tracking, documentation, and operational work management
- Strong troubleshooting, documentation, communication, and cross-functional collaboration skills
- Experience with Google Cloud IAP, Cloud Run, Secret Manager, Artifact Registry, Cloud Monitoring, Gemini, or Vertex AI, including systems spanning Google Cloud Platform and AWS
Familiarity with scientific computing, bioinformatics, laboratory automation, multiomics, research data platforms, or scientific SaaS, plus knowledge of container scanning, software bills of materials, supply-chain security, data integrity, and auditability. Relevant cloud, systems engineering, IT service management, or cybersecurity certifications are desirable
Thanks and regards
Sonali Silswal Team Lead- Talent Acquisition
Email:
Contact:
LinkedIn:
Address: 2591 Dallas Pkwy, Ste 300, Frisco, TX 75034
Website: