Software Developer 4

Menlo Park, CA, US • Posted 3 days ago • Updated 10 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

📊 Calculating match score...

Job Details

Skills

  • Reliability Engineering
  • Deployment Engineering
  • IT Management
  • Data Management
  • Promotions
  • Backup & Restore
  • Batch Processing
  • Operational Efficiency
  • Process Improvement
  • Innovation
  • Management
  • Requirements Analysis
  • Debugging
  • Change Control
  • Reporting
  • SAN
  • Mentorship
  • Database
  • Testing
  • Kubernetes
  • Continuous Integration
  • Continuous Delivery
  • Linux
  • Python
  • Shell Scripting
  • Relational Databases
  • PostgreSQL
  • Backup
  • Recovery
  • Storage
  • Interfaces
  • File Systems
  • Grafana
  • Dashboard
  • Software Development
  • Quality Control
  • Change Management
  • Incident Management
  • Documentation
  • Communication
  • Collaboration
  • Apache Kafka
  • DRP
  • Workflow
  • Weka
  • Amazon S3
  • Data Processing

Summary

The Rubin U.S. Data Facility at SLAC seeks a senior Software Developer / Site Reliability Engineer to lead and execute complex software and infrastructure work supporting Rubin Observatory production services. This role will help operate, automate, and improve the systems that support Alert Production, Data Release Production, database-backed services, Kubernetes applications, Weka and object-storage services, data movement, and other USDF-hosted operational platforms.

This position combines senior software development, site reliability engineering, production operations, deployment engineering, and cross-team technical leadership. The successful candidate will work across Rubin Data Management, SQuaRE, Prompt Processing, DRP, database, storage, networking, and USDF infrastructure teams to ensure that critical services can be deployed, monitored, debugged, upgraded, recovered, and scaled safely.

The role is aligned with the Software Developer 4 level: leading and executing difficult or complex programming and analysis work, contributing broad technical responsibility across multiple functions, and interfacing with other complex systems and programs.

Key responsibilities
  • Lead and execute complex software, automation, and operational engineering projects for Rubin production services at USDF, including Alert Production, DRP, data access, databases, and service infrastructure.
  • Support and improve Kubernetes-hosted services deployed through Rubin's Phalanx/GitOps environment, including Helm configuration, secrets management, environment promotion, release coordination, and operational rollback.
  • Maintain and troubleshoot PostgreSQL-backed services, including backup and restore, performance, schema-change coordination, monitoring, and incident response.
  • Help operate and debug services that depend on Weka storage, S3 gateways, object storage access, shared filesystems, Butler repositories, batch processing, and high-throughput data movement.
  • Develop strategies, methods, tools, and procedures that improve reliability, automation, reproducibility, observability, and operational efficiency across USDF production services, consistent with Software Developer 4 expectations for process improvement and innovation.
  • Direct or drive all phases of selected technical projects, including requirements analysis, design, implementation, testing, deployment, documentation, and operational handoff.
  • Lead testing, debugging, change control, reporting, and documentation for major service and infrastructure changes.
  • Build and maintain monitoring, logging, alerting, dashboards, runbooks, and operational documentation using tools such as Prometheus, Grafana, Loki, and related systems.
  • Provide subject-matter expertise for complex technical problems involving Kubernetes, databases, storage, networking, deployment automation, and distributed production services.
  • Recognize opportunities for high-impact, long-term reliability improvements that span team or organizational boundaries, and recommend or implement actions to resolve them.
  • Mentor technical staff and collaborate with developers, operators, database engineers, storage engineers, and scientists to improve service reliability and operability.

Desired qualifications and experience
  • Bachelor's degree and ten years of relevant experience, or a combination of education and relevant experience, consistent with the Software Developer 4 profile.
  • Demonstrated experience designing, developing, testing, deploying, operating, and maintaining complex applications or infrastructure services.
  • Strong experience with Kubernetes, Helm, GitOps, CI/CD, Linux systems, Python, shell scripting, and production service automation.
  • Strong understanding of relational databases, especially PostgreSQL, including operational monitoring, performance troubleshooting, backup/restore, and schema-change coordination.
  • Experience with large-scale storage or data systems, ideally including Weka, S3/object-storage interfaces, shared filesystems, high-throughput data transfer, or scientific data repositories.
  • Experience with observability systems such as Prometheus, Grafana, Loki, alerting systems, logs, metrics, and operational dashboards.
  • Strong understanding of software development life cycle, quality-control practices, change management, incident response, and operational documentation.
  • Exceptional written and oral communication skills for technical and non-technical audiences, including the ability to work across distributed teams and organizational boundaries.
  • Experience with Rubin Observatory software, Phalanx, Butler, Qserv, Kafka, Prompt Processing, DRP workflows, Weka, S3DF/USDF, or large-scale astronomical data processing would be especially valuable.

Given the nature of this position, SLAC is open to on-site, hybrid, and remote work options
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: RTX169eef
  • Position Id: dde851317317b84980ddef901792c4ca
  • Posted 3 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Palo Alto, California

Today

Full-time

USD 135,000.00 - 155,000.00 per year

Sunnyvale, California

Today

Full-time

USD 138,000.00 - 207,000.00 per year

Menlo Park, California

Today

Full-time

USD 200,000.00 - 287,500.00 per year

Santa Clara, California

Today

Full-time

USD 190,900.00 - 334,100.00 per year

Search all similar jobs