Job Title - Data Engineer
Location - Glendale, CA (Hybrid – 2- 4 Days Onsite)
Duration: Full Time
Job Summary
Mandatory skills
• Databricks experience (primary requirement)
• Apache Airflow
• Advanced SQL skills
• Python
• Spark / PySpark
• Scala
• Experience building and maintaining data pipelines and workflows
Key Responsibilities:
● Design, write, test, and deploy data pipelines using PySpark, Scala, SQL, Python
● Meet with stakeholders to gather requirements and translate them into scalable data platform solutions
● Understanding of Databricks platform and developer tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations
● Ability to explain Spark architecture and pipeline behavior to stakeholders to diagnose root causes and recommend solutions
● Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross-platform integrations
● Manage Databricks platform governance, including Unity Catalog, ACLs, lineage, and data discovery and privacy tooling
● Build and maintain Kubernetes containers and containerized utilities supporting deployed data platform services
● Apply networking knowledge to troubleshoot connectivity and integration errors across platform components
● Perform platform administration: provision and remove access, assess resource utilization, monitor platform health and cost, and evaluate stakeholder requests
● Collaborate with engineers, architects, and product managers to drive Core Data platform success; participate in agile/scrum ceremonies
● Maintain documentation of platform changes, standards, and pipeline configurations to support data quality and governance
Qualifications:
● 5+ years of data engineering experience developing and operating large-scale data pipelines
● Deep hands-on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala
● Strong understanding of Spark architecture—executors, stages, partitioning, shuffle, and performance tuning—with ability to explain tradeoffs to technical and non-technical stakeholders
● Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting
● Proficient in SQL with advanced performance tuning capabilities
● Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines
● Experience managing Databricks platform governance: ACLs, Unity Catalog, lineage, and access provisioning
● Proficiency in Python and at least one additional language (Scala, Kotlin, or SQL-driven pipeline tooling)
● Experience designing and optimizing scalable ETL/ELT pipelines integrating diverse structured and unstructured data sources
● AWS-primary experience (compute, storage, networking, IAM); experience with other cloud providers is transferable
● Proficiency with Docker and Kubernetes for building and maintaining containerized data platform services
● Working knowledge of networking concepts to diagnose cross-platform integration and connectivity issues
● Familiarity with Snowflake and comparable tooling relative to the Databricks ecosystem
● Experience designing and implementing CI/CD and DevOps practices (Git-based workflows)
● Experience implementing data quality checks, monitoring, and logging for pipeline reliability
● Self-starting problem solver with strong analytical and communication skills; willingness to learn new tooling and trends
● Familiar with Scrum and Agile methodologies
● Experience with Snowflake is a plus
● Bachelor’s degree in computer science, Information Systems, or a related field, or equivalent work experience, Master’s Degree is a plus.