Data Engineer
Primary Responsibilities:
Integrate data from multiple on-premises and cloud sources and systems, handling ingestion, transformation, and consolidation to create a unified and reliable data foundation for analysis and reporting
Develop data transformation routines to clean, normalize, and aggregate data, including handling complex data structures and missing or inconsistent data
Implement data de-identification and data masking in line with company standards
Monitor data pipelines and data systems to detect and resolve issues promptly
Develop monitoring tools and automate error-handling mechanisms to ensure data integrity and system reliability
Use data quality tools such as Great Expectations or Soda to support the accuracy, reliability, and integrity of data throughout its lifecycle
Create and maintain data pipelines using Airflow and Snowflake as primary tools
Develop SQL stored procedures to perform complex transformations
Understand data requirements and design optimal pipelines to fulfill business use cases
Create logical and physical data models to maintain data integrity
Support CI/CD pipeline creation and automation using Git and Git Actions
Tune and optimize data processes
Collaborate with data scientists and AI teams to operationalize machine learning models and integrate them into data workflows
Build and maintain feature stores, model input pipelines, and real-time data processing frameworks
Ensure data quality, governance, and security across all stages of the data lifecycle
Participate in Agile development processes, including sprint planning, code reviews, and CI/CD using GitHub or Azure DevOps
Required Qualifications:
Bachelor''''s degree in Computer Science or a related field
Hands-on experience developing data pipelines in Snowflake and writing complex SQL queries
Proven hands-on experience as a Data Engineer
Experience with CI/CD pipelines using Git and Git Actions
Experience building ETL/ELT data pipelines
Proficiency in SQL, with practical experience in PostgreSQL or similar relational databases
Experience building queries in SQL and PSQL
Experience working in an Agile development environment
Experience building unit tests, executions
Experience using a source control repository, preferably GIT
Experience in using GenAI tools (Github Co-pilot etc.)
Experience with related open-source platforms and languages such as Scala, Python, Java, and Linux
Experience with relational and non-relational databases
Experience working on Agile/Scrum projects with high-performing teams
Knowledge of data modeling techniques, including star schema, dimensional models, and Data Vault
Knowledge of data warehousing principles, architecture, and implementation
Knowledge of mesh compliance is desirable but not essential
Solid knowledge of Python
In-depth knowledge of Snowflake architecture, features, and best practices
Good understanding of access control, data masking, and row access policies
Exposure to DevOps methodology
Proficiency in SQL, including window functions and advanced features
Proven solid written and verbal communication skills
Proven solid analytical and problem-solving skills applied to large data sets
Preferred Qualifications:
Bachelor''''s degree or higher in Database Management, Information Technology, Computer Science, or a related field
Experience in Data Engineering
Experience orchestrating data tasks in Airflow to run on Kubernetes for data ingestion, processing, and cleaning
Experience with Databricks
Familiarity with Azure services such as Blob Storage, Functions, Azure Data Factory, Service Principal, Containers, and Key Vault
Expertise in designing and implementing data pipelines to process high volumes of data
Motivated self-starter who excels at managing tasks independently and takes ownership
Proven ability to create Docker images for applications to run on Kubernetes