-
Candidates must possess 5 8+ years of experience in technical support, systems engineering, production support, site reliability, or application support.
-
Hands-on experience supporting or administering the Databricks platform is required.
-
Experience troubleshooting enterprise software or cloud-based platforms is required.
-
Strong experience troubleshooting infrastructure, application, and platform-related issues using systematic root cause analysis.
-
Experience working with Databricks notebooks, clusters, jobs, SQL Warehouses, and workspace administration.
-
Experience troubleshooting Spark job failures and distributed computing environments.
-
Working knowledge of SQL for investigating and resolving technical issues.
-
Experience with cloud platforms such as Azure, AWS, or Google Cloud Platform.
-
Experience with monitoring and observability tools such as Splunk, Grafana, Datadog, or Azure Monitor.
-
Strong understanding of system logs, monitoring tools, and technical troubleshooting methodologies.
-
Basic understanding of networking fundamentals, authentication, and identity management.
-
Familiarity with Linux environments and command-line troubleshooting.
-
Experience working with ticketing and incident management systems such as ServiceNow or Jira Service Management.
-
Strong analytical and problem-solving skills with the ability to independently investigate and resolve technical issues.
-
Excellent verbal and written communication skills with a customer-first mindset.
-
A Bachelor's degree in Systems Engineering, Computer Science, Information Technology, or a related technical field is required.
-
Serve as the primary technical point of contact for customer-reported Databricks platform and application issues.
-
Engage directly with customers to understand business impact, gather technical information, and accurately identify issues.
-
Troubleshoot infrastructure, application, and platform-related problems through systematic investigation and root cause analysis.
-
Investigate system logs, error messages, monitoring data, and platform behavior to identify and resolve technical issues.
-
Resolve Tier 1 support incidents independently whenever possible.
-
Troubleshoot Databricks notebooks, clusters, jobs, SQL Warehouses, and other platform components.
-
Investigate and resolve Spark job failures and distributed computing issues.
-
Use SQL and available monitoring tools to diagnose customer-reported issues and identify potential causes.
-
Monitor platform health and application behavior using observability tools such as Splunk, Grafana, Datadog, or Azure Monitor.
-
Identify whether issues are related to applications, infrastructure, networking, authentication, identity, or the Databricks platform.
-
Escalate complex technical issues to Tier 2 engineering with detailed troubleshooting notes, findings, and supporting documentation.
-
Document troubleshooting activities, root cause analysis, findings, and resolution steps within the incident management system.
-
Track support cases through resolution while maintaining consistent and professional communication with customers.
-
Work with cross-functional technical teams to drive complex incidents toward resolution.
-
Provide a consultative and customer-focused approach when diagnosing technical issues and recommending solutions.