Location: Onsite – Houston Employment Type: Full-Time
About the Role
We are seeking a highly skilled, hands-on Databricks Lead with 8+ years of experience in data engineering, including deep expertise in Azure Databricks, PySpark, and Structured Streaming. This role is ideal for a senior engineer with a programming-first mindset, a strong understanding of distributed systems, and the ability to build high-performance, cost-optimized data solutions.
This is not a traditional ETL role. It requires extensive knowledge of Spark internals, declarative pipeline development, and real-time data processing. The successful candidate will also provide technical leadership and collaborate closely with cross-functional teams to deliver robust, scalable data solutions.
Key Responsibilities
Lead the design and development of scalable, high-performance data pipelines, streaming tables, and Delta Live Tables (DLT) using Azure Databricks and PySpark
Drive Spark performance tuning and implement cost-optimization strategies within the Databricks environment
Build and manage real-time and batch workflows using Structured Streaming
Leverage Delta Live Tables (DLT) and Lakehouse Declarative Pipelines (LDP) to build scalable, reliable, and maintainable data pipelines
Collaborate with data architects, analysts, and business stakeholders to deliver robust data solutions
Provide technical leadership through mentorship, code reviews, and architectural guidance
Ensure system reliability and performance through proactive monitoring and engineering best practices
Must-Have Qualifications
8+ years of hands-on experience in data engineering, big data development, or a related field
Extensive hands-on experience with Azure Databricks and PySpark
Strong programming background and in-depth knowledge of distributed data processing beyond traditional ETL tools
Proven expertise in:
Spark performance tuning and optimization
Cost management within Databricks
Structured Streaming for real-time data processing
Lakehouse Declarative Pipelines (LDP) and Delta Live Tables (DLT)
Familiarity with the Azure data ecosystem, including ADLS, Azure Data Factory, and Synapse
Excellent communication skills and the ability to collaborate effectively with cross-functional teams
Willingness and ability to work onsite in Houston
Nice-to-Have Qualifications
Industry experience in energy, utilities, or heavy industry
Knowledge of data governance, security, and monitoring