
Our customer is a global energy and technology company operating across multiple markets worldwide. With a strong focus on digital transformation, engineering, and innovation, the company uses advanced data and cloud technologies to improve business operations and decision-making at scale.
The project focuses on building and scaling a modern cloud-based data platform on Microsoft Azure.
The platform brings together data from multiple sources and uses Azure Databricks, Data Factory, Data Lake Storage Gen2, PySpark, Spark SQL, and Delta Lake to create reliable, scalable data pipelines and transformation processes.
The solution is designed for large-scale data processing, with a strong focus on performance, data quality, security, governance, and automation.
You will work as part of a cross-functional technology and client team, interacting with technical consultants, application specialists, senior IT professionals, engineers, account managers, business stakeholders, and customer leadership.
Design develop and maintain scalable data pipelines using Azure Databricks and Azure Data Factory.
Develop ETLELT solutions for ingesting transforming and loading data from multiple sources.
Build and optimize PySpark applications and Spark SQL transformations for largescale data processing.
Create and manage Databricks notebooks workflows jobs and clusters.
Implement Delta Lakebased data solutions ensuring data quality reliability and performance.
Develop data ingestion frameworks using Azure Data Lake Storage ADLS Gen2.
Monitor troubleshoot and optimize data pipelines and Databricks workloads.
Implement CICD pipelines for Databricks artifacts using Azure DevOps or GitHub.
Collaborate with data architects business analysts and stakeholders to deliver enterprise data solutions.
Ensure security governance and compliance using Unity Catalog Azure Key Vault and related Azure services.