Senior Data Engineer (Azure Databricks / PySpark)

Nextjob
Full-timeColombo, Sri Lanka

The Company

We are an award-winning software product engineering company that has delivered bespoke web, enterprise, mobile, and cloud solutions to clients around the world for over 25 years. The company has recently expanded into AI-driven innovation, offering AI development, consulting and implementation, data engineering and master data management, and cloud architecture alongside its core software development services. We are a Microsoft Silver Partner and AWS partner, ISO certified, and a SLASSCOM member, and in 2025 won gold at every major national technology award ceremony in Sri Lanka.

With long-standing partnerships spanning 5 to 15+ years with global clients across healthcare, logistics, and enterprise software, we combine deep technical expertise with a client-first approach — helping start-ups and enterprises alike scale through digital transformation. The company prides itself on dedication, continuous improvement, and adaptability, building teams and solutions that grow and evolve alongside its clients' changing needs.

Employment Details

  • Employment Type: 2 Year Contract
  • Work Arrangement: Hybrid – 3 to 4 days per week required onsite
  • Working Days: Monday to Friday
  • Working Hours: 9:00 a.m. to 6:00 p.m.
  • Experience: Minimum Overall 7+ years of experience and 5+ years data engineering with strong production PySpark/Spark experience.
  • Communication: Good English communication skills, both written and verbal (minimum 7/10 proficiency)

Role Overview

Responsible for supporting scalable data pipelines that connect enterprise source systems with MDM and governed downstream platforms. The role focuses on data ingestion, standardisation, transformation, data quality, and ensuring reliable and consistent data is available for MDM and downstream consumption.

Key Responsibilities

  • Develop ingestion pipelines for REST APIs, database/CDC and file-based sources.
  • Build PySpark/Spark transformations and Delta Lake processing on Azure Databricks.
  • Implement Bronze/raw, standardised/conformed and governed data patterns as defined by architecture.
  • Develop reusable standardisation, validation and DQ processing before MDM.
  • Implement incremental loading, schema validation/evolution and restart/recovery patterns.
  • Create source-to-canonical transformations while retaining lineage/provenance.
  • Optimise Spark jobs for performance, cost and operational reliability.
  • Write automated unit/data-quality tests and participate in peer review.
  • Integrate Databricks outputs/inputs with Profisee and downstream Azure services.
  • Support production monitoring, troubleshooting and runbook creation.

Must-Have Experience & Skills

  • Overall 7+ years of experience and 5+ years data engineering with strong production PySpark/Spark experience.
  • Strong Azure Databricks, Delta Lake, SQL and ADLS Gen2.
  • REST/API ingestion and database integration experience.
  • Incremental/CDC processing and schema-evolution experience.
  • Strong data-quality/testing practices and Git/CI/CD.
  • Experience handling large, messy enterprise datasets

Preferred / Nice-to-Have

  • Databricks certification.
  • MDM/entity-resolution project exposure.
  • Azure Data Factory or equivalent orchestration.
  • Exposure to ML-based entity resolution / fuzzy matching (PySpark MLlib), including deterministic and probabilistic matching design
  • Purview/Unity Catalog governance exposure.

Apply for this job

Resume/CV*

Click or drag file to this area to upload your Resume

Please make sure to upload a PDF

First Name*
Last Name*
Email*
Phone Number*
The hiring team may use this number to contact you about this job.

By clicking 'Submit Application', you agree to receive job application updates from Nextjob via text and/or WhatsApp. Message frequency may vary. Reply STOP to unsubscribe at any time. Message & data rates may apply.