Job Description
Experience: 1–4 years in data engineering, backend data pipelines, or closely related software/data roles
Education: Bachelor’s Degree in Computer Science
We are looking for a Data Engineer to support the development and operation of an on-premises analytical data platform that integrates clinical and operational data from multiple hospital sites into a governed lakehouse environment.
The role focuses on building reliable data pipelines, ensuring data quality, supporting platform operations, and transforming captured data into trusted, linked, and analysis-ready datasets.
The successful candidate will work within established architecture and quality standards, implementing and documenting ingestion and transformation processes while gradually taking on greater responsibility and independence.
Key Responsibilities
Data Ingestion & Platform Pipelines
- Design, build, and maintain batch and change-capture data ingestion pipelines from operational databases across multiple sites.
- Develop orchestrated workflows for extraction, landing, transformation, and promotion across lakehouse layers.
- Ensure pipelines are idempotent, observable, and recoverable through retries, checkpoints, and replay mechanisms where appropriate.
- Collaborate with database administrators on read-only access, change-log readiness, and safe data extraction windows.
Lakehouse Layers
- Manage immutable raw landing and curated datasets using a medallion-style architecture.
- Implement data validation, cleansing, typing, and conformed data models using transformation-as-code practices.
- Support slowly changing history for key demographic and reference entities.
- Develop configuration-driven pipeline definitions using declarative specifications maintained under version control.
Cross-Site Identity & Data Linking
- Develop logic to unify records across different hospital sites, including patients, providers, and facilities.
- Apply defined matching rules and support steward review for ambiguous records.
- Maintain bridge and reference structures that map source identifiers to enterprise identifiers with complete audit trails.
Data Quality, Reconciliation & Operations
- Automate reconciliation between source systems and platform copies using counts, keys, samples, and freshness checks.
- Troubleshoot pipeline failures and data drift using established runbooks and escalation procedures.
- Contribute to schema change management, including identification of additive and breaking changes.
- Maintain documentation and alerts related to schema changes.
- Monitor service levels such as pipeline lag, success rates, and data freshness through operational dashboards.
Documentation & Collaboration
- Maintain technical runbooks, pipeline documentation, and change records alongside implementation changes.
- Collaborate with Analytics Engineers on data catalog entries, data dictionary fields, and lineage metadata.
- Provide technical support to clinical and operational data stewards.
- Respect steward ownership of business decisions regarding duplicates, definitions, and data interpretation.
Ideal Candidate
The ideal candidate should have hands-on experience with data engineering, backend data pipelines, database integration, data quality, and analytical data platforms, along with strong problem-solving and documentation skills.
🔎 Source: Official Careers Portal
⚠️ We only share verified job listings and are not the hiring authority.