Global Health EDCTP3Funded by the European Union
The project is supported by Global Health EDCTP3 and its members.
EpiScientia Network
FoundationHealth Informatics

Applied Data Harmonization, Quality, Analytics and Federated Node Deployment

OMOP Data Engineer Certificate

10-week blended professional certificate that develops data engineers able to lead local health-data harmonization at African data-owner institutions: source profiling, structural and semantic mapping, reproducible ETL into the OMOP CDM, data-quality validation, ATLAS cohorts and analytics, and a Dockerised Vantage6 federated data-owner node. Rwanda pilot with extension to African data-owner institutions.

This Certificate is a prerequisite for enrolment in the Postgraduate Certificate programme planned through the EpiScientia Network Project at the University of Rwanda – African Centre of Excellence in Data Science. The Certificate also serves as a prerequisite for joining the Heritage Alliance Ltd technical team supporting African health-data owners in data harmonization and OMOP CDM implementation. This includes contributing to source-data profiling, semantic and structural mapping, ETL development and validation, OMOP CDM deployment, data-quality assessment, and the preparation of local data-owner nodes for participation in federated data networks.

1 · Programme purpose and positioning

Why this certificate

The EpiScientia OMOP Data Engineer Certificate is an applied competency programme designed to develop data engineers who can lead local health-data harmonization at African data-owner institutions. The Rwanda implementation is the initial training setting, while the curriculum is intentionally portable to institutions using OpenMRS, OpenClinic GA, eFiche, eBuzima, MediSoft and registry datasets. The programme follows the source-to-target sequence: source inspection and profiling, structural mapping, semantic standardization, ETL implementation, data-quality validation, cohort construction, analytics, containerization and Vantage6 node operation. Certification is based on demonstrated engineering competence rather than attendance alone: a participant must be able to explain mapping decisions, implement and troubleshoot an ETL, evaluate data quality, construct an executable cohort, and demonstrate a controlled federated computation.

Programme elementDesign
Online instruction10 weekly live sessions; each 60 minutes (25 min theory + 35 min supervised practical)
Asynchronous assessment10 weekly quizzes/artefacts; approximately 20 minutes each
Face-to-face immersionTwo 2-day workshops in Weeks 5 and 10
Core outcomeIndependent conversion of a local EHR/registry source into a validated OMOP CDM data-owner environment capable of participating in Vantage6-mediated federated analytics
AssessmentContinuous assessment 60%; final practical examination 40%
Core referenceThe Book of OHDSI, supplemented by current OMOP CDM and Vantage6 documentation

Teaching staff

Course directors & assistants

Course director

Prof Dr Marc Twagirumukiza, MD, PhD

Course director

Assistant course directors

Carine Umulisa

Assistant course director

Daniel Nsanzabandi

Assistant course director

EpiScientia Mentor

Assistant course director

Topic facilitators

Carine Umulisa

Facilitator

  • · Week 1 — From source EHR to OMOP CDM
  • · Week 2 — Source-data profiling & ETL planning
  • · Week 3 — Structural mapping to OMOP
  • · Week 4 — Semantic harmonization & standardized vocabularies
  • · Week 5 — ETL implementation & reproducibility
  • · Week 6 — Data quality, validation & ETL improvement
  • · Week 7 — ATLAS, WebAPI & cohort construction
  • · Week 8 — Analytics on OMOP & AI-assisted mapping
  • · Week 9 — Docker & Vantage6 data-owner node
  • · Week 10 — Federated OMOP network operations
  • · Workshop I — Build the OMOP Node

Daniel Nsanzabandi

Facilitator

  • · Workshop II — From Local OMOP to Federated Node

The course director leads the programme and appoints the assistant course directors, who oversee the trainees' progress and the course operations. Facilitators lead individual weekly topics and workshops.

2 · Award and programme specifications

Specifications

Award titleEpiScientia Certificate in OMOP Data Engineering
Target audienceData engineers, database administrators, health-informatics specialists, software/data developers, data managers and technically oriented analysts working with clinical, surveillance or registry data
Delivery modeBlended: weekly synchronous online instruction + asynchronous quizzes/artefacts + two in-person workshops
Duration10 consecutive teaching weeks
Live online contact10 hours total
Asynchronous quiz timeApproximately 3 h 20 min total
In-person workshops4 full training days
Recommended additional self-study2–4 hours/week, increasing around Weeks 4–6 and the capstone
Training dataInstitution-owned data where governance permits; otherwise de-identified or synthetic source-like extracts
Reference CDMCurrent supported OMOP CDM release selected by the programme technical lead and frozen for each cohort

3 · Entry prerequisites

Who can join

RequiredWorking knowledge of relational databases and SQL (SELECT, JOIN, GROUP BY, keys and basic data types).
RequiredAbility to navigate a Linux or Windows development environment, install approved software, and manage files/directories.
RequiredAccess to a laptop capable of running the training toolchain or to an EpiScientia-provided remote training environment.
Required for site-based participantsAuthorization to inspect the structure of their institutional source system and use an approved training extract.
RecommendedExperience with PostgreSQL, MySQL/MariaDB or another relational DBMS; Git; Python or R; and basic Docker concepts.
Not requiredPrior OMOP/OHDSI or Vantage6 experience.

4 · Programme-level learning outcomes

On successful completion, the certificate holder will be able to

  1. LO1Explain the OHDSI ecosystem, OMOP CDM architecture, domains, conventions and standardized vocabulary model.
  2. LO2Inspect an unfamiliar EHR or registry schema and produce a source-data inventory and profiling report.
  3. LO3Design structural source-to-target mappings and document transformation logic before coding.
  4. LO4Map local clinical codes and terms to appropriate standard concepts, documenting uncertainty and expert review.
  5. LO5Implement a reproducible ETL that preserves source provenance and populates a selected OMOP CDM release.
  6. LO6Run and interpret OMOP data-quality and characterization outputs and iteratively correct the ETL.
  7. LO7Configure/use WebAPI and ATLAS sufficiently to create concept sets, define cohorts and execute basic standardized analyses.
  8. LO8Critically use alternative and AI-assisted mapping approaches without delegating semantic accountability to automation.
  9. LO9Containerize/configure a local data-owner environment and connect an approved Vantage6 node to a training central server.
  10. LO10Independently demonstrate the end-to-end pathway: source → profile → map → ETL → OMOP → quality → cohort/analytics → federated node.

Teaching model — Each online week uses a fixed 60-minute live pattern: 25 minutes of focused theory followed immediately by 35 minutes of supervised practice. The practical is performed on the learner's own approved data where feasible, or on a common source-like training dataset. A 20-minute asynchronous quiz is linked to the same week's competency and is completed later at the learner's pace. Weekly assessment combines short knowledge items with evidence from the practical exercise so the certificate does not become a knowledge-only award.

6 · Ten-week curriculum map

Week by week

25-minute lecture

  • 0–5 min: Why common data models and distributed evidence networks.
  • 5–10 min: OHDSI ecosystem and role of OMOP CDM.
  • 10–17 min: Core clinical tables and domains: PERSON, VISIT, CONDITION, DRUG, MEASUREMENT, OBSERVATION.
  • 17–22 min: Source concepts, standard concepts and preservation of source values.
  • 22–25 min: End-to-end architecture from local source to analytics/federation.

35-minute supervised laboratory

  • 0–5 min: Open assigned OpenMRS/OpenClinic/eFiche/eBuzima/MediSoft/registry extract and inspect schema.
  • 5–15 min: Identify patient, encounter, diagnosis, observation/lab and medication structures and keys.
  • 15–25 min: Classify representative variables into likely OMOP domains/tables.
  • 25–32 min: Trace one patient journey from source records to candidate OMOP events.
  • 32–35 min: Save a one-page source-to-OMOP hypothesis for instructor review.

Practical focus: Explore source schemas; identify patients, encounters, diagnoses, observations, labs, medications; propose OMOP domains

Quiz / artefact (~20 min): CDM domains, source vs standard concepts, data types and a short source-to-target classification exercise. — OHDSI/CDM architecture and source→target reasoning

8–9 · In-person workshops

Two engineering sprints in Kigali

Workshop I · Week 5

Build the OMOP Node

Convert the first four weeks of mapping work into a working, documented OMOP ETL v1. The workshop is an engineering sprint, not a repetition of the online lectures.

Mon, 2 Nov 2026 – Tue, 3 Nov 2026 Kigali, Rwanda (in person · 2 days)
TimeActivityRequired output
08:30–09:00Registration, environment check, dataset/governance confirmationWorking environments confirmed
09:00–09:45Architecture recap and team ETL planTeam implementation plan
09:45–10:45Source schema clinic: WhiteRabbit findings and unresolved source questionsReviewed source inventory
10:45–11:00Break
11:00–12:30Rabbit-in-a-Hat structural mapping studioReviewed source-to-target design
12:30–13:30Lunch
13:30–15:00Vocabulary/Usagi mapping studioVersioned mapping file
15:00–15:15Break
15:15–16:30Mapping peer review and clinical/semantic escalationMapping decisions + issue log
16:30–17:00Day 1 checkpoint and ETL build planDay 2 task board

Milestone package: OMOP ETL v1; ETL specification; vocabulary mapping file; source inventory/profile; Git repository or equivalent version-controlled package; populated local/training OMOP database; issue log.

Workshop II · Week 10

From Local OMOP to Federated Node

Close the quality loop, validate the cohort, connect the Vantage6 node and sit the final practical examination (Parts A–C) with a multi-node federated capstone.

Tue, 22 Dec 2026 – Wed, 23 Dec 2026 Kigali, Rwanda (in person · 2 days)
TimeActivityRequired output
08:30–09:00Readiness check and capstone briefingIndividual/team capstone plan
09:00–10:15DQD/Achilles rerun and quality triageUpdated quality report
10:15–10:30Break
10:30–11:45ETL correction sprint and mapping closurePriority defects resolved
11:45–12:30ATLAS concept-set and cohort validationValidated cohort definition
12:30–13:30Lunch
13:30–14:30Docker stack verification and configuration reviewReproducible local stack
14:30–15:30Vantage6 node configuration and central registrationConnected training node
15:30–15:45Break
15:45–16:45Approved test computation and troubleshootingSuccessful node execution
16:45–17:00Production-readiness checklistCapstone readiness decision

10 · Assessment framework

Continuous assessment 60% · final practical examination 40%

ComponentWeightEvidence
Weekly assessment 1–1060% total (6% each)Each week: short knowledge component plus practical artefact/result. The exact knowledge/artefact split may vary by week, but practical evidence must be present.
Final practical examination40%Observed end-to-end engineering tasks completed in a controlled training environment during Week 10.

Final practical examination rubric (40%)

DomainWeightCompetent performance
Source profiling & structural mapping10%Accurate source interpretation; correct OMOP target selection; documented assumptions and transformations; traceability from source to target.
ETL implementation & reproducibility10%Executable transformations; correct IDs/keys/dates; provenance; repeatability; logging/error handling; version-controlled implementation.
Vocabulary & semantic mapping8%Appropriate standard concepts/domains; source preservation; evidence-based acceptance/rejection of candidates; explicit handling of uncertainty.
OMOP quality + cohort validation7%Correct use/interpretation of quality outputs; remediation of material defects; executable concept set/cohort with defensible temporal logic.
Vantage6 deployment & federated execution5%Correct node configuration; secure/approved data interface; successful connection and task execution; ability to interpret logs/results and troubleshoot.

Performance descriptors

4 — Independent/ExcellentCompletes the task correctly and reproducibly; explains design choices; detects edge cases; documentation is sufficient for another engineer to reproduce the work.
3 — CompetentCompletes the essential task correctly with minor non-material errors or limited prompting; decisions are technically defensible.
2 — DevelopingPartial completion; material errors, weak traceability or repeated prompting; output is not yet reliable for production use.
1 — InsufficientCannot complete the essential task or makes unsafe/invalid mapping, ETL, quality or federation decisions.

11 · Certification criteria

EpiScientia Certificate in OMOP Data Engineering

  • Overall weighted mark ≥ 60%.
  • Final practical examination mark ≥ 50% (competency gate).
  • Submission of all mandatory capstone artefacts or approved equivalents.
  • Participation in both in-person workshops unless a formally approved equivalent supervised practical assessment is provided.
  • No critical failure in data-governance, security or semantic-mapping practice that would make the demonstrated node unsafe or analytically invalid.
  • Participants who meet the overall mark but fail the practical gate receive a completion statement rather than the OMOP Data Engineer Certificate and may undertake a supervised reassessment.

Statement of competence — A successful EpiScientia OMOP Data Engineer Certificate holder has demonstrated the ability to independently take an approved clinical or registry source through source profiling, structural and semantic mapping, reproducible ETL implementation, OMOP quality validation, cohort construction and local analytics, and to configure the resulting data-owner environment for controlled participation in Vantage6-mediated federated analytics.

Trainer and assessor requirements

RoleMinimum profileRecommended responsibility
Programme scientific leadSenior health informatics/OHDSI practitioner with demonstrable OMOP implementation experienceCurriculum oversight, standards, assessment moderation
Lead OMOP data engineerHands-on ETL engineer experienced with relational databases, OMOP conventions, WhiteRabbit/Rabbit-in-a-Hat and vocabulary mappingWeeks 1–6; workshop engineering supervision
Vocabulary/clinical terminology expertExperience with OMOP standardized vocabularies and clinical terminologies; clinical domain input strongly preferredWeek 4, mapping review, semantic adjudication
OHDSI analytics trainerPractical ATLAS/WebAPI/cohort experienceWeeks 6–8; cohort/analytics assessment
Vantage6/DevOps trainerDocker/Linux/networking plus practical Vantage6 deployment experienceWeeks 9–10; node configuration and troubleshooting
Teaching assistantsSQL/ETL competent; ideally 1 TA per 6–8 learners during workshopsHands-on support without completing assessed work for learners
Independent assessor/moderatorNot the sole trainer responsible for the learner's capstoneFinal practical moderation and borderline decisions

13 · Technical and training infrastructure

Software stack & environments

LayerRequired / recommended tools
Source profilingWhiteRabbit
ETL designRabbit-in-a-Hat plus source-to-target specification templates
Semantic mappingUsagi + OHDSI standardized vocabularies/Athena-derived vocabulary files as institutionally authorized
ETL executionSQL as core; Python/R optional according to source and local engineering standards
OMOP quality/characterizationDataQualityDashboard; Achilles; optional ARES visualization
AnalyticsWebAPI; ATLAS; SQL; selected OHDSI/HADES packages where relevant
ContainerizationDocker
Federated executionVantage6 node + central training server + approved algorithm containers
Engineering governanceGit/version control, issue log, mapping decision log, release notes
Learner workstationModern 64-bit laptop; recommended ≥4 CPU cores, ≥16 GB RAM and ≥100 GB free SSD space (32 GB RAM preferable for heavier local Docker/database work). Administrative rights or a managed preconfigured image/remote desktop. Stable broadband (wired recommended for workshops). Git client; database client; approved code editor/IDE; Docker Engine/Desktop; Java runtime where required by the selected OHDSI tools; access to a PostgreSQL training database or equivalent DBMS.
Central training environmentTraining OMOP database(s) and standardized vocabulary release frozen for the cohort. OHDSI WebAPI + ATLAS training deployment. Central Vantage6 training server with pre-created organizations/users/roles and a controlled algorithm image/registry workflow. Version-control repository for course ETLs, mappings, cohort definitions and assessment submissions. Backup/snapshot capability before each workshop and before the final examination. Separate training/sandbox environment from any production national data node.
Data-owner / site environmentWhere learners work on institutional data, the source system is not modified by training activities: a read-only replica, approved extract or dedicated local data-engineering node is used. Patient-level data remains within the authorized institutional environment. Cross-site exercises use approved aggregate outputs or synthetic/de-identified training data according to local governance.

Data governance, security and responsible AI

  • Use only data for which the participant and training team have explicit institutional authorization.
  • Prefer synthetic or de-identified datasets for cross-country teaching demonstrations; real patient-level data must not be copied into shared course infrastructure without a lawful and approved basis.
  • Maintain separation between training and production credentials, servers and Vantage6 organizations.
  • Record source-to-standard mapping decisions, including rejected and uncertain mappings, so that semantic choices are auditable.
  • AI/LLM tools may propose mappings, transformation code or documentation, but their outputs must be validated against the selected OMOP vocabulary/CDM release and reviewed by a competent human.
  • Patient-level data must not be entered into external generative-AI services unless explicitly authorized under applicable governance.
  • Federated execution does not by itself guarantee privacy: algorithms, output disclosure rules, permissions and node configuration require governance and technical review.

15 · Mandatory learner portfolio

What every certificate holder hands over

  1. Source-system inventory and WhiteRabbit profiling report.
  2. Rabbit-in-a-Hat/source-to-target ETL design.
  3. Versioned vocabulary mapping file with review status.
  4. Executable ETL package and documentation.
  5. OMOP data-quality/characterization report and remediation log.
  6. Saved/exported ATLAS concept set and cohort definition.
  7. Docker/Vantage6 node configuration evidence with secrets removed.
  8. Federated test execution evidence and troubleshooting log.
  9. Final technical handover note describing how the local institution can maintain and rerun the pipeline.

Quality assurance — Each cohort freezes the OMOP CDM release, vocabulary release, tool versions and assessment dataset/environment before teaching begins. Assessment rubrics are calibrated by trainers before Week 5 and moderated after Week 10. Mapping disputes are escalated to a designated terminology/clinical reviewer rather than resolved by convenience. Course feedback distinguishes teaching quality from technical environment failures so that infrastructure problems do not distort competency judgments.

17 · Reference framework

Core references

  • The Book of OHDSI — core conceptual and methodological reference, particularly the Extract-Transform-Load chapter.
  • OHDSI OMOP Common Data Model documentation — current CDM specification and conventions.
  • OHDSI standardized vocabularies and associated mapping tooling/documentation.
  • DataQualityDashboard and Achilles documentation for quality and characterization.
  • ATLAS/WebAPI documentation for cohort and analytical workflows.
  • Vantage6 documentation for server, node, client, Dockerized algorithm execution and node configuration.

Ready to become an OMOP data engineer?

Enrol in the Certification Programme to follow this course with a scholarship, or browse the other courses.

Cookies: policy