Applied Data Harmonization, Quality, Analytics and Federated Node Deployment
OMOP Data Engineer Certificate
10-week blended professional certificate that develops data engineers able to lead local health-data harmonization at African data-owner institutions: source profiling, structural and semantic mapping, reproducible ETL into the OMOP CDM, data-quality validation, ATLAS cohorts and analytics, and a Dockerised Vantage6 federated data-owner node. Rwanda pilot with extension to African data-owner institutions.
This Certificate is a prerequisite for enrolment in the Postgraduate Certificate programme planned through the EpiScientia Network Project at the University of Rwanda – African Centre of Excellence in Data Science. The Certificate also serves as a prerequisite for joining the Heritage Alliance Ltd technical team supporting African health-data owners in data harmonization and OMOP CDM implementation. This includes contributing to source-data profiling, semantic and structural mapping, ETL development and validation, OMOP CDM deployment, data-quality assessment, and the preparation of local data-owner nodes for participation in federated data networks.
1 · Programme purpose and positioning
Why this certificate
The EpiScientia OMOP Data Engineer Certificate is an applied competency programme designed to develop data engineers who can lead local health-data harmonization at African data-owner institutions. The Rwanda implementation is the initial training setting, while the curriculum is intentionally portable to institutions using OpenMRS, OpenClinic GA, eFiche, eBuzima, MediSoft and registry datasets. The programme follows the source-to-target sequence: source inspection and profiling, structural mapping, semantic standardization, ETL implementation, data-quality validation, cohort construction, analytics, containerization and Vantage6 node operation. Certification is based on demonstrated engineering competence rather than attendance alone: a participant must be able to explain mapping decisions, implement and troubleshoot an ETL, evaluate data quality, construct an executable cohort, and demonstrate a controlled federated computation.
| Programme element | Design |
|---|---|
| Online instruction | 10 weekly live sessions; each 60 minutes (25 min theory + 35 min supervised practical) |
| Asynchronous assessment | 10 weekly quizzes/artefacts; approximately 20 minutes each |
| Face-to-face immersion | Two 2-day workshops in Weeks 5 and 10 |
| Core outcome | Independent conversion of a local EHR/registry source into a validated OMOP CDM data-owner environment capable of participating in Vantage6-mediated federated analytics |
| Assessment | Continuous assessment 60%; final practical examination 40% |
| Core reference | The Book of OHDSI, supplemented by current OMOP CDM and Vantage6 documentation |
Teaching staff
Course directors & assistants
Course director
Prof Dr Marc Twagirumukiza, MD, PhD
Course director
Assistant course directors
Carine Umulisa
Assistant course director
Daniel Nsanzabandi
Assistant course director
EpiScientia Mentor
Assistant course director
Topic facilitators
Carine Umulisa
Facilitator
- · Week 1 — From source EHR to OMOP CDM
- · Week 2 — Source-data profiling & ETL planning
- · Week 3 — Structural mapping to OMOP
- · Week 4 — Semantic harmonization & standardized vocabularies
- · Week 5 — ETL implementation & reproducibility
- · Week 6 — Data quality, validation & ETL improvement
- · Week 7 — ATLAS, WebAPI & cohort construction
- · Week 8 — Analytics on OMOP & AI-assisted mapping
- · Week 9 — Docker & Vantage6 data-owner node
- · Week 10 — Federated OMOP network operations
- · Workshop I — Build the OMOP Node
Daniel Nsanzabandi
Facilitator
- · Workshop II — From Local OMOP to Federated Node
The course director leads the programme and appoints the assistant course directors, who oversee the trainees' progress and the course operations. Facilitators lead individual weekly topics and workshops.
2 · Award and programme specifications
Specifications
| Award title | EpiScientia Certificate in OMOP Data Engineering |
| Target audience | Data engineers, database administrators, health-informatics specialists, software/data developers, data managers and technically oriented analysts working with clinical, surveillance or registry data |
| Delivery mode | Blended: weekly synchronous online instruction + asynchronous quizzes/artefacts + two in-person workshops |
| Duration | 10 consecutive teaching weeks |
| Live online contact | 10 hours total |
| Asynchronous quiz time | Approximately 3 h 20 min total |
| In-person workshops | 4 full training days |
| Recommended additional self-study | 2–4 hours/week, increasing around Weeks 4–6 and the capstone |
| Training data | Institution-owned data where governance permits; otherwise de-identified or synthetic source-like extracts |
| Reference CDM | Current supported OMOP CDM release selected by the programme technical lead and frozen for each cohort |
3 · Entry prerequisites
Who can join
| Required | Working knowledge of relational databases and SQL (SELECT, JOIN, GROUP BY, keys and basic data types). |
| Required | Ability to navigate a Linux or Windows development environment, install approved software, and manage files/directories. |
| Required | Access to a laptop capable of running the training toolchain or to an EpiScientia-provided remote training environment. |
| Required for site-based participants | Authorization to inspect the structure of their institutional source system and use an approved training extract. |
| Recommended | Experience with PostgreSQL, MySQL/MariaDB or another relational DBMS; Git; Python or R; and basic Docker concepts. |
| Not required | Prior OMOP/OHDSI or Vantage6 experience. |
4 · Programme-level learning outcomes
On successful completion, the certificate holder will be able to
- LO1Explain the OHDSI ecosystem, OMOP CDM architecture, domains, conventions and standardized vocabulary model.
- LO2Inspect an unfamiliar EHR or registry schema and produce a source-data inventory and profiling report.
- LO3Design structural source-to-target mappings and document transformation logic before coding.
- LO4Map local clinical codes and terms to appropriate standard concepts, documenting uncertainty and expert review.
- LO5Implement a reproducible ETL that preserves source provenance and populates a selected OMOP CDM release.
- LO6Run and interpret OMOP data-quality and characterization outputs and iteratively correct the ETL.
- LO7Configure/use WebAPI and ATLAS sufficiently to create concept sets, define cohorts and execute basic standardized analyses.
- LO8Critically use alternative and AI-assisted mapping approaches without delegating semantic accountability to automation.
- LO9Containerize/configure a local data-owner environment and connect an approved Vantage6 node to a training central server.
- LO10Independently demonstrate the end-to-end pathway: source → profile → map → ETL → OMOP → quality → cohort/analytics → federated node.
Teaching model — Each online week uses a fixed 60-minute live pattern: 25 minutes of focused theory followed immediately by 35 minutes of supervised practice. The practical is performed on the learner's own approved data where feasible, or on a common source-like training dataset. A 20-minute asynchronous quiz is linked to the same week's competency and is completed later at the learner's pace. Weekly assessment combines short knowledge items with evidence from the practical exercise so the certificate does not become a knowledge-only award.
6 · Ten-week curriculum map
Week by week
25-minute lecture
- 0–5 min: Why common data models and distributed evidence networks.
- 5–10 min: OHDSI ecosystem and role of OMOP CDM.
- 10–17 min: Core clinical tables and domains: PERSON, VISIT, CONDITION, DRUG, MEASUREMENT, OBSERVATION.
- 17–22 min: Source concepts, standard concepts and preservation of source values.
- 22–25 min: End-to-end architecture from local source to analytics/federation.
35-minute supervised laboratory
- 0–5 min: Open assigned OpenMRS/OpenClinic/eFiche/eBuzima/MediSoft/registry extract and inspect schema.
- 5–15 min: Identify patient, encounter, diagnosis, observation/lab and medication structures and keys.
- 15–25 min: Classify representative variables into likely OMOP domains/tables.
- 25–32 min: Trace one patient journey from source records to candidate OMOP events.
- 32–35 min: Save a one-page source-to-OMOP hypothesis for instructor review.
Practical focus: Explore source schemas; identify patients, encounters, diagnoses, observations, labs, medications; propose OMOP domains
Quiz / artefact (~20 min): CDM domains, source vs standard concepts, data types and a short source-to-target classification exercise. — OHDSI/CDM architecture and source→target reasoning
8–9 · In-person workshops
Two engineering sprints in Kigali
Workshop I · Week 5
Build the OMOP Node
Convert the first four weeks of mapping work into a working, documented OMOP ETL v1. The workshop is an engineering sprint, not a repetition of the online lectures.
| Time | Activity | Required output |
|---|---|---|
| 08:30–09:00 | Registration, environment check, dataset/governance confirmation | Working environments confirmed |
| 09:00–09:45 | Architecture recap and team ETL plan | Team implementation plan |
| 09:45–10:45 | Source schema clinic: WhiteRabbit findings and unresolved source questions | Reviewed source inventory |
| 10:45–11:00 | Break | |
| 11:00–12:30 | Rabbit-in-a-Hat structural mapping studio | Reviewed source-to-target design |
| 12:30–13:30 | Lunch | |
| 13:30–15:00 | Vocabulary/Usagi mapping studio | Versioned mapping file |
| 15:00–15:15 | Break | |
| 15:15–16:30 | Mapping peer review and clinical/semantic escalation | Mapping decisions + issue log |
| 16:30–17:00 | Day 1 checkpoint and ETL build plan | Day 2 task board |
Milestone package: OMOP ETL v1; ETL specification; vocabulary mapping file; source inventory/profile; Git repository or equivalent version-controlled package; populated local/training OMOP database; issue log.
Workshop II · Week 10
From Local OMOP to Federated Node
Close the quality loop, validate the cohort, connect the Vantage6 node and sit the final practical examination (Parts A–C) with a multi-node federated capstone.
| Time | Activity | Required output |
|---|---|---|
| 08:30–09:00 | Readiness check and capstone briefing | Individual/team capstone plan |
| 09:00–10:15 | DQD/Achilles rerun and quality triage | Updated quality report |
| 10:15–10:30 | Break | |
| 10:30–11:45 | ETL correction sprint and mapping closure | Priority defects resolved |
| 11:45–12:30 | ATLAS concept-set and cohort validation | Validated cohort definition |
| 12:30–13:30 | Lunch | |
| 13:30–14:30 | Docker stack verification and configuration review | Reproducible local stack |
| 14:30–15:30 | Vantage6 node configuration and central registration | Connected training node |
| 15:30–15:45 | Break | |
| 15:45–16:45 | Approved test computation and troubleshooting | Successful node execution |
| 16:45–17:00 | Production-readiness checklist | Capstone readiness decision |
10 · Assessment framework
Continuous assessment 60% · final practical examination 40%
| Component | Weight | Evidence |
|---|---|---|
| Weekly assessment 1–10 | 60% total (6% each) | Each week: short knowledge component plus practical artefact/result. The exact knowledge/artefact split may vary by week, but practical evidence must be present. |
| Final practical examination | 40% | Observed end-to-end engineering tasks completed in a controlled training environment during Week 10. |
Final practical examination rubric (40%)
| Domain | Weight | Competent performance |
|---|---|---|
| Source profiling & structural mapping | 10% | Accurate source interpretation; correct OMOP target selection; documented assumptions and transformations; traceability from source to target. |
| ETL implementation & reproducibility | 10% | Executable transformations; correct IDs/keys/dates; provenance; repeatability; logging/error handling; version-controlled implementation. |
| Vocabulary & semantic mapping | 8% | Appropriate standard concepts/domains; source preservation; evidence-based acceptance/rejection of candidates; explicit handling of uncertainty. |
| OMOP quality + cohort validation | 7% | Correct use/interpretation of quality outputs; remediation of material defects; executable concept set/cohort with defensible temporal logic. |
| Vantage6 deployment & federated execution | 5% | Correct node configuration; secure/approved data interface; successful connection and task execution; ability to interpret logs/results and troubleshoot. |
Performance descriptors
| 4 — Independent/Excellent | Completes the task correctly and reproducibly; explains design choices; detects edge cases; documentation is sufficient for another engineer to reproduce the work. |
| 3 — Competent | Completes the essential task correctly with minor non-material errors or limited prompting; decisions are technically defensible. |
| 2 — Developing | Partial completion; material errors, weak traceability or repeated prompting; output is not yet reliable for production use. |
| 1 — Insufficient | Cannot complete the essential task or makes unsafe/invalid mapping, ETL, quality or federation decisions. |
11 · Certification criteria
EpiScientia Certificate in OMOP Data Engineering
- Overall weighted mark ≥ 60%.
- Final practical examination mark ≥ 50% (competency gate).
- Submission of all mandatory capstone artefacts or approved equivalents.
- Participation in both in-person workshops unless a formally approved equivalent supervised practical assessment is provided.
- No critical failure in data-governance, security or semantic-mapping practice that would make the demonstrated node unsafe or analytically invalid.
- Participants who meet the overall mark but fail the practical gate receive a completion statement rather than the OMOP Data Engineer Certificate and may undertake a supervised reassessment.
Statement of competence — A successful EpiScientia OMOP Data Engineer Certificate holder has demonstrated the ability to independently take an approved clinical or registry source through source profiling, structural and semantic mapping, reproducible ETL implementation, OMOP quality validation, cohort construction and local analytics, and to configure the resulting data-owner environment for controlled participation in Vantage6-mediated federated analytics.
Trainer and assessor requirements
| Role | Minimum profile | Recommended responsibility |
|---|---|---|
| Programme scientific lead | Senior health informatics/OHDSI practitioner with demonstrable OMOP implementation experience | Curriculum oversight, standards, assessment moderation |
| Lead OMOP data engineer | Hands-on ETL engineer experienced with relational databases, OMOP conventions, WhiteRabbit/Rabbit-in-a-Hat and vocabulary mapping | Weeks 1–6; workshop engineering supervision |
| Vocabulary/clinical terminology expert | Experience with OMOP standardized vocabularies and clinical terminologies; clinical domain input strongly preferred | Week 4, mapping review, semantic adjudication |
| OHDSI analytics trainer | Practical ATLAS/WebAPI/cohort experience | Weeks 6–8; cohort/analytics assessment |
| Vantage6/DevOps trainer | Docker/Linux/networking plus practical Vantage6 deployment experience | Weeks 9–10; node configuration and troubleshooting |
| Teaching assistants | SQL/ETL competent; ideally 1 TA per 6–8 learners during workshops | Hands-on support without completing assessed work for learners |
| Independent assessor/moderator | Not the sole trainer responsible for the learner's capstone | Final practical moderation and borderline decisions |
13 · Technical and training infrastructure
Software stack & environments
| Layer | Required / recommended tools |
|---|---|
| Source profiling | WhiteRabbit |
| ETL design | Rabbit-in-a-Hat plus source-to-target specification templates |
| Semantic mapping | Usagi + OHDSI standardized vocabularies/Athena-derived vocabulary files as institutionally authorized |
| ETL execution | SQL as core; Python/R optional according to source and local engineering standards |
| OMOP quality/characterization | DataQualityDashboard; Achilles; optional ARES visualization |
| Analytics | WebAPI; ATLAS; SQL; selected OHDSI/HADES packages where relevant |
| Containerization | Docker |
| Federated execution | Vantage6 node + central training server + approved algorithm containers |
| Engineering governance | Git/version control, issue log, mapping decision log, release notes |
| Learner workstation | Modern 64-bit laptop; recommended ≥4 CPU cores, ≥16 GB RAM and ≥100 GB free SSD space (32 GB RAM preferable for heavier local Docker/database work). Administrative rights or a managed preconfigured image/remote desktop. Stable broadband (wired recommended for workshops). Git client; database client; approved code editor/IDE; Docker Engine/Desktop; Java runtime where required by the selected OHDSI tools; access to a PostgreSQL training database or equivalent DBMS. |
| Central training environment | Training OMOP database(s) and standardized vocabulary release frozen for the cohort. OHDSI WebAPI + ATLAS training deployment. Central Vantage6 training server with pre-created organizations/users/roles and a controlled algorithm image/registry workflow. Version-control repository for course ETLs, mappings, cohort definitions and assessment submissions. Backup/snapshot capability before each workshop and before the final examination. Separate training/sandbox environment from any production national data node. |
| Data-owner / site environment | Where learners work on institutional data, the source system is not modified by training activities: a read-only replica, approved extract or dedicated local data-engineering node is used. Patient-level data remains within the authorized institutional environment. Cross-site exercises use approved aggregate outputs or synthetic/de-identified training data according to local governance. |
Data governance, security and responsible AI
- Use only data for which the participant and training team have explicit institutional authorization.
- Prefer synthetic or de-identified datasets for cross-country teaching demonstrations; real patient-level data must not be copied into shared course infrastructure without a lawful and approved basis.
- Maintain separation between training and production credentials, servers and Vantage6 organizations.
- Record source-to-standard mapping decisions, including rejected and uncertain mappings, so that semantic choices are auditable.
- AI/LLM tools may propose mappings, transformation code or documentation, but their outputs must be validated against the selected OMOP vocabulary/CDM release and reviewed by a competent human.
- Patient-level data must not be entered into external generative-AI services unless explicitly authorized under applicable governance.
- Federated execution does not by itself guarantee privacy: algorithms, output disclosure rules, permissions and node configuration require governance and technical review.
15 · Mandatory learner portfolio
What every certificate holder hands over
- Source-system inventory and WhiteRabbit profiling report.
- Rabbit-in-a-Hat/source-to-target ETL design.
- Versioned vocabulary mapping file with review status.
- Executable ETL package and documentation.
- OMOP data-quality/characterization report and remediation log.
- Saved/exported ATLAS concept set and cohort definition.
- Docker/Vantage6 node configuration evidence with secrets removed.
- Federated test execution evidence and troubleshooting log.
- Final technical handover note describing how the local institution can maintain and rerun the pipeline.
Quality assurance — Each cohort freezes the OMOP CDM release, vocabulary release, tool versions and assessment dataset/environment before teaching begins. Assessment rubrics are calibrated by trainers before Week 5 and moderated after Week 10. Mapping disputes are escalated to a designated terminology/clinical reviewer rather than resolved by convenience. Course feedback distinguishes teaching quality from technical environment failures so that infrastructure problems do not distort competency judgments.
17 · Reference framework
Core references
- The Book of OHDSI — core conceptual and methodological reference, particularly the Extract-Transform-Load chapter.
- OHDSI OMOP Common Data Model documentation — current CDM specification and conventions.
- OHDSI standardized vocabularies and associated mapping tooling/documentation.
- DataQualityDashboard and Achilles documentation for quality and characterization.
- ATLAS/WebAPI documentation for cohort and analytical workflows.
- Vantage6 documentation for server, node, client, Dockerized algorithm execution and node configuration.
Ready to become an OMOP data engineer?
Enrol in the Certification Programme to follow this course with a scholarship, or browse the other courses.


