Data Scientist
Headwater Science · Research Triangle Park, North Carolina · United States · Hybrid
Posted Aug 31, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About Headwater Science
Headwater Science (formerly NoviSci) is a data science and methods company specializing in principled, reproducible evidence generation for complex clinical and regulatory challenges. With deep expertise in comparative effectiveness, causal inference, healthcare utilization and expenditure research, and regulatory-grade analytical software, Headwater Science provides the methodological foundation that delivers reproducible analytic pipelines, novel epidemiologic and statistical methods, and regulatory-grade software validated to hold up under the most demanding scrutiny. The company works with life sciences organizations as a long-term scientific partner. Headwater Science is a Highlander Health company. Learn more at headwaterscience.com .
The Role
In this role, you will support our research projects by helping to transform source data from healthcare databases into analytic-ready data sets, performing statistical analyses, and generating reports of the results. You will work with small teams of epidemiologists and statisticians whose responsibilities span study design and execution. Your work will focus primarily on executing specifications outlined in study protocols/statistical analysis plans (SAP) and building reproducible analytical pipelines that are understandable, well documented, and compliant with quality control standards. Typical projects include studies of natural history, treatment patterns, and comparative effectiveness/safety, often incorporating negative controls to evaluate treatment group comparability.
What You’ll Do
How We Work: You'll spend most of your time writing R code to execute research projects — building cohorts and running the analyses that produce results. You'll work alongside team members, coordinating through GitLab and staying in close contact as the code comes together. Project code runs on remote servers where the data live — usually ours, but sometimes a client or data provider environment. A research project moves through two main phases — cohort building, then analysis and reporting — and you'll work in both.
Cohort Building: This is the largest piece of the role – turning raw healthcare databases into analytic-ready data sets. You will work with deidentified administrative claims and electronic health record (EHR) data, which come as relational tables covering patient demographics, diagnoses, procedures, medications, and lab results. Each data source has its own conventions, and learning them is critical to being successful in this role. In practice, you will:
Draw on domain knowledge of how each database represents clinical events to make appropriate choices when building study variables.
Use dbplyr to query source databases from R, generating SQL against tables that are often too large to hold in memory.
Apply eligibility criteria from a study protocol and/or SAP to identify the study population.
Determine the index date and start of follow-up for each…