JobRaahGet matched free

Jobs

Principal Data Scientist (NLP + Applied AI)

wiley · Remote, GBR · Remote

Posted Sep 17, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Job Description: We believe in bold ideas, diverse perspectives, and the drive to transform knowledge into impact. Here, your curiosity fuels progress, your voice shapes innovation, and your ambition helps redefine what’s possible within science and learning. We are a culture that obsesses over impact, challenges, and drives what’s next to power infinite possibilities for our customers, colleagues and society at large. About the Role: About the role   We're   building the systems that turn one of the world's largest scientific corpora into   research   intelligence. That means   production   NLP pipelines running over millions of journal articles, extracting entities, classifications, claim tuples, and summaries   optimized   for use by downstream agentic applications .   We're   looking for a senior data scientist to own   domain-specific   content   modeling work end to end, from the eval set through the pipeline stage that ships it.   You'll   join a small, senior team where data scientists own their models in production.   You'll   write the code, own the evaluations, ship the changes, and stay accountable for the outcomes. This is a hands-on role for someone who wants to see their models through to real users   in a rapidly evolving market .   What   you'll   do   Design and build NLP enrichment pipelines that extract entities, classifications, claims, and summaries from scientific   full-text   at scale.   Compare NLP approaches to extraction and enrichment against LLM-based   approaches, and   pick the right tool for each task. That means putting traditional NLP (NER, sequence labeling, classification), embedding-based retrieval, LLM prompting, and fine-tuned smaller models on the same table, and defending each choice with evaluation, cost, and operational tradeoffs. This is a core part of the job, not an occasional exercise.   Own evaluation. Build the golden sets   in consultation with SMEs and vendors, choose the metrics, and make productive tradeoffs between speed, quality, and cost.   Write production-quality Python. Manage concurrency and cost for high-volume LLM workloads. Structure code that engineers can   ship   and other data scientists can extend.   Collaborate with a team of data engineers to o rchestrate work in data   pipeline   and data build tools like Airflow and   Dagster . Design idempotent,   retryable , evaluable pipeline stages that stay reliable when a run fails at scale.   Contribute to agentic AI application work: tool-using systems that reason over the enriched corpus, where your NLP and evaluation background will shape how the agent grounds and defends its answers.   Work directly with editors, product managers, and engineers. Bring the modeling perspective into product   decisions, and   translate stakeholder   pushback   into concrete modeling work.   What   you'll   bring   Deep Python.   You've   written it in production, at scale, for years. You know when to…