Data Scientist, Bayesian Statistics and Causal Inference (MS/PhD)
Company: Qureator Inc. / Curiochips
Location: Seoul, South Korea
Employment Type: Full-time (Alternative Military Service Eligible)
About Qureator
Qureator is a TechBio company developing AI and next-generation drug discovery platforms based on our proprietary 3D organ-on-a-chip technology, Curiochips®. Leveraging advanced Microphysiological Systems (MPS), we build highly predictive human disease models that closely recapitulate human physiology and pathology, enabling the evaluation of drug efficacy, mechanism of action, and disease biology in a physiologically relevant setting.
Our platform generates petabyte-scale, 3D fluorescence imaging data from human-derived tissues. By combining AI, computer vision, machine learning, and quantitative biology, we transform these datasets into actionable insights for target discovery, drug response prediction, and patient stratification. Through collaborations with leading pharmaceutical companies, biotechnology firms, hospitals, and research institutions worldwide, we are accelerating the development of transformative therapies for cancer and rare diseases.
The Role
This position is for a Data Scientist who will own the statistical and analytical layer of our platform. Our image analysis pipelines produce large numbers of quantitative and deep learning based phenotypes, and these arrive alongside the functional readouts measured in the same experiments. Your job is to participate in the process of process improvement, experiment design, and data analysis to arrive at sound scientific conclusions with quantified uncertainty.
You will build the models — hierarchical, Bayesian, frequentist — and you will also build the infrastructure those models run on, this is a full-stack role. Expect to stand up and maintain the R and Python analysis stacks, work with our data lake and cloud infrastructure, write production-quality pipeline code, and make analyses reproducible. Our other data scientists work primarily on computer vision, so this person will provide expertise in experimental design, statistical rigor, and everything adjacent to it: data wrangling, QC, visualization, reporting, and the occasional problem that doesn't fit a category.
You will work side by side with computer vision / AI engineers, wet-lab scientists, across Korea and US-based teams.
This is a unique opportunity to build the analytical foundation of a drug discovery platform, at the intersection of biology, imaging, and inference.
Key Responsibilities
Statistical modeling and analysis
- Analyze biological experiments end to end — from study design through estimation of treatment effects to the conclusions reported to project teams.
- Design and fit hierarchical models for image-derived phenotypes and the assay readouts collected alongside them, choosing Bayesian or frequentist approaches based on the question, the data, and the audience.
- Integrate imaging and non-imaging measurements into unified models, so morphological and functional evidence support a single interpretation.
- Quantify and propagate uncertainty from raw measurement through to efficacy and toxicity readouts.
- Apply causal reasoning to perturbation and screening data: adjust for confounding, distinguish genuine treatment effects from technical artifacts, and advise the team on what a given design can and cannot support.
- Maintain a principled modeling workflow — prior and posterior predictive checks, MCMC/variational diagnostics, model comparison, calibration, and appropriate handling of multiplicity and power.
- Establish assay quality statistics, batch-effect correction strategies, and reproducibility standards across the platform.
Engineering and infrastructure
- Build and maintain the R and Python analysis stacks, including reproducible pipeline tooling ({targets}, Kedro) and versioned, documented analysis code.
- Work with our data lake and cloud environment (GCP, object storage, orchestration via Airflow) to get experimental and imaging data into an analyzable state.
- Turn recurring analyses into pipelines and dashboards the rest of the company can use without you in the loop.
- Partner with computer vision / AI engineers to characterize measurement error in segmentation and tracking outputs, and to fold that error into downstream inference.
Collaboration
- Collaborate closely with wet-lab scientists to shape experimental design and data acquisition before experiments are run, not after.
- Communicate uncertainty and statistical reasoning clearly to biologists, engineers, and external pharma partners.
- Stay current on advances in statistical methods, causal inference, and quantitative bioimage analysis.
-
Qualifications
- MS or PhD in Statistics, Biostatistics, Applied Mathematics, Physics, Computer Science, Bioengineering, Computational Biology, or a related quantitative discipline.
- Strong foundation in applied statistics, including experimental design, mixed-effects / hierarchical modeling, and inference under multiple sources of variation. Comfortable working in both Bayesian and frequentist frameworks and choosing between them deliberately.
- Practical Bayesian experience: prior specification, MCMC or variational inference, model checking and comparison, using at least one probabilistic programming framework (Stan, PyMC, NumPyro/JAX, brms, or similar).
- Familiarity with causal inference concepts — DAGs, confounding and identification, potential outcomes — and the judgment to know when an observational analysis supports a causal claim.
- Proficiency in Python and R and their scientific computing stacks (NumPy, SciPy, pandas, scikit-learn, ArviZ, tidyverse). Deep expertise in R and working fluency in Python is acceptable.
- Software engineering fundamentals: git, code review, writing analysis code others can run and reuse.
- Comfort working in cloud environments (GCP or AWS) and with data stored outside a local filesystem — object storage, data lakes, databases. You should be willing to build and maintain your own infrastructure rather than wait for someone else to hand you clean data.
- Working familiarity with machine learning and computer vision concepts — enough to reason about what image analysis pipelines produce and where their measurements go wrong.
- Ability to work independently on ambiguous scientific and technical problems in a fast-paced research environment.
- Strong communication and collaboration skills in English, including the ability to explain statistical reasoning to non-statisticians.
Preferred Skills
- Experience with microscopy, high-content screening, or bioimage analysis data.
- Experience modeling high-dimensional biological data.
- Experience with workflow orchestration and reproducible-analysis tooling — Airflow, Nextflow, Kedro, {targets}, Snakemake, or similar.
- Advanced causal methods: causal discovery, causal representation learning, or causal analysis of large-scale perturbation screens.
- Background in Bayesian optimization or active learning for experimental design.
- Exposure to preclinical study design or nonlinear mixed-effects modeling.
- Experience with MLOps practices, containerization, or CI/CD for analysis pipelines.
- Publications or a project portfolio demonstrating relevant research or technical contributions.
-
Benefits and Culture
We are committed to creating an environment where you can do your best work. We offer a competitive salary and a comprehensive benefits package, including:
- Hybrid and flexible work models.
- Support for professional development, including conferences and education.
- An open, collaborative, and horizontal organizational structure.
- Close collaboration with US-based teams in a global working environment.
- Alternative Military Service Program Available.
-