At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley!
Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation, we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us!
About The Lilly And NVIDIA Partnership
Lilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.
Position Summary
The Chemical Reactivity Landscape project is building a quantitative, predictive map of how reaction outcome depends on substrate, catalyst, and conditions across the chemistry Lilly runs. We are looking for a scientist to own the modeling and analysis layer of that map: turning high-throughput experimentation (HTE) and reaction condition data into models that tell project chemists which conditions to run next, and why.
This is a hands-on individual contributor role reporting to the Scientific Project Leader for the Chemical Reactivity Landscape project. You will work at the interface of quantum chemistry, machine learning, and experimental reaction data, computing descriptors that encode steric and electronic effects, fitting and validating models against real plate data, and closing the loop between what can be calculated about a molecule and what is observed in the lab. You will be collocated with chemists who run the automation platform to quickly assess the quality of your models and how the analytical readouts behind them are produced.
The role is primarily computational and lab-adjacent, based on-site in South San Francisco. You will own significant components of the reactivity modeling stack and be trusted to make technical decisions within them.
Key Responsibilities
Reaction Data Modeling & Analysis
- Build and validate models that predict reaction outcome, yield, selectivity, conversion, impurity profile from substrate structure, catalyst and ligand identity, and condition variables
- Analyze HTE datasets end to end: representation choice, feature engineering, cross-validation design, uncertainty quantification, and honest out-of-domain assessment
- Turn plate-level output into structured, model-ready reaction records with consistent condition encodings, so that campaigns compound into a reusable reactivity dataset rather than isolated screens
- Characterize what the models do not know: identify under-sampled regions of condition space and specify the experiments that would most reduce uncertainty
- Detect and diagnose reactivity cliffs and mechanistic regime changes in the data rather than smoothing them away
Quantum Mechanical & Physical Organic Modeling
- Run DFT and semi-empirical calculations on substrates, catalysts, and intermediatesto quantify steric and electronic effects and to assess whether a transformation is energetically favorable
- Generate atom-, bond-, and molecule-level QM descriptors and combine them with structural representations so that models are grounded in physical organic chemistry rather than correlation alone
- Handle conformational sampling, barrier and transition-state estimation, and solvation and counter-ion effects at the level of theory the question actually warrants
- Bridge from what is calculated about a molecule to what is measured in the plate, and treat the mismatch between the two as information about the model, the mechanism, or the measurement
- Automate QM workflows so descriptor generation scales to thousands of structures without hand-holding
Analytical Chemistry & Data Quality
- Work fluently with the analytical readouts behind reaction data, UPLC/LC-MS, HPLC-UV, NMR, GC, and understand how integration, response factors, ionization behavior, and internal standards shape the numbers a model is being fit to
- Build automated processing for analytical output: peak assignment, calibration, yield and purity calculation, and QC flags that catch bad wells before they reach a training set
- Quantify measurement uncertainty and reproducibility, and propagate them into model training and downstream decisions
- Partner with analytical scientists to improve how reaction data is captured at source, so quality is designed in rather than corrected afterward
Optimization, Benchmarking & Evaluation
- Design and run optimization campaigns over categorical and mixed condition spaces using design of experiments, Bayesian optimization, active learning, and LLM-guided approaches — and choose the right method for the problem rather than defaulting to one
- Build evaluation harnesses for the models and agents in use, including chemistry reasoning and prediction tasks before they are allowed to influence real experiments
Tooling, Agents & Delivery
- Implement version control on datasets and models, log campaign trajectories, and keep results traceable and reproducible by someone else
- Apply appropriate guardrails, human-in-the-loop checkpoints, and safety review to any system that proposes or triggers experimental work
Collaboration & Scientific Engagement
- Partner with the Synthetic Innovation Team, medicinal chemistry, HTE, analytical, and computational chemistry colleagues to turn portfolio bottlenecks into well-posed modeling questions
- Present results and their limitations clearly to synthetic chemists who will act on them and to computational colleagues who will scrutinize them
- Represent the work externally through publications, preprints, conferences, and academic collaborations, and scout emerging methods worth adopting internally
- Contribute to the shared reactivity datasets, descriptor conventions, and modeling standards that other LSMD projects build on
What Success Looks Like
- Lilly has a queryable reactivity landscape that project chemists consult by default when selecting conditions, and that they trust because its uncertainty estimates hold up
- Optimization campaigns reach target conditions in measurably fewer experiments than expert-designed baselines
- QM-derived descriptors demonstrably improve prediction over structure-only models on internal data, with the improvement explainable in chemical terms
- HTE data from routine campaigns flows into the landscape automatically, with QC catching bad data before it trains anything
- Benchmarks are reproducible, baselines are fair, and failure modes are documented rather than discovered by someone else
Basic Qualifications
- Ph.D. in Chemistry, Chemical Engineering, Computational or Physical Organic Chemistry, Cheminformatics, or a related field (or M.S. + 3 years / B.S. + 5 years of equivalent experience)
- Demonstrated experience modeling and analyzing reaction data — HTE, reaction condition, or reaction outcome datasets with rigorous validation and clear treatment of uncertainty
Preferred Qualifications
- Hands-on experience applying quantum chemical methods (DFT and/or semi-empirical) to organic reactivity, including descriptor generation for downstream modeling
- Working understanding of analytical chemistry as it applies to reaction data (LC-MS, HPLC, NMR) and of how analytical practice affects data quality
- Strong scientific Python: RDKit, scikit-learn, pandas, and a deep learning framework such as PyTorch, with version control and reproducible workflows
- Effective English-language written and verbal communication skills, with the ability to explain modeling results to experimental chemists
- Experience with Bayesian optimization, active learning, or LLM-guided optimization applied to reaction conditions, supported by published or internal benchmarks
- Experience building LLM agent or multi-agent systems that call real tools — literature and documentation retrieval, code execution, calculation dispatch
- Familiarity with HTE platforms, automated parallel synthesis, and ELN/LIMS data structures
- Experience with graph neural networks or learned molecular and reaction representations
- Experience with model evaluation and benchmarking methodology, including safety review and failure-mode analysis
- Physical organic chemistry intuition: linear free energy relationships, steric and electronic parameterization, and mechanism-driven hypothesis generation
- Familiarity with catalysis, metal-mediated, organocatalytic, or photoredox, and with the condition variables that matter in each
- Publications or preprints in leading chemistry, cheminformatics, or machine learning venues
- Experience working in multidisciplinary teams spanning experimental and computational science
Additional Information
- Willing and able to work onsite in South San Francisco, CA
- Travel: 0–10%. This is a primarily computational, lab-adjacent role. You will work alongside laboratory teams and may spend time in research laboratory spaces to observe and understand experimental workflows; standard personal protective equipment is required in those settings. The majority of the work is performed in an office and computing environment.
- This job description is intended to provide a general overview of the job requirements at the time it was prepared. Job requirements may change over time and may include additional responsibilities not specifically described in the job description.
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.
Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.
Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).
Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is
$168,000 - $268,400
Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.
#WeAreLilly