Bio-Curation Scientist
Oncology Data, AI & Knowledge Systems
Location: Bengaluru, India
Experience: 0–3 years
Employment type: Full-time
Work model: Hybrid / Remote
Build the data foundation for better cancer intelligence
We are looking for a curious, detail-oriented Bio-Curation Scientist to help transform complex oncology evidence into high-quality, structured, machine-readable data.
In this role, you will work at the intersection of cancer biology, scientific literature, clinical research, and applied AI. You will extract and validate evidence from publications, clinical trials, and regulatory drug labels; standardize it using biomedical ontologies and controlled vocabularies; and help build trusted datasets that support search, analytics, and AI-enabled knowledge systems.
This is an excellent opportunity for an early-career scientist who enjoys both biology and technology and wants to build practical skills in Python, data curation, APIs, LLM-assisted workflows, and biomedical data standards.
What you’ll work on
· Curate oncology, molecular biology, biomarker, drug, and clinical-trial information from peer-reviewed literature, clinical-trial registries, regulatory labels, and trusted biomedical databases.
· Convert unstructured scientific evidence into structured, traceable, high-quality records for our internal knowledge platform.
· Review and validate AI/LLM-assisted extraction outputs, ensuring scientific accuracy, completeness, consistency, and source-level provenance.
· Apply controlled vocabularies, ontologies, metadata standards, and FAIR data principles to make information findable, interoperable, reusable, and ready for analytics and machine-learning applications.
· Standardize entities and relationships across genes, variants, biomarkers, diseases, drugs, indications, mechanisms of action, clinical trials, and evidence sources.
· Use Python notebooks, APIs, JSON/CSV files, and data-quality checks to support scalable curation workflows.
· Identify inconsistencies, missing information, and ambiguous evidence; work with subject-matter experts and technical teams to resolve them.
· Partner with data scientists, data engineers, and product or pipeline owners to improve curation tools, schemas, automation, and data quality.
· Document curation decisions, data definitions, and workflow improvements so that the knowledge base remains transparent and reproducible.
What you’ll bring
Required
· Bachelor’s or Master’s degree in Bioinformatics, Biotechnology, Molecular Biology, Biomedical Sciences, Pharmacology, Life Sciences, Computational Biology, or a related discipline.
· Strong fundamentals in oncology, cancer biology, cell biology, molecular biology, genetics/genomics, pharmacology, or drug discovery.
· 0–3 years of relevant experience through industry work, academic research, internships, dissertation/thesis work, scientific writing, or data-curation projects.
· Strong ability to read, interpret, and summarize scientific publications accurately.
· Basic-to-intermediate Python skills, including working with tabular data and structured files.
· Familiarity with Jupyter Notebooks and Git/GitHub or another version-control system.
· Comfort working with structured and unstructured data, including CSV, Excel, JSON, text documents, and scientific literature.
· High attention to detail, strong ownership, and the ability to follow data standards consistently.
· Clear written and verbal communication skills in English.
Nice to have
· Experience with biomedical or clinical databases such as PubMed, ClinicalTrials.gov, DrugBank, ChEMBL, UniProt, COSMIC, CIViC, OncoKB, or similar resources.
· Familiarity with biomedical ontologies and identifiers, such as MeSH, HGNC, NCBI Gene, Disease Ontology, ICD, SNOMED CT, RxNorm, or MONDO.
· Experience using REST APIs, web scraping responsibly, or automating repetitive data-processing tasks.
· Exposure to SQL, regular expressions, named-entity recognition, knowledge graphs, or data-quality frameworks.
· Familiarity with LLMs, prompt design, retrieval-augmented generation, or human-in-the-loop AI evaluation.
· A publication, thesis, internship, GitHub project, or portfolio demonstrating scientific analysis, curation, or Python-based data work.
Why this role is exciting
· Work on real-world oncology and precision-medicine data—not generic data-labeling tasks.
· Build a rare hybrid skill set spanning life sciences, structured data, AI-assisted research, and data quality.
· Learn directly from interdisciplinary teammates across biology, data science, engineering, and product.
· Gain hands-on exposure to practical tools used in modern data teams: Python, APIs, version control, notebooks, data standards, and AI-assisted curation.
· See your work become part of a reusable knowledge platform that helps make complex scientific evidence easier to discover and use.
· Take ownership early: your ideas on improving extraction, validation, documentation, and workflow automation will be valued.
What success looks like
In your first 3–6 months, you will be able to independently curate defined oncology evidence types, apply the team’s data standards reliably, use Python and APIs to support your work, and contribute ideas that improve data quality or reduce manual effort.
How to apply
Please fill out this form and submit your resume here: https://bit.ly/46gJFia