Research Informatics Software Engineer
Level : MS Level / Early Career
Location : New York City — primarily onsite with periodic remote flexibility.
About the Role
We’re looking for an MS-level Computer Science candidate to join a Research Informatics / R&D IT team at the intersection of scientific software, data engineering, cloud infrastructure, and AI.
This is an opportunity to help build the digital foundation supporting modern drug discovery — from cloud-native data platforms and laboratory informatics systems to data pipelines, LLM/RAG applications, and emerging agentic AI capabilities. The ideal candidate is technically strong, hands-on, and excited to apply modern software engineering and AI technologies to real-world scientific problems.
What You’ll Do
- Build, integrate, and support LIMS, ELN, and analytical informatics platforms, working across data models, APIs, workflows, and scientific data flows.
- Design scalable data pipelines and APIs that make scientific data FAIR, high-quality, and machine-actionable for researchers and AI systems.
- Develop and support cloud-native infrastructure, primarily in AWS, using technologies such as Docker, Kubernetes, CI/CD, and workflow orchestration.
- Prototype and productionize GenAI and agentic AI applications, including LLM agents, RAG/GraphRAG, retrieval pipelines, and multi-agent workflows.
- Partner with scientists and engineering teams to translate research needs into reliable software, data products, and AI capabilities.
- Apply strong software engineering practices around testing, schema design, performance, observability, and production deployment.
- Help continuously improve the digital laboratory environment and its readiness to support increasingly AI-driven workflows.
What We’re Looking For
- Master’s degree in Computer Science or a closely related technical field.
- Strong programming skills in Python and SQL.
- Experience building data pipelines, scientific workflows, or backend/data applications.
- Hands-on experience with AWS, Docker/Kubernetes, PostgreSQL, and modern data engineering tools.
- Exposure to AI/ML and LLM technologies, including RAG, retrieval, or multi-agent systems through coursework, projects, research, or professional experience.
- Experience using modern AI coding assistants / coding agents such as Cursor, Claude Code, GitHub Copilot, or similar tools.
- Strong interest in learning and working with scientific applications, laboratory data, and commercial informatics platforms.
- Ability to work collaboratively across software engineering, data, and scientific teams.
Nice to Have
- Hands-on experience with LLM/agentic AI systems, RAG, GraphRAG, knowledge graphs, or multi-agent architectures.
- Experience with PyTorch, Airflow, Prefect, or related ML/data tooling.
- Experience with high-performance ML or scientific computing.
- Familiarity with Neo4j or other knowledge-graph technologies.
- Exposure to commercial LIMS/ELN or analytical platforms such as Genedata, CDD Vault, Virscidian Analytical Studio, or similar.
- Experience with production-grade software engineering, including CI/CD, testing, API development, schema design, and performance optimization.
- Interest in applying modern AI and data engineering to laboratory workflows and drug discovery.