Requirements
Must-haves
- 4+ years of data research, applied analytics, or evaluation experience
- Experience with LLMs, prompting, RAG, and retrieval concepts
- Proficiency with Python
- Proficiency with SQL
- Experience in maintaining high-quality datasets or evaluation programs
- Ability to navigate ambiguous problems independently
- Strong communication skills in both spoken and written English
What makes you stand out
Fully own and drive the quality of the evaluation system
Nice-to-haves
- Startup experience
- Experience with LLM-as-a-judge evaluation
- Familiarity with (LangSmith, Ragas, DeepEval, Braintrust, Evidently, etc.)
- Knowledge of embeddings, search, reranking, or hybrid retrieval
- Experience creating evaluation rubrics, dashboards, or lightweight automation
- Bachelor's Degree in Computer Engineering, Computer Science, or equivalent
What they'll work on
- Own the quality of the LLM and Retrieval-Augmented Generation (RAG) evaluation system
- Run structured LLM and RAG evaluations and perform root-cause analysis
- Distinguish between retrieval, ranking, source, prompt, model, and citation failures
- Maintain evaluation datasets, including quality reviews, versioning, coverage, and new production cases
- Refine prompts, judge rubrics, and LLM-as-a-judge workflows
- Design controlled experiments with clear baselines, metrics, and success criteria
- Investigate practical solutions to recurring RAG challenges
- Automate repetitive analysis (e.g., Python, SQL, notebooks, APIs)
- Communicate findings and recommended actions to Product, Engineering, and leadership
AI Data Researcher | Python, LLM, RAG
3k-5k USD/month | 4+ years of exp. | C1 English
necessary to complete the registration
https://app.onstrider.com/r/michelecipriano?job=YWktZGF0YS1yZXNlYXJjaGVyLTNjMGE3ZDgxP3JlZmVycmFsPW1pY2hlbGVjaXByaWFubw==