AI Jobs Map

Quik Hire Staffing · India

Software Engineer – Evaluation (Remote)

Remotemid_levelfull timePosted today
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmgithubpythonjavascriptjavarustc++c#rubygitdockerunit-testing

- Role: Software Engineer – Evaluation (Remote)

- Location: Remote (Work from Anywhere)

Role Overview:

We are hiring for one of our clients, seeking a Senior Software Engineer – LLM Evaluation & Repository Validation to work on a Full-Time basis. The role focuses on building evaluation and training datasets that enable large language models to solve realistic software engineering problems. The work involves creating verifiable software engineering tasks from public repository histories using a synthetic approach with human‑in‑the‑loop and expanding coverage across programming languages and difficulty levels.

Key Responsibilities:

• Analyze and triage GitHub issues across trending open‑source libraries with 500+ stars.

• Set up and configure code repositories, including Dockerization and environment provisioning.

• Evaluate unit test coverage and quality for selected repositories.

• Modify and run codebases locally to assess large language model performance in bug‑fixing scenarios.

• Collaborate with AI researchers to design and select repositories and issues that challenge model capabilities.

• Lead junior engineers on project tasks, providing guidance on repository validation workflows.

• Automate development environment setup and maintain reproducible pipelines.

Required Skills & Qualifications:

• Strong experience in at least one language such as Python, JavaScript, Java, Go, Rust, C/C++, C# or Ruby, demonstrated through professional software development work.

• Proficiency with Git for source control, including branching, merging, and issue management.

• Hands‑on experience with Docker and basic software pipeline configuration to create reproducible environments.

• Ability to navigate complex, well‑maintained codebases and understand project architecture.

• Experience running, modifying, and testing real‑world projects locally to evaluate code behavior.

• Prior contribution to or evaluation of open‑source projects is considered a plus.

• Familiarity with unit testing frameworks and assessment of test coverage quality.

• Exposure to large language model research or evaluation projects is desirable but not required.

More About the Opportunity:

This role supports a global leader in the software development and AI research space, contributing directly to the creation of datasets that improve how language models interact with real code. The work combines practical software engineering with AI evaluation, influencing future AI‑assisted development tools.

Equal Opportunity Employer:

We hire based on skills and expertise. All qualified candidates are welcome regardless of background, experience, or prior employment history. Applications are reviewed solely on demonstrated technical ability and qualifications.

Apply Now!

More jobs at Quik Hire Staffing