AI Jobs Map

wellplayd · Hamburg, Germany

AI Dev

Hybridentry_levelfull timePosted 2 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

typescriptllmpostgresqlraggitkubernetesdockergdpropenai

AI Developer (m/f/d) — wellplayd, Hamburg (Hybrid)

We are building production AI into a compliance platform used by sports clubs, associations, and NGOs — and an internal agent fleet that already ships real tickets. Not a research lab. Not a chatbot demo. Code that has to hold.

About wellplayd

wellplayd is a growing startup in good governance and compliance. We help sports clubs, associations, and NGOs implement social standards in a digital, practical way: background checks, reporting channels, case management, submissions, digital signatures, and more. Our vision is that through safe and ethical sports structures, people can experience fairness, tolerance, and solidarity in and through sport.

The role

You will design, build, and operate AI systems that sit next to a live product: TypeScript/Node in the monorepo, Go for the agent fleet, LLMs in production. You own the path from prompt and eval to a merged change — including the cases where the model is wrong and the product still has to be right.

What you will do

- Ship AI features into the product (document understanding, case/triage assistance, structured extraction) with clear evals, not vibes

- Build and harden our agent stack: planning, tool use, review loops, blame/retry, deploy gates

- Write production TypeScript and Go; wire models to APIs, Postgres, and existing services

- Treat prompts, tools, and traces as code: versioned, tested, reviewable

- Keep humans in control on anything that touches people, cases, or compliance data

- Work in a small team: pick up a ticket, ask what’s unclear, land exactly what was agreed

Requirements (must-have)

- 3+ years shipping production software (backend or full-stack), not only notebooks or prototypes

- Hands-on LLM work in a real product: agents, tool calling, structured output, or RAG — you can show the repo or the feature

- Strong TypeScript and/or Go; you can read and change both

- Solid APIs, Postgres, git, and CI; you debug a failing run before you rewrite the prompt

- You evaluate models: fixtures, golden cases, regression checks, cost/latency, failure modes

- You write for production: types, tests, logs, no silent “the model said so”

- English for engineering; German is a plus (our users and partners are in DACH)

- You can work hybrid from Hamburg (Schanzenviertel)

Nice to have

- Built a multi-agent or ticket-driven agent pipeline (Linear/Jira + git + review)

- Kubernetes / Docker / image-level deploys

- Experience with sensitive data (GDPR, case files, background checks) or compliance products

- Prompt/eval tooling (LangSmith, Braintrust, or your own harness)

- Cursor / Claude / similar coding agents as part of how you ship, not as a substitute for judgment

- Interest in sport, NGOs, or social impact

What this is not

- Fine-tuning research with no product deadline

- “Add ChatGPT to the app” with no evals

- A role where you only write prompts and someone else owns the code

What we offer

- Real ownership in a small impact startup — your changes hit users and the backlog, not a slide deck

- Hybrid work, flexible hours, office in Hamburg’s Schanzenviertel

- A product that has to be careful: people, cases, and trust, not engagement metrics

- Room to shape how AI is used here: stack, evals, and the line we will not cross

How to apply

- Send a short note — and especially why wellplayd and this role — to [email protected]. Link a repo, PR, or feature where you put an LLM into production and had to make it reliable.