In a nutshell
LangWatch builds tools that help AI teams understand, debug, and improve their LLM products. We're growing fast and looking for an AI Engineer Intern to own our benchmarking work: measuring how models and agents actually perform, on real tasks, and publishing what we find.This is a paid internship (Amsterdam-based) for 6 months with a strong chance of conversion for outstanding performance.\
About LangWatch
LangWatch is an LLMOps platform for teams building with large language models. We help companies understand how users engage with their LLM features, what's working, and where to improve, enabling faster iteration and better user experiences. Our platform makes it easier to monitor, evaluate, and optimize AI products, closing the gap between proof-of-concept and reliable production.Our core is open source, we're backed by great VCs, and thousands of developers use what we ship. Now we're opening a hands-on internship for someone who wants to answer the question every AI team is asking: which model, which prompt, which setup, and how do we know?What you will be working onNew models ship every week and every vendor claims their own benchmark. Teams building real products still can't answer whether a switch would help them. That gap is your project.
You'll work closely with our CTO and the engineering team. Expect real ownership, code review that makes you better, and results that get read outside the company.
- Benchmark design: Build task suites that reflect what our users actually do, including agentic tool use, structured output, retrieval, and long-context work, rather than what looks good on a leaderboard.
- Harness engineering: Build and maintain the infrastructure to run benchmarks reproducibly across model providers, track cost and latency alongside quality, and re-run everything when a new model drops.
- Evaluation methodology: Work on the hard part, which is scoring. LLM-as-judge calibration, inter-rater agreement, variance across runs, and knowing when a difference is real.
- Analysis: Turn raw runs into findings that survive scrutiny, with error bars and honest caveats.
- Publishing: Write up results as reports, posts, and open datasets or repos that the community can check and reproduce.
- Product feedback: Feed what you learn back into LangWatch's evaluators, our gateway's model routing, and the guidance we give customers on model selection.
You'll learn how to:
- Design an evaluation that measures the thing you care about instead of a proxy for it.
- Tell a real improvement from noise, and defend the difference.
- Benchmark non-deterministic systems reproducibly, and version everything so results stay comparable.
- Reason about the quality, cost, and latency trade-off the way a team shipping to production has to.
- Publish technical results in the open and handle the feedback that comes with it.
Who should apply
- You're studying CS/AI/ (or self-taught with strong projects) and can commit 6 months (full-time preferred; part-time 24 to 32h/week possible).
- You can write code. Python is essential here, and you're comfortable with data analysis in it.
- You have a statistical instinct: sample sizes, variance, and significance are not new words.
- You've built something with LLMs, even if it was small or broken.