In a nutshell LangWatch builds tools that help AI teams understand, debug, and improve their LLM products. We're growing fast and looking for an AI Engineer Intern to own our benchmarking work: measuring how models and agents actually perform, on real tasks, and publishing what we find.This is a paid internship (Amsterdam-based) for 6 months with a strong chance of conversion for outstanding performance.\

About LangWatch LangWatch is an LLMOps platform for teams building with large language models. We help companies understand how users engage with their LLM features, what's working, and where to improve, enabling faster iteration and better user experiences. Our platform makes it easier to monitor, evaluate, and optimize AI products, closing the gap between proof-of-concept and reliable production.Our core is open source, we're backed by great VCs, and thousands of developers use what we ship. Now we're opening a hands-on internship for someone who wants to answer the question every AI team is asking: which model, which prompt, which setup, and how do we know?What you will be working onNew models ship every week and every vendor claims their own benchmark. Teams building real products still can't answer whether a switch would help them. That gap is your project.

You'll work closely with our CTO and the engineering team. Expect real ownership, code review that makes you better, and results that get read outside the company.

You'll learn how to:

Who should apply