Design and run evaluations for your product area: reference tests, heuristics, model-graded checks tailored to search relevance, chat quality, document understanding, or audio performance.
Define and track metrics that matter: task success, helpfulness, hallucination proxies, safety flags, latency, cost.
Own prompt and orchestration design: write, test, and iterate on prompts and system prompts as a core part of your work.
Run A/B tests on prompts, models, and configurations; analyze results; make rollout or rollback decisions from data.
Set up observability for LLM calls: structured logging, tracing, dashboards, alerts.
Operate model releases: canary and shadow traffic, sign-offs, SLO-based rollback criteria, regression detection.
Improve core behaviors in your product area, whether that's memory policies, intent classification, routing, tool-call reliability, or retrieval quality.
Create templates and documentation so other teams can author evals and ship safely.
Partner with Science to diagnose regressions and lead post-mortems.
About you
3-4 years of experience; backgrounds that fit well include ML engineers moving closer to product, or software engineers with real AI/ML production experience.
Strong TypeScript or Python skills - we have both tracks depending on team fit.
Production LLM experience: prompts, tool/function calling, system prompts.
Hands-on with evals and A/B testing; you can design metrics, not just run them.
Comfortable implementing directly in product code, not only notebooks.