Own the eval strategy for North: define what we need to measure across agent workflows, tool use, enterprise knowledge work such as deep research, document creation or editing, and other human-AI interactions.
Build high-quality evals from the realities of the product: user feedback, production failures, privacy-preserving usage logs, internal dogfooding, customer needs, and forward-looking product goals.
Create systems that continuously turn what North is learning from users, customers, and feature teams into evals, so measurement keeps pace with the product rather than becoming a static benchmark.