#evaluation
-
2026-03-17
Your Golden Dataset Is Lying To YouYou defined 12 user archetypes. Your data contains 17. How hand-crafted personas fail at tens of millions of users, and what happens when you let BigQuery ML K-Means and diversity-aware sampling tell you the truth.
-
2026-03-12
Optimising Recall and Precision in LangSmith ExperimentsYour retrieval pipeline returns results. But does it return the right results? How New Computer used LangSmith's experiment framework to achieve 50% higher recall and 40% higher precision in agentic memory retrieval — and what you can steal from their approach.