Evaluation
Benchmarks should match the job.
How we connect retrieval research, task-specific checks, and business milestones.
A benchmark is useful when it helps answer a practical question: can this system do the work we are asking it to do?
Start with evidence
We publish memory-retrieval results with the dataset, metric, and evaluation context. Those results describe a defined test. A business workflow also needs examples drawn from the work it will actually encounter.
For a tailored evaluation, we assemble representative questions, source material, expected results, and failure cases. That might include finding a customer commitment, reconciling a document with a source, or checking whether an agent completed an authorized action.
Check the system, then the outcome
Compare the harness and the model baseline on the same tasks, inputs, permissions, and budget. Record completion, accuracy, elapsed time, and cost. Inspect execution traces and verify outputs against their sources; use human review where the expected result requires judgment.
The business milestone remains a separate measure. Retrieval accuracy, a correct tool call, and revenue growth describe different things. We choose the measures that fit the engagement, record a baseline, and use the results to decide what to improve or automate next.