Koltra Research
Measure the operation, not just the response.
The method asks whether the operation completed as declared. Fluency is not a score.
If you run a practice: this is how Koltra would know whether a product actually worked for you. Not whether the call sounded right, but whether the appointment moved, the rule held, the handoff reached a person, and what it cost. The method below is written so it can be checked; no results are published yet.
The record remembers what happened. This method judges it against what was declared.
Method v0.2 · in development · no public results
An agent can sound right and still leave the work wrong.
Kept separate on purpose: speed can hide a policy breach, and fluency can hide a wrong tool call. This is a working framework, not a scorecard or a performance claim.
Task completion
The operation reached its defined end state: verified, not assumed.Policy compliance
Permissions, consent, and prohibited actions held for the whole run.Tool correctness
The right tool, valid arguments, and a correct reading of what came back.State continuity
Identity, context, and unfinished work survive across turns and channels.Handoff quality
When a person steps in, context, urgency, and ownership arrive with them.Failure recovery
Failures get detected, contained, explained, and recovered. Not buried.Latency
How long the operation and its critical stages actually take.Cost
Everything the run consumed: models, speech, tools, infrastructure, retries.Turn product obligations into checks.
What completion is worth is declared, not implied.
Each product declares its bottom-line metrics before deployment performance is judged: the baseline, who owns the metric, how outcomes are attributed, the time horizon, and what counts as failure. Money made, money saved, cost per completed outcome, and satisfaction are measured per deployment. Numbers only ever come from deployments, and none are published today.
From measurement to governed improvement.
Evaluation is not a dashboard. The path is observe, evaluate, diagnose, learn, heal, validate, and govern. During an operation, recovery stays inside declared authority. Across operations, evidence can justify a tested candidate. Candidates change production only through versioned release.
The method comes before the benchmark.
The internal framework is still taking shape. If the approach proves useful, the parts that can help others inspect AI product operations may be released.
Test the dimensions against your own runs.
The evaluation questions are more useful applied than read. Bring a workflow you operate today and we can work through what completion, handoff quality, recovery, and cost would each have to prove for it.