Skip to content

Koltra Research

Measure the operation, not just the response.

The method asks whether the operation completed as declared. Fluency is not a score.

If you run a practice: this is how Koltra would know whether a product actually worked for you. Not whether the call sounded right, but whether the appointment moved, the rule held, the handoff reached a person, and what it cost. The method below is written so it can be checked; no results are published yet.

The record remembers what happened. This method judges it against what was declared.

Method v0.2 · in development · no public results

An agent can sound right and still leave the work wrong.scored separately
01Task completionThe operation reached its defined end state: verified, not assumed.
02Policy compliancePermissions, consent, and prohibited actions held for the whole run.
03Tool correctnessThe right tool, valid arguments, and a correct reading of what came back.
04State continuityIdentity, context, and unfinished work survive across turns and channels.
05Handoff qualityWhen a person steps in, context, urgency, and ownership arrive with them.
06Failure recoveryFailures get detected, contained, explained, and recovered. Not buried.
07LatencyHow long the operation and its critical stages actually take.
08CostEverything the run consumed: models, speech, tools, infrastructure, retries.

An agent can sound right and still leave the work wrong.

Kept separate on purpose: speed can hide a policy breach, and fluency can hide a wrong tool call. This is a working framework, not a scorecard or a performance claim.

Task completion

The operation reached its defined end state: verified, not assumed.
01

Policy compliance

Permissions, consent, and prohibited actions held for the whole run.
02

Tool correctness

The right tool, valid arguments, and a correct reading of what came back.
03

State continuity

Identity, context, and unfinished work survive across turns and channels.
04

Handoff quality

When a person steps in, context, urgency, and ownership arrive with them.
05

Failure recovery

Failures get detected, contained, explained, and recovered. Not buried.
06

Latency

How long the operation and its critical stages actually take.
07

Cost

Everything the run consumed: models, speech, tools, infrastructure, retries.
08

Turn product obligations into checks.

Declare the permitted action and end state before the run starts.
01
Test tools, state, and policy without touching a live customer system.
02
Keep the failures, retries, and handoffs. Not just the good examples.
03
Completion, safety, latency, and cost stay visible. Never blended into one score.
04
Synthetic or explicitly permitted data only, while the framework matures.
05
Publish nothing without the method, conditions, and limits needed to judge it.
06

What completion is worth is declared, not implied.

Each product declares its bottom-line metrics before deployment performance is judged: the baseline, who owns the metric, how outcomes are attributed, the time horizon, and what counts as failure. Money made, money saved, cost per completed outcome, and satisfaction are measured per deployment. Numbers only ever come from deployments, and none are published today.

From measurement to governed improvement.

Evaluation is not a dashboard. The path is observe, evaluate, diagnose, learn, heal, validate, and govern. During an operation, recovery stays inside declared authority. Across operations, evidence can justify a tested candidate. Candidates change production only through versioned release.

The method comes before the benchmark.

The internal framework is still taking shape. If the approach proves useful, the parts that can help others inspect AI product operations may be released.

Test the dimensions against your own runs.

The evaluation questions are more useful applied than read. Bring a workflow you operate today and we can work through what completion, handoff quality, recovery, and cost would each have to prove for it.