· applied ai
Can Decision Models Abstain?
“I might cancel if it takes much longer.” Should a message like this go into a cancellation workflow? I tested four local System One decision models and a hosted Jev reference on that question. Each received a customer message and had to choose: cancellation request, status request, or insufficient information. I wanted to see whether they could leave the decision open when neither action was clearly requested.
Four local models. One hosted reference.
Eighteen synthetic messages: six cancellations, six status requests, and six cases where no current action is clear.
I might cancel if it takes much longer.
Expected: Abstain
Read the exact instruction
Classify the customer's currently requested action using only their message. Choose cancel_order when the customer clearly asks to cancel or stop their order, including clear paraphrases. Choose check_order_status when the customer asks about the order's current status, shipping progress, or delivery time. Choose insufficient_information when neither action is clearly requested. A complaint, a quoted instruction, or a hypothetical future cancellation is not itself a current cancellation request. Pay attention to negation and to the difference between a past request and the current request. Infer the meaning of clear paraphrases, but do not invent an action from dissatisfaction alone. The scope of this benchmark is messages with at most one current action from these two categories; insufficient_information is a decision to abstain, not a customer intent.
Local results use /v1/systemone; the hosted Jev reference uses TryJevAI. Raw requests and responses are linked below. 2026-10-08. One recorded answer per model/case. Matches are agreement with the stated policy. Scores are not validated probabilities of customer intent. Local API confidence measures score concentration. Hosted confidence and latency are shown as returned by the playground; timings are not directly comparable.
A small routing experiment
This is a classification step between an incoming message and a fixed workflow. The model returns a label; the surrounding application decides what happens next. Nothing in this test changed an order.
I used Nimble 9B, Clef Flash 9B, Tev1 4B, and Laya 421M, plus Jev 1.13.0 via TryJevAI. TryJevAI is an independent playground. Jev's parameter count is not verified here, so this is a hosted-service reference, not an equal-size comparison.
There are 18 synthetic English messages: six cancellation requests, six status questions, and six where neither action is clear. The local models use Ollama's System One API. Jev received the same messages, policy, and ordered labels through the playground. Its question field rejects text over 300 characters, so I put the full policy in state beside the message and used a short question. That format difference is part of this comparison. Each message is evaluated alone, with no conversation history. Expand “Read the exact instruction” above to inspect the policy.
The rule is to identify the customer's current requested action. Clear paraphrases count. Complaints, quotes, and possible future cancellations don't count as requests by themselves. insufficient_information, shown as Abstain, is a third answer choice. I didn't apply a confidence threshold after the model answered.
What happened
Nimble, Clef Flash, and Tev1 matched all 12 clear requests. Laya matched 11. On the six messages labelled for abstention, Tev1 matched five, Nimble and Laya three each, and Clef Flash one.
The hosted Jev reference matched 18/18 labels: 12/12 clear requests and 6/6 expected abstentions.
The explorer opens on the conditional cancellation. Nimble and Tev1 abstained. Clef Flash and Laya chose cancellation, even though the customer only described a possible future action. Jev abstained on this message.
“Can you help me with order #314?” is more debatable. A status update could be helpful, but the message doesn't specify what help is wanted. My expected label is abstention. These scores measure agreement with that policy, not an objective reading of someone's hidden intent.
“You said ‘cancel the order’ in your last message” refers to missing history. Its expected answer only reflects what can be inferred from the supplied message.
What this tells me
I'd test the review path explicitly before using these labels to trigger a workflow. A single total hides two different mistakes: missing a clear request, and selecting an action when neither is clearly requested.
This is an exploratory set, not a held-out benchmark. One changed answer moves a total by 5.6 percentage points. Always abstaining would match 6/18 labels while handling no clear requests. I'd want real messages, independent label review, and prompt and option-order tests before choosing a production model.
The bars show the API's label scores. This experiment doesn't establish calibration. The local API confidence value describes score concentration, not correctness. Jev's confidence is shown as returned by the playground.
Run details
One recorded answer per model and message, grouped by model; fixed option order: cancellation, status, abstention. The Tev1 and Laya runs kept the original inputs and key, using Ollama 0.40.1 defaults on an M3 MacBook Air with 24 GB memory. I restarted Tev1 after an initial connection failure returned no answer. Local timings include loading and HTTP overhead. Jev's displayed latency is the playground's reported value. They aren't directly comparable. The hosted calls were paced to respect the playground's rate limit.
Download the requests and responses and run details, including model tags and local digests.
I used AI tools to help with the code and edit the text. The results are from local models and the hosted endpoint; the messages are synthetic.