Laya is the open decision engine from Convai Innovations and one of September 2026’s most‑starred ML repositories. It is a non‑autoregressive System 1: a 421‑million‑parameter encoder reads text and typed questions and returns option probabilities in one forward pass with zero output tokens. Here, we test those claims on real labeled data: the banking domain of CLINC150. We measure zero‑shot versus a simple classifier, the effect of option wording and order, probability honesty, temperature fitting, abstention gates, out‑of‑scope traffic, biased yes/no, and pydantic‑typed outputs. It is the open answer to TypeSafe’s Jev — but how does it hold up in production?
What is Laya and how does single‑pass decisioning work?
In short: Laya reads a message and typed questions, and returns option probabilities in a single forward pass. It never generates text and logs zero output tokens. You provide a set of labels, a scale, or a yes/no — and you get a distribution. That is its pitch: speed and calibrated confidence.
We install laya 0.3.27 and load the English checkpoint at the author‑reviewed revision via laya.PINNED_REVISIONS. On CUDA, Laya autocasts to half; we switch it off to keep all devices in fp32. This lets a GPU run reproduce the CPU numbers we report. Reproducibility matters.
Calibration shows a first snag before inference. The entry for choice questions with 11+ options is 0.10, outside the valid range; the loader clamps it to 0.5 and warns. A temperature below one sharpens probabilities, so answers with that many options will look about twice as certain as the raw model is. That matters for any decision built on top.
One call to predict answers three typed questions about a support ticket: department as a choice, urgency as a 0–2 score, and churn risk as yes/no. The result includes a probability for each option and two easy‑to‑confuse confidence fields. answer_confidence is the probability of the reported answer. confidence is one minus normalized entropy, whose scale depends on how many options the question has. Metrics and gates use answer_confidence.
What does a pass cost, and how does that shape question design?
The rule is simple: each question becomes a row in the batch. Sixteen yes/no questions take about eight times as long as one yes/no. All options in one choice share a row and its head budget. So a forty‑option choice costs barely twice a three‑option one, and roughly a quarter of sixteen yes/no on our CPU. Prefer one multi‑option choice over many binaries.
Before building on that rule, we test it on labeled data. We load CLINC150 from the Hugging Face Hub as parquet and focus on banking: 15 intents with 100/20/30 train/val/test. We ask Laya to route 450 test queries zero‑shot, giving each intent its name and a one‑line description a developer might write.
The result is strong for zero labeled examples: 0.804 accuracy zero‑shot. For scale, a TF‑IDF plus logistic‑regression baseline reaches 0.651 with three examples per intent, 0.848 with ten, and 0.904 with thirty. One Laya pass gets close to a stronger baseline that needs material labeling.
And yes, it is still one pass with zero generated tokens. Isn’t that what you wanted from a decision engine? But what happens if we change only the words in the options or their order? Do we move answers by accident?
Why do option wording and order change the answers?
Direct answer: bare intent names beat descriptions. Dropping our one‑line descriptions and keeping only the names lifts accuracy from 0.804 to 0.878. Time almost halves too, because options use less than half the tokens. Fewer words, less confusion.
Descriptions blurred intents that names keep apart. account_blocked got routed to freeze_account ten times. Interest‑rate questions went to balance. These small phrasing choices cost percentage points of accuracy and extra milliseconds. The lesson is simple: test the exact wording you will deploy.
Option order also matters. Reversing the bare‑name list changes 4.2 percent of individual answers, though overall accuracy barely moves. That signals a position prior. Do not trust intuition — trust labeled data and freeze your order at validation time. Then keep it unchanged in production.
From here on we route on bare names. This is not a hunch — it is measured on our split. Want to save tokens and time? Say it briefly. Want stability? Lock the order you tested. Obvious after seeing the numbers, right?
How honest are shipped probabilities, and how to calibrate and abstain?
Short answer: they are overconfident for our 15‑option questions. That falls into the choice:11+ bucket, where the shipped temperature is clamped to 0.5. On the 450 test queries, 92 percent of answers claim at least 0.9 confidence, but 91.1 percent of those are right. Mean confidence is 0.974 versus 0.878 accuracy. Expected calibration error is 0.102. Measure it on your own labels.
We then fit temperature on validation data. We build 300 records of raw logits and targets and call fit_temperatures, which fits a choice temperature of 1.258. The fitter replaces the entire table: each bucket disappears without 2,000 records, and score and yes/no reset to 1.0 because none were seen. One fit for a choice silently changes calibration for other types. We restore shipped values and install only the measured bucket.
On the test set, calibration error falls from 0.102 to 0.059 with accuracy unchanged, since temperature never changes the winner. save_calibration writes a JSON file that laya.load can read back. This is a careful, local fix that makes probabilities useful for thresholds and error budgets. Not magic — just hygiene.
What about an abstention gate? fit_abstention_thresholds on the same validation records, with a 5 percent error target, picks 0.602. It keeps 95.7 percent of validation queries at 4.5 percent error, and predict_batch applies it via min_confidence, marking 35 of 450 test answers as abstained. On the test set, though, the gate keeps 92.2 percent at 9.2 percent error — nearly twice the budget. The 2 percent target realizes 5.3 percent. The ranking is sound: the most confident half of test answers are 97.8 percent correct. An error target just holds only for traffic that looks like validation. Thresholds stay per option‑count bucket.
Out‑of‑scope: a gate or an explicit ‘other’ option?
Production traffic includes queries outside the domain. We add 150 out‑of‑scope CLINC queries and 150 from other CLINC domains. Calibrated confidence separates them sharply: banking queries average 0.912; the others, about 0.25. The 5 percent gate stops 89.3 percent of other‑domain queries and 93.3 percent of out‑of‑scope ones, while abstaining on 7.8 percent of banking.
The alternative needs no labels: a sixteenth option, not a banking request. It catches 80.0 percent of other‑domain and 90.0 percent of out‑of‑scope. The cost: 2.0 percent of banking queries are routed away, and banking accuracy drops from 0.878 to 0.864. Adding an option changes what every intent is scored against. Remember that before launch.
Which is better in production? The gate yields stronger blocking of foreign traffic on our split. The ‘other’ option is simpler, but eats into core accuracy. Your choice depends on risk metrics and human‑in‑the‑loop costs. Have labels? Measure both and lock your rules.
In every case, honest confidence is the key. It draws the line between in‑scope and out‑of‑scope and prevents confusing safe abstention with unsafe error. Without calibration, that line is missing.
Yes/no, schema, and practical takeaways: what temperature cannot fix
An in‑scope check as a dedicated yes/no feels natural. We ask whether each message is about the user’s bank account, bills, or payments. It ranks well, with an AUROC of 0.945, but is biased towards no: banking queries average a probability of only 0.361, and at the default 0.5 cut it recognizes just 28.9 percent of them. Temperature cannot repair this.
We fit temperature on 550 validation answers, and it hits the ceiling of 5.0. At the 0.5 cut, nothing changes, because dividing two logits by any temperature never changes which is larger. Temperature scaling fixes a scale, not an offset. What works is treating probability as a score and picking the cut on labels. A cut of 0.09 chosen on validation yields 92.0 percent recall and 83.0 percent specificity on test.
Finally, we wire this into application code via a pydantic schema. decide_batch turns a Literal field into a choice over bare names, and a bool into a yes/no, then projects answers back into a validated model. Passing the gate as min_confidence turns an uncertain intent into None, which Optional accepts — an explicit ask‑a‑human signal. The two out‑of‑scope queries come back as intent=None.
However, the boolean inside decide is cut at a fixed 0.5, so a fraud report is marked in_scope=False at a probability of 0.29. Reading the probability from the result details and applying the 0.09 cut gives the right answer. That is how to handle binary questions in Laya.
“A calibrated decision model is one you have calibrated, on your own labels, for your own questions.”
The bottom line. Laya delivers most of what it promises: one pass answers several typed questions with no generated tokens, the bare‑name router reaches 0.878 on fifteen banking intents with no training data, and calibrated confidence separates in‑scope from out‑of‑scope well enough to stop more than nine in ten out‑of‑scope queries. But the value of its probabilities depends on work you must do: the shipped temperature for 11+ options sharpens, the one‑line calibration call wipes other question types, a validation‑fitted error budget does not transfer without margin, and a yes/no can sit on the wrong side of 0.5 that no temperature can fix — which schema projection then hard‑codes. Each has a few‑line fix, if you measure on your own traffic.
Based on the Laya tutorial.