All posts
    Engineering

    Speed or Rules? What Caching a Decision Model Costs

    AltSlate LabsOctober 7, 20268 min read

    Many decisions inside a business are not open-ended writing. A support system routes a message to one of a fixed set of teams; a workflow engine picks the one action a policy allows. For these, a small model that simply scores the options is faster and easier to calibrate than a chatbot that writes a paragraph first. The most accurate way to score reads the situation, the rules and each option together, so it gets slower with every option you add. Storing the options ahead of time makes it fast, but we found it also makes the model much worse at picking the action a rule requires. We trained much of that back on synthetic rules. On real rules, the benefit is not established. Here is what we measured, including what did not work.

    1.00 → 0.24
    Choosing the rule-required action, before and after storing the options
    joint scorer vs best cacheable scorer, 17 options
    5×
    Faster at a 77-option menu with stored options
    242 ms vs 46 ms, median, one GPU
    98.7%
    Counterfactual pairs answered correctly on both sides after targeted training
    synthetic rules; 30% before training

    01 — the trade-off

    Two ways to score a menu of options

    Certo is our small decision model. It reads a state — the situation plus the rules that govern it — and returns a probability for each option in a single pass, with no text generated.

    The reference design is a joint scorer: the state, the rules and one option go into the model together, and it scores that option. Reading them together is what lets the model apply a rule to an option. It is also what makes this design expensive: the work for every option is redone for every request, so the cost grows with the size of the menu.

    The alternative is a cacheable scorer, the design behind modern search: encode the state once (turn it into a list of numbers), encode each option once, and compare them cheaply. An option's encoding no longer depends on the request, so it is computed once and reused everywhere. The catch is that the state and the option never meet inside the model, and applying a rule may be exactly the interaction that separation removes.

    Joint scorerCacheable scorerstate + rulesoption 1scorestate + rulesoption 2scorestate + rulesoption 3scoreredone for every option, every requeststate + rulesoption 1storedoption 2storedoption 3storedcompareoptions encoded once, reused for every request
    Where the state and the option meet. The joint scorer (left) reads state, rules and each option together, and repeats that work for every option on every request. The cacheable scorer (right) encodes the options once and reuses them; only the state is encoded per request, and the comparison is a cheap similarity score.

    02 — why caching is worth wanting

    The joint scorer slows down as the menu grows

    On our setup, the joint scorer took 52 ms with five options, 135 ms with twenty and 242 ms with seventy-seven. The cacheable scorer stayed at about 46 ms across the same range, because the options are already encoded. At seventy-seven options that is roughly five times faster. These are median timings for a single query on one GPU: they show the trend, not a tuned production deployment.

    latency (ms)01002005 options20 options77 options52 ms135 ms242 ms≈ 46 msjoint (re-read per option)cacheable (stored options)
    The serving case for caching. Median latency as the number of options grows. The joint scorer re-reads everything for each option; the cacheable scorer reuses stored option encodings and stays nearly flat. Our joint baseline did not reuse the shared state between options, which a tuned deployment could, so the real gap may be narrower.

    03 — what breaks

    Caching the options breaks rule-sensitive choice

    We trained four scorers from the same base model (Qwen3-4B) with the same budget, and tested two kinds of decision: routing, which means picking the option whose meaning matches a message, and rule-sensitive choice, which means picking the action a stated rule requires even when another option looks more similar.

    The joint scorer was perfect on rule-sensitive choice. The best cacheable scorer, which stores one summary per option, kept some of its routing ability but scored 0.24 on rules at seventeen options, against the joint scorer's 1.00. A variant that stores one summary per word collapsed to near chance, no better than guessing, on both.

    jointcacheable, one vectorcacheable, many vectors0.650.420.01pick by meaning (routing, 77 options)1.000.240.08follow the stated rule (17 options)
    Accuracy by design. Share of decisions where the top-scored option was the right one. Both cacheable designs score far below the joint scorer on rule-sensitive choice; the single-vector one keeps part of its routing ability. Routing is shown at the largest menu, 77 options; the rule-sensitive set has 17.

    We also tried the obvious rescue: use the fast cacheable scorer to shortlist a few options, then let the joint scorer pick among them. It failed. The goal was to keep the right answer in the shortlist at least 98% of the time. Among the shortlist sizes we tested, only the full menu met that bar, which saves nothing. The options the fast scorer dropped were the rule-compliant ones. That is consistent with it ranking by surface similarity, though it does not by itself prove the cause.

    04 — training it back

    Targeted examples restore it on synthetic rules

    Rather than change the design, we changed the training data. We continued training the one-summary cacheable scorer on counterfactual pairs: two situations that differ in exactly one condition, so the correct action flips. The tempting-but-wrong option is always on the menu, and the rule stays in the state, so the options remain cacheable.

    Policy: release the payment. Exception: if the request is past the deadline, deny it.Situation: the request is within the deadline✓ release_payment✗ deny_out_of_windowSituation: the request is past the deadline✓ deny_out_of_window✗ release_payment
    One counterfactual pair. The policy and the menu are identical; one clause in the situation changes, and so does the correct action. The option that was right becomes the tempting wrong answer.

    On held-out rules, the retrained scorer handled reworded rules almost perfectly (0.99–1.00). It answered both sides of a counterfactual pair correctly 98.7% of the time, against 30% before training. It combined two rule types it had only seen separately (0.73), and partly handled rule types outside its training. Every gain held against both the original model and a control trained on unrelated data, and the result reproduced with a second training run. Routing did not change measurably, though a drop of up to about 8 points cannot be ruled out.

    beforecontrolafter counterfactual training0.510.270.99reworded exceptions0.550.560.99counterfactual0.420.480.73new combination0.640.600.73precedence rules
    Held-out synthetic rules. Accuracy before training, for a control trained on unrelated data, and after counterfactual training. Shown: reworded exception rules (reworded threshold rules: 0.55, 0.59, 1.00) and precedence, one of two untrained rule types (the other, multi-condition, reached 1.00).

    We stop short of calling this rule-following. These pairs change a fact in the situation, which shows the model is sensitive to the relevant facts, but does not rule out a learned, template-specific pattern. The decisive test — keep the facts fixed and change only the stated rule — has not been run yet.

    05 — real rules

    On real rules, the gain is not established

    Synthetic rules are templated. So we scored the models on real, human-written decisions from an internal benchmark of insurance, policy and contract questions. The first pass looked like no gain at all on the hardest questions, but that turned out to be an artefact of our own setup. The joint scorer appends each option after the document, and on long documents the input was cut off at the end, removing the option entirely. Only 29 of the 67 hard questions fitted in full.

    Rescoring only the questions that fitted told a clearer story. On the shorter questions the joint scorer had a significant edge: 0.861 against 0.500 for the retrained cacheable scorer. On the 29 hard questions it was ahead too, 0.655 against 0.483, but that gap is too small to rule out chance. And the synthetic training gave no convincing improvement over the cacheable scorer before training. Transfer to real rules is not established, which is weaker than saying it was disproved.

    One more attempt: we replaced 40% of the synthetic training mix with real privacy-policy decisions and tested on contract questions from different sources. In a single run, it made things worse, cutting contract accuracy by 9.3 and 16.2 points on two test sets. Because the change both removed synthetic examples and added privacy ones, we cannot say which caused the drop.

    06 — the takeaway

    Choose the design by the decision, not by default

    Where menus are small and following the rule matters, the joint scorer is the safer choice. Where menus are large, the cacheable scorer's speed is real — and so is its accuracy cost.

    This is a trade-off that depends on the application, not a general recommendation. Even on plain routing at seventy-seven options, the cacheable scorer was less accurate (0.42 against 0.65), so the choice comes down to how much accuracy a decision can give up for speed. Targeted training recovers a great deal on rules that look like the training data; we could not show that it carries over to real rules from other sources.

    There is also a lesson about measurement. The apparent failure on hard real questions came from how we cut off long inputs, not from the model. Reporting results separately for inputs that fitted and inputs that were cut off is cheap insurance against that kind of mistake.

    The next experiments are the rule-only test (same facts, different rule), training and testing on the same family of real documents, and a joint scorer that reuses the shared state between options, which would narrow the speed gap. The full paper, with every interval and table, is on arXiv.