Ask typed questions about any state and get back labels, levels and calibrated probabilities your code can branch on. Every question in one forward pass, on your own machine, with nothing to parse and nothing to hallucinate.
bundle add ruby-laya
Ruby 3.3+. No Python, no LibTorch, no compiler. Two prebuilt gems come with it, and the first call downloads the checkpoint.
Questions are data, so they live in your code, your database or your config. The answer carries the full distribution, so a threshold is a decision you make rather than one the model makes for you.
One label from a set you define, with a probability for each option.
A position on ordered levels you describe, weighted across the whole rubric.
The calibrated probability that a statement holds. No separate confidence to reconcile.
# Declare the questions once. They are code, so they diff and they test. class TicketTriage < Laya::Decision choice :department, "Which team should handle this?", billing: "invoices, payments, refunds", technical: "bugs, outages, system errors", other: "everything else" score :urgency, "How urgent is this?", levels: ["not urgent", "soon", "critical deadline"] noul :churn_risk, "Does the customer threaten to cancel?" end # One forward pass answers all three. triage = TicketTriage.decide(email) triage.department # => #<Choice billing 95.9%> triage.department == :billing # => true triage.department.billing? # => true triage.urgency.score # => 1.36 triage.urgency.label # => "soon" triage.churn_risk? # => true triage.churn_risk.probability # => 0.827 # Gate on confidence, because the probabilities are trained to mean something. if triage.department.confidence >= 0.85 route_to triage.department.to_sym else escalate_to_human triage end
# Not worth a class? Ask inline. Laya.ask(email) .noul(:refund, "Do they want money back?") .choice(:tone, "How does this read?", %w[calm annoyed furious]) .decide .refund .probability # => 0.856
Script and language detection runs in pure Ruby, before any inference, and sends each request to the checkpoint that can read it. The English model does not degrade on Devanagari or Han, it collapses while staying confident, so the choice has to be made before the forward pass rather than after.
Two checkpoints stay resident by default, and loading is thread-safe while inference is not serialized behind it. Tell it the language when you already know.
Laya.configure do |config| config.preload = true # every checkpoint resident config.device = "coreml" # or "cpu", "cuda", ... end triage = TicketTriage.decide(hindi_ticket) triage.routing.model # => "multilingual" triage.routing.reason # => "non-Latin script (devanagari, 100% of letters); # the English checkpoint cannot read it" # Or say so, when you already know. TicketTriage.decide(ticket, lang: "pt-BR") class InvoiceCheck < Laya::Decision model "typed-decisions" end
Jev 1.13.0 is the hosted decision model whose API this mirrors, so it receives byte-identical questions. Three general models answer the same questions through OpenRouter, constrained to the label set with structured outputs. 500 items per set, drawn with a fixed seed.
| Model | AG News 4 labels | Emotion 6 | Banking77 77 | MASSIVE 17, 12 langs |
|---|---|---|---|---|
| ruby-laya, local | ||||
| jev-1.13.0 | ||||
| qwen3-30b-a3b | ||||
| gemini-3.1-flash-lite | ||||
| gpt-5-nano |
AG News by five points over Jev, and the margin survives a paired test on the same items (p = 0.0005). It answers in 42 to 153 ms against 208 to 217 ms for the hosted model and about a second for the LLMs, costs nothing per call, keeps the data on your machine, and returns the same answer every time. Jev's answers moved between two runs; these did not.
With 77 labels it reaches 39% against Jev's 81%, because the options share one token budget. On multilingual intent it reaches 57% against 87%. Raising the budget buys 4.8 points and telling the router the language buys 2.3, so both gaps belong to the checkpoints. The roadmap says what would move them.
HF_HOME, HF_HUB_OFFLINE and HF_TOKEN.