ruby-laya
System 1 decisions for Ruby

Decisions, not text.

Ask typed questions about any state and get back labels, levels and calibrated probabilities your code can branch on. Every question in one forward pass, on your own machine, with nothing to parse and nothing to hallucinate.

bundle add ruby-laya

Ruby 3.3+. No Python, no LibTorch, no compiler. Two prebuilt gems come with it, and the first call downloads the checkpoint.

TicketTriage.decide(email) 155 ms · local CPU
subject Duplicate charge on invoice #4411
body We were billed twice for March. Please refund the duplicate today or we will cancel our plan.
department billing · 84% confidence
billing95.9%
sales1.5%
technical1.4%
other1.2%
urgency 1.36 of 2 · “soon”
not urgent13.3%
soon37.7%
critical48.9%
churn_risk 82.7% likely
82.7%
refund_requested 85.6% likely
85.6%
Three question types

Every answer is a type, and a distribution.

Questions are data, so they live in your code, your database or your config. The answer carries the full distribution, so a threshold is a decision you make rather than one the model makes for you.

choice

One label from a set you define, with a probability for each option.

triage.department == :billing
triage.department.confidence

score

A position on ordered levels you describe, weighted across the whole rubric.

triage.urgency.score # => 1.36
triage.urgency.label # => "soon"

noul

The calibrated probability that a statement holds. No separate confidence to reconcile.

triage.churn_risk?  # => true
triage.churn_risk.probability
# Declare the questions once. They are code, so they diff and they test.
class TicketTriage < Laya::Decision
  choice :department, "Which team should handle this?",
         billing:   "invoices, payments, refunds",
         technical: "bugs, outages, system errors",
         other:     "everything else"

  score  :urgency, "How urgent is this?",
         levels: ["not urgent", "soon", "critical deadline"]

  noul   :churn_risk, "Does the customer threaten to cancel?"
end

# One forward pass answers all three.
triage = TicketTriage.decide(email)

triage.department              # => #<Choice billing 95.9%>
triage.department == :billing   # => true
triage.department.billing?     # => true
triage.urgency.score           # => 1.36
triage.urgency.label           # => "soon"
triage.churn_risk?             # => true
triage.churn_risk.probability  # => 0.827

# Gate on confidence, because the probabilities are trained to mean something.
if triage.department.confidence >= 0.85
  route_to triage.department.to_sym
else
  escalate_to_human triage
end
# Not worth a class? Ask inline.
Laya.ask(email)
    .noul(:refund, "Do they want money back?")
    .choice(:tone, "How does this read?", %w[calm annoyed furious])
    .decide
    .refund
    .probability                # => 0.856
Routing

Three checkpoints. It picks.

Script and language detection runs in pure Ruby, before any inference, and sends each request to the checkpoint that can read it. The English model does not degrade on Devanagari or Han, it collapses while staying confident, so the choice has to be made before the forward pass rather than after.

Two checkpoints stay resident by default, and loading is thread-safe while inference is not serialized behind it. Tell it the language when you already know.

Laya.configure do |config|
  config.preload = true      # every checkpoint resident
  config.device  = "coreml"  # or "cpu", "cuda", ...
end

triage = TicketTriage.decide(hindi_ticket)

triage.routing.model    # => "multilingual"
triage.routing.reason
# => "non-Latin script (devanagari, 100% of letters);
#     the English checkpoint cannot read it"

# Or say so, when you already know.
TicketTriage.decide(ticket, lang: "pt-BR")
class InvoiceCheck < Laya::Decision
  model "typed-decisions"
end
Benchmarks

Measured against the hosted alternative, including where it loses.

Jev 1.13.0 is the hosted decision model whose API this mirrors, so it receives byte-identical questions. Three general models answer the same questions through OpenRouter, constrained to the label set with structured outputs. 500 items per set, drawn with a fixed seed.

Accuracy, 2026-09-23 · ruby-laya 0.1.0 · Apple M5 Max, CPU only · every bar on one scale, 0 to 100%
Model AG News 4 labels Emotion 6 Banking77 77 MASSIVE 17, 12 langs
ruby-laya, local 91.8% 57.4% 39.0% 57.5%
jev-1.13.0 86.6% 57.4% 80.8% 86.8%
qwen3-30b-a3b 85.6% 54.6% 72.4% 84.8%
gemini-3.1-flash-lite 82.6% 55.0% 78.4% 84.8%
gpt-5-nano 71.0% 56.6% 71.0% 79.9%
What it wins

AG News by five points over Jev, and the margin survives a paired test on the same items (p = 0.0005). It answers in 42 to 153 ms against 208 to 217 ms for the hosted model and about a second for the LLMs, costs nothing per call, keeps the data on your machine, and returns the same answer every time. Jev's answers moved between two runs; these did not.

What it loses

With 77 labels it reaches 39% against Jev's 81%, because the options share one token budget. On multilingual intent it reaches 57% against 87%. Raising the budget buys 4.8 points and telling the router the language buys 2.3, so both gaps belong to the checkpoints. The roadmap says what would move them.

What you are installing

Two prebuilt gems and a download.

dependencies
ONNX Runtime and Hugging Face tokenizers, both shipping prebuilt binaries. Nothing compiles.
first call
Fetches the English checkpoint, about 820 MB, into the standard Hugging Face cache. Respects HF_HOME, HF_HUB_OFFLINE and HF_TOKEN.
weights
ONNX exports of the checkpoints Convai Innovations published, verified against PyTorch before they were written.
without a model
Routing, script and language detection, email cleaning, the question presets and the shortlist are pure Ruby and load without ONNX Runtime.
faithfulness
Over 3,000 assertions compare this port with fixtures recorded from upstream Python, and an opt-in suite replays 43 calls across nine languages against the real checkpoints.