Open source project by Mert Cobanov (
ollaya-dev/ollaya, Apache-2.0).
Why use it?
Agent workflows are full of small judgements with a known set of outcomes: is this ticket a refund request, is this prompt an injection attempt, which team should handle this. Asking an LLM to reason each one out in text costs tokens and time, and the answer comes back as prose you then have to parse.
A decision model reads a state (a message, an email, a ticket, any JSON) plus typed questions (choice, score, noul) and returns calibrated probabilities in a single forward pass, in milliseconds. It never generates text. Ollaya pulls these models by name, serves them from a local daemon, and speaks TypeSafe’s /v1/systemone wire format.
What you can do with it
The README’s first example runs the triage preset on a single customer message:
ollaya run winnow:e4b --preset triage "Third time this year you've double-charged me. Refund it today or I'm cancelling and moving to a competitor."
It returns intent: refund (0.91), is_urgent: yes (0.92), refund_requested: yes (0.99), and churn_risk: yes (0.99).
The agent docs list six built-in presets:
| Preset | What it decides |
|---|---|
triage | intent, is_urgent, frustration, refund_requested, churn_risk for a customer message |
email | category, is_spam, is_phishing, urgency, needs_reply for an email body |
guard | jailbreak, prompt_injection, sensitive_data, harm_severity, topic for a prompt |
moderation | toxic, harassment, threat, spam, severity for a post |
router | difficulty, domain, needs_tools, is_sensitive for a request |
agent | action (run, ask, block), on_task, risk, destructive: reviews a command an agent is about to run |
You can also write your own questions. A choice question routes a ticket to billing, engineering, sales, or other; a score question rates urgency from “can wait” to “right now”; a noul question checks “Does the customer ask for their money back?” Each answer comes with a probability you can threshold before you route, label, block, or escalate.
Key features
-
One binary
ollaya serveruns the daemon;ollaya run,pull,list,ps,show,rm,cp,stopandcreatework the way they do in Ollama. If the daemon isn’t running, the CLI starts it. -
For agents
ollaya mcpserves the models to Claude Code, Claude Desktop, Cursor and other MCP clients, with the toolsdecide,list_models,show_model, andpull_model. Theollaya-decisionsskill teaches agents when and how to use them: picking a model, writing questions, and setting confidence thresholds. -
Many open decision models
winnow:e4bis the recommended model with an NVIDIA GPU;layais the fastest and runs well on a CPU. Others includedecider:2b-visionfor questions about an image, theqwen3guardsafety guard, and the zero-shotnliandgliclassclassifiers. All models and their numbers: ollaya.dev/search. -
TypeSafe-compatible
POST /v1/systemone,/v1/decisionsandGET /v1/modelsare wire-identical to TypeSafe. The official SDK works unchanged when you setTYPESAFE_BASE_URL=http://localhost:11435. -
Modelfiles
Bake a question set into your own model, then run
ollaya create triage -f Modelfileandollaya run triage "…". -
Weights come from their authors
Ollaya reads the original weight files from the author’s Hugging Face repository, pinned to a commit and verified by sha256. Ollaya never re-hosts weights.
How to use
On Linux and macOS (Apple silicon):
curl -fsSL https://ollaya.dev/install.sh | sh
When an NVIDIA GPU is present, the installer adds the CUDA runtime. On Windows, run irm https://ollaya.dev/install.ps1 | iex in PowerShell. A desktop app for macOS, Windows and Linux is at ollaya.dev/download.
Connect it to Claude Code and add the skill:
claude mcp add ollaya -- ollaya mcp
npx skills add ollaya-dev/ollaya --skill ollaya-decisions
Or try it from the terminal. The CLI pulls the model on first use:
ollaya run laya --preset triage --format json "I was charged twice and want a refund."
Notes
- When to use it — Reach for it when the answer is one of a known set of outcomes and you will act on it: route, label, block, escalate, pick a template. Keep reasoning in text for open-ended work.
- Hardware matters — With five questions per request, the encoders (
laya,nli,gliclass,von,qwen3guard) take 7 to 35 ms on a GPU and 0.3 to 2.4 s on a CPU; the decoders 0.1 to 0.9 s on a GPU and 1.3 to 25 s on a CPU.nimble,jeeves, andclefneed a 24 GB GPU. - Requirements — Linux needs glibc ≥ 2.38 (e.g. Ubuntu 24.04+). macOS support is for Apple silicon.
- Local by default — The server listens on
127.0.0.1:11435. States and questions are never logged; the network is used only to pull models. - Early project — Ollaya was published in September 2026 and ships new versions every few days. Check ollaya.dev/results for current models and numbers.
- Model licenses — Ollaya is Apache-2.0. Each model keeps its own license.
- Independent project — Ollaya is not affiliated with or endorsed by Ollama or TypeSafe.
- Source — ollaya-dev/ollaya (Apache-2.0). Docs: ollaya.dev/docs.