Ollaya Runs Decision Models Locally With Millisecond Structured Outputs
A new Rust daemon pulls, serves and runs open decision models locally behind a TypeSafe-compatible API, the way Ollama runs LLMs.
- Ollaya is an open-source Rust daemon that runs decision models locally, Ollama-style, at github.com/ollaya-dev/ollaya.
- Serves TypeSafe-compatible endpoints so existing Jev SDK clients work by changing one environment variable.
- Models return calibrated probabilities for typed questions in a single forward pass, no text generation.
- winnow:e4b hits 0.722 accuracy on typed decisions in 89 ms on an RTX 4090.
- Beats Ollama 0.35 on calibration (ECE 0.022 vs 0.122 on identical Nimble weights).
- Apache-2.0, with installers for macOS, Windows, Linux and Docker at ollaya.dev.
Ollaya serves decision models through an Ollama-style local API
Ollaya is a new open-source runtime that downloads and serves decision models through a local HTTP API. Its Ollama-style CLI combines model management and a daemon in one binary, while TypeSafe-compatible endpoints let existing Jev clients switch by changing a base URL. For developers using LLMs to route tickets, screen content, or select tools, the project offers a fast inference layer built around structured probability outputs.
Probabilities in one forward pass
A decision model evaluates state, such as a message, email, ticket, image, or JSON object, against typed questions. Those questions can request categorical choices, scores, or Boolean flags. The model returns probabilities for every question in a single forward pass, often within milliseconds, without generating prose.
A customer-support request could pair a message with fields such as intent, is_urgent, frustration, and churn_risk. The response contains a typed result and associated probabilities for each field. This schema-based exchange removes generated JSON parsing, malformed-output retries, and instructions intended to force a language model into a fixed response shape.
One install, three API routes
curl -fsSL https://ollaya.dev/install.sh | sh
ollaya run winnow:e4b --preset triage "Third time this year you've double-charged me..."Ollaya’s CLI follows familiar Ollama conventions. ollaya serve starts the daemon, while run, pull, list, ps, show, rm, cp, stop, and create manage models and processes. The CLI starts the daemon automatically when a command needs it.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.