Ask everything. Generate nothing.
snap is a decision engine for LLM pipelines. The input is your data plus every question you have about it; the output is a set of typed answers, each with its full probability distribution, from a single pass over the model. There’s no generated text, so there’s nothing to parse, retry or clean up.
$ brew install emnlmn/snap/snap
thensnap serve
Open source (MIT) · runs on your hardware · Jev wire-compatible · all downloads
state
questions
The decisions inside your pipeline.
snap is the shorter path for any label, flag or score you would otherwise prompt a model for and then parse. Below are six common cases, each with a real request and snap’s real response.
Logits, not generated text.
At every step, a language model scores every token in its vocabulary as the possible next one. Those scores are the logits. Chat inference picks a token, appends it and runs the model again, once per token, until there is enough text for you to parse.
snap asks each question so that the answer is a single letter, runs the model once, and reads the scores of those letters directly.
How we measured
snap and Ollama 0.34.3 run the same MiniCPM5-2B weights on one Apple M1 Max, and OpenAI runs gpt-6-luna with reasoning turned off. Every request uses the same prompt and JSON schema, with data no engine has seen before.
Times are the median of 15 requests (5 for OpenAI). The bars show each engine’s measured phases, and the JSON is a real reply from the model, replayed at the speed it was written.
You can reproduce these numbers with eval/vs_ollama.py and eval/vs_openai.py. Every workload and mode is in Benchmarks.
-
1
Letter prompts
Each question becomes a prompt that stops exactly where the answer’s first token goes. Each option gets a letter, so the whole answer is a single token.
A yes/no question (
noul) gets A and B, achoicegets one letter per option, ascoreone per level, and anumericquestion one per anchor value across your range. -
2
One shared prefill
The state goes through the model once. What the model computes from it, the KV cache (its working memory of your input), is shared by every question.
The questions branch off it like a tree, so any text they have in common is computed only once, and prompt openings that repeat across requests stay cached.
-
3
One batched decode
All the questions are decoded together in a single
llama_decodecall, with one KV sequence each. Four questions cost one pass, not four conversations.At startup, snap checks which method the model’s architecture supports and uses it. There’s nothing to configure.
-
4
One logit row
For each question, snap reads a single row of logits, keeps only the option letters, merges the different tokens for the same letter (like “A” and “ A”) and turns them into your distribution with a softmax.
coveragetells you how much of the model’s probability landed on your letters.snap calibratecan fit a temperature for each question type.
These 26 letters are everything snap can output. With more than 26 options, each option gets its own yes/no question over the same shared prefix, up to 256.
No text channel for prompt injection.
Injection works by getting a model to produce text you didn’t intend: a leaked system prompt, a smuggled tool call, a link your UI renders. snap never produces text. What leaves the engine is a probability over options you wrote, in a JSON shape the server builds, not the model.
| Attack | LLM that generates text | snap |
|---|---|---|
| Data exfiltration through the response | Possible. The response is free text. | No text channel. Answers are numbers keyed by your option names. |
| A smuggled tool call or command | Possible. The model writes the call. | Nothing to execute. Your code maps a decided option to an action you wrote. |
| A leaked system prompt or context | Possible. It can be asked to repeat itself. | Nothing is generated for it to leak into. |
| Markup or links injected downstream | Possible. Output gets rendered in a UI or an email. | The response schema is fixed by the server. |
| A broken output format | Handled with validators, repair and retries. | Can’t happen. The model doesn’t write the response. |
| Persuasion toward a different answer | Possible, and invisible in the output. | Still possible, but limited and measurable. Text in the state can only shift probability between your options. That shows up as split confidence or low coverage, and allow_abstain lets the model say it can’t tell. Set a threshold and send those cases to a person. |
On your hardware
A single binary with llama.cpp built in, the same on a laptop, a VM, a Kubernetes pod or an air-gapped rack. There’s also a fully static Linux build, with no Python and no runtime to keep patched.
No data egress
There’s no API key, no telemetry and no third party handling your records. The weights are a GGUF file on disk, and after the first download snap runs offline. A daily check for new releases prints to stderr, and SNAP_NO_UPDATE_CHECK=1 turns it off.
Auditable by design
The same input always gives the same distribution. You can log the full probabilities with every decision, replay it later, and check x_snap to see exactly what the engine did.
Open source
It’s MIT-licensed Rust you can read end to end. Models come from a tested list, and a calibration file won’t load with a different model or prompt version than the one it was fitted on.
Latency and accuracy on the same weights.
All numbers are medians of snap bench --requests 20 on an Apple M1 Max with Metal, and every one is reproducible with snap bench and snap evaluate eval/cases.jsonl.
snap vs Ollama 0.34.3, same weights, same machine
Both engines run MiniCPM5-2B on the same machine, every request uses data neither has seen, and times are the median of 15 runs. The percentage compares snap with Ollama’s fastest option in each row, which is always one letter per question: the quickest Ollama can go, but it gives you text to parse, no probabilities, and one request per question (when we asked for all the letters in one reply, it stopped early). JSON under a schema is what a pipeline actually uses, and with probabilities it carries the same information as a snap answer, written one token at a time. You can reproduce it with eval/vs_ollama.py.
Latency with a shared state
| scenario | minicpm5-2b | spark-4b | qwen3.8-4b |
|---|---|---|---|
| single question | 50 ms | 89 ms | 82 ms |
| 4 questions, shared | 126 ms 32/q | 232 ms 58/q | 227 ms 57/q |
| 8 questions, shared | 252 ms 32/q | 406 ms 51/q | 472 ms 59/q |
| 8 questions, direct | 1170 ms 146/q | 2158 ms 270/q | 2381 ms 298/q |
| 8 questions on a 1.8 KB document | 1887 ms 236/q | 3278 ms 410/q | 3327 ms 416/q |
| 4 KB state, 1 question | 1364 ms | 2476 ms | 2164 ms |
All the prompts in a request are decoded in one batched call that shares their common prefix, on every architecture; the direct row decodes each question on its own, for comparison. Qwen’s hybrid architecture runs 17 parallel sequences instead of 65, and snap picks the right path at startup.
Accuracy · eval/cases.jsonl, 301 cases
| model | accuracy | ms/case |
|---|---|---|
| qwen3.8-4b Q4_K_M | 86.4% | ~300 |
| spark-4b Q8_0 | 85.7% | ~245 |
| minicpm5-2b Q4_K_M | 69.1% | ~135 |
Raw probabilities aren’t calibrated. snap calibrate fits one temperature per question type on your labeled cases, ties it to the model and prompt version, and prints the calibration error (ECE) before and after.
A drop-in replacement for Jev.
Same request and response format as Jev, so a client written for Jev works with a local snap server, unchanged. The additions are optional fields a Jev client can ignore: status, confidence, coverage and the full probabilities on every answer.
noulYes or no, with the probability of yes.choiceOne of 2 to 256 options, with a probability for each.scoreAn ordered level, with an expected score between 0 and 1.numericA number in your range, read from anchor values.allow_abstainAn extra answer for “not enough information”.{
"state": "My payouts have been failing for 3 days.",
"questions": {
"urgent": {
"type": "noul",
"instructions": "SLA breach?"
},
"team": {
"type": "choice",
"criteria": {
"billing": "payments, refunds",
"technical": "bugs, outages"
}
}
}
}
{
"answers": {
"urgent": {
"type": "noul",
"status": "decided",
"noul": 0.7285,
"boolean": true,
"confidence": 0.4571,
"coverage": 0.9993
},
"team": {
"type": "choice",
"status": "decided",
"choice": "billing",
"probabilities": {
"billing": 0.9828,
"technical": 0.0172
}
}
}
}
A real response from minicpm5-2b, without the usage and x_snap fields.
snap next to Jev and SemIf.
The two closest projects are Jev, the hosted original, and SemIf, an open research scorer.
| Capability | snap | Jev | SemIf |
|---|---|---|---|
| Deployment | Local, one Rust binary | Hosted service | Local Python, WebGPU demo |
| Open weights | GGUF | No | Qwen |
| Jev-compatible HTTP API | Yes | The original | CLI only |
| Numeric answers | Yes | No | No |
| Calibrated probabilities | Per question type | Claimed | Per workload |
| TypeSafe public eval | 73% on a 4B model | ~88%, self-reported | Not reported |
The Jev column comes from its public description and the SemIf column from its repository. The TypeSafe row is agreement with the published consensus on its 373 public decisions, reproducible with eval/typesafe_public.py.
Three steps to your first decision.
The first run downloads the model once (1.5 to 4 GB). After that, everything runs offline on your machine.
Install
A single binary: Homebrew on macOS, a tarball everywhere else.
Serve
The default model on port 8018, or any other from the tested set with --model.
$ snap serveAsk
The playground, built into the binary: questions in a form, answers with their full distributions, and the request as cURL.
localhost:8018/playgroundOpen
Prebuilt for macOS on Apple Silicon, Linux x86_64 and aarch64 (including a fully static build and Vulkan) and Windows x86_64. From source: CUDA, ROCm, Vulkan and Metal.