Skip to content
snap
GitHub
Get snap
$ brew install emnlmn/snap/snap

Ask everything. Generate nothing.

snap is a decision engine for LLM pipelines. The input is your data plus every question you have about it; the output is a set of typed answers, each with its full probability distribution, from a single pass over the model. There’s no generated text, so there’s nothing to parse, retry or clean up.

Install for
$ brew install emnlmn/snap/snap

thensnap serve

Open source (MIT) · runs on your hardware · Jev wire-compatible · all downloads

state

questions

    0.0ms minicpm5-2b · Q4_K_M · Apple M1 Max

    The decisions inside your pipeline.

    snap is the shorter path for any label, flag or score you would otherwise prompt a model for and then parse. Below are six common cases, each with a real request and snap’s real response.

    Logits, not generated text.

    At every step, a language model scores every token in its vocabulary as the possible next one. Those scores are the logits. Chat inference picks a token, appends it and runs the model again, once per token, until there is enough text for you to parse.

    snap asks each question so that the answer is a single letter, runs the model once, and reads the scores of those letters directly.

    snap vs
    questions
    Ollama writes
    How we measured

    snap and Ollama 0.34.3 run the same MiniCPM5-2B weights on one Apple M1 Max, and OpenAI runs gpt-6-luna with reasoning turned off. Every request uses the same prompt and JSON schema, with data no engine has seen before.

    Times are the median of 15 requests (5 for OpenAI). The bars show each engine’s measured phases, and the JSON is a real reply from the model, replayed at the speed it was written.

    You can reproduce these numbers with eval/vs_ollama.py and eval/vs_openai.py. Every workload and mode is in Benchmarks.

    1. 1

      Letter prompts

      Each question becomes a prompt that stops exactly where the answer’s first token goes. Each option gets a letter, so the whole answer is a single token.

      A yes/no question (noul) gets A and B, a choice gets one letter per option, a score one per level, and a numeric question one per anchor value across your range.

    2. 2

      One shared prefill

      The state goes through the model once. What the model computes from it, the KV cache (its working memory of your input), is shared by every question.

      The questions branch off it like a tree, so any text they have in common is computed only once, and prompt openings that repeat across requests stay cached.

    3. 3

      One batched decode

      All the questions are decoded together in a single llama_decode call, with one KV sequence each. Four questions cost one pass, not four conversations.

      At startup, snap checks which method the model’s architecture supports and uses it. There’s nothing to configure.

    4. 4

      One logit row

      For each question, snap reads a single row of logits, keeps only the option letters, merges the different tokens for the same letter (like “A” and “ A”) and turns them into your distribution with a softmax.

      coverage tells you how much of the model’s probability landed on your letters. snap calibrate can fit a temperature for each question type.

    ABCDEFGHIJKLMNOPQRSTUVWXYZ

    These 26 letters are everything snap can output. With more than 26 options, each option gets its own yes/no question over the same shared prefix, up to 256.

    No text channel for prompt injection.

    Injection works by getting a model to produce text you didn’t intend: a leaked system prompt, a smuggled tool call, a link your UI renders. snap never produces text. What leaves the engine is a probability over options you wrote, in a JSON shape the server builds, not the model.

    AttackLLM that generates textsnap
    Data exfiltration through the response Possible. The response is free text. No text channel. Answers are numbers keyed by your option names.
    A smuggled tool call or command Possible. The model writes the call. Nothing to execute. Your code maps a decided option to an action you wrote.
    A leaked system prompt or context Possible. It can be asked to repeat itself. Nothing is generated for it to leak into.
    Markup or links injected downstream Possible. Output gets rendered in a UI or an email. The response schema is fixed by the server.
    A broken output format Handled with validators, repair and retries. Can’t happen. The model doesn’t write the response.
    Persuasion toward a different answer Possible, and invisible in the output. Still possible, but limited and measurable. Text in the state can only shift probability between your options. That shows up as split confidence or low coverage, and allow_abstain lets the model say it can’t tell. Set a threshold and send those cases to a person.

    On your hardware

    A single binary with llama.cpp built in, the same on a laptop, a VM, a Kubernetes pod or an air-gapped rack. There’s also a fully static Linux build, with no Python and no runtime to keep patched.

    No data egress

    There’s no API key, no telemetry and no third party handling your records. The weights are a GGUF file on disk, and after the first download snap runs offline. A daily check for new releases prints to stderr, and SNAP_NO_UPDATE_CHECK=1 turns it off.

    Auditable by design

    The same input always gives the same distribution. You can log the full probabilities with every decision, replay it later, and check x_snap to see exactly what the engine did.

    Open source

    It’s MIT-licensed Rust you can read end to end. Models come from a tested list, and a calibration file won’t load with a different model or prompt version than the one it was fitted on.

    Latency and accuracy on the same weights.

    All numbers are medians of snap bench --requests 20 on an Apple M1 Max with Metal, and every one is reproducible with snap bench and snap evaluate eval/cases.jsonl.

    snap vs Ollama 0.34.3, same weights, same machine

    Ollama, JSON + probabilitiesOllama, JSON answersOllama, one letter eachsnap
    1 question
    898 ms · 69 tokens
    163 ms · 9 tokens
    61 ms · 1 token
    54 ms · 0 tokens
    −11%
    4 questions
    3704 ms · 282 tokens
    342 ms · 26 tokens
    257 ms · 4 tokens
    135 ms · 0 tokens
    −48%
    8 questions
    7874 ms · 562 tokens
    732 ms · 50 tokens
    504 ms · 8 tokens
    263 ms · 0 tokens
    −48%
    5 KB state, 1 question
    3872 ms · 73 tokens
    2738 ms · 8 tokens
    2650 ms · 1 token
    2226 ms · 0 tokens
    −16%

    Both engines run MiniCPM5-2B on the same machine, every request uses data neither has seen, and times are the median of 15 runs. The percentage compares snap with Ollama’s fastest option in each row, which is always one letter per question: the quickest Ollama can go, but it gives you text to parse, no probabilities, and one request per question (when we asked for all the letters in one reply, it stopped early). JSON under a schema is what a pipeline actually uses, and with probabilities it carries the same information as a snap answer, written one token at a time. You can reproduce it with eval/vs_ollama.py.

    Latency with a shared state

    scenariominicpm5-2bspark-4bqwen3.8-4b
    single question50 ms89 ms82 ms
    4 questions, shared126 ms 32/q232 ms 58/q227 ms 57/q
    8 questions, shared252 ms 32/q406 ms 51/q472 ms 59/q
    8 questions, direct1170 ms 146/q2158 ms 270/q2381 ms 298/q
    8 questions on a 1.8 KB document1887 ms 236/q3278 ms 410/q3327 ms 416/q
    4 KB state, 1 question1364 ms2476 ms2164 ms

    All the prompts in a request are decoded in one batched call that shares their common prefix, on every architecture; the direct row decodes each question on its own, for comparison. Qwen’s hybrid architecture runs 17 parallel sequences instead of 65, and snap picks the right path at startup.

    Accuracy · eval/cases.jsonl, 301 cases

    modelaccuracyms/case
    qwen3.8-4b Q4_K_M86.4%~300
    spark-4b Q8_085.7%~245
    minicpm5-2b Q4_K_M69.1%~135

    Raw probabilities aren’t calibrated. snap calibrate fits one temperature per question type on your labeled cases, ties it to the model and prompt version, and prints the calibration error (ECE) before and after.

    A drop-in replacement for Jev.

    Same request and response format as Jev, so a client written for Jev works with a local snap server, unchanged. The additions are optional fields a Jev client can ignore: status, confidence, coverage and the full probabilities on every answer.

    noulYes or no, with the probability of yes.
    choiceOne of 2 to 256 options, with a probability for each.
    scoreAn ordered level, with an expected score between 0 and 1.
    numericA number in your range, read from anchor values.
    allow_abstainAn extra answer for “not enough information”.
    POST/v1/systemonerequest
    {
      "state": "My payouts have been failing for 3 days.",
      "questions": {
        "urgent": {
          "type": "noul",
          "instructions": "SLA breach?"
        },
        "team": {
          "type": "choice",
          "criteria": {
            "billing": "payments, refunds",
            "technical": "bugs, outages"
          }
        }
      }
    }
    200 OKresponse
    {
      "answers": {
        "urgent": {
          "type": "noul",
          "status": "decided",
          "noul": 0.7285,
          "boolean": true,
          "confidence": 0.4571,
          "coverage": 0.9993
        },
        "team": {
          "type": "choice",
          "status": "decided",
          "choice": "billing",
          "probabilities": {
            "billing": 0.9828,
            "technical": 0.0172
          }
        }
      }
    }

    A real response from minicpm5-2b, without the usage and x_snap fields.

    snap next to Jev and SemIf.

    The two closest projects are Jev, the hosted original, and SemIf, an open research scorer.

    CapabilitysnapJevSemIf
    DeploymentLocal, one Rust binaryHosted serviceLocal Python, WebGPU demo
    Open weightsGGUFNoQwen
    Jev-compatible HTTP APIYesThe originalCLI only
    Numeric answersYesNoNo
    Calibrated probabilitiesPer question typeClaimedPer workload
    TypeSafe public eval73% on a 4B model~88%, self-reportedNot reported

    The Jev column comes from its public description and the SemIf column from its repository. The TypeSafe row is agreement with the published consensus on its 373 public decisions, reproducible with eval/typesafe_public.py.

    Three steps to your first decision.

    The first run downloads the model once (1.5 to 4 GB). After that, everything runs offline on your machine.

    1

    Install

    A single binary: Homebrew on macOS, a tarball everywhere else.

    2

    Serve

    The default model on port 8018, or any other from the tested set with --model.

    $ snap serve
    3

    Ask

    The playground, built into the binary: questions in a form, answers with their full distributions, and the request as cURL.

    localhost:8018/playground
    Open

    Prebuilt for macOS on Apple Silicon, Linux x86_64 and aarch64 (including a fully static build and Vulkan) and Windows x86_64. From source: CUDA, ROCm, Vulkan and Metal.