Portable local probabilistic inference runtime. Contribute to bulyaki/Credence development by creating an account on GitHub.

1 points•naughtiusmax•4 days ago•1 comment•

1 comment

naughtiusmax4 days ago
Credence is a small C++ runtime on top of llama.cpp that turns a GGUF model into a typed decision function. You give it evidence and a proposition and get back a boolean, the probability the model put on "true", and diagnostics that say whether that probability means anything. No text is generated: the answer is read off the next-token logits over a fixed set of labels, so a decision costs one prompt decode.

I built it because I liked the shape of Jev (TypeSafe AI's decision-model API: booleans, choices and scores with probabilities instead of prose) but not that it is a closed hosted model. Credence is the same interface, local, with a model you pick and can inspect.

Example, Qwen3-4B Q4_K_M on an M4 Pro:

  credence decide --model Qwen3-4B-Q4_K_M.gguf \
    --profile profiles/qwen3/qwen3-decision-v2.json \
    --context "Invoice #4471 was issued on 3 March. The full balance remains unpaid and a reminder was sent." \
    --proposition "The invoice has been paid."

  {
    "value": false,
    "decision_basis": "raw_probability",
    "raw_probability": 8.03059e-11,
    "calibrated_probability": null,
    "entropy": 1.94703e-09,
    "top_two_margin": 1,
    "answer_conformity": 1,
    "decision_profile": "qwen3-decision-v2",
    "top_tokens": [{"piece": "2", "probability": 1}, {"piece": "1", "probability": 8.03059e-11}, ...]
  }
Change the context to "paid in full on 10 March" and you get value true, raw_probability 1. The whole CLI call is about 0.8 s, most of it loading the 2.5 GB model; the decision is one forward pass.

How it works: a JSON profile renders evidence and proposition into the model's native chat format with a fixed answer prefix (for Qwen3, an empty <think> block), states that label 1 means true and label 2 means false, and reads the logits at the first answer position. The two labels are normalised against each other to get P(true). answer_conformity is how much full-vocabulary mass landed on the permitted labels at all; if it is low, the model was not playing the game and the probability is junk. Multi-token labels are scored on copied KV-cache branches.

Things to be upfront about:

- Raw probabilities from a 4B model are absurdly overconfident; 1 - 1e-10 is not a belief anyone should act on. So there is a calibration step: tools/calibrate.py fits Platt scaling (temperature plus bias) on a labelled JSONL set, and credence decide applies it and says so in decision_basis. On the 22 checked-in cases the 4B model gets 21/22 raw; after calibration the "false" above reports 0.04 instead of 1e-10 and the one miss (a double negation) flips. Those are training-set numbers on 22 samples, and the fitter warns you about exactly that.

- Model size matters more than the plumbing. Qwen3-0.6B, the smoke-test model, answers "true" to almost everything: 11/22. Before I fixed the prompt to state which label meant true it was also 11/22 with far worse log loss. A small model will happily pick a label without knowing what it stands for.

- Boolean-only today. Choice and score operations, a C ABI, a resident daemon so you do not pay the model load per call, and an MCP adapter are next. There is no library API yet, just the CLI.

- Untrusted evidence is delimited and flagged in the output, but that is delimiters, not a defence against prompt injection.

C++20, llama.cpp pinned as a submodule, Apache 2.0. CMake presets for macOS (Metal or CPU), Linux and Windows; ctest runs unit tests plus an end-to-end check on the smoke model.

I would like to hear from anyone who has run logit-scored classification with local models at scale, especially how you evaluated calibration with sets larger than mine.

Read the full thread on Hacker News →

Related stories