Portable local probabilistic inference runtime. Contribute to bulyaki/Credence development by creating an account on GitHub.
1 comment
I built it because I liked the shape of Jev (TypeSafe AI's decision-model API: booleans, choices and scores with probabilities instead of prose) but not that it is a closed hosted model. Credence is the same interface, local, with a model you pick and can inspect.
Example, Qwen3-4B Q4_K_M on an M4 Pro:
credence decide --model Qwen3-4B-Q4_K_M.gguf \
--profile profiles/qwen3/qwen3-decision-v2.json \
--context "Invoice #4471 was issued on 3 March. The full balance remains unpaid and a reminder was sent." \
--proposition "The invoice has been paid."
{
"value": false,
"decision_basis": "raw_probability",
"raw_probability": 8.03059e-11,
"calibrated_probability": null,
"entropy": 1.94703e-09,
"top_two_margin": 1,
"answer_conformity": 1,
"decision_profile": "qwen3-decision-v2",
"top_tokens": [{"piece": "2", "probability": 1}, {"piece": "1", "probability": 8.03059e-11}, ...]
}
Change the context to "paid in full on 10 March" and you get value true, raw_probability 1. The whole CLI call is about 0.8 s, most of it loading the 2.5 GB model; the decision is one forward pass.How it works: a JSON profile renders evidence and proposition into the model's native chat format with a fixed answer prefix (for Qwen3, an empty <think> block), states that label 1 means true and label 2 means false, and reads the logits at the first answer position. The two labels are normalised against each other to get P(true). answer_conformity is how much full-vocabulary mass landed on the permitted labels at all; if it is low, the model was not playing the game and the probability is junk. Multi-token labels are scored on copied KV-cache branches.
Things to be upfront about:
- Raw probabilities from a 4B model are absurdly overconfident; 1 - 1e-10 is not a belief anyone should act on. So there is a calibration step: tools/calibrate.py fits Platt scaling (temperature plus bias) on a labelled JSONL set, and credence decide applies it and says so in decision_basis. On the 22 checked-in cases the 4B model gets 21/22 raw; after calibration the "false" above reports 0.04 instead of 1e-10 and the one miss (a double negation) flips. Those are training-set numbers on 22 samples, and the fitter warns you about exactly that.
- Model size matters more than the plumbing. Qwen3-0.6B, the smoke-test model, answers "true" to almost everything: 11/22. Before I fixed the prompt to state which label meant true it was also 11/22 with far worse log loss. A small model will happily pick a label without knowing what it stands for.
- Boolean-only today. Choice and score operations, a C ABI, a resident daemon so you do not pay the model load per call, and an MCP adapter are next. There is no library API yet, just the CLI.
- Untrusted evidence is delimited and flagged in the output, but that is delimiters, not a defence against prompt injection.
C++20, llama.cpp pinned as a submodule, Apache 2.0. CMake presets for macOS (Metal or CPU), Linux and Windows; ctest runs unit tests plus an end-to-end check on the smoke model.
I would like to hear from anyone who has run logit-scored classification with local models at scale, especially how you evaluated calibration with sets larger than mine.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 9 days ago
- Hacker News · 2 points · 4 days ago
- Hacker News · 2 points · 10 days ago
- Hacker News · 1 points · 5 days ago
- Show HN: A local alternative to Jev – 94% on Banking77gist.github.comHacker News · 2 points · 4 days ago
- Hacker News · 3 points · 3 days ago