grep for meaning: ask a question about lines, logs, diffs or JSON and explore the answers live. Runs locally on CPU, one Rust binary. - sfmqrb/gutcheck

5 points•sfmqrb•5 days ago•1 comment•

1 comment

sfmqrb5 days ago
How it works / tradeoffs

gutcheck is a thin CLI around a typed decision model (Laya multilingual). For each line it builds something like [question] [options] [text], runs one forward pass, and reads either:

- noul — P(yes) for a yes/no question (-n filters, -s prints the score), or - choice — softmax over labels you passed with -c.

No autoregressive generation. No embedding index. No chunking/vector DB. Every line is judged independently against your question.

Why CPU is usable on logs anyway: the model is compute-bound (~10 distinct lines/sec on a laptop). Speed comes from asking less: identical lines answered once; --fuzzy shares answers across lines that differ only in numbers/timestamps/ids; --estimate counts forward passes first. Loghub 500-line samples: Apache 57s→4s (99.8% same), HDFS 101s→9s (96.8%), Linux 93s→9s (99.4%).

vs jgrep: closest relative. jgrep (hosted/API) generally wins on raw throughput for distinct free text. gutcheck wins on privacy/offline, CPU-only with no key, repetitive logs + fuzzy collapsing, and Unix pipes.

Not claiming: semantic search, milliseconds on arbitrary corpora, or high zero-shot accuracy out of the box. RAM/cold-start still measuring.

Happy to answer questions about the ONNX packaging, the fuzzy normalizer, or where a decision filter belongs next to regex in a pipeline.

Read the full thread on Hacker News →

Related stories