Open-source contrastive embedding stack for mapping technical CVE disclosures to cybersecurity compliance controls (NIST SP 800-53 / CMMC) - applied-inference-lab/autormf-ml-lab
2 comments
Curious how you set up the eval. Did you hold out whole controls, or whole frameworks, so the encoder had to map things it never saw during training? That's usually where fine-tuned models look great on a random split and then drop off.
Disclosure: I work on Ookami (github.com/KatyarAILabs/Ookami), an open-source layer for serving and fine-tuning models. That comparison is basically what our eval gate automates: a new fine-tuned version only goes live if it beats the current one on your own held-out and audit sets, with confidence intervals rather than one accuracy number. It's built for LLMs today rather than encoders, but "prove the small model beats the incumbent before shipping it" is the same idea.
The problem is a severe vocabulary gap: vulnerability reports describe low-level exploit mechanics (buffer overflows, memory corruption, protocol stacks), while compliance controls describe administrative governance (flaw remediation, boundary protection, update testing).
When we benchmarked standard approaches against an isolated 2,000-sample validation set, we hit unexpected results:
Zero-shot foundation models (BGE-Large, GTE-Qwen2-7B) got less than 1% strict top-1 accuracy out of the box. Adding BM25 hybrid search via Reciprocal Rank Fusion (RRF) caused strict accuracy to collapse from 77.45% down to 13.65% due to lexical rank dilution across mismatched vocabularies. Adding a 7B generative LLM reranker (Qwen2.5-7B-Instruct) downstream degraded top-1 accuracy by over 5 percentage points (77.45% down to 72.00%) because the general LLM was distracted by operational attack details. What worked was a compact 110M parameter bi-encoder (BGE-Base) fine-tuned with Multiple Negatives Ranking Loss (MNRL). It reached 77.45% strict top-1 accuracy and 87.15% parent control roll-up accuracy. We quantized it to GGUF (Q8_0) so it runs in-process with sub-millisecond latency on commodity CPU hardware.
We decided to open source the full stack (training pipelines, LoRA configurations, evaluation suites, GGUF conversion tools, and the 9-page white paper) under Apache 2.0 via Applied Inference Lab rather than keep it proprietary.
The white paper PDF is attached to the v1.0.0 release: https://github.com/applied-inference-lab/autormf-ml-lab/rele...
Happy to answer questions about the contrastive setup, the failure modes, or the evaluation methodology.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 10 days ago
- Hacker News · 3 points · 4 days ago
- Local JEV, fine tuned in < 30 mins on Mac Airshmcsensei.github.ioHacker News · 1 points · 7 days ago
- Ars Technica · 0 points · 13 days ago
- Hacker News · 2 points · 9 days ago
- Lobsters · 17 points · about 4 years ago