Open-source contrastive embedding stack for mapping technical CVE disclosures to cybersecurity compliance controls (NIST SP 800-53 / CMMC) - applied-inference-lab/autormf-ml-lab

2 points•jpusateri•6 days ago•2 comments•

2 comments

katyarailabs6 days ago
Nice result, and it matches what we keep seeing: a small model fine-tuned on a narrow, well-defined task beats a general 7B model that has to reason its way there. Control mapping is about as narrow as it gets.

Curious how you set up the eval. Did you hold out whole controls, or whole frameworks, so the encoder had to map things it never saw during training? That's usually where fine-tuned models look great on a random split and then drop off.

Disclosure: I work on Ookami (github.com/KatyarAILabs/Ookami), an open-source layer for serving and fine-tuning models. That comparison is basically what our eval gate automates: a new fine-tuned version only goes live if it beats the current one on your own held-out and audit sets, with confidence intervals rather than one accuracy number. It's built for LLMs today rather than encoders, but "prove the small model beats the incumbent before shipping it" is the same idea.

jpusateri6 days ago
Author here. In building compliance automation at Middle Coast Software, one of the biggest friction points was mapping technical vulnerability scans (CVEs) to formal security control catalogs (NIST SP 800-53 Rev 5 and CMMC).

The problem is a severe vocabulary gap: vulnerability reports describe low-level exploit mechanics (buffer overflows, memory corruption, protocol stacks), while compliance controls describe administrative governance (flaw remediation, boundary protection, update testing).

When we benchmarked standard approaches against an isolated 2,000-sample validation set, we hit unexpected results:

Zero-shot foundation models (BGE-Large, GTE-Qwen2-7B) got less than 1% strict top-1 accuracy out of the box. Adding BM25 hybrid search via Reciprocal Rank Fusion (RRF) caused strict accuracy to collapse from 77.45% down to 13.65% due to lexical rank dilution across mismatched vocabularies. Adding a 7B generative LLM reranker (Qwen2.5-7B-Instruct) downstream degraded top-1 accuracy by over 5 percentage points (77.45% down to 72.00%) because the general LLM was distracted by operational attack details. What worked was a compact 110M parameter bi-encoder (BGE-Base) fine-tuned with Multiple Negatives Ranking Loss (MNRL). It reached 77.45% strict top-1 accuracy and 87.15% parent control roll-up accuracy. We quantized it to GGUF (Q8_0) so it runs in-process with sub-millisecond latency on commodity CPU hardware.

We decided to open source the full stack (training pipelines, LoRA configurations, evaluation suites, GGUF conversion tools, and the 9-page white paper) under Apache 2.0 via Applied Inference Lab rather than keep it proprietary.

The white paper PDF is attached to the v1.0.0 release: https://github.com/applied-inference-lab/autormf-ml-lab/rele...

Happy to answer questions about the contrastive setup, the failure modes, or the evaluation methodology.

Read the full thread on Hacker News →

Related stories