RLCD is a calibrated, schema-conditioned extension of pairwise reward modeling: Bradley–Terry becomes Plackett–Luce, and the reward model becomes Jev’s typed decision interface.

68 points•tnspacetime•7 days ago•8 comments•

8 comments

firejake3086 days ago
> The operational signal was always relative preference. The scalar merely hid it.

Is this another Claude-ism? "X was always Y. The Z merely hid it." Or am I overcalling it?

dilyevsky6 days ago
You're right to question this and OP shouldn't have done it.
agos6 days ago
not overcalling, it’s rife with claudisms
WalterGR7 days ago
RLCD, not defined in the article, is Reinforcement Learning for Calibrated Decisions.
borgel6 days ago
Ah, so not Reflective LCD [1] then.

[1] https://www.e3displays.com/reflective-lcd-display-monitor/

tnspacetime6 days ago
If you need any background info on Jev read this:

https://software.human-tokens.dev/

daemonk6 days ago
Yeah the calibration is really what makes it useful in practice for quick, small decisions. Asking a LLM to give scores to a problem will yield inconsistently scaled/anchored results that changes at a whim.

The blog is pretty heavy on statistics. I'll have to study it more when I have time. Is it essentially bootstrapping results to statistically normalize the answers?

tnspacetime6 days ago
I have not studied it properly too. Good that it has both code and note though.

Read the full thread on Hacker News →

Related stories