← map
○  learning

LLM-as-judge

not started  ·  evaluation

Using a model to grade model output. Cheap and scalable, and quietly vulnerable to position bias and self-preference.

Nothing written yet. It's on the map because I want to understand it. The node exists so I can see the gap.