G-Eval LLM-as-a-Judge: Grading AI Outputs With Another AI, Reliably If you're building anything with LLMs, at some point you need to answer "is this output actually good?" at a scale where a human can't read every response. The standard technique for this now is LLM-as-a-judge: you use a second LLM