Evaluating an LLM Is Different From Evaluating a RAG System

I recently read Hamel Husain and Shreya Shankar’s article about similarity metrics. Their argument is that metrics like ROUGE and BERTScore often miss the failures that matter in an LLM application. I agree. But I think we need to separate two things: evaluating what a model learns during training and evaluating whether a RAG application works. Both involve an LLM, but the questions are different. If you do not understand the question, choosing a metric will not help you much. ...

October 6, 2026 · 5 min · 982 words · Necati Demir