Evaluating an LLM Is Different From Evaluating a RAG System
I recently read Hamel Husain and Shreya Shankar’s article about similarity metrics. Their argument is that metrics like ROUGE and BERTScore often miss the failures that matter in an LLM application. I agree. But I think we need to separate two things: evaluating what a model learns during training and evaluating whether a RAG application works. Both involve an LLM, but the questions are different. If you do not understand the question, choosing a metric will not help you much. ...