I recently read Hamel Husain and Shreya Shankar’s article about similarity metrics. Their argument is that metrics like ROUGE and BERTScore often miss the failures that matter in an LLM application.

I agree.

But I think we need to separate two things: evaluating what a model learns during training and evaluating whether a RAG application works.

Both involve an LLM, but the questions are different. If you do not understand the question, choosing a metric will not help you much.

What Does a Better Training Result Mean

When you train a language model, you change its parameters. For an autoregressive model, the basic objective is to get better at predicting the next token.

Loss and perplexity help measure that. Evaluating them on held-out text tells you how well the model predicts text it did not train on. You still need consistent measurement conditions; tokenization and context affect perplexity.

That is useful information. But it does not tell you everything about the model’s behavior.

Maybe your model predicts text better, but still ignores an instruction. Maybe fine-tuning improves overall accuracy, but makes an important category worse.

The second problem happened in my Laya experiment. Fine-tuning improved accuracy, but hate-speech recall fell. Whether that tradeoff was acceptable depended on how the application handled those classes.

So even during model development, I want to measure the actual task. For classification, check the classes you care about. For generated code, run it against tests. For instruction following, check whether the instructions were followed.

The training result tells you something specific. You have to understand what that is before using it to make a deployment decision.

RAG Adds More Places for Things to Go Wrong

A RAG system retrieves information from an external source and gives it to the model. Now the answer depends on the documents, retrieval, context assembly, and generation.

There is some overlap with training here. The original RAG paper trained the query encoder and generator together. Fine-tuning and RAG can belong in the same system.

Imagine a store’s support assistant. Its policy says standard items can be returned within 30 days, but clearance items are final sale.

A customer asks whether they can return a clearance item after 20 days.

The assistant retrieves the general return policy and misses the clearance exception. Then it tells the customer they can return the item.

The retrieved document was relevant. The answer was still wrong.

Now imagine that the assistant retrieves the complete policy, including the exception, and gives the same wrong answer. This time, the information reached the model. The model failed to apply it.

Those failures need different fixes.

For this assistant, I would check three things:

  1. Did the necessary evidence reach the model?
  2. Did the model use that evidence correctly?
  3. Did the final answer resolve the customer’s question?

One score cannot explain all three.

A Similar Answer Can Still Be Wrong

This is where similarity metrics can become misleading.

ROUGE measures overlap with reference text. BERTScore uses contextual embeddings to compare candidate and reference tokens. They can be useful if you have checked that their scores reflect quality for your task.

But consider these two answers:

Clearance items are eligible for return.

Clearance items are not eligible for return.

Almost all the words are the same. One word changes what the customer should do.

For our assistant, I want an evaluation that reliably catches that difference. I also want it to catch the missing exception.

Cosine similarity can help a retriever find passages. But a high similarity score does not establish that the passage contains enough information to answer. A document about returns can still leave out the only fact that matters for this customer.

ROUGE and BERTScore are also not the usual next-token training objective. There is no general rule that makes them appropriate for training and inappropriate for RAG. Their usefulness depends on what you need to measure.

Following the Source Is Not Enough

There is another distinction I think matters: correctness and faithfulness.

Faithfulness checks whether the answer’s claims are supported by the supplied context. That is what the Ragas faithfulness metric measures.

Suppose our assistant retrieves an old policy that allowed clearance returns. The model follows it perfectly.

The answer is faithful to that document. It is wrong under the current policy.

The reverse can happen too. The model might give the correct answer from its learned knowledge while the retrieved documents do not support it. If your product requires answers backed by current company documents, that matters.

And when the available information is insufficient, the assistant should recognize that. I would measure both unsupported answers and unnecessary refusals. Refusing every request does not make a useful support assistant.

Evaluate What You Would Actually Fix

I would start with real questions and read the failed answers. Then I would trace where each failure happened.

Was the policy missing from the knowledge base? Did search miss it? Was the exception removed when the context was assembled? Or did the model receive it and still answer incorrectly?

Try giving the same model the complete current policy yourself. If the answer improves, investigate retrieval and context assembly. If it still fails, investigate how the model uses the information. Keep in mind that changing context can also change ordering and distractions.

Use code checks for requirements you can test directly. For judgments about meaning, a model judge can help, but compare its decisions with someone who understands the domain. The judge can miss the same exception you are trying to catch.

If I fine-tune this assistant, I still want to test the clearance question. If I change its retriever, I want to test it again.

The customer’s requirement stays the same. The checks that explain a failure depend on what changed.

Before choosing a metric, be clear about the behavior you need and the mistake you need to catch. Then measure those things.