Pick metrics for foundation-model evaluation — perplexity, ROUGE, BERTScore, RAGAS, LLM-as-judge — and understand what each one quietly fails to measure.
It covers perplexity, ROUGE, BLEU, BERTScore, RAGAS, LLM-as-judge and human evaluation, including what each one quietly fails to capture. It is written for teams who need to prove an LLM feature improved, not just that it shipped.
Every branch states the trade-off that decided it, so the recommendation you end on comes with the reasoning attached — something you can paste into a design note or defend in a review.
Built by Tarek Atwan — twenty years in data and AI, four books, four-time Pluralsight Elite instructor, Fortune 500 engagements across eight countries. Consulting through Ensemble Methods. Source on GitHub.