view post Post 2644 How do I test an LLM for my unique needs?If you work in finance, law, or medicine, generic benchmarks are not enough.This blog post uses Argilla, Distilllabel and 🌤️Lighteval to generate evaluation dataset and evaluate models.https://github.com/argilla-io/argilla-cookbook/blob/main/domain-eval/README.md
benchmarks meituan-longcat/LARYBench Updated Apr 30 • 2.94k • 18 llamaindex/ParseBench Benchmark • Updated Apr 19 • 169k • 14.1k • 101 nvidia/QCalEval Viewer • Updated Apr 13 • 243 • 1.16k • 19 allenai/olmOCR-bench Benchmark • Updated Feb 19 • 7.83k • 256
RULER Datasets Falcon-H1-3B-Base RULER Datasets lighteval/RULER-131072-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 29 lighteval/RULER-65536-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 79 lighteval/RULER-32768-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 16 lighteval/RULER-16384-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 25
benchmarks meituan-longcat/LARYBench Updated Apr 30 • 2.94k • 18 llamaindex/ParseBench Benchmark • Updated Apr 19 • 169k • 14.1k • 101 nvidia/QCalEval Viewer • Updated Apr 13 • 243 • 1.16k • 19 allenai/olmOCR-bench Benchmark • Updated Feb 19 • 7.83k • 256
RULER Datasets Falcon-H1-3B-Base RULER Datasets lighteval/RULER-131072-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 29 lighteval/RULER-65536-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 79 lighteval/RULER-32768-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 16 lighteval/RULER-16384-Falcon-H1-3B-Base Viewer • Updated Jun 18, 2025 • 6.5k • 25