https://datumo.com/
Advanced LLM Evaluation Platform by datumo
Apr 7, 2026 - Discover the power of Datumo Eval for LLM evaluation. Generate industry-specific question datasets and assess question quality effortlessly.
llm evaluationadvancedplatform
https://reactor.microsoft.com/en-us/reactor/events/26291/?ref=berberich.dev
Building Scalable LLM Evaluation Pipelines with Azure Cosmos DB | Microsoft Reactor
Learn new skills, meet new peers, and find career mentorship. Virtual events are running around the clock so join us anytime, anywhere!
azure cosmos dbllm evaluationbuildingscalablepipelines
https://openevals-evaluation-guidebook.hf.space/
The LLM Evaluation Guidebook
Understanding the tips and tricks of evaluating an LLM in 2025
llm evaluationguidebook
https://aws.amazon.com/blogs/machine-learning/operationalize-llm-evaluation-at-scale-using-amazon-sagemaker-clarify-and-mlops-services/
Operationalize LLM Evaluation at Scale using Amazon SageMaker Clarify and MLOps services |...
Feb 5, 2024 - In the last few years Large Language Models (LLMs) have risen to prominence as outstanding tools capable of understanding, generating and manipulating text...
llm evaluation
https://arxiv.org/abs/2511.08598v1
[2511.08598v1] OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open...
Abstract page for arXiv paper 2511.08598v1: OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking
llm evaluation
https://dokimos.dev/
Dokimos | LLM Evaluation Framework for Java
Dokimos is an evaluation framework for LLM applications in Java. It helps you evaluate responses, track quality over time, and catch regressions before they...
llm evaluationframeworkjava
https://aihub.hkuspace.hku.hk/effective-cross-lingual-llm-evaluation-with-amazon-bedrock/
Effective cross-lingual LLM evaluation with Amazon Bedrock - HKU SPACE AI Hub
Evaluating the quality of AI responses across multiple languages presents significant challenges for organizations deploying generative AI solutions globally....
llm evaluation
https://agenta.ai/
Agenta - Prompt Management, Evaluation, and Observability for LLM apps
Agenta is an open-source platform for building robust LLM Application. It provides tools for prompt engineering, evaluation, debugging, and monitoring of...
prompt managementfor llmagentaevaluationobservability
https://openreview.net/forum?id=PuhF0hyDq1
New Evaluation Metrics Capture Quality Degradation due to LLM Watermarking | OpenReview
With the increasing use of large-language models (LLMs) like ChatGPT, watermarking has emerged as a promising approach for tracing machine-generated content....
evaluation metricsdue tonewcapturequality
https://writing-showdown.com/
LLM Writing Evaluation Results
writing evaluationllmresults
https://research.google/pubs/towards-a-human-in-the-loop-framework-for-reliable-patch-evaluation-using-an-llm-as-a-judge/
Towards A Human-in-the-Loop Framework for Reliable Patch Evaluation using an LLM-as-a-Judge
https://llm-eval.github.io/pages/leaderboard/advprompt.html
Adversarial robustness | LLM Evaluation
adversarial robustnessllmevaluation
https://arxiv.org/abs/2410.03608
[2410.03608] TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
Abstract page for arXiv paper 2410.03608: TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
https://gsi.upm.es/en/investigacion/lineas-de-investigacion?view=publication&task=show&id=724
Publication - A Modular LLM-Enhanced Agent-Based System for the Generation and Evaluation of...
https://arxiv.org/html/2504.20612v1?ref=canartuc.com
The Hidden Risks of LLM-Generated Web Application Code: A Security-Centric Evaluation of Code...
https://aws.amazon.com/id/blogs/machine-learning/llm-as-a-judge-on-amazon-bedrock-model-evaluation/
LLM-as-a-judge on Amazon Bedrock Model Evaluation | Artificial Intelligence
Feb 12, 2025 - This blog post explores LLM-as-a-judge on Amazon Bedrock Model Evaluation, providing comprehensive guidance on feature setup, evaluating job initiation through...
llm as a judgeon amazonmodel evaluation
https://waingram.github.io/publications/ingram2025learning/
Learning from LLM Disagreement in Retrieval Evaluation
William A. Ingram is a scientist an academic leader who uses AI to structure and interpret scholarly data, advancing discovery, synthesis, accessibility,...
learning fromllmdisagreementretrievalevaluation
https://gsi.upm.es/en/investigacion/publicaciones?view=publication&task=show&id=728
Publication - Evaluation of Diversity in LLM-Based News Discovery Through an Agent-Based System
https://arxiv.org/abs/2601.03986
[2601.03986] Benchmark^2: Systematic Evaluation of LLM Benchmarks
Abstract page for arXiv paper 2601.03986: Benchmark^2: Systematic Evaluation of LLM Benchmarks
benchmarksystematicevaluationllm
https://arxiv.org/abs/2412.11417v1
[2412.11417v1] RL-LLM-DT: An Automatic Decision Tree Generation Method Based on RL Evaluation and...
Abstract page for arXiv paper 2412.11417v1: RL-LLM-DT: An Automatic Decision Tree Generation Method Based on RL Evaluation and LLM Enhancement
https://arxiv.org/html/2412.11417v2
RL-LLM-DT: An Automatic Decision Tree Generation Method Based on RL Evaluation and LLM Enhancement
https://aws.amazon.com/blogs/machine-learning/llm-as-a-judge-on-amazon-bedrock-model-evaluation/
LLM-as-a-judge on Amazon Bedrock Model Evaluation | Artificial Intelligence
Feb 12, 2025 - This blog post explores LLM-as-a-judge on Amazon Bedrock Model Evaluation, providing comprehensive guidance on feature setup, evaluating job initiation through...
llm as a judgeon amazonmodel evaluation