https://vercel.com/blog/agents-md-outperforms-skills-in-our-agent-evals
AGENTS.md outperforms skills in our agent evals - Vercel
A compressed 8KB docs index in AGENTS.md achieved 100% on Next.js 16 API evals. Skills maxed at 79%. Here's what we learned and how to set it up.
our agentagentsmdskillsevals
https://aligneval.com/
AlignEval: Making Evals Easy, Fun, and Semi-Automated
A prototype tool/game to help you look at your data, label it, evaluate output, and optimize evaluators.
easy funmakingevalssemiautomated
https://github.com/kiln-ai/kiln
GitHub - Kiln-AI/Kiln: Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents,...
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more. - Kiln-AI/Kiln
ai build
https://evals.openai.com/
OpenAI Evals
openaievals
https://support.vectorevaluationsplus.com/s/
Vector Evals+/PD (Formerly Teachpoint)
vectorevalspdformerlyteachpoint
https://commandline.microsoft.com/assert-written-intent-executable-evals/
Turn specs into evals for any agent with ASSERT - Command Line
Jun 27, 2026 - ASSERT is an open-source framework for converting natural language behavior requirements into executable evaluations of AI models and agents.
turnspecsevals
https://github.com/langfuse/langfuse
GitHub - langfuse/langfuse: 🪢 Open source AI engineering platform: LLM evals, observability,...
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain,...
open source aillm evalsgithublangfuse
https://github.com/mclenhard/mcp-evals
GitHub - mclenhard/mcp-evals: A Node.js package and GitHub Action for evaluating MCP (Model Context...
A Node.js package and GitHub Action for evaluating MCP (Model Context Protocol) tool implementations using LLM-based scoring. This helps ensure your MCP...
https://app.evals.net/login
EVALS
evals
https://maven.com/parlance-labs/evals
AI Evals For Engineers & PMs by Hamel Husain and Shreya Shankar on Maven
Learn proven approaches for quickly improving AI applications. Build AI that works better than the competition, regardless of the use-case.
https://visr.dev/
Visr — Integration Evals for Agentic Loops
Visr turns real agent sessions into reusable task evals, score history, and promotion evidence for teams shipping agent workflows.
visrintegrationevalsagenticloops
https://nicknisi.com/posts/writing-my-first-evals/
Writing My First Evals | Nick Nisi
I had no background in evals. I built two very different evaluation systems for two AI-powered developer tools, and they taught me the same lesson: trust isn't...
my firstwritingevalsnicknisi
https://who-to-bother-at.vercel.app/t/nextjs-evals
Next.js Evals Contacts | Who to Bother at Vercel
For questions about Next.js evaluation and testing frameworks
next jsevalscontactsbothervercel
https://apps.law.uci.edu/shib/lawevals/
Law Evals
lawevals
https://archives-manuscripts.dartmouth.edu/repositories/2/archival_objects/53588
OP/2P EVALS BLOCK 1 98/9 PEDI CLERKSHIP SUBJECT FILES | Dartmouth Libraries Archives & Manuscripts
https://drive.google.com/file/d/1nkzn01fqONdhnAUGdWKrMUy_-MjpwYsL/view?usp=sharing
Evals_S19.pdf - Google Drive
evalspdfgoogledrive
https://amycmitchell.substack.com/p/evals-for-product-managers
Think Evals Are Just for AI? Think Again - by Amy Mitchell
Why product managers need evals and how to put them into practice with examples to get started fast
just forthinkevalsaiamy
https://www.timeisnoweducationcenter.com/
Time Is Now Education Center | DWI Classes and Evals | San Antonio
time is noweducation center
https://archives-manuscripts.dartmouth.edu/repositories/2/archival_objects/551119
Course Evals Crewe-Washburn 1970-2009, 1970-01-01 - 2009-12-31 | Dartmouth Libraries Archives &...
course evals
https://www.psychologytoday.com/us/psychiatrists/renewed-adhd-treatment-center-adhd-testing-evals-anaheim-ca/768653
Renewed ADHD Treatment Center-ADHD Testing & Evals, Psychiatric Nurse Practitioner, Anaheim, CA,...
adhd treatmentpsychiatric nurserenewedcentertesting
https://cran.csiro.au/web/packages/pander/vignettes/evals.html
Capturing evaluation information with evals
evaluation informationcapturingevals
https://dev.to/danielsogl/skills-without-evals-are-just-markdown-and-hope-3a71
Skills Without Evals Are Just Markdown and Hope - DEV Community
TL;DR. I built an Anthropic Agent Skill for @ngrx/signals and ran it through the full eval pipeline:... Tagged with claude, ai, angular, ngrx.
skillswithoutevalsmarkdownhope
https://dynamicsgpblogster.blogspot.com/2011/04/microsoft-dynamics-convergence-atlanta_15.html
Microsoft Dynamics Convergence Atlanta 2011: Evals Reminder
Hope you had a great time at Microsoft Dynamics Convergence Atlanta 2011 and that you made it home safely. Now that you have had some time ...
microsoft dynamicsconvergenceatlantaevalsreminder
https://ai-evals.io/
AI-Evals.io
ai evalsio
https://www.foxnews.com/politics/hegseth-incredibly-talented-battle-proven-leader-military-evaluations-show
Trump Defense pick Hegseth performance evals praise a 'battle-proven leader' | Fox News
President-elect Donald Trump's defense secretary nominee Pete Hegseth was praised in military evaluations as an "incredibly talented, battle-proven leader."
https://devblogs.microsoft.com/foundry/build-2026-open-trust-stack-ai-agents/
Build agents you can trust across any framework with open evals and a control standard | Microsoft...
Jun 2, 2026 - Learn how Microsoft helps developers build trustworthy AI agents with open evaluations, portable runtime controls, production observability, and security...
https://archives-manuscripts.dartmouth.edu/repositories/2/archival_objects/563546
Course Schedules/Evals, Cross Listings, Minor Program Info 1997-2003, 1997-01-01 - 2003-12-31 |...
https://rahulgarg.ai/
Rahul Garg | RL Environments & AI Evals
I build RL environments and evals that train AI agents: Dockerized environments, verifiable graders, and DevOps-domain tasks for frontier-model post-training.
rl environmentsrahulgargaievals
https://mobiclass.csc.ncsu.edu/2016/04/team-evals.html
Visual Interfaces for Mobiles @ NCSU: Team evals
Folks, As we approach this semester's crescendo, please make sure that you perform for your team. For those of you who feel one or more ...
visualinterfacesmobilesncsuteam
https://www.langchain.com/langsmith-platform
LangSmith: AI Agent & LLM Observability and Evals Platform
LangSmith is the complete framework agnostic AI agent and LLM observability, evaluation, and deployment platform.
ai agentllm observabilitylangsmithevalsplatform
https://lrs.sog.unc.edu/lrs-subscr-view/bills_summaries/465298/S368
Bill Summaries: S368 PHYSICAL AND PSYCH. EVALS. FOR LEO'S. | Legislative Reporting Service
https://en-us.spreaker.com/episode/the-intel-with-greg-cosell-evals-eagles-free-agents-and-other-nfc-east-additions--64864937
The Intel With Greg Cosell: Evals Eagles Free Agents And Other NFC East Additions
A digital show and podcast featuring free agent and college prospect breakdowns by NFL Films senior producer Greg Cosell, co-host of ESPN's "NFL Matchup Show."...
https://pakodas.substack.com/p/continual-learning-without-evals
Continual Learning Without Evals Is Just Drift
the component that cannot be taken for granted
continual learningwithoutevalsdrift
https://pydantic.dev/jobs/evals-continuous-learning-engineer
Evals & Continuous Learning Engineer | Pydantic
continuous learningevalsengineerpydantic
https://voloshin.net/
Alex Voloshin — Agentic-dev tooling, evals & tri-vendor parity
Helping engineering teams operationalize AI coding agents: tri-vendor parity, eval rubrics, and production patterns for Claude Code, Codex, Windsurf.
dev toolingalexagenticevalstri
https://u22a8.ai/
u22a8.ai — evals without the LLM
Run your AI evals in semantic space, not through an LLM — cheaper, more accurate, and deterministic than LLM-as-a-judge. RAG, safety, quality. API, MCP, CLI.
aievalswithoutllm
https://telehealthevals.com/
TeleHealth Clinical Evals
telehealthclinicalevals
https://gosuevals.com/
Gosu Evals - AI Coding Agent Evaluations
ai coding agentgosuevalsevaluations
https://www.mcpevals.io/
MCP Evals - Evaluate Your MCP Tools with Confidence
A powerful Node.js package and GitHub Action for evaluating Model Context Protocol (MCP) tool implementations using LLM-based scoring.
your toolsmcpevalsevaluateconfidence
https://ednotesonline.blogspot.com/2012/01/will-uftunity-leadership-cave-to.html?showComment=1326675520510
Ed Notes Online: Will the UFT/Unity Leadership Cave to Pressure on Evals While Claiming Victory to...
Ed Notes defends public education and promotes democratic teacher unionism with a focus on the UFT.
https://www.rohitdiwakar.com/
Rohit Diwakar - Multi Agentic AI Developer | LangGraph, Evals, ADK, A2A, MCP, N8N
agentic ai developer
https://dev.to/nazar-boyko/llm-evals-for-developer-tools-useful-correct-safe-33jg
LLM Evals For Developer Tools: Useful, Correct, Safe - DEV Community
Jul 16, 2026 - Someone on your team built an LLM feature. Maybe it's an inline code-suggest. Maybe it's a
llm evalsfor developertoolsusefulcorrect
https://www.ajc.com/news/navy-yard-shooter-navy-evals-career-come-light/uq20e8SLH1pWcB8LsyHLtL/
Navy Yard shooter's Navy evals, career come to light
A recently published story by The Navy Times provides some new details about the Navy career of Aaron Alexis, the man who authorities say killed 12 people at...
navy yardshooterevalscareercome
https://doctorguilford.com/
Capital Psychology Consultants | ADHD Evals | Tallahassee, FL 32301
tallahassee flcapitalpsychologyconsultantsadhd
https://fastforward.utoronto.ca/webtool-tag/course-evals/
course evals | Webtool Tags | fastforward
course evalswebtooltagsfastforward
https://archives-manuscripts.dartmouth.edu/repositories/2/archival_objects/23990
BUDGET/VISA/MJD PRESENTATIONS/EVALS/MENTOR, 1/1/2000-12/31/2004 | Dartmouth Libraries Archives &...
https://www.samaramunizpsychology.com/
Trauma & multiculturalism therapy. Psych evals. Boston
Therapy for Trauma and Multiculturalism, and Psych eval for immigration, accident and restraining orders in Boston, Massachusetts. Bilingual Services.
traumamulticulturalismtherapypsychevals
https://www.paloaltonetworks.com/cortex/cortex-xdr/mitre
Cortex XDR Performance in MITRE Evals - Palo Alto Networks
See how Cortex XDR by Palo Alto Networks performs in MITRE evaluations, showcasing advanced EDR capabilities to detect and prevent cybersecurity threats.
cortex xdrpalo altoperformancemitreevals
https://developers.openai.com/cookbook/examples/evaluation/use-cases/completion-monitoring
Evals API Use-case - Monitoring stored completions
Evals are task-oriented and iterative, they're the best way to check how your LLM integration is doing and improve it. In the following eva
use caseevalsapimonitoringstored
https://deepmind.google/research/evals/
Evals — Google DeepMind
evalsgoogledeepmind
https://www.acc.af.mil/News/Article-Display/Article/2395465/wsep-east-kicks-off-fy21-with-air-to-air-and-air-to-ground-evals/
WSEP East kicks off FY21 with Air-to-Air and Air-to-Ground Evals Air Combat Command Article...
The first Weapons System Evaluation Program-East of Fiscal Year 2021 provided air-to-air and air-to-ground training and evaluation from October 13-23, 2020. ,
https://developers.openai.com/cookbook/examples/evaluation/getting_started_with_openai_evals
Getting Started with OpenAI Evals
**Note: OpenAI now has a hosted evals product with an API! We recommend you use this instead. See Evals** The OpenAI Evals framework consis
getting started withopenaievals
https://www.acesevals.com/
ACES - ATLAS Clinical Evals Software
acesatlasclinicalevalssoftware