AI evaluation and human input: a quick guide

A quick guide to AI evaluation, from model benchmarks and application evals to user evals, human input, and AI experience benchmarking.

There are a huge range of different ways to test AI products, each focused on different aspects of AI and different testing methodologies. At first glance, the choices and their specialized language can feel like an overgrown jungle – so dense that you can’t make sense of it, let alone find a path through. This document untangles the family tree of AI evaluation, and describes where and how human input can help.

A very simplified view of the AI evaluation world

The meaning of “eval”

The first thing we need to untangle is the wording we use to describe AI tests. The most common word used for AI testing is “evaluation,” often shortened to “eval.” In the AI world there are two common meanings for eval: Some people use it to mean any form of AI test, while others use it to refer specifically to custom-designed tests of AI applications. It’s confusing to have the same name for both the whole category and a part of the category, so in this document we call the whole field of AI tests “evaluation,” and we use “application eval” to mean customized tests of AI apps.

The difference between model benchmarks and application evals

There are two broad types of AI evaluations: The benchmark tests used to assess foundational models, and the application-specific evals used to test AI products built on top of those models.

AI model benchmarks are standardized sets of questions and tasks designed to measure the knowledge, reasoning, and capabilities of foundational models (it’s a bit like testing an automobile by measuring its acceleration, handling, and braking power). Because benchmarks are standardized, they can be used to compare AI models. There’s a thriving ecosystem of AI model benchmarks, constantly being updated as AI capabilities improve and AI models learn how to ace the older benchmarks.

There’s also a category of benchmarks focused on safety, including refusing dangerous subjects (such as how to build a bomb) and avoiding dangerous behaviors (like hacking or deception). Those tests also are evolving rapidly.

If you’re not developing your own foundational models (and most companies are not), you’ll probably use AI model benchmarks only when choosing which model you want to build your AI product on top of. 

Application evals are customized tests that measure the performance of the particular AI application you are building. It’s a strong best practice to run application evals during development of an AI application. The most common application evals focus on the accuracy, appropriateness, and completeness of the application’s responses to prompts. 

Technologist Hamel Husain is a thought leader in application evals. Here’s a very simplified outline of the process he advocates, which is commonly called “LLM-as-a-judge”:

  • A trusted domain expert reviews sample interactions with the AI application and grades them pass/fail with critiques explaining the reasoning.
  • Those human-labeled examples are used to develop and calibrate an LLM judge. Some examples can be included in the judge prompt, while separate examples are used to test how well its judgments align with the expert’s. 
  • Once the judge can reliably reproduce the domain expert’s judgments, it can be used to evaluate much larger volumes of AI interactions as the application is iterated.
  • Because everything is now automated, the AI eval process can run more frequently and at a greater scale.
  • Human review continues periodically and when material changes are made, helping ensure the judge remains aligned with expert judgment. 

As an enhancement to the LLM-as-a-judge model, some companies advocate Human-in-the-Loop Evals (HITL). HITL keeps human experts involved alongside automated evaluation. Subject matter experts can review outputs flagged by an automated judge, inspect edge cases or ambiguous results, and periodically review samples of production interactions. This lets teams use automated evals at scale while applying human judgment where it adds the most value.  (There’s a detailed description of the process here).

There are also application evals for other attributes of AI apps, such as latency and cost efficiency.

The application eval process never ends. AI applications operate in changing environments, where user behavior, prompts, underlying models, and system components can evolve over time. Those changes can affect performance, even if the application was well tuned at launch. To monitor performance in production, teams often use online evals that score samples of real user interactions on an ongoing basis. These can be combined with other methods, including production monitoring, A/B testing, and periodic human review, to identify changes in performance and opportunities for improvement. 

Measuring the overall effectiveness of AI applications

There’s an important gap in AI application evals. Although we have many ways to measure applications’ efficiency and accuracy, there’s been comparatively little work on measuring the subjective experience of using those applications. This is especially difficult prior to launch, but it’s also an issue after launch. Data management education firm TDWI wrote a good overview of this problem, which it described as “the gap between benchmark performance and deployment performance” (source).

We many measurements of the car’s performance, but we struggle to evaluate how it feels to drive it

In terms of our automobile analogy, although we have many measurements of the car’s performance, we struggle to evaluate how it feels to drive it. Since the success of most AI applications depends on user satisfaction and adoption, this is an important issue, and companies are starting to address it in different ways.

One emerging approach is User Evals, which use conventional user tests as a separate layer in the eval process. This approach is being explored in several places, most prominently by AI teams at Microsoft and Meta. Their approach, which they sometimes label UXR Evals, was presented at the 2026 Research Week conference, and in a series of posts by Pooja Dhaka, a UX researcher at Microsoft.

User Evals evaluate AI applications through user tests in first-person, multi-turn interactions, conducted with enough scale to produce reliable comparisons between model versions.

Automated evals can measure many aspects of system performance efficiently, but they can struggle with highly subjective qualities that depend on human context and judgment. As Giuliano Morse of Meta Superintelligence Labs observed at Research Week, measures such as accuracy and latency are relatively straightforward to evaluate, while qualities like humor are much harder. Human SMEs can give some feedback, but because they are experts they can’t reproduce the reactions of a typical user. As Morse put it, “they struggle with the highly subjective components of quality.” 

Chuck Kwong, a principal UX researcher for Microsoft AI, gave an example. He said that in a shopping-assistant study, an LLM model recommended stiletto shoes for a beach wedding. The response looked correct and complete, but real users quickly surfaced the problem: "I can't be wearing stilettos at a beach wedding. I will literally sink in the sand." The user's context and understanding were missing from every other form of eval.

Another emerging approach is AI experience benchmarking. This approach combines the accuracy of numerical benchmarks with inputs from videos of users. In an AI experience benchmark, the application is tested on the five attributes that research has shown drive human engagement with an AI application: understanding, trust, control, outcome, and emotional affinity. The application receives a numerical score, which can be used to track improvement over time and to benchmark against competing solutions. Video clips of users help to identify the cause of problems and give insight on how to fix them.

UserTesting’s ARQ™ assessment is an AI experience benchmarking service.

How human input strengthens AI evaluation

Human input can strengthen the AI evaluation process at several points: 

  • Specialized recruitment can identify subject matter experts who help define evaluation criteria, provide high-quality human judgments, and calibrate or validate LLM-as-a-judge and Human-in-the-Loop eval processes. 
  • User evals can be planned, recruited, and managed through a human insight system. Multiple testing methods can be used. AI-moderated interviews are often used for this testing, but any user test, including unmoderated self-interviews, can work. 
  • AI experience benchmarks can also be executed through a human insight system. These benchmarks are very new and require customization to the AI application being tested, so it can be helpful to purchase them as a service rather than trying to create them on your own.

Related Articles