How to be sure your AI experience is excellent: Introducing ARQ

Learn how ARQ™ measures the human-AI relationship across trust, control, understanding, outcomes, and affinity to help improve AI experiences.

Summary

There’s a blind spot in the way we evaluate AI experiences. 

A thriving ecosystem of evals helps us improve the accuracy, appropriateness, and completeness of AI responses, but they don’t cover the full human experience of using an AI product. People using AI often judge it the same way they judge a relationship with a human being: Do they trust it?  Do they understand its motivations and actions? Is it personable?

Those subjective reactions to AI often drive customer preference and adoption. To fully optimize conversational and agentic AI products, we have to evaluate how people relate to them. 

AI Relationship Quality (ARQ) is the first comprehensive benchmark and assessment of the human-AI relationship. It measures the five elements of an AI experience that drive human reactions, and identifies strengths and opportunities for improvement.

Filling the gap in AI evaluation

Traditional AI evals focus primarily on system performance, including accuracy, efficiency, safety, and whether the system responds appropriately.  Many of these tests can be automated, and run continuously during and after development.  Evals have been effective at improving AI system performance.

But those technical measures are only one part of the human experience of using an AI product. To fully optimize AI quality, we need to understand how people relate to an LLM: how they feel about it as a thought partner when sharing ideas, and as an assistant when performing tasks. 

Leaders in the AI space are starting to recognize this gap in the way we evaluate AI experiences:

  • "Are the experiences I have with this technology good?…. It could be small things like the last time I chatted with an AI, was it delightful? Or was it annoying?" --Bret Taylor, Chairman, OpenAI, in Masters of Scale
  • "How do you prove to yourself which one of those (AI experiences) is the best? I don't have a confident answer to that. I expect this is where the good old-fashioned usability testing comes in." --Simon Willison, AI engineering leader, in Lenny’s Podcast
  • "How do we know if our AI product is good?…It's hard, and it's rapidly evolving and most people are not doing a very good job….We're just learning how do we evaluate an AI product." --Teresa Torres, product discovery leader, in All Things Product

Leading design and research professionals take the idea further. They say companies need to design AI specifically for relationship:

  • "Designing for agentic AI is designing for a relationship. This relationship, like any successful partnership, must be built on clear communication, mutual understanding, and established boundaries." --Victor Yocco, UX Researcher at ServiceNow, in Smashing Magazine
  • "We are no longer designing tools. We are designing relationships." --Saranya Gunasekaran, User Experience Architect at Honeywell, in UXmatters

Research on AI experiences and human-AI interaction has identified five elements of the experience that drive the human-AI relationship. The most successful AI products generally excel in these five areas:

  • Understanding. Does the AI feel like a knowledgeable coworker, rather than a black box? Do human users feel they understand the AI’s thinking, intent, and boundaries? An effective AI product shows what it is doing while it works, and states its limits plainly
  • Trust. Do people give the appropriate level of trust – not too much, but also not too little -- to an AI experience? Effective AI needs to give people a calibrated “trust meter,” so they know when they should rely on the AI and when to step in. A successful AI product signals uncertainty at the right times, hedges when it should, and doesn’t overclaim. This is the factor most directly tied to long-term customer retention.
  • Control. Can people steer the AI, correct it, and recover when it goes wrong? A well-controlled AI product feels like a collaborative partner that stays within the boundaries the user set, never taking over or making people feel helpless. Loss of control is one of the strongest drivers of AI rejection, especially for agents. 
  • Outcome. Did the user feel the AI-human partnership delivered better results than the person could alone, and do those results feel relevant and aligned with what they actually wanted? This factor depends in part on AI accuracy, but the focus is on the user’s overall outcome, not just the AI’s answer.
  • Affinity. How does it feel to work with the AI? Does the AI's manner fit the stakes of the task? The AI’s personality, tone, warmth, and emotional response all live here, as does its fit with the brand. Our research has shown that when two AI products produce similar answers, the AI’s personality often drives user preference.

Introducing ARQ: a comprehensive view of the human-AI relationship

ARQ (AI Relationship Quality) is a comprehensive benchmark and analysis of how people relate to an AI experience. It features quantitative ratings of the five elements that drive AI adoption, plus an overall score. The assessment also includes video user actions and thoughts to understand the “why” behind the ratings and to identify opportunities for improvement.

Benefits of ARQ include:

  • An AI team can use it to quickly evaluate the quality of an AI experience and identify opportunities for improvement
  • You can measure your company’s AI experiences against competitors
  • You can objectively track improvements in your AI solutions over time
  • The ARQ score makes the AI relationship tangible for VP and C-level executives

ARQ is an ideal accompaniment to AI evals. Evals typically measure accuracy and other technical factors, while ARQ measures the subjective user relationship.

Details on the ARQ methodology

ARQ is a customized assessment of the human-AI relationship, conducted by Auros' professional services team. The process starts with a scoping meeting to identify the exact tasks to be tested and the right user audience for the test. Because AI experiences are so diverse, both the tasks and target users are customized for each AI product. UserTesting then runs the quantitative and qualitative tests, analyzes the results, and delivers a report with findings, insights, and recommendations. ARQ reports can be run as a one-time evaluation, or an ongoing series of scheduled assessments.

ARQ is rooted in more than a decade of software usability assessments run by Auros's UserTesting. Thousands of usability assessments have been conducted, creating a very rich experience base. ARQ was built on that experience, along with published academic research on AI experiences, and the company’s own AI studies. The ARQ process combines a quantitative interaction test, for rigor and repeatability, with qualitative user videos to give insights on causation and to help communicate findings in a compelling manner to stakeholders and executives.

‍The quantitative interaction test. Participants use an AI feature or product to accomplish three tasks with varying levels of difficulty. After each task, they are asked to agree or disagree with up to 21 statements, grouped into the five elements of the AI experience (the number of questions can vary depending on whether they were able to complete each task). This produces a 1-100 score for each element, and an overall 1-100 score for the full experience.

‍The qualitative interviews. Participants complete the same three tasks, speaking their thoughts out loud as they proceed. Their screens, faces, and voices are recorded during the tasks. They are also asked a series of questions about the experience. These videos are paired with the quantitative scores to give insights on what caused the scores, and identify opportunities to improve them. The videos are also extremely helpful to help communicate customer feelings to stakeholders and executives, which helps drive group agreement on priorities.

Analysis of the findings

Because people interact with all elements of an AI experience at the same time, the scores tend to move together. For example, a particularly bad rating in one area can pull down the other scores. So rather than hyper-focusing on a single element score, it’s more useful to look at patterns in how the scores relate to each other. Auros’ pro services team will help you with this analysis, but here are some examples of results we’ve seen in testing:

All five low together This indicates a capability problem. Focus first on improving Outcome and model quality.
Outcome high, Control low The LLM gets there but takes the wheel. This is a common problem in agents. Look at the user’s ability to correct, undo, and hand off.
Outcome high, Understanding low It works but nobody knows why. This is a trust time-bomb: the relationship is fine until the first error, then confidence collapses.
Everything decent, Affinity low This indicates a tone and personality problem. Fixes here are often high-impact.
Everything decent, Trust low People do not believe it yet. Look for overclaiming and missing uncertainty signals. Treat this as a retention risk.

A single element with a low score is an indication to dig deeper. Something is going wrong, but the survey often won’t capture why. It’s important to watch the video sessions for a low-rated element.

Sample ARQ output

ARQ reports are delivered as interactive HTML documents plus video highlight reels. Here are anonymized sample images from ARQ reports.

Q&A

‍How does ARQ relate to AI evals?

They are complementary. Evals usually measure performance and accuracy, and are run automatically at scale. ARQ measures the human-AI relationship, and is run periodically to assess progress and check for problems. You should be doing both.

‍I have a very specialized user base. Can ARQ work for my AI experience?

Yes. The assessment can be customized to almost any audience or experience.

‍How often can we run an ARQ assessment?

As often as you want, but let’s be clear that ARQ is not designed to be automatically repeated constantly the way an automated eval is. This is about bringing unique human insights and human judgment into the AI development process, and that can’t be fully automated.

‍What research was ARQ based on?

We reviewed peer-reviewed research on the factors driving AI adoption, and also current software assessment scales. Based on that, we ran experiments on different question sets and scoring systems. The design of ARQ was also influenced by our own AI tests, research we’ve run for our customers, and the company’s decade-plus experience assessing traditional software experiences.


Related Articles