Technical evals tell you how the system performs. ARQ shows how people experience it — whether they feel understood, trust the interaction, stay in control, achieve the outcome they wanted, and want to come back.

Technical performance is only part of AI quality. ARQ adds a structured measure of how people relate to the experience, helping teams understand how the interaction lands with people and where the experience needs improvement.
Technical evals are essential for measuring system performance. ARQ complements them by showing whether people understand the experience, trust it, feel in control, and want to use it again.
Conversational and agentic AI introduce human factors that technical metrics do not fully explain. ARQ helps identify where understanding, trust, control, outcomes, or affinity are creating friction.
A shared measure gives product, research, engineering, and leadership a common way to discuss the human experience and compare how it changes across releases or experiences.
The benchmark provides an overall score for the AI experience and a score for each of the five dimensions, supported by participant evidence that helps explain what is driving the result.
Repeat the benchmark and the score is comparable, so you can track a release, compare two experiences, or set it against competitors.
Every engagement is built around your AI product, target audience, and business goals. You receive a clear view of overall performance, the five dimensions, supporting participant evidence, and recommendations for where to focus next.
One number for the AI experience, on a 1 to 100 scale, so the result travels into a leadership deck without needing a translation layer.
Understanding, trust, control, outcome, and affinity are each scored separately, helping teams see which dimensions are strengthening the experience and which need attention.
Participant clips and supporting evidence help explain the scores, so teams can see the moments and behaviors behind the findings instead of relying on a summary alone.
We recruit the people your AI is built for from our own network, whether that is a consumer segment, a profession, or your existing product users.
A written read on where the relationship is strong, where it is weak, and which changes are worth making first. Delivered as an interactive report.
Run it again after a release with the same audience and the same test design. The second wave moves faster because the scoping is already done.
ARQ follows a standardized methodology, with each engagement tailored to the AI experience and audience being evaluated. This structure makes the benchmark repeatable while leaving room for custom scoping when the use case requires it.
A mixed method design: ten think-out-loud sessions plus two hundred interaction tests.
Three tasks against one AI product, such as a chatbot, an agent, or a single AI feature.
General population, recruited from our network, screened to the audience your AI experience is built to serve.
An interactive report with the scores, the findings and the clips, delivered to you rather than through a platform login.
Two to three weeks. Repeat waves usually run faster, because the audience and test design are already validated.
The five dimensions draw on published research on AI adoption together with our own studies. The methodology is documented so research teams can understand how the benchmark is structured and compare results consistently over time.
Elements scored separately
Participants per engagement
Weeks to delivery
ARQ is designed to make the methodology clear to the teams using the results. The sampling approach, scoring model, task design, and participant criteria are documented as part of the engagement.
The scoring model, task design, and screening criteria are documented so teams can understand how the benchmark was run.
People agree to the specific session, including the recording, before they take part. Consent covers the AI experience they are about to use.
Your prompts, your product and your participants' sessions stay inside your engagement. They are never pooled into a shared corpus or resold.

Choose the AI experience you want to understand. We align on the audience, tasks, and goals, run the benchmark, and deliver the score, supporting evidence, and recommendations for what to improve next.
Can’t find the answer you’re looking for? Talk to our team.
ARQ complements technical evals rather than replacing them. Evals measure system performance; ARQ measures how people experience the interaction across understanding, trust, control, outcome, and affinity.
Both are benchmarks, and they measure different things. The scoring model here was designed from the ground up on published AI adoption research. What carried over from QXscore is what we learned running it.
No. Technical performance stays with your evals. ARQ focuses on the human experience across understanding, trust, control, outcome, and affinity, giving teams a complementary view of AI quality.
Human-in-the-loop review evaluates whether AI responses are accurate or appropriate. ARQ asks the intended audience how the interaction is experienced. The two approaches answer different, complementary questions.
An interactive report with the overall score, scores across the five dimensions, key findings, recommendations, and supporting participant evidence. Standard engagements are delivered without requiring you to provision a separate platform environment.
Not as part of the standard engagement. If you need the work run in your own environment, that can be discussed as a custom scope based on your requirements.