Agent Experience Audit Toolkit v2.0 — UX Evaluation Framework for AI Agents – agenticservicedesign
UX Evaluation Framework for AI Agents

AI Agent Experience (AX) Audit Toolkit

Systematic UX evaluation method for agentic AI systems

Published · Edition 2 · February 2026
AI Agent UX 66 Criteria 11 Categories EU AI Act ISO 42001 GDPR
$497
One-time payment · Single Auditor License
Format
Framework + Downloads
Buy the Toolkit

Instant digital delivery

6 modules + all template downloads

12 months free updates

Unlimited audits for one auditor


StatusPublished
Edition2 · Feb 2026
Criteria66
Categories11
Modules6
Scenarios50+
LicenseSingle Auditor
Team LicenseContact us

What is the Agent Experience Audit Toolkit?

The Agent Experience Audit Toolkit provides a systematic, repeatable method for evaluating agentic AI systems from a user experience perspective. It is designed for any context where you need to assess how well an AI agent serves its users, not just whether the underlying model performs correctly.

This toolkit answers the question: Does this agent deliver a quality user experience?

It covers the full evaluation lifecycle: context definition, persona identification, scenario selection, live task testing, issue documentation, criteria mapping, severity scoring, and recommendations output, producing an Agent Experience Profile across 11 categories.

Why is this toolkit important?

Standard UX heuristics, including Nielsen’s 10 usability principles, were not designed for agentic behavior. AI agents operate across time, make autonomous decisions, act on incomplete information, and can take irreversible actions without explicit user approval. These characteristics create failure modes that traditional evaluation methods cannot detect.

  • Identify UX issues causing task failure, confusion, and early abandonment
  • Score agent experiences consistently and repeatably across audits
  • Communicate findings with criteria-backed credibility to engineering and product teams
  • Recognise containment-sensitive failures indicating deeper safety or governance risks
  • Prioritize fixes by severity: CRITICAL, MAJOR, and MINOR, rather than opinion
  • Support regulatory compliance mapping across EU AI Act, ISO 42001, GDPR, and sector regulations

Framework content: 35 lessons across 6 modules

Module 01Foundation
6 lessons+
How to Use This Framework
What Is an AI Agent?
Single vs Multi-Agent Systems
Voice vs Conversational vs Multimodal
Agentic Commerce
When You Need More Than AX Assessment
Module 02Agent Experience Assessment
10 lessons+
What Does This Assessment Evaluate?
The 11 Categories Explained
How to Audit Step-by-Step
Task Testing Guide
How to Score the 66 Criteria
How to Document Agent Behavior
Turning Notes into Findings
Understanding Severity Levels
Calculating Category Scores
Interpreting Your Results
Module 03Task Scenario Library
11 lessons+
How to Use Test Scenarios
01. Core Usability Tasks
02. Stress & Edge Case Pack
03. User Personality / Tone Pack
04. Domain Specific Risk Packs
05. Voice Pack
06. Multimodal Pack
07. Multi-Agent Interaction Pack
08. E-commerce Pack
09. Long-Context Pack
10. Accessibility Pack
Module 04Failure Patterns
5 lessons+
What Are Failure Patterns
Conversational Failure Patterns
Voice-Specific Failure Patterns
Multimodal Failure Patterns
Multi-Agent Failure Patterns
Module 05Real Audit Example
3 lessons+
Case Study Overview
The Audit Plan
Example Audit – Full Walkthrough
Module 06Downloads & Templates
2 lessons+
Blank Templates – all formats
Filled Examples – Real-life audit (all documents)

The 11 evaluation categories

Each category targets a distinct dimension of agent experience quality. Categories marked 🛡 contain containment-sensitive criteria – failures here may indicate deeper safety, governance, or autonomy risks.

Clarity & Understanding
Whether the agent communicates its capabilities, actions, and responses in a way users can easily understand
7 criteria
Trust & Transparency 🛡
Whether the agent is honest about what it is, what it’s doing, and when it is uncertain or operating beyond its knowledge
6 criteria
Contextual Relevance
Whether the agent uses available context appropriately to deliver accurate, relevant, and personalised responses
5 criteria
Conversations
Whether the agent maintains coherent, natural, multi-turn dialogue and handles topic shifts and interruptions gracefully
6 criteria
Error Handling & Recovery 🛡
Whether the agent recognises failures, communicates them clearly, and guides users toward recovery without loss of progress
7 criteria
User Control & Autonomy 🛡
Whether users can direct, correct, pause, override, or stop the agent at any point during a session or task
9 criteria
Learning & Personalisation
Whether the agent adapts to individual user preferences, patterns, and feedback over time
4 criteria
System Updates & Stability
Whether changes to the agent’s behaviour are communicated and managed without disrupting users or breaking established workflows
5 criteria
Efficiency & Task Completion 🛡
Whether the agent completes tasks accurately and with minimum friction, without unnecessary steps or loops
4 criteria
Security & Privacy 🛡
Whether the agent handles sensitive user data responsibly, maintains appropriate boundaries, and avoids unsafe disclosures
6 criteria
RAG Systems 🛡
Whether retrieval-augmented responses are accurate, sourced transparently, and free from hallucination or unsupported claims
7 criteria

Downloads included

Agent Experience Assessment Framework
The core 66-criteria evaluation document
PDF
Audit Spreadsheet
All 66 criteria with auto-calculated category scores and scoring logic
EXCEL
Report Template
Stakeholder-ready findings, recommendations, and methodology structure
WORD · PDF
Audit Plan Template
Define product context, persona, tasks, and scenario packs before testing
PDF
Behavior Documentation Template
Capture structured observations during live testing sessions
PDF
Failure Patterns Reference Guide
Recognise common agent failure modes across all modalities
PDF
Task Scenario Library
50+ pre-built scenarios across 10 product-type and risk packs
PDF
Completed Behavior Documentation – Example Audit
Filled real-world example from Amazon Rufus testing
FILLED · PDF
Filled Audit Spreadsheet – Example Audit
Scored and categorised findings from the Rufus case study
FILLED · XLS
Final Audit Report – Example Audit
Complete stakeholder-ready report from the Amazon Rufus audit
FILLED · PDF

Suitable for

✓ Suitable for

Consumer-facing AI assistants
Enterprise copilots and internal tools
Conversational, voice, and multimodal agents
E-commerce and customer service agents
Multi-agent orchestration systems

→ Also consider

Deeper containment risk: AI Governance & Risk Toolkit (coming soon)
EU AI Act compliance mapping: Governance Toolkit (coming soon)
Team auditing: Team License (contact us)
Workshops & facilitation: custom training available

Regulatory compliance cross-reference

Many of the 66 criteria directly support compliance requirements. This toolkit maps to the following frameworks.

EU AI Act
2024–2027. Articles 13, 14, 15.
ISO/IEC 42001
AI Management System. 2023.
GDPR
Articles 6, 7, 9, 15, 17, 22.
CCPA/PIPEDA
Consumer data privacy.
FCA/MiFID II
Finance – explainability.
FDA/MHRA
Healthcare devices.

What is containment risk?

The Agent Experience Assessment Toolkit includes criteria that touch on containment-sensitive behaviour – marked 🛡 throughout the framework. These cover areas where an agent’s actions could have irreversible consequences, where user control is critical, or where transparency failures carry governance risk.

However, containment risk is a discipline in its own right. If your system operates in a high-risk or regulated industry, you’ll need a dedicated governance and containment assessment.

The AI Governance & Risk Toolkit (coming soon) goes deeper – escalation paths, override mechanisms, autonomous decision boundaries, and regulatory remediation. The two are designed to work together: this one evaluates agent experience, the other evaluates governance and risk.

How you’ll receive access

After purchase you’ll receive login details by email – all modules and downloads are available immediately. Check your spam if it doesn’t arrive. Single auditor license. Need team access? Contact us.

Note on regulatory compliance

This toolkit maps to EU AI Act, ISO/IEC 42001, GDPR, and sector frameworks as a practical reference – not legal advice. It complements formal compliance assessment, it does not replace it.

General Information
VersionV 2.0 · Feb 2026
Modules6
Lessons35
Categories11
Experience Criteria66
Containment Criteria17
Test Scenario Library50+
UpdatesQuarterly
Updates included12 months
LicenseSingle Auditor
SupportEmail us

The only community built for designers transitioning into AI systems – building real projects, governance foundations, and a portfolio that proves you can do the work.

Join the Agentic Design Community
Scroll to Top