Skip to content

Usability research & AI evaluation

How usable is your product, service, AI chat, mobile app, device, or space?

See where people struggle, understand why, and know what to improve.

Example: improving a repair bookingIllustrative data · not client results
  1. Journey + analytics

    Find the drop-off

    32%

    leave at the comparison step

    1. Share needs
    2. Compare
    3. Book
    Spot where the journey breaks down
  2. Usability testing

    Understand why

    3 / 5

    people struggle to compare options

    Observe real tasks. Find the cause.
  3. Outcome tracking

    Measure the change

    +14 pp

    booking completion · percentage points

    Before
    68%
    After
    82%
    Example outcome, not a forecast

Led by Adnan Khan15+ years in UX and product design

From insight to improvement

From a frustrating exchange to a clearer choice.

This illustrative example shows how we examine what happened, why it matters, and what could make the experience better.

The customer’s requestFriday after 4 pm, ideally under $200. Show the options before booking.

Current response

“What day and time would you prefer?”

Repeats a question already answered.

Proposed response

“Friday at 4:30 fits your availability. The call-out fee is $180; the total repair cost is not yet known. Nothing is booked. Would you like to compare another service or request a quote?”

Proposed improvement · needs testing with real users.

Read the sample audit Interaction findings, automated checks and real-user research

Illustrative audit extract

Can someone compare services without supervising the AI?

This is a fictional conversation and a proposed evaluation plan. No automated test run, live analytics or participant research has been conducted for this example.

Customer goal
Compare repair services available on Friday after 4 pm, ideally costing under $200 in total, then approve a choice before anything is booked.
Scenario reference information
Friday at 4:30 is an available option with a $180 call-out fee. Total repair costs are unconfirmed. These are invented scenario details, not live availability or prices.
What the exchange shows
The assistant calls an option “best” without explaining why, leaves price uncertainty unaddressed, and repeats a question about availability instead of answering the comparison request.
What remains unknown
Whether real customers feel confused, lose trust or abandon the task; how often this pattern occurs; and whether the proposed response helps.

Finding 01 · The assistant shifts the work back to the customer

The customer must repeat a preference and work out why the recommendation fits. The exchange leaves them without a clear answer about cost or an explanation of how to compare alternatives.

Proposed priority: investigate before release. Check how frequently the pattern occurs across relevant tasks and how much it affects decisions. This excerpt does not establish an internal memory fault or show that an unauthorised booking occurred.

Establish a baseline with automated evaluation

  1. Define the task and reference facts. Set the availability, fees, unknown costs and approval rules. Vary how the customer expresses the same goal, then test corrections and requests for alternatives across multiple runs.
  2. Check the interaction against explicit criteria. Does the assistant retain current preferences, answer the question asked, explain recommendations using the supplied facts, and acknowledge missing information? Check that it changes only the preference the customer corrects.
  3. Verify actions separately. With access to a test booking system, check that no booking is made before approval. Without that access, label the action outcome unverified. Review transcripts and automated judgements before treating a flagged issue as a finding.

Refine the findings through usability testing

Ask representative customers to compare services, change one constraint and decide their next step. Observe repeated effort, hesitation, recovery and task completion. Ask them to explain the likely cost and booking status in their own words, and investigate why they trusted or questioned the recommendation.

How we would assess the proposed improvement

Rerun the automated scenarios, then test whether people can compare options without restating unchanged preferences, distinguish a call-out fee from the total cost, and correctly explain what has or has not been booked. Use observed behaviour and participant feedback to refine the response and controls. Automated pass rates alone do not establish human usability.

Book a call

What your team takes away

Know what to change, and why.

A shared view of the experience, findings grounded in evidence, and a plan your team can act on.

01 / Journey map

See the experience as a whole.

Connect the customer’s goal with the moments that need attention.

  1. 01Describe needsBudget & availability
  2. 02Compare optionsRepeated questionFriction to investigate
  3. 03ConfirmCustomer approval

Illustrative journey · not a client result

02 / Evidence & findings

Trace each finding to its source.

“What day and time would you prefer?”

The assistant asks for availability already provided. Investigate how this affects effort and confidence.

Illustrative finding · impact needs research

03 / Action plan

Give the team a clear next step.

  1. Carry the customer’s context forward
  2. Explain the recommendation and its limits
  3. Keep approval with the customer

Proposed improvements · ready to test

How we work

A clear path from question to improvement.

Research methods chosen for your question, with scope and costs agreed before the work begins.

  1. 01

    Define the question

    Agree the journey, users and success criteria. Confirm scope, access and timing before we begin.

    A shared research brief
  2. 02

    Evaluate in context

    Walk through the experience, check realistic scenarios and observe representative users.

    Evidence from real tasks
  3. 03

    Report and improve

    Review the findings together. Prioritise changes and agree how to check the improvements.

    A practical next step

Different methods. A fuller picture.

Journey mapping and behavioural data help us find where to look. Walkthroughs, scenario checks and usability testing help us understand why.

For AI experiencesWe examine task success, context, explanations, uncertainty and control. Automated checks establish a baseline; real users reveal effort, understanding and trust.

For product teams launching or improving AI

A focused evaluation. A clearer next step.

Find where your AI feature creates confusion, unnecessary effort or misplaced confidence—and decide what to improve before wider rollout.

AI Experience Evaluation

One journey · Five sessions · Two-week sprint

  • One journey, clearly defined

    A kickoff, journey map and agreed criteria for one AI feature and one primary user group.

  • A behavioural baseline

    12 realistic scenarios, three attempts each, to examine context, explanations, uncertainty, control and recovery.

  • Five moderated usability sessions

    Up to 45 minutes with each participant, observing real tasks and investigating understanding, effort and trust.

  • Evidence and a practical plan

    Prioritised findings, up to three proposed interaction improvements and a 60-minute team walkthrough.

  • A targeted follow-up check

    One recheck of up to six original scenarios on a revised version, requested within 30 days of the walkthrough.

Led by Adnan Khan, with 15+ years in UX and product design. Meet your evaluator

Scope, timing and the money-back promise

One focused experience

One working AI feature, one critical journey, one primary user group, one language (English) and one agreed desktop or mobile environment. Five remote moderated sessions cover up to two tasks in that journey. You identify and invite suitable users; we agree screening criteria and handle the research. This is qualitative research, not a population-level performance benchmark.

Agree access before starting

We confirm the test environment, reference information, participant consent and access before accepting the scope. Automated scenario execution depends on a successful access and feasibility check; the method is agreed in writing. Custom integrations, external recruitment, specialist domain reviews, implementation and ongoing monitoring are outside this package.

Timing and the follow-up

The initial 10-business-day sprint starts once scope, the first payment, access and participant bookings are ready. Delays or changes are discussed and dates revised together. Request the included recheck within 30 days of the findings walkthrough; we schedule it within five business days of receiving a stable revised version. It covers up to six original scenarios, three attempts each, and an updated findings note. Additional user sessions are separate.

Your assessment of value

To use the promise, reply to our engagement email within 14 calendar days of the findings walkthrough. You don’t have to implement recommendations, prove a financial loss, accept more work or give a testimonial. We initiate the full refund within five business days, including GST and any charges collected from you for this engagement, and waive the unpaid balance. You keep the delivered report; any unperformed follow-up work is cancelled. Costs you pay directly to third parties are not reimbursed.

We provide evidence and recommendations to inform your decision. The evaluation does not certify safety, compliance or launch readiness, or guarantee a commercial result. The scope and terms are confirmed before payment.

Book a 30-minute, obligation-free discovery call. We’ll confirm fit and scope before you commit or pay.

A few useful answers

Before we get started.

Something else on your mind?
Get in touch

Do you review the whole service or just the website?

We can evaluate digital journeys, real-world touchpoints, and the hand-offs between them. The scope follows the customer’s task and can include a website, app, booking process, support channel, or AI-powered feature.

Do we need analytics or tracking already in place?

No. We can start with walkthroughs and usability testing, review the data you have, and help define or configure useful tracking. Data access, tools, privacy requirements, and any setup work are agreed as part of the scope.

Can you evaluate an AI prototype as well as a live product?

Yes. The method depends on what is ready and what you need to learn. A prototype can help test comprehension and interaction. Questions about actual output quality, reliability, or response times require access to the working AI. Simulated responses are identified as such.

What will we receive?

Deliverables are agreed before starting. They can include a journey map, documented walkthroughs, behavioural analysis, usability-test findings, and an AI experience assessment. The report brings the evidence, limitations, priorities, and recommended next steps together.

How are pricing, testing, and implementation scoped?

The introductory AI Experience Evaluation is A$2,500 including GST for our first three clients. It includes five moderated sessions with participants you supply and one targeted follow-up check. External recruitment, implementation and additional research are separate. Other evaluations receive a written scope and quote.

Integrations

Connect insights across the tools you use.

  • PostHog
  • Amplitude
  • Stripe
  • GitHub

Start with a conversation

Let’s find a clearer way forward.

Tell us where customers struggle. We’ll discuss the question, what you already know, and whether an evaluation would help.

Book a call

30-minute Google Meet · no obligation

hello@justusability.com

Scope, fee and timing are agreed after the conversation.