AI Assistant reports

How your AI is performing: its speed, its honesty about its own answers, the questions it struggles with, and what it costs.

AI Assistant reports

AI Confidence Scores New

What you see. The AI's average self-rating, the low vs high split, the 1–5 distribution, a trend, a per-model breakdown, and the 50 lowest-rated conversations, each with the AI's own reasoning and the model used.

How to use it. This is the AI telling you where it's unsure before a customer complains. Read the low-rated conversations and their reasoning to find topics to add to your knowledge base.

How it's calculated: averages and distribution of the AI's per-conversation self-ratings, grouped by model for the per-model breakdown.

Unanswered Questions

What you see. A paginated list of customer questions the AI erred on or answered with low confidence, showing the question, model, confidence score, time, and reason (error vs low confidence).

How to use it. Your knowledge-gap backlog. Every row is a real customer question your AI couldn't answer well, the highest-value list for deciding what content to add next.

How it's calculated: requests that returned an error, or whose best knowledge-base match scored below 0.5.

AI ROI New

What you see. Hours saved by automation, cost per conversation, automation rate, and total cost; an automation-rate trend; a per-model breakdown; and a short knowledge gaps list.

How to use it. The business case for the AI in one screen. "Hours saved" and "cost per conversation" are the numbers to put in front of leadership; the automation-rate trend shows whether the AI is taking on more over time.

How it's calculated: hours saved = auto-handled replies × average review time ÷ 60 (if there's no review-time data it assumes ~5 minutes per manual review). Costs are estimated from token usage × a fixed per-model price table.

AI Activity Log

What you see. Request count, error rate, total tokens, and latency percentiles (p50–p99); a 24-hour hourly request sparkline; and model cost & token usage (total tokens, estimated cost, average tokens per request, a per-model table, a daily token trend, and a usage heatmap).

How to use it. The technical health and cost of the AI. Rising error rate or p95 latency means customers are waiting or getting failures; the per-model cost table shows where your AI spend goes.

How it's calculated: percentiles are computed over request durations; error rate = failed requests ÷ all requests; cost = tokens ÷ 1000 × per-model rate (estimate).

Common questions

What does the AI Confidence Scores report show?

It shows the AI's average self-rating, the low-vs-high split, the 1 to 5 distribution, a trend, and the 50 lowest-rated conversations, each with the AI's own reasoning and the model used. It tells you where the AI is unsure before a customer complains.

Where do I find customer questions the AI couldn't answer?

Use the Unanswered Questions report. It lists customer questions the AI erred on or answered with low confidence (best knowledge-base match below 0.5), showing the question, model, confidence score, time and reason. It is your knowledge-gap backlog, feed it into your documents.

Are the AI cost numbers exact?

No. AI and token costs are estimates calculated from a fixed per-model price table (tokens divided by 1000, times a per-model rate), so treat them as close estimates for budgeting rather than an invoice.