AI assistant for reporting

Data Explorer is a cybersecurity reporting tool: pick your columns, mix and match filters, and pull the data you need. But what you get back is raw data — numbers, not answers. This explores an AI assistant layered on top: ask a plain-language question, and it turns that raw data into an executive summary, targeted recommendations, and a trend prediction.

Due to confidentiality, the data and several of the screens below are recreated or placeholder rather than pulled from the live product.

The problem

Data Explorer has everything. That's the problem.

This is where every bit of data about AI Exercises actually lives — pick a column, add a filter, cross-reference exercise types, export. It's genuinely capable. It's also a lot: thirteen fields, five filters, and no help figuring out which combination answers the question you actually walked in with.

The goal wasn't another chatbot. It was an assistant people could actually trust with a decision.
Data Explorer with the new AI Report button added next to Save

The button next to Save is where it starts.

The principles

Three principles I designed against

AI reporting can go wrong fast — confident answers nobody can verify, or complexity mistaken for insight. These are the principles I designed against to keep it trustworthy.

01
Showing How AI Reaches Its Conclusions

If an assistant tells a manager their team is falling behind, the next question is "says who, based on what." So nothing here gets to just assert a number. Every generated report comes with a "sources & method" note listing exactly which data it queried, whether a plain rule did the work or a model interpreted something, and which filters were live.

It starts with one question — describe what you want, or start from a suggestion if you don't know the field names yet.

Data Explorer
Explore and download data from the platform
✦ AI Report
What data would you like to see?
Token usage and cost per team over the last quarter, sorted by spend, shown as a bar chart.
Try asking
Which teams spend the most tokens per completed objective?
Score trend for SOC Triage Agent over the last month
Compare cost per attempt across AI Agent Tuning vs AI Ranges
What needs my attention this week?
Or insert a report section
▤ Token Usage & Cost
▤ Team Readiness Summary
▤ Exercise Completion Rates
Cancel
Generate Report →

Generate Report hands off to the assistant. It lands you on a real, editable report that leads with a plain-language summary — and because this principle isn't optional, that summary opens straight into exactly which data it queried and how. That's what "Sources & method," expanded below, actually contains.

AI Assistant
Y
Token usage and cost per team over the last quarter, sorted by spend, shown as a bar chart.
Pulled token usage and cost across every AI Agent Tuning and AI Range exercise from the last quarter, grouped by team and sorted by total spend. Three teams account for 68% of it — SOC Team A is well out in front.
Y
Only Agent Tuning, not Ranges
Done — filtered out AI Range exercises. Updated the report on the right.
Executive summary

SOC Team A accounts for 41% of total spend this quarter — more than the next two teams combined. At the other end, IT Risk & Compliance has spent the least of any team, which is worth a second look before reading it as efficiency.

Sources & method

Queried 3 tables — sessions, exercises, teams — filtered to Q4, Agent Tuning only. Every number above is a deterministic aggregation, not a figure the model estimated.

Ask a follow-up or request a change…
Token Usage & Cost by Team — Q4, Agent Tuning only
📈
Save Report
TEAM×TOKEN_USED×COST×EXERCISE_TYPE = Agent Lab×+ Add field
SOC Team A
$18,240
SOC Team B
$12,110
Security Engineering
$8,760
Platform Engineering
$5,930
IT Risk & Compliance
$3,480
Team ↕Token UsedCostExercise Type
SOC Team A2,140,000$18,240Agent Lab
SOC Team B1,410,000$12,110Agent Lab
Security Engineering1,020,000$8,760Agent Lab
02
Designing For Calibrated Trust
The goal was never to get people to trust the AI completely. It was to help them know when to trust it, and when not to.

A confidence score only matters if it changes what happens next. Below, two teams are flagged as likely to miss their Q1 target, each showing the trend and the AI's confidence in it. High confidence: that's the read. Low confidence: the assessment says so, and hands back control — check the trend, or the data, yourself.

AI Assistant
Y
Which teams are on track to miss their Q1 readiness target?
Projecting current trend against the 75% target. Two teams look likely to land short.
SOC Team B
Projected to miss Q1 readiness target
Six straight weeks of decline, 81% → 58%.
At risk
AI Assessment
Classified asAt risk of missing target
90%
High confidence

Why: Readiness has dropped every week for 6 straight weeks, 81% → 58%. A trend that consistent is safe to project forward.

Platform Engineering
Projected to miss Q1 readiness target
Only two weeks of data, moving in opposite directions.
At risk
AI Assessment
Classified asAt risk of missing target
53%
Low confidence

Why: Only two weeks of data so far this quarter, and they moved in opposite directions, 62% → 68%. Could recover or keep sliding — not enough history to say which.

Low confidence. Worth a manual check-in before treating this as a real trend.
Ask a follow-up or request a change…
03
Earning Trust Through Progressive Autonomy
AI should earn the right to do more over time.

Start with recommendations, guided actions, and automation as confidence grows. Users adopt AI faster when they can see value, build trust gradually, and decide when they're ready to hand over more responsibility.

The first stage looks like this: the assistant checks every session against your thresholds and surfaces what's worth a look, paired with a suggested next step — "Assign Exercises" — that you still have to click yourself.

AI Assistant
Y
What needs my attention this week?
Checked every Agent Tuning session against your usual thresholds. Two people crossed one.
High Cost spiked 4× with no score gain
Riley Nguyen · SOC Triage Agent
Why this was flagged

Cost rose from a $0.94 average to $3.80 on the latest attempt, while score held flat (33 → 33). Flagged when cost more than triples without a 5+ point score gain.

Plateau No improvement across 3 sessions
Lucia Martinez · Threat Intel Analyst
Why this was flagged

Score moved by less than 3 points across the last three attempts (47 → 49 → 47). Flagged after 3 consecutive low-delta sessions.

Ask a follow-up or request a change…
Token Usage & Cost by Team — Q4, Agent Tuning only
✓ Saved
TEAM×TOKEN_USED×COST×EXERCISE_TYPE = Agent Lab×
SOC Team A
$18,240
SOC Team B
$12,110
Security Engineering
$8,760
Platform Engineering
$5,930
IT Risk & Compliance
$3,480
Team ↕Token UsedCostExercise Type
SOC Team A2,140,000$18,240Agent Lab
SOC Team B1,410,000$12,110Agent Lab
Security Engineering1,020,000$8,760Agent Lab

Once the same recommendation repeats often enough to stop being a coincidence, the offer changes — from "here's what to do" to "want me to just start doing this automatically?" Below, that shows up as the assistant noticing you've assigned the same exercise to every person with the same gap, and asking, not assuming, whether it should do that automatically next time.

AI Assistant
Noticed a pattern worth automating — same response, four times in a row.
IF Plateau flag on a Threat Intel Analyst session
THEN assign Threat Intel Refresher
Seen 4 times this month — same flag, same fix, every time
Ask a follow-up or request a change…
Token Usage & Cost by Team — Q4, Agent Tuning only
✓ Saved
TEAM×TOKEN_USED×COST×EXERCISE_TYPE = Agent Lab×
SOC Team A
$18,240
SOC Team B
$12,110
Security Engineering
$8,760
Platform Engineering
$5,930
IT Risk & Compliance
$3,480
Team ↕Token UsedCostExercise Type
SOC Team A2,140,000$18,240Agent Lab
SOC Team B1,410,000$12,110Agent Lab
Security Engineering1,020,000$8,760Agent Lab
And it doesn't stop once you save

A saved report isn't a snapshot. It stays something you can question later, including questions that reach beyond it into everything else the assistant can see — still governed by the first principle: every answer cites where it came from.

AI Assistant
Y
Which teams are most at risk of falling behind on Agent Tuning, looking at both spend and readiness?
AI Summary

IT Risk & Compliance and Platform Engineering are the ones to watch — not because either is over-spending, but because they're barely spending at all. Both sit at the bottom of token usage this quarter, and both already show up in Capability by Area's limited-coverage list for exercises like Threat Intel Analyst. Low spend here isn't efficiency — it's absence.

Sources: this saved report · Capability by Area — Agent Tuning
Ask a follow-up or request a change…
Token Usage & Cost by Team — Q4, Agent Tuning only
✓ Saved
TEAM×TOKEN_USED×COST×EXERCISE_TYPE = Agent Lab×
SOC Team A
$18,240
SOC Team B
$12,110
Security Engineering
$8,760
Platform Engineering
$5,930
IT Risk & Compliance
$3,480
Team ↕Token UsedCostExercise Type
SOC Team A2,140,000$18,240Agent Lab
SOC Team B1,410,000$12,110Agent Lab
Security Engineering1,020,000$8,760Agent Lab
AskOne field: what data would you like to see — typed, suggested, or tapped in from an existing section.
Generate & explainA real, editable report, leading with a cited summary and a sources note nobody had to ask for.
Earn moreRecommendations become guided actions. Repeated, approved patterns get offered as automation — never assumed.