We assessed our own AI.
Here is what we found.
We sell one question: can you evidence your AI controls operating, or are they only asserted? Before asking anyone else, we ran the same method on The Agility Doctor. Ten gaps, closed in two days, each one now backed by evidence we can show on request, and tested so it stays closed.
Inventory and tiering
A small firm still runs more AI than it thinks. We found five systems, two of them high risk: the one that touches client data, and the agent that can change the systems holding it.
Asserted or evidenced?
For every control that mattered, one question. Most of what we had was real in practice and invisible on paper. That is the normal finding, and it is exactly what a customer’s security review will catch.
Asserted. We knew what we ran, but nothing written down said which data each system touched or who approved its output.
An AI register: purpose, data in and out, model, human approver, known failure modes and risk tier for every system.
The register itself, reviewed line by line with the owner.
Asserted. Reports were reviewed by a person in practice, but a single configuration change would have let them go straight to a client, and nothing would have noticed.
Skipping review now needs two separate switches, one of which must be set to an explicit approval value.
An automated check every hour alerts on any report produced without review.
Not in place. Diagnostic data would have been kept indefinitely.
Automatic deletion 12 months after a report is delivered, unless the client has an active engagement.
A daily job runs the deletion, and the hourly check alerts if anything is older than the limit.
Not in place. Deletion would have meant hand-written database work under time pressure.
One-step deletion of a client’s answers, metrics, report and contact details, with a 7-day public commitment.
Every deletion is logged, without keeping any of the deleted data.
Asserted verbally only.
A published client data statement and a public AI Use Policy.
Model provider terms on file; both pages public.
True today, not protected tomorrow. Raw tracker exports were never stored, but no control guaranteed it would stay that way.
A check that runs before every release and blocks it if any code starts storing raw files.
Build logs for every release.
Backup tables sat behind the same interface as live data.
Moved into a private archive that no public interface can reach. Nothing was deleted.
Access rules on the archive, verified during the assessment.
Deleting a client’s data would have left any login they had to our client portal in place.
Deletion now removes the client’s portal logins too. Our own admin accounts are never touched.
Tested every day against planted client, admin and unrelated accounts.
Each control was checked when it was built. Nothing proved it still worked a month later.
Known-case tests: six planted cases run against the live data controls every day, and 183 checks run on the report pipeline before every release.
A failing daily case raises an alert within the hour, and a failing release check blocks the deploy. We proved both by breaking the controls on purpose and watching the tests catch it.
Found while building the tests above. Client text was already stripped of line breaks and length-capped, but a company or team name containing a closing tag could end the protected data block early.
Tag characters in client text are neutralised before any report is drafted.
52 known attack cases across all 13 client text fields, run before every release.
Closed once is not closed
A fix nobody re-checks is an assertion again within a quarter. This is what keeps these ten evidenced.
- Every dayPlanted test cases run against deletion and retention. Each one proves a control still does what the client data statement promises, then removes itself.
- Every hourA monitoring check reads the results, and alerts if a report goes out without review, client data is kept past its limit, or a daily test fails or stops running.
- Every releaseThe report pipeline’s guardrails are tested before the site can deploy: invented numbers, altered scores, and injection attempts in every client field.
What the tests do not cover: the judgment in a report. Automated tests guard the numbers, the scores and the boundaries around the model. Whether the analysis is right is still decided by a person, on every report.
How it maps
Structured around the four functions of the NIST AI Risk Management Framework, with the register and policy organized to match ISO/IEC 42001. Readiness, not certification.
Two days, one owner, one AI agent
An AI security and governance agent ran the audit and built the fixes under tiered approval: reversible changes with no client data could go ahead, anything touching client data needed the owner’s sign-off, and some changes were never automated at all. Every change was logged with a way to roll it back.
The agent is on our own register as a high risk system, with the same oversight rules as everything else. Governing the governor is part of the job.
Want the same read on your organization?
Forty-five minutes with whoever owns your delivery process. You receive a one-page sketch of where your controls are asserted rather than evidenced. Our full AI Use Policy is public.
Request a readiness read →