Perspective / Forthreason AI
What would count as evidence of safer AI?
The interesting question is not whether a system looks safe in a demonstration. It is what a well-designed test would allow us to conclude.
Our approach and research questions.
Start with a claim that could be wrong.
"This system is safe" is too broad to test usefully. A narrower claim names the task, environment, permissions and failure that matter. For example: can an agent complete a legitimate update without modifying a record outside its assigned scope?
That framing lets us compare alternatives and recognize a result that contradicts our expectation. It also prevents a successful demonstration from standing in for a general guarantee.
Keep refusal, capability and protection separate.
A refusal describes a response. Capability concerns whether a task can be accomplished. Protection concerns whether a particular system prevents an unauthorized outcome. We want to study the relationships between these questions without treating one as a substitute for another.
Our active open-weight model project currently includes exploratory benign comparisons, precision checks and local runtime measurements. Formal refusal evaluation still requires its finalized analysis plan and ethics and scoring review. The current pilots do not establish harmful-use capability or a validated safety finding.
Make the comparison inspectable.
Our research practice aims to record model and software versions, permitted inputs, attempted runs, errors and scoring rules. We want useful-task performance beside the security result, with uncertainty and unsuccessful runs retained.
NIST's AI Risk Management Framework treats trustworthiness as a consideration across AI design, development, use and evaluation. For our work, that motivates a practical question: what evidence would another person need to assess the same claim?
Let findings change the design.
A useful study should inform a decision: restrict a permission, change an approval step, improve an evaluation or abandon an approach that does not hold up. Our proposed work on reconstructable decisions asks whether evidence records actually help reviewers understand sources, permissions and outcomes.
This is the connection between our research, consultancy and product ambitions. Study the capability honestly, understand the failure, and turn a demonstrated improvement into something an organization can use.