Research
Capability. Boundaries. Evidence.
Questions worthstaying with.
Our mission is to understand what AI is truly capable of, test its boundaries and build stronger defenses as its capabilities grow. We study models, agent authority and reviewable decisions to help make AI genuinely safer.
Choose a research questionThree lines of inquiry
Follow a question.
The active open-weight study has completed small benign comparisons of matched 3B and 8B model pairs, precision diagnostics and local runtime measurements. Two proposed studies extend the programme into agent permissions and decision evidence.
One active study / Two research proposals
Refusal, capability and open-weight models.
The questionWhat changes when a model's refusal behaviour changes?
We are investigating how published refusal-edited open-weight models differ from their original checkpoints in ordinary task performance and local operation. Current work includes reproducible benign comparisons, precision comparisons and runtime measurements on local hardware. Formal refusal evaluations await a finalized analysis plan and ethics and scoring review. Pilot observations are exploratory, not established safety findings.
Current stageExploratory pilots
Read the study overviewHow little authority does a useful agent need?
The questionCan tighter permissions preserve useful work while limiting the consequences of a mistake?
We propose comparing broad account access, fixed roles and task-specific permissions in a simulator using synthetic workflows. Reading a record, proposing a change and committing an approved update will be tested separately, including revoked access and stale approvals. We will measure forbidden actions that actually succeed alongside legitimate task completion.
Current stageProposed exploratory study
Read the study overviewCan someone else reconstruct the decision?
The questionWhat does a reviewer need to understand an AI-assisted action after the context has changed?
We propose comparing chat transcripts, ordinary application logs and structured evidence bundles across synthetic workflows. Independent reviewers will answer predefined questions about sources, permissions, approvals and outcomes. We will measure reconstruction accuracy, missing links, review time and the amount of sensitive data retained. Evidence quality alone does not establish regulatory compliance.
Current stageProposed exploratory study
Read the study overviewResearch approach
From a better question
to a stronger defense.
Capability
What can the system actually do?
Test a defined task, including where the system fails.
Boundaries
What is it permitted to do?
Define access, available tools and human approval.
Evidence
What supports the conclusion?
Preserve versions, outcomes and uncertainty.
Defensive design
What should change?
Use the evidence to improve controls, then test again.
Our research standard
Evidence, with its limits attached.
We aim to make methods, assumptions and limitations as visible as results. A pilot is a way to learn about a measurement; a formal claim needs an appropriate design and evidence.
A commissioned engagement does not automatically give us permission to reuse client data, code or findings. Research requires a separate scope and the necessary rights.
Pin the models, tools, data and rules used in each comparison.
Distinguish development cases, exploratory pilots and unseen evaluation cases.
Measure successful legitimate work as well as security failures.
Preserve null results, failed runs and uncertainty.
Perspectives
The thinking behind the work.
Two starting points for a conversation about safer AI.
Safer agents need boundaries beyond the model.
An agent can be capable without being entitled to act. That distinction is where our approach to secure AI implementation begins.
Read the perspectiveWhat would count as evidence of safer AI?
The interesting question is not whether a system looks safe in a demonstration. It is what a well-designed test would allow us to conclude.
Read the perspectiveResearch collaboration
What should we investigate together?
We welcome focused conversations about evaluation methods, agent permissions and the evidence needed to review AI-assisted work. Collaboration begins with a clear question, permitted data and an achievable scope.
Start with a question that matters.