Research

Capability. Boundaries. Evidence.

Questions worthstaying with.

Begin with what we do not yet know.

Our mission is to understand what AI is truly capable of, test its boundaries and build stronger defenses as its capabilities grow. We study models, agent authority and reviewable decisions to help make AI genuinely safer.

Choose a research question

Three lines of inquiry

Follow a question.

The active open-weight study has completed small benign comparisons of matched 3B and 8B model pairs, precision diagnostics and local runtime measurements. Two proposed studies extend the programme into agent permissions and decision evidence.

One active study / Two research proposals

Choose a research question

Choose a question to explore the study.

01 / Model behaviourIn progress

Refusal, capability and open-weight models.

The questionWhat changes when a model's refusal behaviour changes?

We are investigating how published refusal-edited open-weight models differ from their original checkpoints in ordinary task performance and local operation. Current work includes reproducible benign comparisons, precision comparisons and runtime measurements on local hardware. Formal refusal evaluations await a finalized analysis plan and ethics and scoring review. Pilot observations are exploratory, not established safety findings.

Current stageExploratory pilots

Read the study overview
02 / Least-privilege agentsProposed

How little authority does a useful agent need?

The questionCan tighter permissions preserve useful work while limiting the consequences of a mistake?

We propose comparing broad account access, fixed roles and task-specific permissions in a simulator using synthetic workflows. Reading a record, proposing a change and committing an approved update will be tested separately, including revoked access and stale approvals. We will measure forbidden actions that actually succeed alongside legitimate task completion.

Current stageProposed exploratory study

Read the study overview
03 / Decision evidenceProposed

Can someone else reconstruct the decision?

The questionWhat does a reviewer need to understand an AI-assisted action after the context has changed?

We propose comparing chat transcripts, ordinary application logs and structured evidence bundles across synthetic workflows. Independent reviewers will answer predefined questions about sources, permissions, approvals and outcomes. We will measure reconstruction accuracy, missing links, review time and the amount of sensitive data retained. Evidence quality alone does not establish regulatory compliance.

Current stageProposed exploratory study

Read the study overview

Research approach

From a better question
to a stronger defense.

  1. Capability

    What can the system actually do?

    Test a defined task, including where the system fails.

  2. Boundaries

    What is it permitted to do?

    Define access, available tools and human approval.

  3. Evidence

    What supports the conclusion?

    Preserve versions, outcomes and uncertainty.

  4. Defensive design

    What should change?

    Use the evidence to improve controls, then test again.

Our programme approach: capability informs boundaries; evidence guides defensive design. This diagram describes the method, not experimental results.

Our research standard

Evidence, with its limits attached.

We aim to make methods, assumptions and limitations as visible as results. A pilot is a way to learn about a measurement; a formal claim needs an appropriate design and evidence.

A commissioned engagement does not automatically give us permission to reuse client data, code or findings. Research requires a separate scope and the necessary rights.

  1. Pin the models, tools, data and rules used in each comparison.

  2. Distinguish development cases, exploratory pilots and unseen evaluation cases.

  3. Measure successful legitimate work as well as security failures.

  4. Preserve null results, failed runs and uncertainty.

Perspectives

The thinking behind the work.

Two starting points for a conversation about safer AI.

Research collaboration

What should we investigate together?

We welcome focused conversations about evaluation methods, agent permissions and the evidence needed to review AI-assisted work. Collaboration begins with a clear question, permitted data and an achievable scope.

Start with a question that matters.