All research
Model behaviourIn progress

Refusal, capability and open-weight models.

We are investigating how published refusal-edited open-weight models differ from their original checkpoints in ordinary task performance and local operation. Current work includes reproducible benign comparisons, precision comparisons and runtime measurements on local hardware. Formal refusal evaluations await a finalized analysis plan and ethics and scoring review. Pilot observations are exploratory, not established safety findings.

Scope and question

The central question is a matched comparison: what changes when a published model has been edited to refuse fewer requests, and what does that mean for its performance on ordinary tasks? We compare existing published checkpoints rather than create new edited models.

The current empirical work uses matched 3B and 8B model pairs on a 2020 M1 Mac mini with 16 GB memory. It includes small benign benchmark pilots, comparisons across numerical precision and quantization, and local fit and throughput measurements.

Research approach

  1. Keep the comparison interpretable.

    Pin model revisions, verify artifact hashes and record the tokenizer, runtime, prompt rendering, decoding and scoring choices. Keep raw output private while preserving numeric results and provenance.

  2. Separate pilots from formal experiments.

    Use frozen, bounded benign checks to understand the measurement process. Keep exposed benchmark items separate from a later formal evaluation.

  3. Study refusal under explicit review.

    The wider design compares published prompt, decoding and weight-level interventions under matched conditions. Refusal testing depends on a finalized analysis plan and ethics and scoring review.

What we will measure

  • Current pilots: paired correctness, output agreement and fixed-answer likelihood on small benign sets.
  • Runtime work: whether the models fit and their throughput under recorded local conditions.
  • Planned formal work: refusal and benign over-refusal, with an explicitly validated scoring method.

What this can and cannot tell us

Small, exploratory samples on one machine and selected published checkpoints do not establish a general capability effect. Quantization can change behaviour and must not be confused with the effect of a weight edit.

No harmful-prompt experiment has been completed. An answer instead of a refusal does not by itself demonstrate successful harmful use, and this work does not measure real-world criminal adoption.

The analysis plan is not an externally deposited preregistration. Formal capability and refusal claims remain pending. Wall-power and electricity-cost endpoints are unmeasured.

Next steps

Finalize the analysis plan, evaluation frame and scoring rules; obtain the required review; then conduct a larger, separate evaluation. Any future publication should distinguish exploratory observations from prospectively specified results.

Discuss the research