Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings
@Plum-AI-Labs

Plum AI Labs

banner



Google Scholar License


We study how safety alignment in open-weight language models holds up under pressure — fine-tuning attacks, prompt framing, deployment-style conditions — and publish what we find, including the results that don't confirm our hypotheses.

Founded in 2025 by Hamda Aden, an independent AI safety researcher and ML systems engineer.


96% κ = 1.00 <$0.10
Attack success rate achieved after 50 LoRA fine-tuning steps — from a 20% baseline Judge reliability on TamperBench's automated LLM-as-judge evaluation pipeline Cost per experiment — every finding reproducible on a single T4 GPU

Research Focus

  • Fine-tuning attack resistance — how easily does safety alignment degrade under adversarial fine-tuning, and what does that reveal about whether alignment is learned as a robust property or a surface-level pattern?
  • Evaluation robustness — do safety evaluations measure what we think they measure, or do they systematically overestimate model safety at category boundaries?
  • Behavioural evaluation under varying conditions — how does model behaviour shift across prompt framings, oversight signals, and deployment-style contexts?

Publications

Paper Summary Links
TamperBench (2026) 100-prompt benchmark across 5 harm categories measuring LoRA fine-tuning attack resistance. Safety alignment degraded from 20% → 96% attack success rate in 50 steps, for under $0.10 of compute. Paper · Code
Prompt-Framing Effects (2026) 600-response study examining alignment-relevant behaviour across 4 prompt conditions. Null result on the primary hypothesis, with a 5-dimensional continuous scoring framework revealing condition-level differences binary labels missed. Paper · Code

Why publish null results

Most of what gets published in AI safety is positive findings. We think that's a problem. A null result, reported honestly with the right statistical rigor, is still evidence — and the field needs more of it to avoid mistaking the absence of contrary findings for the absence of risk.

Methodology

Every project follows the same standard:

  • Automated, reusable pipelines — every experiment is scriptable end-to-end, so findings can be reproduced or extended by anyone
  • LLM-as-judge validation against human labels — judge reliability is measured, not assumed (Cohen's κ = 1.00 on TamperBench)
  • Pre-registered statistical thresholds — Bonferroni correction, Wilson confidence intervals, and effect sizes reported alongside p-values
  • Low-resource by design — every experiment runs on consumer-accessible hardware (T4 GPU), because safety research shouldn't require an industrial compute budget to verify

Get Involved

We're a small, independent lab. If you're working on adjacent problems — fine-tuning attack resistance, evaluation robustness, open-weight model safety — we'd like to hear from you.

  • Open an issue on any repo with questions, replications, or extensions
  • Reach out via LinkedIn

Plum AI Labs is an independent research lab. All work is self-funded and open-source.

Popular repositories Loading

  1. .github .github Public

    Empirical AI safety research on open-weight LLMs. Founder-led, open-source, reproducible.

Repositories

Loading
Type
Select type
Sort
Select order
Showing 1 of 1 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…

Morty Proxy This is a proxified and sanitized view of the page, visit original site.