Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings

tombudd/tombudd

Open more actions menu

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

25 Commits
25 Commits
 
 
 
 
 
 
 
 

Repository files navigation

Tom Budd

AI Safety & Governance Engineer · Research-Grounded Systems Builder

Website Research LinkedIn Get Involved


I build small, reproducible evaluation artifacts for researchers and engineers reviewing agentic AI systems.

My public work focuses on:

  • agent authority boundaries
  • tool-use safety
  • human oversight
  • auditable decision evidence
  • deterministic synthetic evaluations

Public repositories contain synthetic, clean-room, or educational artifacts. They do not expose private production systems, logs, prompts, schemas, or customer data.


Start Here

Evidence-Gated Evaluation of Agentic AI SystemsMETHODS PREPRINT / RUNNABLE / TESTED

A bounded framework for testing agent capability without expanding authority. The public release includes the manuscript, JSON Schema, 18 synthetic cases, deterministic gate logic, regression tests, frozen results, and independent-review materials.

Start with:

frozen claim + authority envelope + evidence ledger -> PASS / HOLD / FAIL / INVALID_RUN

A passing result supports only the predeclared claim within the tested environment. It grants no authority to deploy or act.

Evaluation benchmark portfolio

AI Governance BenchmarksRUNNABLE / TESTED

A clean-room benchmark suite using synthetic cases, deterministic scoring, generated reports, and regression tests.

synthetic case -> scorer -> report -> tests

Start with:

What the portfolio demonstrates

  • Runnable synthetic evaluation cases
  • Deterministic scoring and reproducible reports
  • Tests that detect boundary failures and unsupported public claims
  • Explicit separation between capability evidence and operational authority

Additional Public Work

  • Agent Action Audit TemplateRUNNABLE / TESTED — schema-backed synthetic action receipts, blocked-action examples, human-review metadata, and validation tests.
  • Human-AI Governance LabRUNNABLE / TESTED — toy workflow gates for risk classification, human approval, reports, and synthetic audit receipts.
  • Active Inference PrimerRESEARCH_NOTES / UTILITIES — minimal educational free-energy utilities with synthetic numerical tests and explicit limitations.
  • Eudaimonic AlignmentRESEARCH_NOTES — public research notes on human flourishing, agency, and alignment/governance questions.
  • Quantum AI ExperimentsSANDBOX — simulator-first quantum/AI-adjacent experiments with explicit claim boundaries.

Contact

I am open to serious collaborators in AI evaluation, agent safety, governance engineering, red-teaming, and applied research.

tombudd.com · tom@tombudd.com · Get involved

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Morty Proxy This is a proxified and sanitized view of the page, visit original site.