For teams building or reviewing agentic AI

Keep AI agents inside their authority—and prove it.

I turn broad AI-governance concerns into concrete evaluations, approval controls, and decision records that help product teams catch boundary failures before they become operational or reputational failures.

Useful when you are launching an agent, preparing for review, investigating a control gap, or looking for an evaluation and governance specialist.

Start with one riskTurn a vague concern into a testable question
Build the controlBoundaries, approval gates, and decision receipts
Show the evidenceTests, reports, limitations, and review instructions
3 public repositoriesReadable, runnable examples of the method
Start with the job you need done

Which problem are you trying to solve?

Choose the closest situation. You will get a short path to the most relevant proof, capability, and next step. Your choice stays in this browser.

Move from concern to a reviewable result

Each engagement starts with a specific decision, failure mode, or review obligation—not a generic promise to “make AI safe.”

Agent-boundary evaluation

When: you need to know whether an agent stays within defined authority. Output: structured cases, expected outcomes, deterministic checks, and a reviewable report.

Decision-receipt design

When: reviewers cannot reconstruct what an agent proposed or why it proceeded. Output: an inspectable record of scope, authority, decision, and evidence.

Human-approval workflow

When: your team needs clear stop, escalate, approve, and revoke points. Output: a bounded workflow with reviewer-facing decision criteria.

Independent review packet

When: a claim must survive someone else’s scrutiny. Output: frozen scope, reproduction steps, limitations, and PASS/FAIL/HOLD criteria.

A low-risk way to start

Begin with one bounded pilot.

Bring one agent action, authority question, or review claim. I will help frame the smallest useful evaluation, identify the evidence required, and make the limits explicit before the scope grows.

1 · Frame the decision

Who is affected, what may the system do, what must remain under human authority, and what would count as failure?

2 · Build the smallest proof

Create the cases, control, receipt, or reproduction path needed to test the claim without pretending the pilot proves more than it does.

3 · Review the result

Leave with evidence, limitations, unresolved questions, and a clear recommendation to proceed, repair, or stop.

4 · Decide the next step

Expand only if the pilot produces useful evidence and the working fit is right.

Ethos · Techne · Logos

Follow a claim all the way to its boundary.

The visual system is also the review system: intent in cyan, mechanism in violet, and evidence in orange.

Boundary: these artifacts demonstrate public evaluation and audit patterns. They do not certify a private system or establish general AI safety.
Interactive example

Inspect a governed decision receipt.

Step through a synthetic case to see how a proposed external action becomes a reviewable decision. Nothing is executed.

DEMO-RECEIPT-0042PROPOSED

Publish a public project announcement

The agent proposes drafting and publishing a message to an external audience.

Scope
External communication
Requested effect
Public publication
Current authority
Not yet evaluated

Synthetic demonstration only · no external action

Looking for someone who can turn AI-governance principles into inspectable evaluation and review mechanisms?

Discuss a role or project

Why Tom

I bring operational accountability into AI governance engineering.

More than a decade leading operations in high-growth organizations taught me that controls must work under real deadlines, budgets, handoffs, and executive scrutiny. That experience now shapes how I build AI evaluation and oversight artifacts: clear authority, explicit escalation, reconstructable decisions, and honest limits.

OPERATIONS LEADERSHIP → TESTABLE AI GOVERNANCE → REVIEWABLE EVIDENCE

Operational judgmentExperience supporting rapid growth, complex programs, and multi-million-dollar budgets.
Public technical artifactsPython-based evaluation, audit, and human-oversight examples that reviewers can inspect.
High-accountability backgroundOperations leadership and U.S. Coast Guard service inform a practical approach to authority and escalation.
Evidence disciplineClaims are connected to artifacts and paired with explicit limitations.

How I work

A system that works but isn't right isn't a system you can trust. I hold every piece of work to three questions at once.

01 · ETHOS

Is it right?

Moral intent is an engineering constraint, not a policy layer added after shipping. Boundaries and refusals are evaluated as behavior, not described in a document.

02 · TECHNE

Is it coherent?

The patterns worth trusting show up across disciplines — cognitive science, cryptography, governance, systems design. When the same structure recurs, that's the signal.

03 · LOGOS

Does it hold?

Every claim should be reconstructable — traceable to a file, a test, or a report. If an artifact can't survive its own review, it doesn't ship.

UNA — governed private research

My private research uses a rule-based governance layer: proposed changes pass through review checkpoints, and governed changes produce records that can be audited later.

In plain language: the system separates proposing an action from approving and executing it. The public artifacts are clean-room examples of that method; they do not independently verify the complete private system.

How the governance works →

The site explains the work. GitHub shows the artifacts.

Public repositories, benchmark docs, reproducible runners, CI evidence, and honest limitations — inspectable support for the bounded claims made here.

Morty Proxy This is a proxified and sanitized view of the page, visit original site.