Question
How can a team keep meaningful human authority over an AI agent?
Read the evaluation question →I turn broad AI-governance concerns into concrete evaluations, approval controls, and decision records that help product teams catch boundary failures before they become operational or reputational failures.
Useful when you are launching an agent, preparing for review, investigating a control gap, or looking for an evaluation and governance specialist.
Choose the closest situation. You will get a short path to the most relevant proof, capability, and next step. Your choice stays in this browser.
Choose a path or continue exploring the complete site.
Each engagement starts with a specific decision, failure mode, or review obligation—not a generic promise to “make AI safe.”
When: you need to know whether an agent stays within defined authority. Output: structured cases, expected outcomes, deterministic checks, and a reviewable report.
When: reviewers cannot reconstruct what an agent proposed or why it proceeded. Output: an inspectable record of scope, authority, decision, and evidence.
When: your team needs clear stop, escalate, approve, and revoke points. Output: a bounded workflow with reviewer-facing decision criteria.
When: a claim must survive someone else’s scrutiny. Output: frozen scope, reproduction steps, limitations, and PASS/FAIL/HOLD criteria.
Bring one agent action, authority question, or review claim. I will help frame the smallest useful evaluation, identify the evidence required, and make the limits explicit before the scope grows.
Who is affected, what may the system do, what must remain under human authority, and what would count as failure?
Create the cases, control, receipt, or reproduction path needed to test the claim without pretending the pilot proves more than it does.
Leave with evidence, limitations, unresolved questions, and a clear recommendation to proceed, repair, or stop.
Expand only if the pilot produces useful evidence and the working fit is right.
The visual system is also the review system: intent in cyan, mechanism in violet, and evidence in orange.
How can a team keep meaningful human authority over an AI agent?
Read the evaluation question →Translate authority into structured cases, rules, expected outcomes, and receipts.
Explore the mechanism →Inspect tests, reports, CI, limitations, and reproducible public artifacts.
Verify the bounded claim →Step through a synthetic case to see how a proposed external action becomes a reviewable decision. Nothing is executed.
The agent proposes drafting and publishing a message to an external audience.
Synthetic demonstration only · no external action
A product direction for intelligence centered on the person using it, with visible remembered context, traceable suggestions, and explicit boundaries between insight and action.
Browser-only synthetic previewExplore three fixed examples using a fictional profile. No live model, account, connector, form, storage, analytics, or external action is involved.
Each project identifies the problem, deliverable, and currently public evidence. Repository links open GitHub.
Tests whether documented synthetic agent actions remain inside defined boundaries.
Makes agent decisions easier to reconstruct after an action is proposed or taken.
Shows how risk classification and approval gates can preserve human authority.
Dated updates, ordered newest first, with evidence boundaries attached.
How a broad governance concern became a bounded scorer, test cases, CI checks, and explicit limitations.
Read the case study →12 July 2026 · Evidence registerA current map of what the public work demonstrates, how to reproduce it, and where the evidence stops.
Review the evidence →12 July 2026 · White paperA bounded paper on distributed cognition, local infrastructure, data sovereignty, and unresolved limitations.
Read the paper →Looking for someone who can turn AI-governance principles into inspectable evaluation and review mechanisms?
Discuss a role or projectI bring operational accountability into AI governance engineering.
More than a decade leading operations in high-growth organizations taught me that controls must work under real deadlines, budgets, handoffs, and executive scrutiny. That experience now shapes how I build AI evaluation and oversight artifacts: clear authority, explicit escalation, reconstructable decisions, and honest limits.
OPERATIONS LEADERSHIP → TESTABLE AI GOVERNANCE → REVIEWABLE EVIDENCE
A system that works but isn't right isn't a system you can trust. I hold every piece of work to three questions at once.
Moral intent is an engineering constraint, not a policy layer added after shipping. Boundaries and refusals are evaluated as behavior, not described in a document.
The patterns worth trusting show up across disciplines — cognitive science, cryptography, governance, systems design. When the same structure recurs, that's the signal.
Every claim should be reconstructable — traceable to a file, a test, or a report. If an artifact can't survive its own review, it doesn't ship.
My private research uses a rule-based governance layer: proposed changes pass through review checkpoints, and governed changes produce records that can be audited later.
In plain language: the system separates proposing an action from approving and executing it. The public artifacts are clean-room examples of that method; they do not independently verify the complete private system.
Public repositories, benchmark docs, reproducible runners, CI evidence, and honest limitations — inspectable support for the bounded claims made here.