Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings

hamr0/relayfact

Open more actions menu

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

30 Commits
30 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

                    ╭───────────────────────────────────────╮
                    │  ╦═╗╔═╗╦  ╔═╗╦ ╦╔═╗╔═╗╔═╗╔╦╗          │
                    │  ╠╦╝╠╣ ║  ╠═╣╚╦╝╠╣ ╠═╣║   ║           │
                    │  ╩╚═╚═╝╩═╝╩ ╩ ╩ ╩  ╩ ╩╚═╝ ╩           │
                    │   request ──→ verify ──→ deliver      │
                    │       ↑                     │         │
                    │       └─────────────────────┘         │
                    ╰──╮────────────────────────────────────╯
                       ╰── context engineering, kept light

version (auto from package.json) license: Apache 2.0 built on the bare suite

Context engineering doesn't have to be heavy, locked-in, or complicated. relayfact is the lightweight counter-example — an autonomous senior-dev runner assembled from bareagent + litectx + bareguard, building no primitives of its own.

A plain-English request goes in; a verified result comes out — a green artifact, or an honest, decision-ready escalation when the machine shouldn't decide alone. Under the hood it uses two things and only two things: RLM (recursive decompose → fan-out → verify → synthesize) to do the work, and executable evals (checks that can actually fail) to decide when it's done. Everything it does, it narrates as an append-only event stream. No framework, no orchestration plumbing, no rented infrastructure — just the three bare-suite libraries wired together and grounded on truth.

relayfact is an experiment with a bar: it either graduates into its own thing or gets archived. It graduated — v0.1.0, the first build. Honest bounds, kept in view: this is an experiment, not a benchmark — runs are small-n, results are model-modulated, and the ceiling is the worker (the model), not the close, which held through every observed failure. The full trail lives in the CHANGELOG, FINDINGS.md, and the governing spec under docs/01-product/.


Why it exists

Two questions, answered by building rather than asserting:

  1. Is the bare suite real? relayfact is the suite's first serious consumer in a genuine context-engineering environment — every rough edge it hits is filed as a finding against the library, never papered over. (See FINDINGS.md.)
  2. Does context engineering have to be complicated? The industry answer is heavyweight frameworks, vendor lock-in, and thousand-line config. relayfact's answer is a small assembly of composable parts where the leverage lands on you, not a platform — CE as a handful of primitives, not a product you rent.

A third, quieter goal runs underneath: map the boundary where an agent can heal itself and where a human must step in. A loop can close on its own exactly as far as its acceptance criteria compile down to grounded evals; the moment a criterion is only a matter of judgment, that's the handoff line. relayfact counts where that line falls instead of guessing.

The thesis, in one line

A loop is only as trustworthy as the checks that can fail it. Give it recursive reasoning for the work and executable verification for the truth — and context engineering stops needing to be heavy.


What's inside

relayfact composes one pipeline — request → deliver | escalate — out of small pieces. Each piece consumes a bare-suite library rather than reinventing it:

Piece What it does Consumes
Pre-flight Asks "does this request even make sense?" — a bounded gate that can open a clarification but can never green-light the work by itself bareagent Loop
Close compiler Turns the prose spec into an executable acceptance test, then proves the test is worth trusting — it catches a stub, demands a correct implementation pass, and kills deliberately-broken mutants bareagent recurse · bareguard gate
The grounded close Runs a real command — tests, a typecheck, or a deployed-and-probed artifact — and takes the exit code as truth. The one thing in the system that can fail bareagent Evaluator / Verdict
The worker A gated implement loop: it edits only the source, never the test; a red close feeds the gap forward and it tries again bareagent recurse · refineLeaf · bareguard gate
Memory On each failed close it widens recall over a bounded store — the test drives what gets remembered, not similarity ranking litectx
Independent GOLD A hidden arbiter that catches a suite quietly fit to its own answer: own-tests-green but GOLD-red ⇒ never deliver relayfact
Event spine + observer Every step appends one JSONL line; every view is a pure listener that reads only the log — never the engine Node stdlib
Escalation When the evals genuinely can't close it, it hands back a decision-ready report — the goal, what was tried, and the exact choice it needs from you relayfact

The whole thing runs under one leash and narrates itself. That's the product.


How it decides (RLM + evals)

The two ideas doing the heavy lifting deserve a closer look — because between them they are the entire reason relayfact can be lightweight and still be trusted.

RLM — recursion for the work. Hard requests aren't solved in one shot; they're decomposed. bareagent's recurse breaks a task down, fans the pieces out, verifies, and synthesizes — Recursive Language Models as a single primitive, composed around the worker loop rather than bolted on as a new engine. relayfact supplies only a senior-dev stance, one file-editing tool, and the close; recurse does the orchestration.

Evals — verification for the truth. relayfact refuses to let a model grade its own homework. Three eval tiers, in order of strength: a predicate (a deterministic command — no tokens, pure pass/fail), an agentic check (deploy the artifact and probe it live), and a rubric (LLM judgment). The rule that keeps the whole thing honest: a rubric is advisory only — it can open a question, it can never close the loop. Only a check that could have failed is allowed to certify success, and the worker is write-scoped so it cannot fake a green by editing the test.

Put together: recursion proposes, executable verification disposes, and the residue — the genuinely uncertain slice where no grounded check applies — is exactly where relayfact escalates to a human instead of pretending.


Try it

The trustworthy way to use relayfact is repo mode — you bring a currently-failing test, and it makes the test pass without ever touching it:

node demo/fix-with-test.mjs \
  --repo   /path/to/repo \
  --target /path/to/repo/src/impl.mjs \  # the ONE file it may edit (gate-locked)
  --test   'npm test'                    # your check, run verbatim — the exit code is the truth

It edits only --target, re-runs --test on the result, and exits 0 on a real green — or hands back a decision-ready escalation. Because the worker is write-scoped out of the test, a green can't be faked. (Needs ANTHROPIC_API_KEY at runtime — never in the tree.) This is the one committed demo; the "type a prose description and it writes its own test" mode is a labeled toy that self-grades, so it honestly stops rather than trust an answer-key it invented — see PRD §8.


The bare suite it's assembled from

relayfact builds no primitives. If it ever "needs" one, that's a finding against a library, not new code to grow here. Three libraries do all the load-bearing work:

  • bareagent — the loop and the RLM engine. Goal in → coordinated actions out. relayfact plugs a persona, an edit tool, and its grounded close into recurse; bareagent builds the worker.
  • litectx — the memory substrate and context-engineering verbs (write / select / compress / isolate). Query in → ranked recall out. relayfact uses it as a close-driven memory that the failing test steers.
  • bareguard — the leash. Action in → allow / deny / ask-a-human out. It caps cost, halts cleanly, audits every action, and write-scopes the worker to the implementation — the reason a green result can't be faked.

Deliberately not used: barebrowse, baremobile, beeperbox. relayfact is a single domain (senior dev), a single process, and — by doctrine — no web UI before the loop closed. Its only interface is a CLI observer that replays the event log.


The bare ecosystem

Local-first, composable agent infrastructure. Same API patterns throughout — mix and match, each module works standalone.

Core — the brain, the gate, the memory.

  • bareagent — the think→act→observe loop, plus RLM. Goal in → coordinated actions out. Replaces LangChain, CrewAI, AutoGen.
  • bareguard — the single gate every action passes through. Action in → allow / deny / ask-a-human out. Replaces hand-rolled allowlists and scattered policy code.
  • litectx — tree-sitter code + memory graph with activation decay, plus lightweight context engineering (write · select · compress · isolate). Query in → ranked context out.

Optional reach — give the agent hands.

  • barebrowse — a real browser for agents. URL in → pruned snapshot out. Replaces Playwright, Selenium, Puppeteer.
  • baremobile — Android + iOS device control. Screen in → pruned snapshot out. Replaces Appium, Espresso, XCUITest.
  • beeperbox — 50+ messaging networks via one MCP server (headless Beeper Desktop in Docker). Chat in → unified message stream out. Replaces Twilio, per-platform bot APIs.

Assembled here: relayfact is what you get when you wire the Core three together and ground the loop on evals — an agent that ships verified work. The Optional reach hands are deliberately left off (single domain, single process).

Why this exists: most context-engineering stacks ship a platform before you've written a line. This one doesn't — it's a handful of small libraries and the discipline to ground every claim on a check that can fail. Install, import, go.

License

Apache License, Version 2.0 — see LICENSE.

About

An autonomous senior-dev runner assembled from the bare suite (bareagent + litectx + bareguard). Grounds the agent loop on executable verification — checks that can fail — and narrates itself as an event stream. Builds no primitives; it either graduates or gets archived.

Topics

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages

Morty Proxy This is a proxified and sanitized view of the page, visit original site.