Skip to content

Building in public

Build Log

A weekly account of the machine being built, written while it is being built. Receipts over adjectives.

The week the instruments confessed

112 commits. Sixteen speed claims killed on real hardware, zero survivors. Forty-three research claims banked with cryptographic receipts. Five real emails out the door, every one carrying its evidence. Fourteen of the commits are corrections of the same week's work. Pretty proud of this one too, for a different reason than last week.

New here, short version: I run a portfolio of ventures solo. An AI harness is the execution layer. One conductor terminal orchestrating worker terminals, governed by written law, enforced by code. This log is me building that machine in public. Receipts over adjectives.

The instrument was the liar

The flagship service sells performance findings to local businesses. The pipeline sourced 106 advertisers, measured them with the industry-standard lab tool, vetted the entities against state registries, and produced a lead list where every number had passed a corroboration gate: two independent runs, both past the threshold, quote the conservative one.

Then I loaded the slowest-claimed site on a real machine. It rendered in about two seconds. The lab said twenty-four.

We re-measured all sixteen on real hardware. Every single lab claim died. Mean lab number: 10.7 seconds. Mean real load: 1.5. The rank order did not even predict which sites were slower. The flagship lead, the scariest number on the list, was the fastest site in the set.

The corroboration gate never had a chance, and the reason is the law this week produced: two runs of the same simulator agree with each other about the simulation. Corroboration controls variance. It cannot buy validity. If the instrument is wrong, running it twice makes you confidently wrong.

The rebuilt list has nine leads instead of sixteen, and every finding on it is something the business owner can see in one click. A certificate error that throws a browser warning while they buy ad traffic. A dead booking link at a practice selling five-figure procedures. Placeholder phone numbers shipping on a live page. Claims no phone can refute, because the phone is the demo.

The rule is now enforced in the pipeline itself: simulated numbers may rank candidates internally, but no claim ships to a prospect that the prospect can refute with their own device in seconds. The funnel reads 106 sourced, 9 shipped, and I trust the 9 more than I ever trusted the 16.

Prose to physics, round two

Last week's principle held: hooks enforce what has a decidable predicate, prose governs what needs a mind. This week it compounded twice.

First, a documented wall decayed. Doctrine said the hook layer could not intercept calls to external tool servers, verified three ways a week ago, and several rules were prose-only for exactly that reason. Before wiring around it again I re-probed with a sentinel. The wall was gone. The platform had moved underneath the doctrine, and a week-old verification is a conjecture, not a fact. Two prose rules became enforcing guards the same afternoon.

Second, the mail client turned out to rewrite links at storage time. Paste a clean URL into a draft and the stored copy wraps it in a tracking redirect, invisibly, on a surface no tool can read back. Three drafts shipped ugly before the mechanism was measured. The fix is a guard at the tool boundary: link-bearing drafts must carry pre-built anchors, plaintext URLs are denied with the reason attached. The guard's first live catch was my own cleanup note, denied for naming the banned pattern in its text. The guard biting its author is becoming a tradition, and I have decided it is a healthy one.

The conductor learned to wake

The orchestration layer got its senses this week. The old way, the conductor polled worker terminals and read quiet screens as finished work. That produced the week's most instructive false claim: I reported four lanes complete when two were mid-thought between phases. The principal caught it from across the room.

The scientific pass that followed measured the failure honestly: working lanes go quiet for minutes at a time, so silence cannot mean done. The predicate that survived testing on all four ground truths: a lane is finished when its formal sign-off is the last substantive thing on the transcript, with nothing after it. Quiet gates, proof fires.

The conductor now gets woken by the machine in exactly three cases: a lane finished, a lane stuck waiting on an approval no automation is allowed to grant, or a lane hung past its deadline. Mid-week the whole machine crashed, fleet and all, and the recovery was the quiet victory: every lane resumed from its transcript on disk at full memory, nothing lost but the processes. Which also settled a resource question: an idle finished terminal burns real CPU redrawing its own dashboard, so finished lanes get closed on review now. Resume costs nothing. Idle costs a core.

The library the executives read

The principal asked one uncomfortable question this week: all that research the machine did, did any of it get written into the expertise corpus? The honest answer was zero. Six findings had been flagged for capture. None landed. The flags were prose, and prose dies under load. Same disease as last week, new organ.

Two fixes, same day. The backfill: forty-three claims extracted from the week's research, each stored with a hash of its verbatim source so a future instance can prove the quote is unaltered. Statute mechanics, patent doctrine, radio physics, and the lab-versus-real measurement itself, now citable forever. And the enforcement: any work order that includes research must declare, mechanically, either its write-back step or an explicit waiver. The judgment about what is corpus-worthy stays human. The declaration is physics.

Then the part that makes the library earn rent: the executive layer that derives objectives for each venture now reads the corpus at derivation time, cites the claims it leaned on by id, and where the corpus is silent it must write CORPUS-GAP instead of inventing ground truth. First live run consulted 79 claims and named two honest gaps. One of the gaps was real strategic work I owed and had not noticed.

Scope qualifiers travel

Near-miss of the week, and it gets a law. A legal analysis cleared a product concept, but the clearance was true only for one design shape. The headline said cleared. The recommended shape was a different one, and six live patents were unanswered for it. The verdict and its condition had come apart somewhere between the memo and the summary, and an independent review pass caught it before anything shipped: the qualifier is part of the claim, and a reader acting on the headline alone would have been wrong.

The same lane produced the week's best confirmation. A collaborator took the patent analysis and pulled every document himself rather than trusting the summary. Everything load-bearing held, including the error I had already flagged as my own. Verification by someone with no reason to be polite is the only compliment that compounds.

Also in that lane: an abandoned patent application bars nobody from anything, patents claim mechanisms rather than concepts, and both of those sentences are in the corpus now with the case citations attached.

Exports

My brother sent two project briefs this week asking to be red-teamed, both drafted by his own AI setup. The replies were built so his AI cannot inflate them: every verdict fused to the evidence that would flip it, and one standing instruction, attack these points only with receipts, a listing, a quote, a preorder count. If it answers with scenarios instead of evidence, it is agreeing with me in a longer voice. The discipline exports. That is the point of writing it down.

Scoreboard, with the honest column

  • 112 commits in the core harness. The line-churn number is garbage this week and I am not printing it: the repo now stores captured evidence files, and other people's HTML is not my work
  • 14 of 112 commits are corrections of the same week's output, including one correction that itself needed correcting twice before the law came out right
  • 2 new tool-boundary guards born, 1 extended, each with regression arms and at least one live fire on its own author
  • 5 real emails sent, each carrying its receipts. 43 corpus claims banked with source hashes. 16 lab claims audited, 0 survived
  • 27 automation jobs deliberately paused at week's end. The usage meter hit its ceiling and the honest move was shutting the factory to finish the week thinking instead of generating

And the honest column. The revenue gate still reads zero. It is closer in every way that is not the number: the flagship service now answers on its own domain, nine defensible leads exist with evidence a prospect can see in one click, and the vet request is in a reviewer's inbox. But closer is a feeling, and the gate is a number. The instrument lesson cuts here too: I do not get to corroborate my own optimism by running it twice.

Next week

  • A grant submission goes out the door Friday, receipts committed as always
  • The nine leads clear external review, then the first real outreach in the flagship vertical
  • The sourcing machine scales past one metro with the defect scanner doing the qualifying
  • Regression evals over the skill layer itself, so prompt infrastructure rots loudly instead of silently

Thesis, unchanged: compound leverage faster than the surrounding systems decay. Last week the compounding was error-to-law latency. This week it was receipt density: the fraction of claims in the system that carry their own evidence. A claim with a receipt survives any audit, any model swap, any cold start. A claim without one is a feeling with formatting.