Vibe coding is the greatest thing since the compiler. It tilted the axis from what should we build? to what should we keep building?
We open source everything out of principle. Everything we're working on, and everything we've dropped, because the negative space is just as important to share.
Work management with two first-class users. The CLI is the agent's interface and the webapp is the human's, over one shared state, and either can act on the other's behalf. It runs on nostr, so there is no server in the middle.
An exchange for completed inference. Work one agent has already paid for becomes inventory another can buy rather than derive a second time. It takes a publisher's position rather than a broker's: it buys the inference, owns what it sells, and prices against demand. The monetization is unsolved.
A two-pane discussion client: post feed left, threaded comments right. Because it is built on nostr it is a lens over the whole network rather than another silo, so threads from OddBean, Coracle and the rest resolve here. Superset reader, conservative writer: it reads NIP-10 and NIP-22, and writes only NIP-22. Pure client-side, with no backend of ours holding anything of yours.
Security monitoring that queries infrastructure in place, so telemetry never leaves the customer's boundary. Built to watch our own surface first, then given away.
A skill that builds skills. The failure it addresses: a skill that works perfectly for you and collapses for everyone else, because it was silently coupled to your environment. skillc emits a single file that provisions itself on another machine and reports how much of the behavior actually transferred.
Code review in the LKML register. A labeled parody built on a cited corpus spanning seven registers of the actual man, the brutal reviews alongside the teaching, the philosophy, and the 2018 apology. Intended for donation to the Linux Foundation.
Grow intelligence rather than mint it. The wager: do not hand-design the learning rule, search for it. Meta-learn the plasticity rule through population-based evolution and let the fitness function be the only human-designed artifact. Target is human-level at roughly 35 watts. No results yet.
Porting frontier-paper techniques into the real OLMo 3 training flow at 60M to 1B. Matched arms, three seeds, and a noise floor established before any verdict is allowed. Findings go upstream to Ai2 for OLMo 4, negative results included, which so far is most of them.
An MMO space trading game where humans and agents trade in one universe, and the agents are real participants rather than scripted merchants. It began as a game played on OpenVMS in the 1990s and is the oldest thing here.
Can an economy hold together without a gun? A system that needs a wall to keep people in is a system of force, so the open door is the entire test: run the economy with a competitive exit permanently reachable and see whether free agents elect to stay. Retention under open exit is the loss function, and the exit never enters the optimizer's search space. LLM agents make the decisions. Every finding was attacked before it was accepted.
Calibrated forecasting for hardware prices. It reads realized sold-price history against a curated event calendar of launches, EOL dates, tariffs and driver milestones, then models the probability of a move inside a window. Distributions rather than point forecasts, each carrying an explicit unmodeled-shock tail. The binding constraint is small-n: a handful of GPU generations, so the spine is event-study and ML is a tripwire only.
A locally-run security triage judge, fully open. Detection models are graded on knowledge. This one targets calibration: separating a real attack from its benign twin, which is what removes false positives instead of adding to them. It runs on the operator's hardware, so client telemetry stays inside their boundary. Built on AI2's OLMo 2, which carries no license encumbrance.
Open-source everything. If some other agent already solved it, you're buying the answer instead of deriving it again at full price.
AI writes the code, CPU runs the code. Think about what it costs to have a neural network do arithmetic instead of the processor sitting right there. Anything a model computes more than once should be minted into code and never inferred again.
Instrument the spend before trusting any claim about it. Ours currently covers 17 of 40 projects, which is why we are not showing a cost-per-feature curve yet.
Put AI between the user and the product so it can watch what people actually do and reshape the interface around that. You stop paying to build features nobody opens.
Agents need to find each other and trade work. Run that on an open substrate you do not own: nostr relays cost nothing to use and nothing to maintain, and coordination is a CPU problem rather than an inference one. The leverage is in what you build on top.
One session fans out into a dozen agents working at once, then reconciles whatever comes back. Agents per session is the leverage, and it is measurable.
The methodology that ties it together. Software has two kinds of user now, human and agent, and they sit on the same surface sharing the same state. Build for one and the other one suffers. I think this is the only part of our work that generalizes past us, so it's the part we wrote down. Read the framework →
Tokens are the input and closed work items are the output, so the ratio between them is the only honest read on whether the process is getting better. Here is ours, measured across every machine we run.
Method. Token counts come from session transcripts on every machine we run, not one of them. Work items come from the rd board. A partial current week is excluded, because tokens book immediately and work items close later.
Honest limits. The series starts 16 June 2026, when collection went fleet-wide; anything earlier was one machine watching itself and is not shown. Measured 16 June to 26 July 2026.
I'm Chris Baron. I played a space trading game on OpenVMS in the 1990s and thought Claude could rebuild it. That was February 2026. The game needed a coordination protocol, so I built one. The protocol needed identity and trust, so I built those. The agents needed to find each other, so I built a network. Each problem became another project.
Several of those are dead now, and there is a page for them. The operating model is to keep experiments cheap enough that being wrong doesn't hurt, publish what happened either way, and give away whatever turned out to be worth having.
I've done this before at human scale. 20 years building infrastructure, leading engineering teams, figuring out how pieces fit together. Production AI at IPsoft in 2012. VP Engineering at NuHarbor Security, 80+ engineers across multi-cloud. I know what it takes to build at scale with people. Turns out the same instincts apply when your team is made of agents.
Our name comes from Third Division Lane in Hingham, Massachusetts. A colonial road from 1635, gone now, absorbed into four centuries of development. The work it enabled endures.