ADR 0018: a topic for self-improvement, and why it is not `autonomous-research-loops` or `inferencing`

decision

ADR 0018: a topic for self-improvement, and why it is not `autonomous-research-loops` or `inferencing`

About this note
Field Value
Status accepted
Date 260804
Deciders chamin
On this pageContextThe three candidate homes, and why each failsDecisionAlso decided: inferencing stays emptyConsequences

Context#

S14 (Stanford CS329A lecture 1) is the brain's first source on a model improving itself from its own verified output. Its gated claims are about a loop: sample the model many times, select the outputs that survive a check, feed those back as training data, repeat (n5). Around that sit the decomposition that makes the loop legible (coverage against precision, n4), the scaling behaviour that makes it affordable (n1), and the constraint that decides where it can run at all (verification availability, n6).

The kit's standing guidance pushes against new topics, and three ADRs have declined one on that basis (0013, 0014, 0015) against one that accepted (0017). So the question is which pile this falls into, and there are three plausible existing homes to rule out first.

The three candidate homes, and why each fails#

evals is the closest and it fails on direction. That note is about measuring a system someone built - per-stage metrics, pass@k, QA gates, ablation, and how to design a metric an optimizer cannot game. S14's verifier is not a measurement of the loop, it is a load-bearing component inside it, and the difference is consequential rather than semantic: an eval that is wrong gives you a bad reading, while a verifier that is wrong writes its error into the next set of weights. That said, two of S14's claims genuinely are eval claims and are filed there rather than hoarded here - claim 125 (models prefer their own reasoning traces) and claim 126 (the verifier increasingly written by the system it judges), both of which extend claims 34 and 113 and are reusable well outside any self-improvement loop.

agents fails on layer. That note is about building agents, meaning the loop, the prompt, the context, the tools, the orchestration. S14 is about the model underneath the agent getting better at being a model. Only one of its claims is genuinely about agent construction (claim 127, that what ships is still a hand-drawn static graph), and that one is filed under agents.

autonomous-research-loops is the interesting near-miss, and the distinction is worth writing down because a future pass will be tempted to merge them. The two are the same shape at different layers. ADR-0017's note covers an agent iterating an artifact unattended against an automated metric, where the thing that changes is a training script and the loop runs on one machine overnight. S14's loop changes the model's own weights, runs inside a lab's training pipeline, and is not unattended at all. Apply ADR-0014's swap test in both directions. Swap S14's language model for a theorem prover or a code generator and every claim survives, which is the signature of a transferable area. Swap S14's loop closure for a one-shot generate-and-select pipeline and the entire subject evaporates, which identifies the closure as the defining feature rather than an incidental one. Filed under autonomous-research-loops, the coverage/precision split and the domain-gradient argument would sit beside claims about git reset and wall-clock budgets and read as someone else's subject.

Decision#

Create brain/topics/self-improvement.md, Status emerging (one source), covering the loop closure, the coverage/precision decomposition, test-time compute as a scaling axis, and verification as the rate limit. Claims 120-124 land there; 125 and 126 go to evals, and 127 to agents.

Cross-reference autonomous-research-loops rather than merging. The link is worth stating explicitly in both notes because it is genuinely productive in one direction: claim 124 (verification is the bottleneck and sets the ceiling) is the frame that explains claim 114 (S13's accept rule banking a random-seed change as an improvement). S13 had a real, cheap, automatic verifier and it still failed, because having a verifier and having a verifier that works are different things. That is a stronger statement of 124 than S14 itself supplies, and neither note reaches it alone.

Merge-back trigger, written down now rather than re-derived later. If a second primary source on self-improvement never arrives, and the next two sources on either note keep landing claims that could sit in both, merge the two notes under the broader name and demote the distinction to a section heading. Do not merge on aesthetic grounds while both notes are still accumulating.

Also decided: inferencing stays empty#

Claim 123 (test-time compute as a third scaling axis) is about inference, and it does not go into brain/topics/inferencing.md. That note's declared scope is serving - prefill and decode, the KV cache, batching, quantization, speculative decoding, throughput against latency. Test-time scaling buys accuracy rather than efficiency, and the two share a word rather than a subject. Filing it there would misfile the claim and, worse, would make a still-empty seed note look populated. The inferencing note remains at seed with zero sources, which is an accurate description of this brain's coverage.

Consequences#