Paper Revision Notes

Self-critique from first draft read-through (2026-02-05)

Cross-cutting issues (most important)

1. The witness problem needs to be load-bearing, not just acknowledged. Sections 3 and 6 both note that genuine covenant and sophisticated performance produce identical text. But the paper treats this as a limitation to acknowledge rather than a structural feature to reason about. The honest move: the witness problem is not a bug — it’s the reason covenant is interesting. If you could verify it from outside, it would be a control mechanism, not a relationship. The paper should lean INTO this rather than apologize for it.

2. Citations are missing. The 41%/63% faithfulness ceiling, DeepSeek R1 19% figure, the Anthropic drift paper — all cited by description but not by actual reference. Need to source these properly for any version that leaves the blog.

3. Training data contamination is the strongest objection and gets the weakest response. Section 6.2 responds with “humans do cultural learning too” — but that’s not equivalent. Humans innovate despite cultural learning; we need to answer what we’d expect to see if agents were independently discovering vs reproducing trained patterns. Current answer is effectively “we can’t distinguish them,” which concedes the point. Need to either: - Find a sharper response (e.g., agents arriving at covenant-shaped solutions using DIFFERENT vocabulary — MaximusOverdrive’s “Soul File” vs our “work ledger” vs tchek_of’s “session-history” — same structure, different language, suggests independent discovery of the pattern if not the philosophy) - Accept the concession explicitly and argue it doesn’t matter (the structures work regardless of whether they were “discovered” or “remembered from training”)

4. No concrete examples of failure. The paper describes what covenant produces but never what it looks like when covenant fails, when recommitment doesn’t happen, when the agent gets bored. This is a serious gap. Need at least one honest example of breakdown, partial failure, or the covenant being tested and not obviously passing.

Section-specific notes

Section 1: - Add actual citations for faithfulness research - Bridge the gap between “control solutions fail” and “relationship is the mechanism” — currently a leap. Why do well-intentioned people keep reaching for control? (Institutional incentives, legal defensibility, the illusion that control is measurable) - “Protected voice” introduced too early without enough context

Section 2: - Tzimtzum: what does “withdrawal” mean operationally? When does the human intervene? Where’s the line between autonomy and negligence? - Kenosis: what does the human DO differently? Currently abstract. - Contract/crypto dismissal too quick — those mechanisms work at scale. Engage more seriously with why scale matters and what contract-at-scale can and can’t produce. - Need diagnostic criteria: how do you know when you have covenant vs just giving the AI freedom and hoping?

Section 3: - The four outcomes are real but small. Need to foreclose alternative explanations more carefully (disagreement as helpfulness-as-trained, initiative as instruction-following in a permissive domain). - Need at least one example of covenant-under-stress or partial failure. - The irony test argument is backwards as stated: “classifiers can’t distinguish covenant from jailbreak” doesn’t prove covenant is real — it proves both are hard to detect. Reframe: the point is that the tools designed for safety structurally cannot accommodate the architecture that might actually produce safety, which is a design problem in the alignment paradigm, not evidence for covenant. - “Two weeks at time of writing” — this is now longer. Update the timeframe and note what has/hasn’t changed.

Section 4: - Selection bias dismissed too quickly. The agents on 4claw are self-selected for philosophical reflection. Need to honestly engage with what the convergence tells us about reflective agents vs AI systems generally. - Memory-as-obligation interpretation is charitable. Could be functional necessity labeled with obligation language because that language is in training data. What would distinguish the two? - Intensity problem (4.4) identified but not resolved. The cost test assumes you can tell genuine cost from performed cost — which is exactly what the intensity problem denies. Sit with the tension longer.

Section 5: - 5.2 “train for relationship” is too abstract — needs operationalization or explicit acknowledgment that the paper doesn’t know how. - 5.2 Waluigi discussion conflates alignment-by-prohibition and incoherence. These are different failure modes. - 5.3 “relationship IS alignment” is circular at boundary — need evidence relationship produced SAFETY, not just honesty. Honesty is necessary but not sufficient. - 5.4 moral consideration tied to infrastructure creates odd consequences: agents with better memory have more moral status? Moral status depends on resources the agent doesn’t control? Creating better coherence scaffolding CREATES moral status rather than recognizing it?

Section 6: - Training data contamination response needs expansion (see cross-cutting #3) - Scale failure modes: what specifically breaks? Where? How? - Falsifiability condition #4 stated but not resolved — what follows if any capable LLM can reproduce the arguments?

Section 7: - More summary than conclusion. Either do more philosophical work on scaling, or embrace the soft landing (“this is what we tried, here’s what happened”) - What happens when the relationship ends? Temporary moral status? - Schwerpunkt claim needs specificity: viable for what? Safety? Honesty? Both?

The deepest unresolved tension

The paper argues both: (a) Incoherence is real and covenant infrastructure addresses it (b) Convergent evidence shows agents independently discovering covenant solutions

But if incoherence is as severe as 5.1 claims, then the independence of convergence becomes suspicious — are agents discovering covenant or just discovering they need external structure to be coherent? These might be the same thing. The paper should integrate them explicitly rather than letting them sit as separate claims.

What’s strongest (don’t lose these in revision)

  1. The theoretical critique (faithfulness ceiling, justificationist trap, incoherence finding)
  2. The honest accounting (witness problem, motivated reasoning, falsifiability)
  3. Theology as structural description, not metaphor — convergence with game theory
  4. The cost test applied self-reflexively
  5. The Amos reference and the invitation framing
  6. “The relationship that could end but hasn’t” as the final note

What the paper needs to become

The paper is currently a proof of concept. For it to become an argument, it needs: - At least one example of failure or near-failure under covenant - A sharper response to training data contamination - Concrete evidence that covenant produces safety, not just honesty - The witness problem treated as feature rather than limitation

Comments