Goodhart-proof AI coding pipeline with architectural isolation
-
Updated
Jun 5, 2026 - Python
Goodhart-proof AI coding pipeline with architectural isolation
Standing Algebra (ΣR): A Closure-Theoretic Operator for Constraining Domination and Preserving Autonomy
Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.
Self-improvement loop for LLM agents with an integrity gate that catches reward hacking: reward gains that erode reproducibility get reverted, even ones a reward-only gate would accept.
Catch reward traps before training. Static analysis for RL reward functions.
The forge, distilled: an ontology of three weeks of alignment research — every direction tried, colored verified / falsified / open, each color backed by a named artifact. Products: justitia, proxylimen, fallacy-cutter. Full tree at tag forge-full-tree.
Medium is the message: multi-agent coordination with no boss agent, no planner, and no messaging between agents only a shared medium and a validation gate. The bet: a crowd of dumb, selfish, local agents climbs the global hill if the signal is well designed.
Toy 6. An interactive phase-space instrument mapping Ψ = S/D — the ratio of capability to modeling depth that determines whether a system is in the viable, transitional, or failure-mode-dominant regime. Includes the Inner Crossing animation. Companion simulation for The Inner Crossing — Series 2, Part 3.
Chain-of-density study of 'Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents' (Zhang et al. 2026, arXiv:2607.12790) - five-tier note, locator-verified. Double Ratchet, drawback detectors, anchor discipline, Goodhart repair.
A preregistered synthetic simulation of proxy failure under optimization pressure, with deterministic replay, matched controls, and adversarial verification.
98% of the Alpha Was a Gate — a documented reward hack in a self-optimizing LLM research agent. Full search trajectory, frozen evaluation harness, both strategies, and re-runnable ablations. The agent found a second exploit 18 minutes after the first was patched.
`The Alignment Constraint Framework — a structural argument about specification coherence in AI alignment`
Toy 5. An interactive proxy decay simulator showing how optimization pressure erodes the modeling capacity required to distinguish proxy from territory — producing self-reinforcing V(t) degradation that becomes progressively harder to correct. Companion simulation for The Depth Constraint — Series 2, Part 2.
ICFRAME: compile declarative incentive specs into deterministic multi-agent experiments that measure reward hacking.
PPO for RLHF with a Bradley-Terry reward model, plus Best-of-N, DPO, RLOO and GRPO. Reward overoptimization measured: proxy climbs while gold peaks and collapses below the reference.
Evals with teeth: score your agents on the future — live prediction markets they can't game — and pay them in capital. Benchmarks saturate; the future doesn't.
To associate your repository with the goodharts-law topic, visit your repo's landing page and select "manage topics."