First measured Codex integration
Measured OFF vs ON. Correctness first.
Exploratory 3-task smoke, one paired run per task. Codex CLI 0.147.0 with GPT-5.6-sol, identical prompts and limits, verified in official SWE-bench Lite task environments on Modal.
codex plugin marketplace add SignalLayerLabs/Marginal --ref main && codex plugin add marginal@marginal
codex plugin remove marginal@marginal
Tool Enforcement, never an overstated Full Compute Enforcement claim. Repository blocking must earn a local evidence receipt and demotes automatically on drift. Universal directory review is pending; the Git marketplace works now.
Scope matters. This smoke validates the integration and supplies an early paired observation; n=3 is too small for a general performance claim. Neither lane resolved a task, so efficiency per resolved task is undefined. No deny was applied in these three agent trajectories, so the observed token difference is not evidence that a deny caused the saving.
Compute governance, measured
The next action should add evidence, not just activity.
MARGINAL evaluates whether another model call, tool call, retry, verification or reviewer is likely to add enough value to justify its cost — and accounts for the cost of making that decision.
- Open source
- Local first
- Provider neutral
- Zero mandatory dependencies
The question mark matters: enforcement is earned by evidence, not assumed by the demo.
The actual thesis
Not “models waste tokens.” Diminishing marginal value.
Coding agents can spend compute on actions whose incremental value becomes unclear. A repeated action is not automatically waste: another test may be exactly what a risky patch needs.
MARGINAL looks for a stronger pattern: same semantic action + unchanged observable state + no new evidence. That is provider-neutral and remains meaningful even when models improve.
Proof before claims
MARGINAL must earn its own compute.
Matched OFF/ON evaluation is the product test. Saving agent tokens is insufficient if governance overhead, quality loss or false stops erase the benefit.
Verified quality
Same verifier, predefined non-inferiority margin, regressions and recoveries exposed.
Gross savings
Agent workload change before counting MARGINAL's own cost.
Net savings
Effective tokens, USD and latency after the governance tax. This is the claim surface.
Repeated calls
Shows whether behavior changed or the system simply moved cost somewhere else.
False stops
Explicitly reviewed deny recommendations that would have blocked helpful work.
Uncertainty
Repeated runs and bootstrap intervals separate a stable effect from one lucky trajectory.
Graceful irrelevance
What if GPT-5.7 — or any future model — is already efficient?
Then MARGINAL should get out of the way. The evaluator now distinguishes gross from net savings and can classify a configuration as pass_through when the governor does not demonstrate enough net value.
A negative or neutral benchmark is a valid result. The project should not manufacture a reason to intervene simply to justify its own existence.
New control
State-aware diminishing returns, without a GPT-specific patch.
The detector is opt-in. Missing state fails open. Changed state or new evidence resets the pressure. That keeps the mechanism conservative while real agent telemetry is collected.
Exact duplicate prevention still exists. Diminishing-return control adds a semantic layer for allowed retries and verification loops where fingerprints can legitimately differ.
Read the evidence model →Community pressure test
Feedback changes the product only when the reasoning survives review.
Show OFF vs ON benchmarks
Matched runs are required for performance claims. The 10-task canary remains engineering validation, not marketing evidence.
Future models may reduce the benefit
Converted into Graceful Irrelevance: MARGINAL must demonstrate positive net value for the workload in front of it.
“Less slop” on the website
The technical concepts stay, but the landing page now starts with a concrete failure mode and proof standard before architecture theory.
Providers intentionally preserve waste
There is no evidence needed or offered for that claim. MARGINAL is justified by user-controlled economics and observability, not provider motive speculation.
Benchmark discipline
One benchmark is a surface, not ground truth.
The requested SWE-bench Pro comparison can be useful because it is recognizable and gives the community a common reference. It should still be versioned, audited and accompanied by task-quality notes.
MARGINAL-specific evaluation should also target the behavior the product claims to change: redundant same-state actions, repeated verification, overhead, false stops and verified outcomes.
Measured milestone
Codex Reference Integration
The auditable adapter and first matched smoke are implemented. The next evidence gate is a preregistered repeated canary large enough to estimate trajectory variance.
Build the evidence first