First measured Codex integration

Measured OFF vs ON. Correctness first.

Exploratory 3-task smoke, one paired run per task. Codex CLI 0.147.0 with GPT-5.6-sol, identical prompts and limits, verified in official SWE-bench Lite task environments on Modal.

Native Codex plugin · Shadow Mode by defaultEarned Enforcement, one-command install.
codex plugin marketplace add SignalLayerLabs/Marginal --ref main && codex plugin add marginal@marginal codex plugin remove marginal@marginal

Tool Enforcement, never an overstated Full Compute Enforcement claim. Repository blocking must earn a local evidence receipt and demotes automatically on drift. Universal directory review is pending; the Git marketplace works now.

Verified quality0/3 → 0/3OFF → ON resolved
Effective tokens24.93% fewer1,098,747 → 824,839
Tool calls3.03% fewer33 → 32
Governance latency7.06 s0 external tokens · $0
Decisionpass_throughnot support-eligible

Scope matters. This smoke validates the integration and supplies an early paired observation; n=3 is too small for a general performance claim. Neither lane resolved a task, so efficiency per resolved task is undefined. No deny was applied in these three agent trajectories, so the observed token difference is not evidence that a deny caused the saving.

Compute governance, measured

The next action should add evidence, not just activity.

MARGINAL evaluates whether another model call, tool call, retry, verification or reviewer is likely to add enough value to justify its cost — and accounts for the cost of making that decision.

  • Open source
  • Local first
  • Provider neutral
  • Zero mandatory dependencies
Illustrative traceNot a benchmark
01
Read README.mdnew information acquired
RUN
02
Verify README.mdnew evidence acquired
RUN
03
Verify README.mdsame state · no new evidence
DISCOUNT
04
Verify README.mddiminishing return threshold reached
STOP?

The question mark matters: enforcement is earned by evidence, not assumed by the demo.

The actual thesis

Not “models waste tokens.” Diminishing marginal value.

Coding agents can spend compute on actions whose incremental value becomes unclear. A repeated action is not automatically waste: another test may be exactly what a risky patch needs.

MARGINAL looks for a stronger pattern: same semantic action + unchanged observable state + no new evidence. That is provider-neutral and remains meaningful even when models improve.

Proof before claims

MARGINAL must earn its own compute.

Matched OFF/ON evaluation is the product test. Saving agent tokens is insufficient if governance overhead, quality loss or false stops erase the benefit.

01

Verified quality

Same verifier, predefined non-inferiority margin, regressions and recoveries exposed.

02

Gross savings

Agent workload change before counting MARGINAL's own cost.

03

Net savings

Effective tokens, USD and latency after the governance tax. This is the claim surface.

04

Repeated calls

Shows whether behavior changed or the system simply moved cost somewhere else.

05

False stops

Explicitly reviewed deny recommendations that would have blocked helpful work.

06

Uncertainty

Repeated runs and bootstrap intervals separate a stable effect from one lucky trajectory.

Primary target Minimize effective compute per verified successful task, subject to quality and false-stop constraints.

Graceful irrelevance

What if GPT-5.7 — or any future model — is already efficient?

Then MARGINAL should get out of the way. The evaluator now distinguishes gross from net savings and can classify a configuration as pass_through when the governor does not demonstrate enough net value.

$ marginal public-eval baseline.jsonl marginal.jsonlintervention.status: pass_throughNo positive net intervention value demonstrated.

A negative or neutral benchmark is a valid result. The project should not manufacture a reason to intervene simply to justify its own existence.

New control

State-aware diminishing returns, without a GPT-specific patch.

Same semantic keye.g. verify the same artifact for the same purpose
Same state hashthe underlying workspace has not changed
No new evidencethe previous pass did not change the evidence state
Repeat count risesexpected gain decays until a configured stop threshold

The detector is opt-in. Missing state fails open. Changed state or new evidence resets the pressure. That keeps the mechanism conservative while real agent telemetry is collected.

Exact duplicate prevention still exists. Diminishing-return control adds a semantic layer for allowed retries and verification loops where fingerprints can legitimately differ.

Read the evidence model →

Community pressure test

Feedback changes the product only when the reasoning survives review.

Accepted

Show OFF vs ON benchmarks

Matched runs are required for performance claims. The 10-task canary remains engineering validation, not marketing evidence.

Accepted

Future models may reduce the benefit

Converted into Graceful Irrelevance: MARGINAL must demonstrate positive net value for the workload in front of it.

Partially accepted

“Less slop” on the website

The technical concepts stay, but the landing page now starts with a concrete failure mode and proof standard before architecture theory.

Rejected

Providers intentionally preserve waste

There is no evidence needed or offered for that claim. MARGINAL is justified by user-controlled economics and observability, not provider motive speculation.

Benchmark discipline

One benchmark is a surface, not ground truth.

The requested SWE-bench Pro comparison can be useful because it is recognizable and gives the community a common reference. It should still be versioned, audited and accompanied by task-quality notes.

MARGINAL-specific evaluation should also target the behavior the product claims to change: redundant same-state actions, repeated verification, overhead, false stops and verified outcomes.

Same modelSame promptSame toolsSame limitsSame task orderSame verifierRaw paired JSONLPreregistered gates

Measured milestone

Codex Reference Integration

The auditable adapter and first matched smoke are implemented. The next evidence gate is a preregistered repeated canary large enough to estimate trajectory variance.

v0.2Learning Loop Foundationuniversal protocol · ledger · privacy · replay
HardeningNet-value evidence layergovernance tax · false stops · diminishing returns
v0.3 foundationCodex integration smokepinned runtime · telemetry · matched OFF/ON · Modal verification

Build the evidence first

Observe. Measure. Let intervention earn enforcement.