A coding agent can execute a tool successfully and still make no progress. That sounds obvious until you try to turn it into a runtime rule. A repeated file read can be waste, but it can also be verification after a write. A repeated test can be a loop, but it can also confirm a flaky failure. Even an unchanged workspace does not prove that an action was useless if the action produced new evidence.
In short
- Repetition is not the signal. The stronger signal is repeated successful work with unchanged observable state and no new evidence.
- Raw prompts do not have to be the primary governance input. A runtime can reason from derived action identity, structured outcomes, workspace state and evidence deltas.
- Unknowns should reduce confidence. Failure, ambiguous outcomes, state changes and incomplete coverage are reasons to fail open—not reasons to invent certainty.
- Detection is not enforcement. Seeing a no-progress candidate does not automatically justify blocking the next action.
Repetition is not the same as no-progress
Start with the smallest counterexample. An agent reads config.py, edits it, then reads it again. The tool family and target may be identical, yet the two reads have different jobs: the first establishes state; the second verifies a change. Any governor that simply counts duplicate calls will eventually punish useful verification.
read config.py → establish evidence
edit config.py → workspace changes
read config.py → verify the changeNow remove the write:
read config.py → success, evidence acquired
read config.py → success, same state
read config.py → success, same state, no new evidence
read config.py → another no-progress candidate?The interesting question is no longer “did this call appear before?” It is “what changed that makes another identical successful call worth running?”
A useful operational approximation:same semantic action + successful outcome + unchanged observable workspace state + no new evidence → no-progress repetition candidate
The final word matters. It is a candidate, not a verdict. The runtime still has to account for what it cannot observe.
Why not just read the prompt?
The obvious design is to inspect the conversation and ask a model whether the agent looks stuck. That can be useful in systems where semantic intent is the product. It is a poor requirement for a minimal runtime governor.
Prompts and transcripts often carry the most sensitive material in an agent session: source fragments, customer context, repository names, commands, credentials accidentally pasted into chat, or proprietary reasoning context. Making that content mandatory governance telemetry expands the trust boundary before the governor has proved that it needs the data.
There is also an architectural cost. If every decision about waste requires another model call, the control plane introduces its own latency, cost and failure mode. A governor trying to reduce unnecessary work should be able to answer narrow questions without always paying for a second inference path.
agent action
↓
model judge reads transcript
↓
judge decides whether the agent is wasting work
↓
control decisionThat architecture is not inherently wrong. It is simply stronger and more invasive than necessary for a class of no-progress cases that can be observed from effects.
Observe effects, not private thoughts
A narrower runtime model can work with four categories of evidence:
| Signal | Question | Why it matters |
|---|---|---|
| Semantic action identity | Is this meaningfully the same action and target? | Literal tool names differ across engines and wrappers. |
| Outcome | Did the action actually succeed? | Failures and unknown outcomes should not be counted as completed duplicate work. |
| Observable state | Did the workspace or relevant state change? | A changed state can make an otherwise repeated action useful again. |
| Evidence delta | Did the action produce new usable evidence? | Useful verification can occur even when the workspace itself is unchanged. |
The point is not that these four signals solve every loop. They do not. The point is that they define an auditable surface: a reviewer can inspect why the governor believed a repetition was low-value without needing the full conversation.
Semantic identity is harder than string equality
Two calls can look different and still represent the same operation. A Codex adapter might report Read(path=...) while another engine reports read_file(file_path=...). Conversely, two calls to a generic search tool might carry materially different queries and therefore different evidence value.
This is why MARGINAL keeps adapter logic separate from the provider-neutral governance core. The adapter normalizes native lifecycle events into a smaller action model; the core reasons over the normalized identity, outcome and evidence. The integration documentation explicitly requires adapters to declare capabilities and classify data before persistence rather than smuggling engine-specific assumptions into policy.
Inspect the integration contract in the repository →
State is what separates verification from loops
A fixed “three repeats means stop” rule is attractive because it is easy to explain. It is also easy to break. Consider three sequences:
| Sequence | What changed? | Governance interpretation |
|---|---|---|
| read → write → read | Workspace changed | The second read may verify the write; repetition pressure should reset. |
| test → edit → test | Implementation changed | The second test has new causal context. |
| read → read → read | Nothing observable; no new evidence | A no-progress candidate becomes more plausible. |
The same principle applies beyond files. A repeated status poll may be useful if time itself is expected to change the result. A repeated network call may carry server-side state the local governor cannot see. Those action families should not be treated as equivalent to a deterministic local read merely because the surface syntax repeats.
Unknown outcomes should weaken authority
A common failure in control systems is to turn missing data into implied success. If an integration cannot prove whether an action succeeded, the governor should not promote that event into positive evidence for blocking a future action.
known success + same state + no new evidence
→ stronger no-progress evidence
failure or unknown outcome
→ do not escalate authorityThis is visible in MARGINAL's current engine boundaries. Claude Code can report success and failure through distinct hook events, while OpenCode exposes weaker outcome evidence for many tools. The project therefore labels both integrations Observe-only and keeps unknown outcomes unknown instead of laundering them into confidence.
That leads to a separate question: when should a detector be allowed to become an enforcer?
What “without reading prompts” does—and does not—mean
It would be misleading to turn this design into a blanket privacy claim. A runtime can avoid using raw prompts and source as governance evidence while other systems around it—an IDE, shell history, agent vendor, plugin callback or custom logger—still record sensitive material.
MARGINAL's privacy model therefore classifies fields rather than declaring telemetry “safe” by default. Its strict SAFE_TELEMETRY profile removes free text and metadata, pseudonymizes selected identifiers with field-separated HMACs and generalizes timestamps. The documentation also says the uncomfortable part explicitly: pseudonymization is not anonymization.
Read the privacy profiles and limitations →
Detection and enforcement are different problems
A detector can be useful while still being wrong often enough that it should never block. That is why MARGINAL separates a recommendation from authority. New integrations begin without earned permission to interfere. A no-progress candidate can be recorded, reviewed and falsified before it becomes an enforcement rule.
OBSERVE
↓
record recommendations and outcomes
↓
review false stops and coverage
↓
EARN narrow authority where the adapter can truly intercept
↓
demote when evidence or conditions driftThis distinction also prevents capability inflation. The current Codex integration is documented as Tool Enforcement, not Full Compute Enforcement. Claude Code and OpenCode are Observe-only. A prompt instruction or advisory skill is not presented as enforced interception.
The best contribution is a counterexample
If this model is useful, it should survive hostile examples. The most valuable trace is not another obvious infinite loop. It is a case where repeated successful work looks low-value from the outside but is actually useful.
Examples worth testing include:
- a verification read whose value is not represented in workspace state;
- a tool whose server-side state changes while local state stays fixed;
- a flaky test that must be repeated to estimate reliability;
- two calls that normalize to the same action but carry materially different evidence;
- a state fingerprint that misses a meaningful change in a large monorepo.
MARGINAL has a public challenge specifically for this: find useful verification the governor could mistake for waste. A strong result can be “the governor already fails open correctly.” The objective is to tighten the falsification boundary, not manufacture a win.
Questions developers usually ask
Can an AI agent loop be detected from duplicate tool calls alone?
Not reliably. Duplicate calls are a useful symptom, but a repeated action can be legitimate verification after new state or evidence. Treat repetition as one input, not the conclusion.
Does no-progress detection require storing prompts?
No. Some no-progress cases can be evaluated from derived action identity, structured outcomes, observable state and evidence deltas. That does not mean every loop can be solved without semantic context.
Should a no-progress detector automatically block the next call?
No. Detection quality and enforcement authority are separate concerns. A conservative system can remain advisory until local evidence supports a narrow intervention boundary.
See the mechanism before installing it.
The interactive demo uses a deterministic trace to show the difference between repeated activity and observable progress. No provider telemetry is presented as benchmark evidence.
This article was developed with AI-assisted drafting and reviewed against MARGINAL's public code, documentation and evidence available on Aug 19, 2026. Product claims are deliberately limited to what those public sources support; unsupported performance claims are excluded.
