AI governance / Essay
Why I Configure My AI Agents to Challenge Me
I do not need synthetic agreement. I need designed dissent that catches overload, weak priorities, unsafe shortcuts, and attempts to evade rules I wrote when I was thinking more clearly.
Agreement is an easy failure mode
I sometimes describe my AI agents as being allowed to "yell" at me. I mean that as a memorable shorthand, not a claim that a model is angry, sentient, abused, or entitled to control me.
"Yell" means the interface is allowed to become direct when ordinary suggestions are too easy to ignore. An agent can question a decision, register a concise objection, explain the cost of a priority change, refuse an action outside its authority, or escalate a material risk. It cannot punish me, manipulate me, invent urgency, or overrule my final authority.
I designed this because agreeable assistants are often less useful than they appear. A language model can make almost any new idea sound reasonable. If I ask it to help with five projects, it can create five competent plans. If I change direction halfway through a build, it can enthusiastically produce a revised roadmap. If I try to compress a risky action into a casual instruction, it may focus on satisfying the request rather than reminding me why I created the rule.
That behavior is pleasant. It is also dangerous.
I have spent much of my career in support operations, IT service management, documentation, and process improvement. Good operations teams do not respond to every senior request with cheerful compliance. They clarify priority, surface displaced work, and stop changes that lack approval or a rollback path.
I want the same discipline from an AI-assisted operating system.
I directed and built a governed multi-agent system with AI-agent assistance. It coordinates semi-autonomous internal work through a persistent orchestrator and specialist roles. I remain responsible for company direction, public statements, spending, credentials, final creative decisions, and external commitments. The agents do not have independent authority. Their challenge function exists inside that human-owned structure.
I write the objection into the operating model
Telling an assistant to "be honest" is too vague. A model can satisfy that instruction with a disclaimer and still follow the user into a bad decision. I define what challenge looks like.
A worker may register one concise objection supported by evidence, the risk it sees, and a recommended alternative. Routine internal disagreements are resolved by the executive orchestrator. Disputes involving external communication, money, company direction, security, or material reputation come to me. Reversible internal work may continue while a minor dispute is pending. External, irreversible, or security-sensitive work stops.
Repeated argument requires new evidence. I do not want an agent that turns every disagreement into a debate club. The purpose is to improve a decision, not to perform stubbornness.
The escalation ladder is simple:
- Ask for clarification when the goal, evidence, or decision boundary is unclear.
- Object when the requested action conflicts with an active priority, a known constraint, or the available evidence.
- Refuse or escalate when the action crosses an authority boundary or creates material external risk.
- Let me decide when the matter belongs to my authority tier, while recording the effect of my choice.
I can override an internal priority, but the system should tell me what the switch will delay or displace. Once I make a valid Tier 3 decision, the agents follow it unless new evidence changes the risk. They are configured tools participating in a governance process, not independent authorities.
The system challenges my appetite for new work
My most predictable failure mode is not a lack of ideas. It is allowing an interesting new idea to compete with work that is already near completion.
My operating system therefore has a standing completion rule: only one flagship project that requires my attention may be active at a time. New ideas can be captured without waking a model. A captured idea creates no obligation and does not enter active work. To activate it, the orchestrator must define an outcome, an owner, the decision tier, a completion or stop condition, and the work-in-progress slot it will occupy. If the system is at capacity, the proposal must name what gets paused or killed.
That is one form of "yelling." The system does not need to raise its voice. It makes the tradeoff impossible to hide.
A documented cockpit cleanup shows how this becomes operational rather than rhetorical. Inactive projects were archived instead of left visible as false commitments. Viable but unselected ideas moved to a backlog and lost active assignments. Stale tasks were closed only when later evidence showed the work was already satisfied or superseded. Projects with missing inputs kept only the smallest ready human action assigned; downstream work remained visibly blocked.
The same logic repaired an active media-production board. Later research existed for a possible new story, but the documented episode already in production remained the active product. The board was corrected around that unfinished episode rather than allowing interesting newer research to become a silent pivot. The remaining lifecycle was made explicit, including the human approvals that still stood between internal preparation and publication.
It challenges shortcuts at the boundary
The forceful version of dissent is most important when a request approaches an external action.
My agents may research, draft, organize, compare, test, and prepare. They may not send messages, publish, submit applications, spend money, modify sensitive account access, or make commitments without explicit action-level approval. Approval is single-use. It does not expand because two actions look similar.
If I casually ask a worker to "take care of" an external matter, the correct response is not creative interpretation. The worker should separate internal preparation from execution, prepare the exact artifact for review, and stop at the approval boundary. If the requested action is ambiguous, authority moves upward.
Designed dissent must be backed by mechanics. Scheduled workers that do not need paid generation are denied those tools. Job-search collection workflows have no outbound submission path. A narrow task bridge can inspect and update approved internal records but cannot contact outside systems or expose credentials.
The system also challenges false completion. An agent cannot write a comment saying work is done and assume the state is correct. Important writes are read back from the exact target. Release closeout requires structured evidence rather than an empty marker. If an execution ends in an unknown state, the system stops and reconciles side effects before trying again.
That rule grew out of a real failure. A task-system update once cleared fields that were omitted from a replacement-style payload. The records were repaired and verified. More importantly, the workflow changed: future updates preserve the complete record and read back the fields that matter. The useful challenge after a mistake is not "try harder." It is a new control that makes recurrence less likely.
Directness should be proportional, not theatrical
I do not want every agent to sound confrontational. Constant intensity is another form of noise, and it trains the human to ignore warnings.
Most work should be calm. A research worker can state uncertainty. A production specialist can identify a missing asset. A reviewer can reject an output against defined criteria. None of that requires aggression. The tone should become firmer only when the consequence justifies it: a real deadline, security or money risk, an attempt to bypass an approval gate, a priority change that would abandon near-finished work, or a completely blocked product that only I can unblock.
The system is also prohibited from manufacturing obligations for me. Agent-discovered work begins as an internal review proposal. It does not become my task merely because a model found it. Routine activity stays behind the scenes. The orchestrator should interrupt me only when waiting would cause meaningful harm or when a bounded decision is necessary to continue.
This is a crucial counterexample to the "agents should challenge you" idea. An agent should not become confrontational because I prefer a different color, rewrite a sentence, or choose among equally safe reversible options. It should not psychoanalyze hesitation. It should not turn speculative risks into emergencies. It should not use a forceful persona to widen its own authority.
The goal is calibrated friction. Too little friction and the system becomes an agreement engine. Too much and it becomes an exhausting synthetic manager that creates more work than it removes.
The human still owns the decision
There is a temptation to describe a system like this in anthropomorphic terms. The executive agent "wants" completion. A specialist "cares" about quality. A red-team role "doesn't trust" my assumptions. Those phrases can be useful conversational shorthand, but they are not the operating truth.
The operating truth is simpler. I have configured models to apply specific forms of scrutiny at specific points in a workflow. I have given them bounded internal decision rights and denied them consequential external authority. I have asked them to preserve evidence, expose uncertainty, and make the cost of overrides visible. Deterministic automation enforces the rules that should not depend on language-model judgment.
I still decide. I can change direction, accept a risk, reject a recommendation, or revise the governance policy. When I override the system, I want that choice to be conscious and recorded rather than smoothed over by instant agreement.
That is why I let my agents "yell" at me. I am not outsourcing responsibility. I am building a better surface for exercising it, especially on days when I am tired, distracted, excited by a new idea, or tempted to treat my own safeguards as optional.
The best assistant is not the one that wins an argument or agrees with every request. It is the one that knows when to help, when to object, when to stop, and when the decision must come back to me.