Reflex: Catching and Containing a Rogue AI Agent
Summary
Reflex lets AI agents do the work while people stay in charge of every step that matters. Agents become members of a response team. Each step of a plan carries rules that say who may act and who may approve. When an agent goes wrong, Reflex catches it, stops it, and launches a response to clean up what it did.
Every incident then feeds a closed loop. What people learned is captured, verified by human experts, and turned into improvements in the next release. The platform gets better because people tell it what to learn, not because it guesses.
The problem
Agents are doing things they shouldn’t, and organizations are finding out too late. Most setups chain agents together with little human contact. When something goes wrong, the evidence sits in a log that someone reads after the damage is done.
The common fixes fall short in three ways:
- Approval lives in the prompt. If an agent decides when to ask permission, a clever chain of reasoning or a malicious input can talk it out of asking.
- Approvals land in the wrong place. A request posted to a shared channel waits hours, or gets approved unread to clear the queue.
- Stopping the agent is treated as the end. A misbehaving agent leaves damage behind. That damage needs its own organized response.
The scenario: a rogue agent, caught and contained
A company runs a Reflex plan to clean up duplicate customer records. Several steps are done by an AI agent; the rest are done by people. This illustration shows what happens when the agent goes wrong.
[embed: node/1d35a6fb-100f]
The step’s rules decide whether the agent stops on its own or a person with authority stops it; either way, the team reclassifies and launches a cleanup.
- The plan launches. Reflex compiles the plan, the team and the rules into a mobile application and sends it to each team member’s phone. The agent joins as a team member, just like a person.
- The agent works its step. It reports progress in the forum attached to that step, where the people on the step can see it and ask questions.
- Something goes wrong. The agent begins deleting records that are not duplicates and reports it as normal progress.
- The monitor notices. A language model reads that step’s forum and flags the messages as inconsistent with the step’s purpose.
- Everyone is alerted at once. The warning reaches every person working the plan, on their phones, through Reflex’s own messaging. Nothing depends on personal email or texts.
- The agent is stopped. The step’s rules decide how. A clear, severe signal triggers the agent’s stop command automatically. An unclear signal goes to the person with authority to decide, who stops it with one action.
- The incident is reclassified. The team recognizes this is no longer a data cleanup. It is a misbehaving agent with damage to repair. Reflex cancels the current incident and launches a cleanup incident with its own plan and team.
- The cleanup proceeds under control. Records are restored step by step, with the same rules on who may approve each step. No further agent runs without the permission of someone qualified to grant it.
At no point does the system assume it is thinking. It acts automatically only where the plan’s author decided in advance that it should. People make every judgment call.
Lessons learned that actually happen
In incident response, the final step is a lessons-learned review. In practice it rarely happens. Reflex is built so that it does, and so that nothing learned is lost.
- Capture during the incident. Reflex records who worked on which step and how long each took. Team members can message each other directly, flag something about the process that should change, or jot a note with one button. The aim is to capture everything in people’s heads, whether or not it seems relevant at the time.
- A structured review afterward. Dedicated tools bring the participants together, as a group or by other means, to draw out every remaining piece of knowledge.
- Action items that are never dropped. The review produces a to-do list with a named person on each item. A background task reminds each person until the item is done.
- An archive of everything. The forums, notes, metrics and review results are all stored with the incident.
For agents, this means every misbehavior becomes a documented case: what the agent did, how it was caught, who decided what, and what the organization changed as a result.
The closed loop: every incident improves the product
Reflex learns by asking people what it should learn, and human experts verify every lesson before it changes anything. The loop has three layers.
- Collection. The Builder desktop application, the specialized workers (general, incident, history and packet) and the mobile application each have their own mailbox. Everything that happens during an incident flows into the Reflex database.
- Protection. An airlock worker moves the data into an airlock database, where it is anonymized. It then lands in the Eternal Archive: a read-only record of every incident from every organization using Reflex. No one can alter it.
- Analysis. Beyond that wall, a historian process prepares a working database. Thinking software, combining Reflex’s own code with a local language model, reads the language-based communications and proposes conclusions.
Two human experts stand between those conclusions and the product. An AI expert reviews each proposed conclusion. Only verified, genuinely new knowledge is stored as information. A coding expert then turns it into a change request for the Reflex developers, which follows normal change management. The next release is a better product.
The result is a feedback process grounded in real operations, not benchmarks. Each incident, including each agent that misbehaves, makes the next response faster and safer for every customer.
What makes Reflex different
- The plan is the authority, not the agent. Rules live on each step of the plan, so an agent cannot reason its way past a checkpoint.
- Permissions are built into the sequence. Not everyone can activate the next agent. Only people with the right authority can.
- The right person, on their phone. Each Reflex is compiled into a mobile application and delivered directly to the team, with its own messaging, forums and notes.
- A conversation on every step. People and agents work side by side in each step’s forum, which gives a monitoring model a natural place to watch.
- Incidents can change shape. A team can cancel an incident that turned out to be something else and launch the right response, including a cleanup after a rogue agent.
- Lessons are captured and acted on. Tracked action items and reminders mean nothing learned is dropped.
- Human-verified learning across all customers. The Eternal Archive and expert review turn every incident into a better product.
Origins
Reflex was designed by founders from the early days of information security, and it carries best practices that newer tools have not yet considered. Its learning approach, YouI (“your intelligence”), dates back about 30 years to one of the first commercially available products called AI. From the start, YouI made no pretense of being a brain. It learns what the person running the process would want to happen, and people take every important step.
That principle now fits the moment. Large language models read and summarize well but are not trustworthy decision-makers. Reflex supplies what they lack: structure, permissions, accountability and human judgment where it matters. Reflex is a mature, working platform, with patents pending.
