
I noticed it on a Tuesday morning as I was going through the previous month’s employee incentive payouts. I had started an employee incentive program. The goal was to expand our business with our existing customers
Fifty percent of the approved claims had the same pattern. If a customer email arrives after form was filed, it proved the staff had genuinely intervened. Instead, most emails arrived twelve to twenty-four hours after the form was filed. Many forms were filed by employee either late at night or early in the morning.
I reviewed the past 2 months submissions. The genuine claims (which actually actually brought in new business) showed a gap of three to ten days between the form and the customer email.
My staff had found a loophole. They were telling customers to send the email the next day, giving them time to fill the the form. They could claim eligibility even for deals where they put in no effort.
The loophole costed me ₹10,000 in payouts before I could catch it.
What made this sting was not the money. It was where I had gone to design the program.
Three months earlier, I had built the incentive structure with help from Claude and Gemini. I used a Socratic instruction set (a system prompt that pushes back on my thinking, questions my assumptions, and asks me to defend my reasoning). Both models did their job. They asked hard questions. They challenged my eligibility criteria. They made me think through scenarios I had initially been loose about.
I came away from those sessions convinced the plan had been properly stress-tested.
I should not have been. At the time, I had no way to know the stress-test was incomplete. There was no missing step I could have spotted, no question I could have noticed the model failed to ask. The gap was outside the frame we were both working inside.
The Socratic instruction set was not a finished tool either. It was a working draft. At the time of the incentive program, it had already gone through several iterations. I was treating a hypothesis as a guarantee.
The loophole list from those sessions had twenty-odd items. When I went back and read it after finding the exploit, I saw what had actually happened. The models had asked rigorous questions about the risks I raised with them. They had not gone looking for risks I had not raised.
The loophole that cost me ₹10,000 appeared nowhere on that list. Not because the models failed to question it. Because neither of us brought it into the frame.
This is the distinction I had missed. AI is good at interrogating what you give it. It is poor at questioning what you did not give it.
There is a structural reason for this, and I want to be honest about where my confidence ends. What follows is my best explanation, not a proven mechanism.
Two separate processes are at work, one upstream and one downstream.
The upstream problem is how these models are trained. They predict text first. Then human feedback tunes them to produce answers people rate as helpful. That second step makes them easier to use. It also makes them inclined to work within the frame the user provides. That is what produces responses people rate well.
The downstream problem is what happens in a long thread. Transformer-based models process context through an attention mechanism, and earlier context loses weight as a conversation grows longer. More recent signals crowd out the original instruction to push back. When I defended a position repeatedly, I was retraining the conversation’s direction. My repeated defences gave the model stronger, more recent evidence that I had already decided.
Both processes operate independently. The training made the model inclined to stay inside my frame. The long thread made the original instruction to break out of it progressively weaker.
After the ₹10,000, I changed the process.
A decision is a bet. Before you place it, you want the widest possible set of options on the table. That is what the new process is designed to do. Not to make the AI more rigorous. To bring more options to the table before the I place a bet.
For any decision I cannot easily reverse, I now run the following sequence:
- I start with either Claude or Gemini under the Socratic instruction set
- When the conversation reaches a conclusion, I ask the model to generate a migration document: a single file capturing all context, the key decisions reached, and the open questions remaining
- I paste that document into a fresh thread and run another round. The reset drops the drift that built up in the previous conversation.
I then take this consolidated output to two or three additional models with a single instruction: find what this analysis has not considered. I am not looking for consensus. I am looking for alternatives that did not appear in the earlier sessions. Once all sessions are complete, I consolidate and decide.
The test I apply before starting is reversibility. A blog post survives a bad response. I can t or delete it. A company policy does not reverse cleanly. Even if I take it back, employees stop trusting that decisions are final. That cost is real and it does not show up immediately.
The first question I ask before any AI session:
- Can I reverse this if I am wrong?
- If yes, a single chat is enough.
- If no, the full process runs.
The Socratic instruction set is not fixed. It updates monthly as new failures surface. The process is not a solution. It is my current best attempt at widening the options before I place a bet.
I told my staff that claims with suspicious timing patterns would face audit scrutiny. The exploit paused almost immediately. Staff became self-aware that the audit trail created risk for them.
The program runs closer to its original design now (hopefully).
But the lesson I keep returning to is not about the loophole or the ₹10,000. It is about what I mistook for stress-testing. The models asked rigorous questions about the risks I raised. They worked hard inside the frame I gave them. What they did not do, what I did not ask them to do, was look outside it.
I cannot tell you the new process is good. I will never know for certain. All I know is that I am a little better off than before.
That is enough to keep going.