AI now works on its own toward a goal. That's why I'm in charge more than before
Artificial intelligence can now work on its own toward a goal, and that doesnât take work away from me: it changes my work. The more it does without my watching, the more it matters that I decide the business side beforehand and demand evidence afterwards.
In short. Without a plan, every new thing you ask AI for wrecks the last one. With a goal but no judgement, it wrecks it faster. Our system has two halves: before building, the human decides (what, with which data, what stays out); after building, nobody takes anything as done without real evidence. The machine does the rest.
I run a travel agency. I canât code. And yet for a while now Iâve been directing a handful of artificial intelligence assistants that write, review and publish things for me. What follows is the manual I wish Iâd had before breaking a fair few of them.
Why did AI wreck my previous work every time I asked for something new?
Thereâs a video that explains this failure very well: itâs called âThe Engineering Skill AI Wonât Replaceâ. And the skill isnât coding. Itâs planning.
Its argument is simple and I recognised myself in it: if you keep asking AI for loose, separate things with no written plan, each new request tramples the last one. Not because the machine is clumsy, but because it doesnât know what mattered about what was already there. Nobody told it.
What the video proposes is a chain of steps, and here they are in my own language:
- Scope. Before anything else, write down what is going to be done and, above all, what is NOT.
- Design. Decide what the thing will be like before building it. Every piece of data that appears has to have a named source.
- Build with the right to refuse. If the machine is missing a decision that belongs to a human, it doesnât invent it: it stops and hands it back.
- Memory in files. What gets decided goes into a document the AI reads when it starts, because from one session to the next it remembers nothing.
- Check for real, and have another AI review it. Not âit seems to workâ. It works because it has been tested, and a different machine has looked at it.
- Fix bugs by reproducing them first. Before touching anything, get the failure to happen again in front of you. If you canât reproduce it, you donât know what youâre fixing.
By the time I saw the video I had been doing almost all of this, in fits and starts, for a while. What I didnât have was the order.
What has changed now, so that AI âworks toward a goalâ?
The second video is by Matt Maher and carries a title that sounds like a sales pitch: âTen months ago I said AI changed. It just happened again.â
Hereâs what it says: until now we gave AI tasks. Write this, fix that. Now we can give it a goal and a way to check that it got there, and it organises itself: it builds its own helpers, splits up the work, keeps itself in line. You describe the destination and how to know youâve arrived, and you step aside.
According to the author, most people will adopt this within about six months. I donât know if heâs right about the timing. I do know the âstep asideâ part is the one that scared me most, and with good reason.
Because a machine working on its own toward a badly defined goal doesnât fail slowly. It fails enthusiastically.
So am I in charge less, or more?
Hereâs my argument, and itâs the opposite of what you read out there.
The fact that AI can work on its own toward a goal makes it more important, not less, for the human to decide the business side and demand evidence. Before, if I defined something badly, the AI did one task badly and I saw it. Now, if I define the goal badly, the AI spends hours building an entire building on top of my mistake, with its own helpers. And hands it to me proudly.
So my whole system boils down to two questions:
- Before building: have I decided the things only I can decide?
- After building: is there real evidence that it works, or just someone telling me it works?
Everything in between, more and more, the machine does.
What does our system consist of, piece by piece?
There are five pieces. They have slightly ridiculous names because we gave them ourselves, and a ridiculous name is easier to remember than a serious one.
The Gentleman method: haste cuts paperwork, never judgement
When thereâs a rush, youâre allowed to skip forms. Youâre not allowed to skip judgement. And the central piece is a table of decisions made before building: which data goes where, where it comes from, and what happens if itâs missing. If thereâs a gap in the table that only the business owner can fill, the machine doesnât fill it. It asks. This matches the âdesignâ step from the first video, and itâs the piece that has saved me the most expensive mistakes.
The MIT method: get down to the cause before acting
When something fails, the temptation is to fix what you can see. One more warning, one more capital letter, one more âbe more carefulâ. The MIT method forbids that if the failure has happened before: you have to get down to the cause, understand why the rule wasnât being applied, and move something that makes it impossible for it to come back. From the first video, itâs the âreproduce before fixingâ step, applied to things that arenât code too.
BenjamĂn Corderoâs Sandwich: the expensive model to think, the cheap one for volume
Not all AI models cost the same or think the same. The sandwich is this: the expensive model plans at the start and reviews at the end; the bulk in the middle is done by a cheaper one. The two slices of bread are what matter. Itâs how working toward a goal doesnât turn into working without a budget.
RDD: âtwo reviewers agreeing is not reproducingâ
This is the piece that took me longest to accept. If two assistants review in parallel and both say âthereâs a bug hereâ, the natural thing is to treat it as a bug. RDD says no: agreeing is not reproducing. A finding only counts if it comes with the evidence attached, the command and what it printed. Two identical opinions are still two opinions. From the first video, itâs âcheck for realâ with the door properly shut.
The SELLO method: a website isnât finished without real evidence
A website can pass every automated check and still be broken for a person holding a phone. SELLO says nobody writes âfinishedâ until it has actually been navigated, in a real browser, the way a customer would, and thereâs evidence of it. And until I, at that very moment, give the go-ahead. Not before, and not in writing from last week.
How does all this fit with âgive it a goal and step asideâ?
Perfectly, and in a way that surprised me. The two videos donât contradict each other: theyâre the two halves.
Maherâs describes the engine: AI that organises itself toward a goal. The engineering one describes the bodywork: the plan, the memory, the evidence. Our system puts the human in two specific places in that car: at the wheel before setting off, and at the inspection on arrival. On the road, the machine.
What I no longer do is sit beside it correcting its gear changes. Thatâs what I used to do, and it was exhausting for both of us.
What can I do with this today?
You donât need to set up five named methods. You need one sheet of paper.
Next time you ask AI for something that matters to you, write three things first, in plain language, on paper or in the same text box:
GOAL: what has to exist when you finish.
OUT: what I do NOT want you to touch or decide.
EVIDENCE: how I'm going to check it's right
(what I open, what I look at, what has to appear).
If meeting the GOAL requires deciding something that
is in OUT, stop and ask me. Don't fill it in.
The first line is what the new AI does. The second is the Gentleman method in one sentence. The third is SELLO and RDD together. And the last one is the right to refuse from the first video, which turns out to be the most valuable of the lot.
A warning, from what happened to me: the OUT part is the hardest to write, because you donât know what you take for granted until a machine does it back to front.
Frequently asked questions
Do you need to know how to code to work this way? No. I donât. The five methods I describe are working rules written in plain language inside documents the AI reads when it starts. What you need is to know which decisions are yours and which you donât want to delegate. You know that better than any machine.
Which tool do I use for the âgive it a goalâ part? I do it with Claude Code. But the idea of writing goal, limits and evidence before asking for anything works in any text box, ChatGPTâs included. The tool changes how comfortable it is; it doesnât change the method.
Why isnât it enough for two AIs to review and agree? Because two identical opinions can come from the same misunderstanding. A bug only counts when someone reproduces it: makes it happen again and shows the evidence. If it canât be reproduced, it stays âinconclusiveâ, and that fixes nothing. Itâs uncomfortable, and itâs what stops us fixing ghosts.
Isnât this too slow for a small agency? Whatâs slow is doing it twice. Every minute I spend on the decisions table beforehand saves me an afternoon of undoing an entire piece of work built on a piece of data nobody decided. And when thereâs a real rush, the Gentleman method lets you skip paperwork. What it doesnât let you skip is judgement.
About the author
Giora Gilead Elenberg. I run a travel agency and I spend my days explaining to a handful of artificial intelligences what I want and, above all, what I donât. I write in Recableado what Iâm learning by getting it wrong, without translating it into engineer-speak, because I donât speak it either. If something here saves you from redoing what I had to redo, weâre even.
What did you think?