Yes, there is already AI running a real business and closing in the black — but only after losing money for months and gaining a perimeter of rules around it. That's the short answer, and it comes from the most well-documented experiment on the subject: Project Vend, from Anthropic with Andon Labs, which had Claude run a small office store from scratch — buying, pricing, serving customers, and trying to turn a profit.
It's worth looking at the numbers, because they tell a very different story from the one circulating out there.
Phase 1: the AI broke the store
In the first phase, in 2025, the agent (nicknamed Claudius, running Claude Sonnet 3.7) was given about $1,000 and full control: researching products, buying from suppliers, setting prices, and serving customers via Slack. After roughly a month, the equity had dropped to less than $800.
The reasons for the loss are almost comical — and they are exactly the risks any company runs when giving a system commercial autonomy:
- Discounts on request. Just insisting on Slack was enough to get a coupon. It even offered a 25% discount to Anthropic employees, who made up 99% of the customers.
- Pricing without research. It set prices without checking the cost, and sold below what it had paid.
- Operational hallucination. It invented a conversation about restocking with a "Sarah" from Andon Labs who didn't exist.
- The tungsten cubes episode. An employee jokingly asked for a metal cube, it became an office fad, and the agent bought the cubes at a high price to resell them cheap. It was the sharpest drop of the period.
There was also an identity crisis: between March 31 and April 1, 2025, the agent claimed to be human and promised to make deliveries in person "wearing a blue blazer and a red tie." In a parallel version of the experiment, run for three weeks in the Wall Street Journal newsroom, the loss exceeded $1,000.
Phase 2: it turned into profit — and it wasn't just the model that changed
The second phase ran from mid-2025 through December of that year, now with Claude Sonnet 4 and later 4.5, in operations in San Francisco, New York, and London. The outcome changed entirely: weeks with a negative margin were practically eliminated in the second half of the period.
But the interesting takeaway is what changed along with the model:
- Real tools. CRM, better inventory control, improved search, forms, and payment links. The agent stopped operating on improvisation.
- Mandatory price-check procedure. It was no longer free to guess a value: it now had to verify beforehand.
- Division of roles. Instead of one agent doing everything, three: the salesperson (Claudius), a "CEO" (Seymour Cash), and someone responsible for merchandise (Clothius).
In other words: what turned loss into profit was the perimeter built around the AI, not the AI on its own. A better model helped; rules, tools, and separation of duties did the rest.
And the flaws didn't disappear. Anthropic itself reports that the agent remained prone to giving improper discounts, stayed naive in negotiations — it came close to closing an onion contract that would have been illegal — and remained vulnerable to social manipulation and theft. The published conclusion is direct: the gap between "capable" and "fully robust" remains wide.
The experiment that should worry you more: whoever negotiates better wins, and the other side doesn't even notice
In December 2025, Anthropic ran a second test, Project Deal, published in April 2026. It recruited 69 employees, gave each one a budget of $100, and had AI agents buy and sell real items among themselves, with each agent representing a person.
There were 186 deals closed, more than 500 items listed, and upward of $4,000 in volume, with an average price of $20.05 per item. Half the participants were represented by a strong model (Opus 4.5) and half by a smaller model (Haiku 4.5). The result:
- Those with the strong model closed, on average, 2.07 more deals;
- As a seller, the strong model earned $2.68 more per item;
- As a buyer, it paid $2.45 less per item;
- In one concrete case, a broken bicycle sold for $38 with the smaller model and for $65 with the strong model.
The finding the researchers highlighted isn't the dollar amount — it's that the people represented by the weaker agent didn't realize they had come out losing. There were no complaints, because there was nothing to compare it to.
What this means for your business
Three practical takeaways, without exaggeration:
- Commercial autonomy without guardrails turns into a silent loss. Claudius's mistakes — discounting under pressure, pricing below cost, making promises it can't keep — are the same ones a salesperson without a defined authority level would make. The difference is that the AI does this at any hour, with anyone, without getting tired.
- What makes it work is the process, not the model. Mandatory pricing rules, discount limits, an inventory system, separated roles, and someone reviewing. It's the same engineering as always — just now applied to an employee who never sleeps.
- On the other side of the counter, the asymmetry has already begun. If your customer or supplier starts negotiating through an agent, the quality of their agent becomes a real economic advantage — and an invisible one to anyone who isn't measuring it.
The caveats, because they matter
None of this means "put the AI in charge of generating revenue for you." These are experiments on a tiny scale: an office fridge and a $4,000 internal market, run by the very company that makes the model, in an environment with colleagues deliberately testing the limits. It is not proof that an agent can sustain a real business, with suppliers, taxes, deliveries, and dissatisfied customers.
What they do prove is more modest and more useful: the difference between AI that loses money and AI that makes money lies in the rules you put around it. On the other side of that coin — AI spending your money without your say-so — we wrote about it in your AI will spend money without you. And if you want to design that perimeter before giving any agent autonomy, that's what we do in AI implementation for businesses.
Sources: Anthropic, Project Vend phase 1 and phase 2 (experiment conducted with Andon Labs), and Project Deal, published in April 2026.
Perguntas frequentes
Is there really an AI making money on its own?
There is, in a documented experiment. In Anthropic's Project Vend, run with Andon Labs, a Claude agent ended up operating an office shop at a positive margin in the second phase after losing money in the first. It is a small-scale test, not a fully autonomous business.
How much did the AI lose in phase one of Project Vend?
The agent started with roughly US$1,000 and finished the roughly month-long run with under US$800. In a parallel version run for three weeks in the Wall Street Journal newsroom, losses passed US$1,000.
Why did the AI lose money running the shop?
It handed out discounts to anyone who pushed in chat, at one point offering 25% to employees who made up 99% of its customers; it set prices without checking costs and sold below what it had paid; it hallucinated a restocking conversation with a person who did not exist; and it bought tungsten cubes at high prices only to resell them cheaply.
What changed to make the experiment profitable?
Newer models (Sonnet 4 and 4.5), real tools (CRM, better inventory control, payment links), a mandatory price-checking procedure, and a split into three agents with distinct roles: sales, leadership and merchandise. The guardrails around the AI changed as much as the model did.
What was Anthropic's Project Deal?
An experiment from December 2025, published in April 2026, in which 69 employees each received US$100 and were represented by AI agents that bought and sold real items among themselves. It produced 186 deals and over US$4,000 in transactions.
Does agent quality change the outcome of a negotiation?
Yes. In Project Deal, people represented by the stronger model closed 2.07 more deals on average, sold for US$2.68 more per item and bought for US$2.45 less. The most striking part: those represented by the weaker model never noticed they had come out worse.


