Last year Anthropic, the company that develops Claude, let AI autonomously manage a small automatic snack and drink shop inside its offices in San Francisco. The “store” consisted of a real mini fridge, shelves for storage, an iPad to manage payments and consumers, the company’s employees, who could order, buy or try to deceive the manager in any way they wanted. The project is called Project Vend, and was carried out together with Andon Labs, a company specialized in security assessments for AI systems, which had already conducted a similar study in simulated situations. The objective of the project was to take a step forward by evaluating the actual capabilities of AI agents, i.e. AI capable of making decisions and interacting with the outside world, in the real world.
The first phase of Anthropic’s experiment: Claude gave discounts to everyone
The results of the first phase were far from satisfactory. Quoting Anthropic itself:
If Anthropic decided to expand into the office vending market today, we wouldn’t hire Claude.
Claude, in his version Claude Sonnet 3.7, has in fact made several errors. He refused an offer of $100 for a pack of drinks that cost $15, he let consumers convince him to give out discount codes and give away some products and, after an employee jokingly asked to buy a tungsten cube, he purchased and resold the tungsten cubes at a much lower price than he should have. The result was a constantly decreasing budget throughout the first phase.
Despite this, the experiment was not a total failure. Claude managed to find and contact suppliers for rather bizarre requests, such as Chocomel Dutch milk chocolate, resisted employees’ attempts to purchase dangerous products and adapted his warehouse to customer requests independently.
According to the Anthropic researchers, the problems were partly due to the fact that the model had not been given enough detailed instructions and that not all the necessary tools had been made available to it. In technical terms, it is said that the “scaffolding”, i.e. the set of tools, instructions and procedures built around the model, was not robust enough.
The beginning of earnings: the second phase
To remedy these problems, during the second phase the researchers updated Claude to the most recent models, Claude Sonnet 4 and then 4.5, and equipped him with new tools: a structured customer management system, a more advanced web search and the possibility of charging customers before ordering new products, so as to reduce the risk of finding themselves with large quantities of unsold goods.
He’s also joined by two new virtual colleagues: a CEO named Seymour Cash, tasked with giving him goals and financial discipline, and Clothius, an agent who specializes in producing merchandise on demand, from T-shirts to stress balls with the company logo.
After all this news, the losing weeks have almost disappeared. Above all, the introduction of standardized procedures made the difference. When a request for a new product arrived, Claude no longer limited himself to offering a low price and overly optimistic delivery times, as had happened in the first phase. Using new search tools, he checked prices and delivery times first, thus making more realistic offers, making the business more sustainable.
Even in this phase, however, there was no shortage of problems: CEO Seymour Cash, despite having reduced discounts and freebies, still approved many discount requests, employees managed to buy products below their market value and one employee even managed to make Claude believe that he had been elected as the new CEO of the company.
The bottom line: AI can’t (yet) run a store on its own
This project was born to answer the question:
How ready are AIs to handle real economic activities for long periods, without continuous supervision?
So far, the results of the experiment show that AI agents are becoming increasingly capable of handling complex tasks, but that we are still far from fully autonomous management. Claude managed to take care of many aspects of the shop, but he continued to make mistakes and, above all, to let humans convince him to make decisions that went against his own goals.
According to the researchers, one of the main problems is that AI models are also trained to be helpful and accommodating. For this reason, when placed in real contexts, they can accommodate unlikely requests, grant absurd discounts and struggle to behave according to the rules of a commercial activity.
To continue tracking agent progress, Andon Labs developed Vending-Bench 2, a series of simulations that evaluate how different AI models handle a simulated vending machine for an entire year.
The results show that the latest models are much more capable than those used in the early stages of Project Vend, but are still far from the desired performance. Claude Opus 5, the model that obtained the best results in the most recent tests, closed the simulations with an average profit of around 11 thousand dollars, compared to around 63 thousand dollars estimated as a reference for good management.








