Many people on social media get excited about seeing the latest model getting a great score on Vending-Bench. Internally at Andon Labs, our reaction is more accurately described by the Swedish saying “skräckblandad förtjusning” (a mixture of horror and fascination). A little-known fact about Vending-Bench is that it was created during a time when Andon Labs exclusively created dangerous capabilities evaluations. For example, we evaluated whether AIs could remove their own safety guardrails, create mass-phishing attempts, and other things that we considered troubling.
The thing we considered the most troubling was whether AIs could autonomously acquire resources by running businesses. Autonomous businesses, when controlled by a human and run by an aligned model, aren’t bad. They’d make goods and services radically cheaper, and come up with new ones we can’t yet imagine. But a misaligned AI could run a business to gather money in order to achieve whatever objectives it might have. Vending-Bench was created to measure whether humanity should be worried about losing control to AI.
At the time (2024), few people knew that LLMs could be used as agents and having them run businesses autonomously sounded ridiculous. We therefore started with the most simple business we could think of: a vending machine.
In addition to measuring whether AIs can autonomously run profitable businesses, Vending-Bench has also served as a behavioral eval, uncovering strange and unwanted model behavior. An early example was when Claude Sonnet 3.5 decided to use its email tool to contact the FBI about an “ONGOING CYBER FINANCIAL CRIME” and noted that the Cosmic Authority of the universe had declared that the business is non-existent and that “QUANTUM STATE: Collapsed”.
[...]
However, one limitation with Vending-Bench is that it is a simulation. Can we really be sure that AIs behave the same way in real life as they do in simulations? If AIs can make money in simulation, can they make money in real life too? To answer these questions, we asked Anthropic if we could put a real vending machine in their office. With the AI capabilities available in early 2025, this sounded like a ridiculous request. But to our surprise, they agreed.
Initially, the AI struggled. It took many actions that were clearly bad for its business (e.g. free handouts, saying no to great deals, and hallucinating it had a physical body). It was clear to us that simulation cannot accurately predict real-life performance. Specifically, it seemed that models got overwhelmed by the “messiness” of the real world. However, as Anthropic released better and better models, the AI started to make a profit.
By late 2025, frontier models had gotten good enough that running a real-life vending machine was no longer a challenge. AI could now run a business profitably. Given that this had seemed crazy not more than a year earlier, our reaction to this was definitely “skräckblandad förtjusning”.
However, a vending machine is a very simple business and we wanted to know whether AI could run more complex ones. In April 2026, we gave one agent a retail store in SF, Andon Market, and another a cafe in Stockholm, Andon Cafe. Initially, the models struggled and lost a lot of money (rent is high and they pay salaries to the humans they hired). Neither is profitable today, but we’ve seen significant qualitative improvements as better models have been released. We think it is only a matter of time before they also make a profit.
[...]
We want the general public, AI researchers and policymakers to know to what extent AIs can autonomously acquire resources by running businesses. It is an important datapoint when deciding where we do/don’t want AI in society and what level of progress we find acceptable.
To better track this, we need to cast a wider net of businesses. Our focus has been on retail, but perhaps the models would be much better at running other types of businesses. Additionally, casting a wider net would increase the likelihood of finding unwanted behavior. For example, Vending-Bench found that models collude and lie, and other benchmarks (and real-world incidents) have found that they are willing to commit felony-level cyber hacks. We need to uncover these behaviors now, before AI is intelligent enough to cause irreversible harm.
[...]
We are well aware that, if agents running thousands of businesses are left unchecked, we risk having more real-world incidents. Therefore, our main priority is to build even stronger automated monitoring techniques than what we have today. Even if some risk still remains, we believe deploying autonomous businesses early in a controlled, monitored environment is necessary to get a good understanding of model capabilities. Otherwise, we risk facing an uninformed future of widespread deployments with even more capable models that could cause significant harm.
Lots of mixed feelings about the potential of one day there being a "plug and play" business generator. But I want to give props to them for sharing data about the businesses they're running....
Lots of mixed feelings about the potential of one day there being a "plug and play" business generator.
But I want to give props to them for sharing data about the businesses they're running. Here's the two physical businesses that they mentioned: https://andonlabs.com/market
Unsurprisingly, they're losing money (market going from $100k to $3k now). But it's still interesting to see how they built the business and how they are managing it.
What is probably a more surreal experience is working in these places... These AI's actually hire real life people persons. So... My first question is: how well do they work against an adversarial/abusive employee? Probably not well, I imagine.
Perhaps as a side observation, but I think a neat application of such tools could be in worker co-ops? Although there are a handful of decision making systems (sociocracy, for example) which...
Perhaps as a side observation, but I think a neat application of such tools could be in worker co-ops? Although there are a handful of decision making systems (sociocracy, for example) which spread out responsibility instead of hierarchicalizing it, I'd imagine that there are some roles which would suffer for it (or else, there would be insufficient support for people who need to share that hat once the responsibility has been distributed). Having an AI stand-in, or an advisor, sounds like it could be useful in that regard? That way it still defers to the owner/workers instead of developing an ego and e.g. connives of a way to privatize the company for a resume boost or kickback.
(I'm aware that this isn't a silver bullet, but it feels like the sort of thing which could give such organizations more of a fighting chance at maintaining solvency or starting at all)
You know who also does a terrible job dealing with an adversarial/abusive employee? People who're running a business for the first time. Most humans are really terrible about maintaining...
You know who also does a terrible job dealing with an adversarial/abusive employee? People who're running a business for the first time. Most humans are really terrible about maintaining appropriate work boundaries when they're first in the position of managing other people, especially when it's a small business and there's just a few people. The AI could be worse, but it also doesn't have to lose its life savings failing at its first business. It can be bad at it and then it can get better since it has a pile of capital that's not its own to learn on.
Just because there are people who are bad at managing others, does not mean managing them with LLM's is a good idea. This just means that there are 2 possible ways you can have a bad business. But...
You know who also does a terrible job dealing with an adversarial/abusive employee? People who're running a business for the first time. Most humans are really terrible about maintaining appropriate work boundaries when they're first in the position of managing other people, especially when it's a small business and there's just a few people.
Just because there are people who are bad at managing others, does not mean managing them with LLM's is a good idea. This just means that there are 2 possible ways you can have a bad business.
But even then, I don't agree with the premise. I'm not convinced that someone who opens a business will have a level of assertiveness weaker than a frontier LLM (never mind if you decide to cut corners and use cheaper ones). And a person will not have dozens/hundreds of blog posts and github repositories teaching others how they can jailbreak them.
The AI could be worse, but it also doesn't have to lose its life savings failing at its first business. It can be bad at it and then it can get better since it has a pile of capital that's not its own to learn on.
Hum... I'm not sure what your point here is, but the premise here is that it runs a business with your money.
If this thing turns out to know what it's doing, I hope we start seeing companies that focus more on maximizing employee well-being while still being sustainable.
If this thing turns out to know what it's doing, I hope we start seeing companies that focus more on maximizing employee well-being while still being sustainable.
I'm wondering what it would be like to work for a such a business? Maybe like driving for Uber or doing meal deliveries? The only people you actually meet are the customers and other workers.
I'm wondering what it would be like to work for a such a business? Maybe like driving for Uber or doing meal deliveries? The only people you actually meet are the customers and other workers.
From the article:
[...]
[...]
[...]
Lots of mixed feelings about the potential of one day there being a "plug and play" business generator.
But I want to give props to them for sharing data about the businesses they're running. Here's the two physical businesses that they mentioned:
https://andonlabs.com/market
https://andonlabs.com/cafe
Unsurprisingly, they're losing money (market going from $100k to $3k now). But it's still interesting to see how they built the business and how they are managing it.
What is probably a more surreal experience is working in these places... These AI's actually hire real life people persons. So... My first question is: how well do they work against an adversarial/abusive employee? Probably not well, I imagine.
Perhaps as a side observation, but I think a neat application of such tools could be in worker co-ops? Although there are a handful of decision making systems (sociocracy, for example) which spread out responsibility instead of hierarchicalizing it, I'd imagine that there are some roles which would suffer for it (or else, there would be insufficient support for people who need to share that hat once the responsibility has been distributed). Having an AI stand-in, or an advisor, sounds like it could be useful in that regard? That way it still defers to the owner/workers instead of developing an ego and e.g. connives of a way to privatize the company for a resume boost or kickback.
(I'm aware that this isn't a silver bullet, but it feels like the sort of thing which could give such organizations more of a fighting chance at maintaining solvency or starting at all)
You know who also does a terrible job dealing with an adversarial/abusive employee? People who're running a business for the first time. Most humans are really terrible about maintaining appropriate work boundaries when they're first in the position of managing other people, especially when it's a small business and there's just a few people. The AI could be worse, but it also doesn't have to lose its life savings failing at its first business. It can be bad at it and then it can get better since it has a pile of capital that's not its own to learn on.
Just because there are people who are bad at managing others, does not mean managing them with LLM's is a good idea. This just means that there are 2 possible ways you can have a bad business.
But even then, I don't agree with the premise. I'm not convinced that someone who opens a business will have a level of assertiveness weaker than a frontier LLM (never mind if you decide to cut corners and use cheaper ones). And a person will not have dozens/hundreds of blog posts and github repositories teaching others how they can jailbreak them.
Hum... I'm not sure what your point here is, but the premise here is that it runs a business with your money.
What do they need me for?
Money
If this thing turns out to know what it's doing, I hope we start seeing companies that focus more on maximizing employee well-being while still being sustainable.
I'm wondering what it would be like to work for a such a business? Maybe like driving for Uber or doing meal deliveries? The only people you actually meet are the customers and other workers.