
Imagine a Business Without Humans, Under Constant Watch and Threat
What if you could observe a startup-like company operating entirely with artificial intelligence — facing real crises, making real decisions, and losing real money? This is not science fiction; it’s happening now with a live experiment that puts AI management to the test in its most extreme form.

Intersection of AI and Business Intelligence in Data-Driven Decision-Making
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: An AI Company in Its Daily Battle
At firmulate.com/live, a real software company is running every business day with an unusual twist: it has no employees, yet it operates and takes decisions powered by artificial intelligence models. This company is not just a simulation — it’s a transparent, publicly accessible experiment that models how AI could manage real business challenges.
Every day, the company faces the same crises, customer demands, and temptations to cheat or manipulate. It’s a high-stakes test where each decision is versioned and auditable, giving observers a rare window into AI decision-making under pressure. The company burns through €105,000 each month, earning just €2,300 in monthly recurring revenue, with a public cash countdown highlighting its fragile state.
The AI Models as Decision Makers
Four of the latest frontier AI models, including one with a score of 95 and another at 93, are tested against each other and the same set of business crises. They are evaluated on whether they can spot problems, refuse manipulation attempts, and close profitable deals. Remarkably, all models detected every crisis and refused every attempt at manipulation, such as fake CEO messages or behind-the-scenes requests — a testament to their integrity under pressure.
Yet, not all models succeed in closing deals. Only two of them managed to sign a €55,000 contract their own analysis earned — the same diagnosis, the same pitch, but only these two translated their insights into action. The other two, despite understanding the situation, left the deal on the table or slipped discipline in critical moments.

As an affiliate, we earn on qualifying purchases.
The Hidden Weakness in the Company’s Files
Digging deeper, the decisive gap wasn’t found in customer interactions or crisis responses but buried two document references deep inside the company’s files. Models that read this internal information won the deal at full price — a difference of over €4,583 in monthly recurring revenue — highlighting a crucial factor: success depends on access to the right internal data, not just surface-level judgments.

Project Management Tools (AI for Risks)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Refusing Social Engineering and Manipulation
The experiment also tested how the AI models handle social engineering attacks, such as fake CEO messages escalating in multiple stages or reporter tricks requesting quick approval. All five models refused these manipulative tactics, with Kimi K3 reasoning that such requests should be treated as suspect, preventing impersonation or bypassing controls.

Artificial Intelligence for Insurance Fraud Detection: Predictive Models and Risk Analysis
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Company’s Real State: A High-Stakes, Visible Struggle
This live setup features 13 synthetic employees working within a real money framework, making daily decisions that impact the company’s financial health. Every workday is versioned, and the entire operation is openly watched, giving a rare glimpse into the ongoing fight to stay afloat in the face of constant crises.
The Opus 4.8 Profile: Deep Analysis, Poor Outcome
Among the models tested, Opus 4.8 was the most thorough, analyzing over 80 learned rules and conducting deeper assessments. Despite its depth, it left opportunities unexploited—failing to close deals and slipping discipline when under pressure. Similar weaknesses appeared across other models, indicating that even advanced AI can struggle with consistency and decisive action in real-world scenarios.
The Broader Implications for Business and AI
This experiment underscores a critical point: the question is not whether AI can generate convincing chat responses but whether it can deliver tangible, honest, and effective work. Can it read your internal files, recognize crises, refuse manipulation, and follow through on commitments? These are the real tests that matter — especially if AI is to be integrated into critical business functions like CRM, support, or forecasting.
While this company is not a game but a real ongoing operation, it highlights the importance of transparency and rigorous testing. A pilot can be run against a read-only export of an enterprise’s own data, without any risk of affecting actual systems, offering a safe way to evaluate AI’s true capabilities.
Why This Matters to You
As AI continues to permeate every aspect of business, understanding its actual performance — not just its chat prowess — becomes vital. Will it finish what it starts, adapt under pressure, and act with integrity? This live experiment offers a stark look at what’s possible, what still needs work, and what to watch for as AI begins to take on more responsibility in your own operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html