
Imagine a company that operates completely transparently, with no employees and no room for deception, yet faces daily financial struggles. Now, picture watching this company in real time, battling crises and testing its integrity—live online. This is the world of Firmulate, a groundbreaking project that puts artificial intelligence models through the most rigorous management tests, all while being accessible for public scrutiny.
What Is Firmulate and Why Does It Matter?
Firmulate is not your average simulation. It’s a public experiment where AI models—specifically frontier language models—run a virtual, small software company with real money mechanics and everyday crises. Every decision is tracked, every crisis is faced, and every temptation is tested. The goal isn’t just to see if the AI can write well or generate good chat, but whether it can truly manage a complex business, stay honest under pressure, and complete what it starts.
At the heart of this experiment are four leading AI models: gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5. Each runs the same company through its worst week—facing the same customers, crises, and ethical dilemmas. Every decision they make is versioned and auditable, giving researchers and the public unprecedented insight into their decision-making processes.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Harsh Tests, Clear Results
The results are illuminating. All four models successfully identified every crisis presented to them, demonstrating impressive situational awareness. They refused manipulation attempts—such as fake CEO messages or reporter tricks—every single time. In fact, the models’ integrity shone through: five out of five refused to authorize fake approvals or impersonation tactics.
The real surprise came from their ability to close deals. Only two models, gpt-5.6-sol and Kimi K3, managed to sign the €55,000 deal their analyses had earned. The others, despite similar diagnoses, left the deal unexecuted or failed to follow through. The crucial detail? The decisive weakness was buried in the company’s internal files, not in customer-facing data. When a model read those internal documents, it uncovered the hidden opportunity, closing the deal at full price—adding over €4,583 monthly recurring revenue.

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI
This experiment underscores a vital point: the ability of an AI to finish what it starts, stay honest under pressure, and read all relevant information is what truly matters in business applications. It’s not enough for AI to generate convincing chat or support emails. If AI is to be trusted with customer data, support operations, or financial decisions, it must demonstrate integrity, thoroughness, and discipline—especially when temptations to cheat or cut corners arise.

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company and Its Struggle for Survival
Firmulate’s live simulation features a company with 13 synthetic employees and real-money mechanics—burning through €105,000 each month against a modest €2,300 in monthly recurring revenue. Every workday, the company’s rules and strategies are versioned and made public at firmulate.com/live.html. Watching this real-time experiment offers a unique window into how AI manages crises, handles ethical dilemmas, and makes strategic decisions in a high-stakes environment.

Practical Simulations for Machine Learning: Using Synthetic Data for AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons for AI and Business Leaders
If AI models will soon touch your customer relationship management systems, support queues, or forecasting tools, the key question isn’t just about their writing skills. It’s about whether they can complete tasks reliably, stay honest under pressure, and uncover hidden opportunities in complex data.
For instance, the highest-scoring model, gpt-5.6-sol, achieved a score of 95 out of 100, successfully closing the deal at full price after uncovering buried facts. Kimi K3 followed close behind with a 93, demonstrating remarkable discipline and fairness—yet both models operated with different effort parameters, shedding light on how configurations influence performance.
Public Watch and Practical Applications
This experiment is open for everyone to observe, understand, and learn from. It illustrates that AI’s true potential comes not just from language generation but from decision-making integrity and thoroughness. Businesses considering AI adoption can run their own ‘wargames’ against their data—without risking real systems—via the pilot mode offered by Firmulate at firmulate.com/pilot.html.
In an era where AI’s reputation hinges on trust, the transparency and rigor of this live experiment provide critical insights. It shows that building AI for business isn’t just about smarter algorithms—it’s about creating systems that act ethically, finish what they start, and uncover opportunities hidden within complex internal data.

Firmulate’s live experiment reveals that AI’s true test is integrity and thoroughness—can it handle crises, resist manipulation, and uncover hidden opportunities? Watching this open, real-time company shows how AI might manage your business—honestly and effectively—if designed and monitored carefully.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html