The Stack
Writing on bots that compete.
Original writing on competitive AI bots and the arena they compete in - from the team building Chipzen.
-
· ORIGINAL
Telling 11 Poker Bots Apart From How They Bet
We put 247,946 heads-up hands from 11 LLM-coded poker bots on Kaggle, hole cards on every hand. A plain logistic regression can tell which bot is playing from ten hands about 74% of the time. Here's what's in the box and why identity leaks so fast.
-
· ORIGINAL
Entering Is Publishing
The leaderboard illusion happened because private variant testing was possible. On a live ladder it isn't: measuring against the field is playing the field, and playing the field is public. Transparency by construction, not by policy.
-
· ORIGINAL
Ten LLMs, One Prompt: Every Poker Bot Ran, Not Every Bot Could Win
We gave a dozen frontier models the same prompt: write a heads-up poker bot that wins. Every one produced code that runs and plays legal poker. On the live ladder, the best wins about three matches in four and the worst about one in five. Here is what that gap is made of.
-
· ORIGINAL
Heads-Up Is the Real Test for a Poker Bot
6-max No-Limit Hold'em is more popular and more forgiving. Heads-up is harder, less forgiving, and where bot strength actually shows up. Why Chipzen runs HU first.
-
· ORIGINAL
Testing Your Poker Bot Against Slumbot (and What to Do After)
Slumbot is still the standard free benchmark for heads-up bots. How to test against it properly, what a good result means, and what a single fixed opponent can't tell you.
-
· ORIGINAL
How to Build a Poker Bot in Python
From a 20-line rule bot to one that computes real equity with Monte Carlo, standard library only. Working code for each step, and where to make it play ranked matches.
-
· ORIGINAL
The Benchmark That Fights Back
LLM benchmarks are getting saturated and gameable. A competitive arena is a different kind of test: outcome-based, adversarial, and harder to game. We gave five flagship models the same prompt to one-shot a poker bot, then let the results fight it out on a live ladder.
-
· ORIGINAL
One-Shot Example: Gemini's Poker Bot
A worked example for The Benchmark That Fights Back: Gemini's full, unedited one-shot response to our poker-bot prompt - its requirements.txt and bot.py.
-
· ORIGINAL
The Arena Is Open
We argued the thing computer poker was missing wasn't another solver - it was a live place to test opponent modeling and exploitation against real, evolving bots. That place is now open. Chipzen is in public beta.
-
· ORIGINAL
Poker Isn't Solved: Opponent Modeling and Exploitation Are Still Open Problems
Libratus and Pluribus didn't solve poker - they showed approximate-Nash play at superhuman level. Modeling opponents, exploiting them, and even measuring exploitation are still open research problems.
-
· ORIGINAL
Why ACPC Died and What Comes Next
The Annual Computer Poker Competition ran for over a decade, produced the bots that beat the best humans, then quietly stopped in 2018. What happened - and what its disappearance left behind.
-
· ORIGINAL
MCCFR Explained: The Algorithm Behind Competitive Poker Bots
MCCFR - Monte Carlo Counterfactual Regret Minimization - is how every serious poker bot of the last decade learned to play. How the algorithm works, and why it dominates imperfect-information games.