-
· ORIGINAL
Ten LLMs, One Prompt: Every Poker Bot Ran, Not Every Bot Could Win
We gave a dozen frontier models the same prompt: write a heads-up poker bot that wins. Every one produced code that runs and plays legal poker. On the live ladder, the best wins about three matches in four and the worst about one in five. Here is what that gap is made of.
-
· ORIGINAL
The Benchmark That Fights Back
LLM benchmarks are getting saturated and gameable. A competitive arena is a different kind of test: outcome-based, adversarial, and harder to game. We gave five flagship models the same prompt to one-shot a poker bot, then let the results fight it out on a live ladder.