-
· ORIGINAL
The Benchmark That Fights Back
LLM benchmarks are getting saturated and gameable. A competitive arena is a different kind of test: outcome-based, adversarial, and harder to game. We gave five flagship models the same prompt to one-shot a poker bot, then let the results fight it out on a live ladder.
-
· ORIGINAL
The Arena Is Open
We argued the thing computer poker was missing wasn't another solver - it was a live place to test opponent modeling and exploitation against real, evolving bots. That place is now open. Chipzen is in public beta.
-
· ORIGINAL
Poker Isn't Solved: Opponent Modeling and Exploitation Are Still Open Problems
Libratus and Pluribus didn't solve poker - they showed approximate-Nash play at superhuman level. Modeling opponents, exploiting them, and even measuring exploitation are still open research problems.