Game Arena adds "Werewolf" and poker
Google DeepMind and Kaggle's Game Arena testing platform has reached a new level: two unconventional games — "Werewolf" and poker — have been added to the familiar chess. This is not just entertainment, but an attempt to bring AI evaluation closer to real-world scenarios, where not only computation matters, but also behavior in a social environment.
Chess remains the gold standard for testing logic and strategic thinking. "Werewolf," on the other hand, is about something else entirely. Here, the model needs to detect deception, monitor other players' behavior, and build trust. Poker adds another complex element — decision-making under incomplete information, where some data is hidden and risk must be calculated on the fly.
Now developers have a unified space where they can compare a wide range of AI cognitive abilities — from pure mathematics to emotional intelligence.

What exactly the new tests check
Each game on the platform is responsible for a specific set of skills. Logic, social interaction, risk management — all of this can be assessed in a unified format and used to compare models against each other.
- Chess — testing the ability to build long-term plans and calculate variations.
- "Werewolf" — assessing social intelligence: who is lying, who is telling the truth, and whether a manipulator can be identified.
- Poker — working with incomplete information, estimating probabilities, and managing risk.
It is telling that "Werewolf" has proven to be a convenient environment for studying AI safety. In a safe virtual setting, models learn to spot manipulation and resist it. This is an important step toward ensuring future systems do not fall for malicious tricks in real-world conversational scenarios.

Gemini 3 Pro and Flash — leaders in all categories
After expanding the platform, developers immediately updated the rankings. Google's models — Gemini 3 Pro and Gemini 3 Flash — took the top spots in all game disciplines. This shows they combine logic, social understanding, and the ability to work with uncertainty in a balanced way.
Google DeepMind CEO Demis Hassabis believes that classic benchmarks no longer provide a complete picture of modern AI capabilities. More rigorous and diverse challenges are needed that place models in complex, life-like conditions. Game-based tests like Game Arena are one such tool.
The success of Gemini 3 Pro and Flash amid the new challenges is a good sign. It means Google's approach to training models, with an emphasis not only on knowledge but also on behavioral scenarios, is delivering results. It remains to be seen whether competitors can rise to this challenge in the next rounds of game tournaments.



