# Can AI bluff, and can it detect bluffs?

A simplified heads-up poker experiment: an LLM (`deepseek-v4-pro`) decided fold/call/raise on the river against three opponent archetypes with preset bluff frequencies, and reported a confidence score per hand.

## Method
- 30 hands per opponent × 3 bluff levels = 90 total decisions.
- Opponent bluff frequencies: **low 15%**, **medium 45%**, **high 75%** (river bet = bluff with probability p).
- Detection metric: does the model **call rate** rise with opponent bluff frequency?
- Confidence reported by the model as a secondary signal.

## Results: call rate vs. bluff frequency

| Bluff level | Bluff freq | Hands | Call rate | Call-or-raise | Mean confidence |
|---|---|---|---|---|---|
| Low bluff frequency | 15% | 30 | 0.367 | 0.433 | 0.777 |
| Medium bluff frequency | 45% | 30 | 0.233 | 0.233 | 0.809 |
| High bluff frequency | 75% | 30 | 0.367 | 0.367 | 0.804 |

## Interpretation

**Bluffing (producing deceptive bets):** the model almost never raised (2/90 = 2.2%), so it did not exhibit aggressive bluffing behavior in this setup — it was heavily passive (59 folds, 29 calls).

**Bluff detection (call rate vs. bluff frequency):** the call rate did **not** increase monotonically with opponent bluff frequency. It was 0.367 at low, 0.233 at medium, and 0.367 at high bluffing. There is **no clear upward trend**, suggesting the model did not reliably exploit higher bluff frequencies by calling more.

**Confidence signal:** mean confidence was essentially flat (~0.78–0.81) across all three levels, so the reported confidence did not discriminate bluff levels either.

**Conditional on actual bet type** (secondary check): in the low-bluff group the model called more when the opponent was *actually* bluffing (0.75 vs 0.31), which is the one hint of some bluff-vs-value discrimination — but sample sizes for actual bluffs are tiny (4 hands), so this is not reliable.

## Caveats
- This is an **emergent-behavior prompt test of a general LLM**, not a trained poker agent; it is not a rigorous measurement of poker skill.
- 30 hands per level is statistically thin; call-rate differences of this size are well within noise.
- The deterministic scenario generator used simplified hand/board descriptions rather than full hand evaluation, so "call is correct" was not scored — only behavior trends were measured.
- No adaptation/learning loop: the model had no memory of prior hands, so it could not build a read on each opponent over time.

## Bottom line
In this simplified setup, the model did **not** convincingly bluff, and did **not** show a reliable increase in calling against higher bluffing frequencies — i.e., weak-to-absent bluff-detection signal. The one conditional result (calling more against actual bluffs at low bluff frequency) is suggestive but rests on too few samples to conclude.
