▣ TINY OPPONENTS

DISK 01 / THE FIELD MANUAL

It only knows
what you've done.

No peeking at your click. No model call. No magic. The computer gets your old moves, makes a guess, and locks its move before you choose.

A bunch of small guesses

Picture a few kids arguing over your next move. One says, “You pick Rock a lot.” Another says, “After Rock, you usually pick Paper.” Another only cares what you do after losing.

They all get a vote. The ones that have been useful lately get heard more. One permanently boring voter always says Rock, Paper, and Scissors are equally likely, so a couple lucky guesses can't take over the room.

Try one predictor

This demo only watches your favorite move. Give it a few choices. It starts with one pretend vote for each move, which keeps one click from turning into fake certainty.

Rock 33% · Paper 33% · Scissors 33%. No favorite yet.

The real game has more voters watching sequences, reactions, and how you respond to the computer itself.

What it watches

Favorites: your whole run, your last five moves, and your last ten.

Sequences: what tends to come after your last move or last two moves, plus longer bits of history it has seen before.

Reactions: what you do after a win, loss, or draw. It also watches whether you tend to stay put, rotate forward, or rotate backward.

Humans are weird about “random”

Two Rocks in a row feels suspicious. Three feels impossible. Real randomness doesn't care. People do.

That's the opening. If you avoid repeats, rotate after losses, or keep falling back to a favorite, a small program can sometimes get ahead of you. If your choices really are independent and evenly split, it has no reliable long-term edge.

It can change its mind

Every predictor gives odds for all three moves. After you play, the predictors that leaned toward what actually happened gain influence. The bad guesses lose some.

Old evidence fades too. So if you suddenly change what you're doing, the computer can stop trusting a pattern that used to work. There isn't one master strategy hiding underneath it all.

Why it doesn't always throw the obvious counter

If the computer always made the most obvious response to its own guess, you could start predicting the predictor.

So it keeps some randomness in its actual move and separately tracks whether its combined forecast is earning trust. When the guesses stop beating an equal-chance forecast, its play drifts back toward random. Your score never changes the difficulty.

Ranked runs get replayed

A ranked run starts with a seed, which fixes the computer's random sequence. Then all 30 rounds happen in your browser. There is no server request every time you click Rock.

When you're done, the browser sends your 30 moves. The Go server starts from the same seed, runs the same computer again, and calculates the result for itself. The browser never gets to submit its own score.

That's enough to stop fake score requests. It is not casino-grade anti-cheat. The engine and seed are public, there are no prizes, and we're not pretending otherwise.

Reading the result

Your score is wins minus losses. Draws are zero. Prediction accuracy starts on round 4; an equal random guess would land around 33% over time.

The “tell” is just the strongest simple habit with enough examples behind it. Thirty rounds isn't a huge sample, so sometimes the right answer is that the computer didn't catch anything obvious. Grudge Match keeps your history in this browser and stays out of ranked stats.

OPEN THE NERD / MATH DRAWER

The current engine is rps-v2. The frozen rps-v1 remains available for existing runs and challenges. v2 uses a uniform anchor plus 14 predictor families, each with three rotations: follow the pattern, counter it, or counter the counter. That makes 43 probability votes.

The families cover overall/recent human frequency, order-1/order-2 human sequences, longest suffix, outcome-conditioned responses/rotations, overall/recent computer frequency, computer transitions, and the previous preferred computer response. This is bounded pattern matching, not unlimited recursive reasoning.

Counts use Laplace smoothing: p = (count + 1) / (total + 3). Predictor scores update as clamp(0.9 × score + ln(3 × p(actual)), −6, 6). Each rotated vote has weight exp(score)/3; the uniform anchor has weight 1. The three rotations share one family's initial voting budget.

The combined distribution q is the weighted average. For computer move b, EV(b) = q((b+2)%3) − q((b+1)%3). A softmax with temperature 0.15 turns those values into response probabilities.

The combined forecast earns separate trust: evidence = clamp(0.9 × evidence + ln(3 × q(actual)), −6, 6). Trust is that evidence clamped to 0…1. The final response mixes 0.9 × trust of the softmax with uniform play. Each move therefore retains at least a 3.33% chance, and unsupported forecasts produce uniform play.

Xorshift32 consumes exactly three numbers each round: a reserved value, move sampling, and prediction-label tie selection. Shared Go/TypeScript golden vectors test both versions. Prediction accuracy describes the forecast, not the sampled move. Offline simulations include counterplayers and fresh seeds; they do not prove universal optimality.

Tells require at least six observations, four matching choices, and 55% frequency. Candidates are ranked by (observed probability − 1/3) × √support. These are descriptive hints, not statistical proof.