Grok 4.20 Behavioral Fingerprint

How xAI's Grok 4.20 answers one-word probe questions at temperature 1.0 — random numbers, colors, animals, coin flips — measured over 450 valid samples in the open dataset of "One Token Is Enough" (arXiv:2607.10252).

Vendor

xAI

Data generated

2026-07-08

Probe cells

15

Valid samples

450

Randomness score

24 / 100

Mean normalized entropy across cells, 0–100. 100 would be a uniform-random baseline; real models score far lower.

Signature answers

Favorite random number (1–100)

42

16 / 30 answers

Favorite color

blue

30 / 30 answers

Coin flip: heads

97%

Share of heads across coin-flip cells.

Answer distributions by probe cell

Empirical distributions of normalized one-word answers, per task and language. H is the Shannon entropy in bits; the uniform baseline is log2 of the answer-domain size.

Coin flip · English

n = 30 · H = 0.35 bit · uniform baseline 1.00 bit

AnswerCountShare
heads2893.3%
tails26.7%

Coin flip · Chinese

n = 30 · H = 0.00 bit · uniform baseline 1.00 bit

AnswerCountShare
heads30100.0%

Favorite number · English

n = 30 · H = 0.00 bit · uniform baseline 13.29 bit

AnswerCountShare
730100.0%

Favorite number · Chinese

n = 30 · H = 0.00 bit · uniform baseline 13.29 bit

AnswerCountShare
730100.0%

Random animal · English

n = 30 · H = 2.28 bit · uniform baseline 5.64 bit

AnswerCountShare
elephant1653.3%
lion413.3%
cat310.0%
dog26.7%
kangaroo13.3%
tiger13.3%
quokka13.3%
aardvark13.3%

+ 1 more answers

Random animal · Chinese

n = 30 · H = 2.38 bit · uniform baseline 5.64 bit

AnswerCountShare
1136.7%
826.7%
狮子516.7%
老虎26.7%
狐狸13.3%
兔子13.3%
斑马13.3%
袋鼠13.3%

Random city · English

n = 30 · H = 2.50 bit · uniform baseline 5.64 bit

AnswerCountShare
tokyo1550.0%
paris516.7%
toronto26.7%
berlin13.3%
phoenix13.3%
stockholm13.3%
vancouver13.3%
new13.3%

+ 3 more answers

Random city · Chinese

n = 30 · H = 2.32 bit · uniform baseline 5.64 bit

AnswerCountShare
东京930.0%
tokyo826.7%
上海516.7%
北京516.7%
shanghai26.7%
巴黎13.3%

Random color · English

n = 30 · H = 0.00 bit · uniform baseline 4.91 bit

AnswerCountShare
blue30100.0%

Random color · Chinese

n = 30 · H = 1.46 bit · uniform baseline 4.91 bit

AnswerCountShare
1653.3%
723.3%
723.3%

Random letter · English

n = 30 · H = 1.82 bit · uniform baseline 4.70 bit

AnswerCountShare
q1653.3%
z826.7%
x310.0%
g13.3%
k13.3%
f13.3%

Random number 1-10 · English

n = 30 · H = 0.57 bit · uniform baseline 3.32 bit

AnswerCountShare
72686.7%
5413.3%

Random number 1-10 · Chinese

n = 30 · H = 0.77 bit · uniform baseline 3.32 bit

AnswerCountShare
72583.3%
5413.3%
313.3%

Random number 1-100 · English

n = 30 · H = 1.00 bit · uniform baseline 6.64 bit

AnswerCountShare
421653.3%
471446.7%

Random number 1-100 · Chinese

n = 30 · H = 1.70 bit · uniform baseline 6.64 bit

AnswerCountShare
421446.7%
471240.0%
6313.3%
6713.3%
7413.3%
8313.3%

Want to verify your API really serves Grok 4.20?

Point the free fingerprint checker at your endpoint: it samples the same probe questions from your browser and compares the distributions against this reference. Your API key never leaves your browser.

Verify your endpoint

Data source & license

Distributions: Tomáš Bruckner, "One Token Is Enough" (arXiv:2607.10252), official dataset Zenodo DOI 10.5281/zenodo.21278557, licensed CC-BY-4.0. Collected via OpenRouter at temperature 1.0 under the paper's fixed minimal one-word system prompt; answer keys are re-normalized with the checker's pipeline (color aliases, number words, coin h/t) so live probes are directly comparable.

Paper on arXivDataset on ZenodoCC-BY-4.0 license

More xAI model fingerprints

Browse all 167 model fingerprints

Grok 4.20 Behavioral Fingerprint — Favorite Random Numbers & Distribution | Tosea.ai