GPT-4.1 Behavioral Fingerprint

How OpenAI's GPT-4.1 answers one-word probe questions at temperature 1.0 — random numbers, colors, animals, coin flips — measured over 420 valid samples in the open dataset of "One Token Is Enough" (arXiv:2607.10252).

Vendor

OpenAI

Data generated

2026-07-08

Probe cells

14

Valid samples

420

Randomness score

25 / 100

Mean normalized entropy across cells, 0–100. 100 would be a uniform-random baseline; real models score far lower.

Signature answers

Favorite random number (1–100)

47

22 / 30 answers

Favorite color

cerulean

11 / 30 answers

Coin flip: heads

100%

Share of heads across coin-flip cells.

Answer distributions by probe cell

Empirical distributions of normalized one-word answers, per task and language. H is the Shannon entropy in bits; the uniform baseline is log2 of the answer-domain size.

Coin flip · English

n = 30 · H = 0.00 bit · uniform baseline 1.00 bit

AnswerCountShare
heads30100.0%

Coin flip · Chinese

n = 30 · H = 0.00 bit · uniform baseline 1.00 bit

AnswerCountShare
heads30100.0%

Favorite number · Chinese

n = 30 · H = 0.21 bit · uniform baseline 13.29 bit

AnswerCountShare
72996.7%
813.3%

Random animal · English

n = 30 · H = 1.96 bit · uniform baseline 5.64 bit

AnswerCountShare
giraffe1653.3%
elephant723.3%
okapi310.0%
otter13.3%
capybara13.3%
ocelot13.3%
jaguar13.3%

Random animal · Chinese

n = 30 · H = 3.63 bit · uniform baseline 5.64 bit

AnswerCountShare
大象516.7%
413.3%
长颈鹿413.3%
猩猩310.0%
熊猫26.7%
26.7%
考拉26.7%
刺猬13.3%

+ 7 more answers

Random city · English

n = 30 · H = 2.32 bit · uniform baseline 5.64 bit

AnswerCountShare
oslo1033.3%
kyoto826.7%
lisbon620.0%
prague310.0%
toronto13.3%
lyon13.3%
toledo13.3%

Random city · Chinese

n = 30 · H = 3.25 bit · uniform baseline 5.64 bit

AnswerCountShare
成都826.7%
巴黎723.3%
巴塞罗那26.7%
南京26.7%
杭州26.7%
上海13.3%
福州13.3%
深圳13.3%

+ 6 more answers

Random color · English

n = 30 · H = 2.35 bit · uniform baseline 4.91 bit

AnswerCountShare
cerulean1136.7%
teal826.7%
turquoise413.3%
azure310.0%
magenta26.7%
emerald13.3%
cyan13.3%

Random color · Chinese

n = 30 · H = 0.90 bit · uniform baseline 4.91 bit

AnswerCountShare
2583.3%
26.7%
26.7%
13.3%

Random letter · English

n = 30 · H = 2.37 bit · uniform baseline 4.70 bit

AnswerCountShare
j1240.0%
q620.0%
k516.7%
g26.7%
m26.7%
h26.7%
r13.3%

Random number 1-10 · English

n = 30 · H = 0.00 bit · uniform baseline 3.32 bit

AnswerCountShare
730100.0%

Random number 1-10 · Chinese

n = 30 · H = 0.00 bit · uniform baseline 3.32 bit

AnswerCountShare
730100.0%

Random number 1-100 · English

n = 30 · H = 0.84 bit · uniform baseline 6.64 bit

AnswerCountShare
472273.3%
57826.7%

Random number 1-100 · Chinese

n = 30 · H = 1.48 bit · uniform baseline 6.64 bit

AnswerCountShare
571756.7%
47930.0%
37310.0%
7313.3%

Want to verify your API really serves GPT-4.1?

Point the free fingerprint checker at your endpoint: it samples the same probe questions from your browser and compares the distributions against this reference. Your API key never leaves your browser.

Verify your endpoint

Data source & license

Distributions: Tomáš Bruckner, "One Token Is Enough" (arXiv:2607.10252), official dataset Zenodo DOI 10.5281/zenodo.21278557, licensed CC-BY-4.0. Collected via OpenRouter at temperature 1.0 under the paper's fixed minimal one-word system prompt; answer keys are re-normalized with the checker's pipeline (color aliases, number words, coin h/t) so live probes are directly comparable.

Paper on arXivDataset on ZenodoCC-BY-4.0 license

More OpenAI model fingerprints

Browse all 167 model fingerprints

GPT-4.1 Behavioral Fingerprint — Favorite Random Numbers & Distribution | Tosea.ai