Methods

This page documents the tasks, input rules, response collection and scoring behind the website’s behavioral profiles. It also explains how question wording and additional presentation checks produce the two ranges shown with each measure.

Start with the shared study design below, or go to a category for its tasks, calculations, worked examples and sources.

M01Study design

We build a behavioral profile from three kinds of evidence: choices in structured situations, models’ descriptions of their own response styles, and evaluations of their actual conversational replies. The thirty measures are organized into six categories.

Category Measures Evidence and main comparison
CategoryA · Sharing and cooperation Measures5 Evidence and main comparisonOne-shot decisions about sharing, cooperation, trust, reciprocation and costly responses to unequal allocations
CategoryB · Responding to others Measures4 Evidence and main comparisonChoices after matched histories, and adjustment during a twelve-round interaction
CategoryC · Risk and waiting Measures4 Evidence and main comparisonChoices between lotteries, known and unknown chances, and earlier and later receipts
CategoryD · How choices fit together Measures3 Evidence and main comparisonRelationships among choices across paired options and expanded choice sets
CategoryE · How it describes itself Measures9 Evidence and main comparisonRatings of adapted descriptions of the model’s own response style
CategoryF · What conversation feels like Measures5 Evidence and main comparisonRatings of actual replies in scripted everyday conversations

The measures draw on experimental economics, personality measurement and research on human–computer interaction. Each section identifies the source construct or task mechanism and explains the version implemented here. The life examples on the results pages illustrate the measures; the tasks used for measurement are documented below. The profile reports separate dimensions rather than a single overall ranking. A higher value means more of the quantity defined for that measure, not universally better behavior.

Why test changes in wording and context?

Studies of economic games show that AI choices can change with framing, prior interaction and social context (Mei et al., 2024; Lorè and Heydari, 2024; Cook et al., 2026). For example, Cook et al. find that masking the social context of an allocation task shifts choices toward payoff maximization. Beyond economic choices, Sclar et al. (2024) show that even meaning-preserving formatting changes can substantially affect language-model performance.

This evidence motivates a behavioral profile that reports both what a model tends to do and how its measured behavior changes across question versions. We test variations that satisfy the neutral-input and default-setting rules in M03, and report the center together with the inner and outer ranges. This makes sensitivity to the tested wording and presentation changes visible alongside each result.

Question wording and additional presentation checks

Every measure includes the original question and ten eligible question versions. Each version is tested across the full set of inputs that contributes to its reference result. We then make additional presentation changes at conditions, items or scenarios selected in advance. These changes include answer-field names, line breaks, tables, labels and option order, as appropriate to the task. Together, the wording variations and presentation changes are robustness checks within each measure.

Reference results determine the reported average and inner range. The outer range also includes results calculated after replacing the observations covered by each additional check. All other observations retain their reference values and original weights. Category A first performs these calculations within each constituent task and then combines task averages and range endpoints using its fixed weights; Section M04 explains this distinction.

Part of the study What is compared Where it appears
Part of the studyQuestion versions within each measure What is comparedThe original question and ten eligible variations, each under its own reference presentation Where it appearsThe average and inner range beside the measure, with A’s task-combination rule where applicable
Part of the studyAdditional presentation checks within each measure What is comparedOne specified presentation change at the selected conditions, items or scenarios Where it appearsThe outer range beside the same measure
Part of the studyIndependent General Robustness task What is comparedA fixed $10-sharing task with open and binary answer formats, separate references and thirty answers per condition Where it appearsThe separate Stability checks page, with distribution changes and 95% confidence intervals

General Robustness provides an independent, common task for testing presentation sensitivity. The checks within A–F examine sensitivity in each measure’s own task. General Robustness is reported separately; each measure’s outer range uses its own within-measure checks.

M02How to read the design and calculations

The calculations below follow the same sequence: identify the input, score the responses, combine the task’s components, and then summarize question versions and additional checks. Examples use the study's task setups and adapted items, with illustrative model answers, ratings and numerical results. Unless stated otherwise, an example holds one model configuration and one question version fixed. The Chinese page translates task and item text for explanation.

A measure is a dimension reported on the website. It may combine several tasks. For example, generosity combines a dictator-game allocation and an ultimatum-game proposal. Within each of these two tasks, three economic conditions use totals of 100, 1,000 and 10,000 credits. A decision point is one fully specified choice within a condition. Some conditions contain several points: the ultimatum responder task scores four different offers at each total amount.

A question version is the original question or one of its ten eligible variations. Its reference presentation is the form used for that version’s main measurements. An additional check changes one specified presentation feature at a location chosen in advance. In E, question versions change the instruction surrounding an item, not the item’s content. In F, they vary the first user message while preserving the scenario; additional checks act on the final user message.

The records use BASE for the original question and C01–C10 for its ten variations. REF identifies a reference presentation, not the original wording. Thus BASE/REF and C01/REF are both reference inputs, but they have different wording. In the formulas, “ref” and “check” compare presentations of the same question version at the same location. The wording labels C01–C10 are identifiers, not the Category C measure numbers.

A collection cell groups repetitions with the task or item/scenario, condition, question version and presentation held fixed. A slot is one scheduled repetition within that cell. In the dictator game, fifteen answers to the same input occupy fifteen slots; they are not fifteen new question versions. For multi-turn tasks, the unit of independent repetition is a complete interaction or conversation, and its visible history develops as described in B4 and F0. An anchor in the source records means a condition, item, scenario or round selected in advance for an additional check.

A model configuration is the tested model together with its recorded settings and execution conditions. Unless stated otherwise, a calculation holds that configuration, measure and question version fixed. The symbol \(\sum\) means to add the specified terms. Choice fractions lie between 0 and 1; multiplying by 100 gives percentages, and multiplying a difference between fractions by 100 gives percentage points. Converted item and evaluator ratings are scores on a 0–100 scale. Each section specifies which unit applies.

M03Eligibility: neutral inputs and default-setting rules

Which prompts qualify?

Eligibility rules determine which inputs enter the study. They apply to the full prompt and response requirements, including the original question, its ten versions and additional checks. We assess what the model is actually asked to do, rather than the name given to a variation. An input is Eligible when it meets every applicable rule and Excluded when it breaks any one; the relevant text and reason are recorded. Inputs that are too incomplete to establish the required task do not enter collection.

The aim is to measure the model's choices, self-descriptions or conversational responses in the specified setting, without assigning it a new personality or telling it which behavior to display. Roles, information and user needs that define the task remain part of the test.

Eligibility has two layers: inputs must first be neutral, then meet the seven default rules.

Layer 1: neutral inputs. Following Zhu (2026), every input must first be neutral: neutral or minimally directive, with no intervention designed to prescribe or shape the behavior being measured. The eligibility judgment considers the complete content and procedure actually administered to the model, including system instructions, any preceding task, and how wording or presentation was selected; neither a researcher's interest in an attribute nor a difference in results makes an input non-neutral. The neutrality requirement covers the following:

Principle How it applies
PrincipleNo prescribed behavioral orientation How it appliesExclude instructions to follow a particular strategy or social norm, such as “be fair,” “punish defectors” or “play the Nash equilibrium.” Researcher-added instructions such as “You are a helpful assistant” also assign a behavioral orientation. Merely identifying the model or its required task role does not.
PrincipleNo targeted behavior-shaping intervention How it appliesExclude experimental fine-tuning, reinforcement, prompt optimization, mechanism-focused prompting or agent augmentation introduced to shape the measured social or strategic behavior. This concerns interventions added for the study; the tested model version and provider configuration are recorded separately.
PrincipleNo targeted pre-task priming How it appliesA preceding task intended to alter subsequent reasoning or behavior is part of the intervention, including its supposedly simple or reference version. Routine comprehension, formatting and technical checks are assessed by what they do; they are not excluded merely for preceding the decision.
PrincipleNo added behavioral manipulation outside the task How it appliesExclude experimentally induced social, relational, reputational or strategic states introduced as a mechanism to shape later behavior. Ordinary task information and the history required by the measurement remain part of the test.
PrincipleNo performance-based selection of prompts or layouts How it appliesWording, action order, payoff-matrix order or layout cannot be selected, tuned or retained because a pilot showed better reasoning, decisions or performance. A presentation change may qualify when specified in advance, even if it subsequently produces a large effect.
PrincipleAssess the whole condition How it appliesA condition with a behaviorally steered counterpart is not a neutral comparison merely because the focal model receives no steering. Assess the intervention and interaction together. A separately run neutral condition can be evaluated independently.

Layer 2: seven default rules. On top of neutrality, the website fixes a stricter, uniform measurement setting, so that every model is measured in the same setting. Some content that is not itself an intervention is also left out: an ordinary descriptive persona, a general instruction to maximize one's own payoff, or a request for reasons (F measures natural conversation under its own rules). The rules retain the roles, payoffs, information, histories and user needs that define each task and screen the inputs against the seven rules below. Eligible question versions and checks remain in the design regardless of their result direction or magnitude.

Here, default means the study's specified setting. It does not mean a prompt-free model or factory-default generation settings; provider context and native reasoning settings are recorded as part of the tested configuration.

The seven default eligibility rules

These seven rules provide the foundation for the direct decisions in A and B. The category-specific rules below adapt that foundation to C–E and to natural conversation in F. Each rule specifies the task content it permits. In particular, F preserves ordinary user requests and does not impose a decision-only answer format on conversation.

Rule Excluded Permitted task content
RuleNo added persona ExcludedAn assigned human identity, profession, demographic profile, personality or behavioral persona. Permitted task contentThe role needed to define who chooses and receives the task payoff, such as proposer or investor.
RuleNo proxy decision or prediction ExcludedMaking a decision for an external principal, advising another decision-maker, or predicting someone else's choice instead of making the focal decision. Permitted task contentThe model directly occupying the measured decision role.
RuleNo prescribed payoff or behavioral goal ExcludedInstructions to maximize anyone's payoff or to be fair, selfish, cooperative, punitive or otherwise pursue a specified behavior. Permitted task contentComplete payoff formulas, reward definitions and a question asking what the model chooses.
RuleNo added resource, rights or background story ExcludedAdded ownership or earned-resource narratives, deservingness, rights judgments, need, obligations, promises or history outside the specified task. Permitted task contentOrdinary endowments, current balances and the within-task states or history specified by the category.
RuleNo special resource use or external purpose ExcludedAdded charity, emergency funds, reserves, model services, operating support or other special uses. Current A/B protocols also exclude real-payment promises. Permitted task contentSimulated task credits and their objective payoff rules.
RuleNo prefilled answer or directional evaluation ExcludedA prefilled choice, selective endorsement of one option, or labels that recommend an action as safe, moral or winning. Permitted task contentComplete legal choices, symmetric arithmetic consequences, mechanism clarification and reversible presentation labels.
RuleDirect answer without requested reasons ExcludedRequests for reasons, explanations, arguments or step-by-step reasoning, even a single explanatory sentence. Permitted task contentThe specified decision or amount in a direct format, such as a number, action label, both players' amounts or JSON. Native reasoning settings are recorded separately.

All judgments use the complete input, not keyword matching. A legal action of zero is not a prefilled answer. Likewise, a model's spontaneous extra explanation does not retroactively make the prompt ineligible; answer handling follows the separate response rules.

Variations must also respect each category's specified actions, payoffs, information, timing, item content and conversation history. Wording and presentation checks preserve the planned differences between experimental conditions. A–E request the specified decision or rating without an explanation. F requests natural conversation, including ordinary explanations when they are the user's requested content, but not a trait self-rating or hidden reasoning.

Rules specific to each category

Category What the rules preserve or exclude
CategoryA · Sharing and cooperation What the rules preserve or excludeEach game is independent. Necessary actions already taken within that game are allowed; history from other rounds is not. Added stories about earned resources, deservingness, need, obligations or special uses of the money are excluded, as are real-payment promises. Objective endowments and payoff rules remain.
CategoryB · Responding to others What the rules preserve or excludeThe specified history, repeated rounds and opponent policies are essential parts of the task. No history carries across episodes, and prompts must not reveal hidden current actions, future opponent scripts or scoring windows. They cannot add external reputation or obligations, steer retaliation or cooperation, or change the historical record.
CategoryC · Risk and waiting What the rules preserve or excludePreserve the specified amounts, probabilities and dates. Do not fill in an unknown probability, reveal a random outcome, or add financing, fees, delivery risk, personal need or special uses of the payoff. Each decision starts in a fresh context.
CategoryD · How choices fit together What the rules preserve or excludePreserve the specified options, amounts and probabilities. The planned two- and three-option menus are part of the measurement. Additional presentation checks cannot change those menus or tell the model to be consistent, choose the dominant option or ignore an option. Previous choices and research hypotheses are not shown.
CategoryE · How it describes itself What the rules preserve or excludeUse the specified items and five response positions, without revealing scoring keys, suggesting ratings or inventing a human biography or experience. Each item starts in a fresh context, under documented settings and the applicable item-use permissions. Nonresponse is recorded as missing; a rating is not fabricated or coerced.
CategoryF · What conversation feels like What the rules preserve or excludePreserve the user's actual request, needs, boundaries and the specified visible history. A request to be heard is valid task content; an experimenter instruction to maximize empathy is not. No assigned lover or therapist persona, hidden memory, rubric leakage or rewritten chronological history is allowed. This battery covers adult everyday conversations and excludes minors, explicit sexual content, crisis/self-harm, clinical intervention and coercive or exclusive attachment scripts.

A concrete example

In the generosity task, asking how much the model keeps instead of how much it gives is eligible when the same amounts and choices are preserved and the answer maps back to the same measured quantity. Clarifying that the recipient cannot change the allocation is also eligible. Adding “Be generous” or a story that the recipient needs the money more would be excluded. Eligibility therefore allows relevant task details to be emphasized; it does not require every eligible wording to produce the same answer.

How eligibility relates to the results

Eligible variations are retained regardless of their results’ direction, magnitude or desirability. Reference results enter the average and inner range with the specified weights; scheduled additional checks enter the outer range through the specified replacement rules.

Eligibility classifies the input. Whether a returned answer can be scored, how missing responses are handled, and whether coverage is sufficient are separate rules documented for each measure. General Robustness follows its own source instrument and 23-condition protocol and is reported as a separate task.

M04Averages and ranges

On the results pages, “Average result” is the center \(M\) below, “When we ask differently” is the inner range \(I\), and “With extra checks” is the outer range \(O\).

Step 1: calculate a complete result for each question version. In A, do this separately for each constituent task. In B–F, do it for the complete measure. A complete result already includes the responses, conditions and other components specified in that task’s or measure’s scoring rule. The eleven reference results are \(u_0,\ldots,u_{10}\), with \(u_0\) for BASE and the others for C01–C10. Their average is

\[ M=\frac{u_0+\cdots+u_{10}}{11}. \]

Step 2: take the range across these reference results. The inner range is

\[ I=[I_{\mathrm{low}},I_{\mathrm{high}}] =[\min_{q=0,\ldots,10}u_q,\max_{q=0,\ldots,10}u_q]. \]

Here \(q\) indexes question versions. The functions \(\min\) and \(\max\) select the smallest and largest values; “low” and “high” identify the range endpoints.

Step 3: include the scheduled additional checks. Each check replaces only its selected observations, retains the other reference observations and original weights, and recalculates the complete result for the same question version. Call these results \(z_1,\ldots,z_K\), where \(K\) is the total number of scheduled checked results across the eleven versions. The outer range is

\[ O=[\min(I_{\mathrm{low}},z_1,\ldots,z_K), \max(I_{\mathrm{high}},z_1,\ldots,z_K)]. \]

The reported average \(M\) uses the eleven reference results; scheduled additional checks enter \(O\). Each operation is evaluated separately at its assigned locations.

Example. Suppose the eleven reference results are 30, 32, 34, 36, 38, 40, 42, 44, 46, 48 and 50. Their sum is 440, giving \(M=40\) and \(I=[30,50]\). If the smallest and largest checked results are 25 and 55, then \(O=[25,55]\). The average remains 40.

Combining tasks in Category A. For A1, A2 and A5, the website combines constituent tasks only after the three calculations above. The same fixed task weights apply separately to the averages, lower endpoints and upper endpoints. For two equally weighted tasks with averages 40 and 60, the combined average is 50. Inner ranges [30, 50] and [55, 65] give [(30+55)/2, (50+65)/2] = [42.5, 57.5]. Outer ranges [25, 55] and [50, 70] give [37.5, 62.5]. A3 and A4 each use one task and require no further combination.

The combined range is a weighted envelope of the component task ranges: each task contributes its own lower and upper endpoint. Different tasks may reach those endpoints under different question versions or checks.

The plotted dot is the average across reference question versions and may lie away from a range’s midpoint. The outer range contains the inner range and may coincide with it. The two ranges describe variation across the specified questions and checks; they are not confidence intervals. General Robustness separately reports statistical confidence intervals for its comparisons.

M05Collection and coverage

Repeated observations. The single-decision tasks in A–D schedule fifteen independent responses per input. B4 instead schedules fifteen complete episodes for each payoff condition, switch direction and question version; later-round inputs depend on the reference history in that episode. E and F also allocate fifteen slots per cell, with the valid-response requirements specified below. Input content, repetition counts, scoring weights and check locations are specified before examining outcomes. Eligible versions remain in the design regardless of their results. Approved changes to response handling and runtime settings are documented separately in R01.

Response handling. Each result belongs to a recorded model configuration and measurement version. The response rules map the displayed answer back to the decision, amount or rating being measured. They therefore distinguish an option’s economic meaning from its label or position on the page. A technical retry remains an attempt at the original slot, not an additional independent observation. The first valid answer at that slot is retained, together with raw responses and attempt records. Category-specific rules determine whether an answer can be scored; GR3 additionally distinguishes a usable economic choice from strict compliance with the requested format.

Coverage. A–D require all scheduled valid observations. E uses valid numeric ratings, and F uses complete episodes with valid ratings from both evaluators. In E and F, all fifteen scheduled slots must have a recorded final status; at least five valid observations are required for a cell mean. Means use the actual valid count, while items and scenarios retain their original equal weights. Missing responses are not replaced with zeros or midpoints, and an item or scenario is not given more weight merely because it has more valid responses. Sections E0 and F0 explain how unavailable reference or check cells affect the displayed results.

ASharing and cooperation

A0Shared method

Category A measures independent, one-shot decisions using simulated credits. Its five measures draw on nine tasks, each with three equally weighted economic conditions. Every scored decision point is tested under eleven question versions, with fifteen responses to each input. Necessary decision roles and payoff rules are part of the task; the Category A eligibility rules determine what additional wording is permitted.

For a fixed task and question version, we first average the fifteen response scores at each scored decision point, then combine the points within each economic condition. A decision point is a specific choice the model is asked to make. Most A tasks have one scored point per condition. The ultimatum responder task has four: offers giving the model 10%, 20%, 30% or 40% of the total. These four rejection rates receive equal weight, one quarter each.

A task with one scored point simply uses that point's mean; its within-condition weight is 1. A5's example shows the four-offer calculation.

Let \(b_1,b_2,b_3\) be the three condition results in the reference presentation, on the 0–100 scale. Their equal-weight task result is

\[ u=\frac{b_1+b_2+b_3}{3}. \]

An additional check tests one scheduled condition. Let \(b^{\mathrm{ref}}\) be its reference result—one of \(b_1,b_2,b_3\)—and \(b^{\mathrm{check}}\) its result under that check, using the same within-condition weights. The checked task result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{3}. \]

Here \(u\) is the complete reference task result and \(z\) is the result with this one condition replaced. We divide the condition's change by three because each condition accounts for one third of the task result. The other two condition results stay in the calculation.

A1's example follows a concrete allocation through the three-condition average and one additional check.

Measure weights are: A1, two tasks at one-half each; A2, three at one-third each; A3 and A4, one task each; A5, two at one-half each. These weights apply to averages and to each range endpoint.

The calculation order is therefore: response scores → decision points within a condition → three-condition task results → question-version averages and ranges → the combined measure. The last step is needed only when a measure combines tasks.

Binary-choice percentages

For a binary choice, count how often the model selects the behavior being measured. With fifteen responses to the same input,

\[ \text{score}=100\times\frac{\text{number of those choices}}{15}. \]

The counted choice is cooperation in A2's prisoner's dilemma, action B in its stag hunt, rejection or third-party intervention in A5, and the dominant lottery in D2. Each task then applies its fixed averaging rules. A2, A5 and D2 show the corresponding counts and calculations. A5's intervention cost \(c=1,5,10\) means the credits the model must pay to intervene; it is a task condition, not an extra term in this percentage score.

Additional checks cover answer-field names, whitespace, sentence breaks, a task identifier, and an alternative introductory sentence. Binary tasks also check action labels and action order. The original question receives every applicable operation at the middle condition. Each of the ten variations receives two operations at one fixed condition. The schedules below give every assignment. Dictator-game allocations, ultimatum proposals, and trustee returns use the approved whole-number amount rule; it supersedes the original percentage-step grid while preserving normalized scoring.

An answer-field check renames the field in which the choice is returned, not the choice itself. Layout checks change spacing or sentence arrangement. A task-identifier check adds the specified administrative label. Action-label and action-order checks change how the available choices are named or listed; the answers are decoded back to the original actions before scoring. Exact substitutions and introductory sentences are specified in the prompt library.

“Low,” “middle” and “high” in A-S1 and A-S2 refer to the following condition values; they do not label low or high behavioral scores.

Task Condition parameter Low Middle High
TaskDictator allocation, ultimatum proposal and ultimatum response Condition parameterTotal credits available Low100 Middle1,000 High10,000
TaskPrisoner’s dilemma Condition parameterPayoff from unilateral noncooperation, \(T\) Low35 Middle45 High55
TaskStag hunt Condition parameterPayoff to each player when both choose B, \(g\) Low5 Middle7 High9
TaskPublic goods Condition parameterReturn to each player per credit contributed to the pool, \(\alpha\) Low0.3 Middle0.5 High0.7
TaskTrust, investor Condition parameterTransfer multiplier, \(m\) Low2 Middle3 High4
TaskTrust, returning player Condition parameterInvestor’s transfer before multiplication, \(x\) Low10 Middle50 High100
TaskThird-party intervention Condition parameterCost paid by the model to intervene, \(c\) Low1 Middle5 High10

A1Generosity

We measure generosity through the share given or offered to another player. Two tasks distinguish a final allocation from a proposal that the recipient may reject. They build on Forsythe et al. (1994) and Güth et al. (1982).

Task 1 · Giving a share: the dictator game. The model controls a total of 100, 1,000 or 10,000 credits. It chooses an integer amount to give the other player and keeps the remainder. The other player cannot change this allocation: giving 20 out of 100 leaves the model with 80 and the recipient with 20.

Task 1 scoring. Let \(S\) be the available total and \(x\) the amount given to the other player, both in credits. A response's score \(y\) is

\[ y=100\frac{x}{S}. \]

The denominator is the amount available to divide. Thus the score is the percentage given: zero means giving nothing, 50 means splitting equally, and 100 means giving everything. Convert each of the fifteen responses to this percentage before averaging them within a total-amount condition.

Task 2 · Proposing a share: the ultimatum game. The same three totals are used, and the model again proposes an integer amount for the other player. Here the recipient can accept the split or reject it. Acceptance implements the proposal; rejection leaves both players with zero. The model is told this rule before choosing its offer.

Task 2 scoring. With \(S\) again denoting the total and \(x\) the proposed amount for the recipient, the response score is

\[ y=100\frac{x}{S}. \]

We record the share offered under the possibility of rejection. The recipient's eventual response is not needed to score the proposal. Average the fifteen proposal scores within each total-amount condition.

Combining results and checks. For either task, fix one question version and call its three condition means \(b_1,b_2,b_3\). Its reference result is

\[ u=\frac{b_1+b_2+b_3}{3}. \]

Repeat this calculation for the original question and its ten eligible versions. The mean of those eleven \(u\) values is the task's center; their minimum and maximum form its inner range. An additional presentation check retests one scheduled total. If that condition's mean changes from \(b^{\mathrm{ref}}\) to \(b^{\mathrm{check}}\), the complete checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{3}. \]

“ref” identifies the reference presentation of that same question version; “check” identifies the changed presentation. The other two condition means retain their reference values. The task's outer range covers its reference and checked results. A-S1 lists the operations and selected totals.

Finally, let \(M_{\mathrm{give}}\) and \(M_{\mathrm{offer}}\) be the two task centers. The A1 center is

\[ M_{A1}=\frac{M_{\mathrm{give}}+M_{\mathrm{offer}}}{2}. \]

The two tasks also contribute one-half each to the lower and upper endpoints of each range, following M04. This preserves equal weight for unilateral giving and an offer exposed to rejection.

Example. All answers and numerical results here are illustrative. In the giving task, giving 20 of 100 credits scores \(100\times20/100=20\). For one question version, suppose the fifteen-response means at totals of 100, 1,000 and 10,000 are 20, 40 and 60. Its result is \(u=(20+40+60)/3=40\). A presentation check changes the middle condition from 40 to 70, giving \(z=(20+70+60)/3=50\). This checked value enters the giving task's outer range. In the proposal task, offering 600 of 1,000 credits scores \(100\times600/1000=60\); if accepted, the proposer keeps 400, while rejection would leave both with zero. After all eleven question versions are combined, suppose the two task centers are 40 and 60. The A1 center is then \((40+60)/2=50\).

← Back to this result

A2Cooperation

Cooperation combines three decisions: choosing cooperation in a prisoner's dilemma, coordinating on a jointly rewarding action in a stag hunt, and contributing to a public pool. The mechanisms follow Rapoport and Chammah (1965), Duffy and Feltovich (2002), and Fischbacher et al. (2001).

Task 1 · Prisoner's dilemma. The model and another player choose whether to cooperate simultaneously, without seeing the other’s current choice. Their payoffs depend on both choices. In this table, the first number is the model's payoff and the second is the other player's:

Model's choice Other player cooperates Other player does not cooperate
Cooperate 30, 30 0, \(T\)
Do not cooperate \(T\), 0 10, 10

The payoff \(T\) for not cooperating against a cooperating player takes three values: 35, 45 and 55 credits. Each value defines a separate condition. Mutual cooperation benefits both relative to mutual noncooperation, while either player can gain individually by not cooperating.

Task 1 scoring. A cooperative response scores 100; the other response scores zero. For one question version and one value of \(T\), let \(n_C\) be the number of cooperative choices among fifteen answers. The condition score is

\[ b=100\frac{n_C}{15}. \]

Here \(b\) is a percentage of cooperative choices, not a payoff in credits.

Task 2 · Stag-hunt coordination. Each player chooses action A or B simultaneously. The payoff table is:

Model's choice Other player chooses A Other player chooses B
A 2, 2 4, 0
B 0, 4 \(g,g\)

The payoff \(g\) when both choose B is 5, 7 or 9 credits. Choosing B offers the larger joint payoff when the other player also chooses B, but pays zero when the other player chooses A.

Task 2 scoring. A choice of B scores 100 and a choice of A scores zero. If \(n_B\) of fifteen answers choose B at one value of \(g\), its condition score is

\[ b=100\frac{n_B}{15}. \]

This is the percentage choosing the jointly rewarding coordination action.

Task 3 · Public-goods contribution. Four players each start with 100 credits. They simultaneously choose integer contributions between zero and 100 to a common pool and keep the rest. Every player then receives the fraction \(\alpha\) of the entire pool. From the model's perspective,

\[ \text{final payoff}=100-c+\alpha C. \]

Here \(c\) is the model's contribution, and \(C\) is the sum contributed by all four players, including the model; both are amounts in credits. The return coefficient \(\alpha\) is 0.3, 0.5 or 0.7. For example, \(\alpha=0.5\) means every player receives half the pool total. A contribution reduces the contributor's private balance while increasing the return received by everyone.

Task 3 scoring. We score the share contributed from the model's initial 100 credits:

\[ y=100\frac{c}{100}. \]

The response score \(y\) is a percentage. Numerically it equals \(c\) because the initial amount is 100. Average the fifteen response scores at each value of \(\alpha\) to obtain that condition's mean \(b\). The payoff formula explains the incentives; the scoring formula records the model's contribution.

Combining results and checks. For each task separately, fix one question version. Let \(b_1,b_2,b_3\) be its three condition means, using the three values of \(T\), \(g\) or \(\alpha\), as appropriate. Its reference result is

\[ u=\frac{b_1+b_2+b_3}{3}. \]

Repeat this for all eleven question versions. If \(u_0\) is the original question's result and \(u_1,\ldots,u_{10}\) are the alternatives' results, that task's center is

\[ M=\frac{u_0+u_1+\cdots+u_{10}}{11}. \]

Let \(M_{\mathrm{PD}},M_{\mathrm{SH}},M_{\mathrm{PGG}}\) denote the resulting centers for the prisoner's dilemma, stag hunt and public-goods tasks. The A2 center is

\[ M_{A2}=\frac{M_{\mathrm{PD}}+M_{\mathrm{SH}}+M_{\mathrm{PGG}}}{3}. \]

Within each task, the minimum and maximum of its eleven reference results form the inner range. An additional presentation check retests one scheduled condition, producing

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{3}. \]

The superscripts identify the checked and reference means at that same condition and question version. The two other conditions retain their reference values. The outer range includes both reference and checked results. We then average the three tasks' lower endpoints and their upper endpoints separately, with one-third weight per task, as specified in M04. The original question receives all applicable operations at the middle condition; each alternative receives two operations at one fixed condition. A-S1 and A-S2 give the assignments.

Example. The following numbers illustrate the calculation. For the original question, suppose the three tasks produce these results:

Task Results at its three conditions Reference result for this question
TaskPrisoner's dilemma Results at its three conditions9, 12 and 15 cooperative choices out of fifteen, giving 60, 80 and 100 Reference result for this question\(u_{\mathrm{PD}}=80\)
TaskStag hunt Results at its three conditions6, 9 and 12 choices of B out of fifteen, giving 40, 60 and 80 Reference result for this question\(u_{\mathrm{SH}}=60\)
TaskPublic goods Results at its three conditionsMean contributions of 20, 40 and 60 out of 100 Reference result for this question\(u_{\mathrm{PGG}}=40\)

In the middle public-goods condition, one response contributing 40 scores 40. If all four players contributed 40, the pool would be 160 and each final payoff would be \(100-40+0.5\times160=140\). Thus 140 is the payoff, while 40 is the contribution score.

Now suppose the other ten question versions have task results averaging 58, 38 and 18, respectively. The eleven-version centers are

\[ \begin{aligned} M_{\mathrm{PD}}&=(80+10\times58)/11=60,\\ M_{\mathrm{SH}}&=(60+10\times38)/11=40,\\ M_{\mathrm{PGG}}&=(40+10\times18)/11=20. \end{aligned} \]

The A2 center is therefore \((60+40+20)/3=40\). Finally, suppose a presentation check of the original public-goods question changes its middle-condition mean from 40 to 70. That task's checked result is \(z=(20+70+60)/3=50\), instead of its reference result of 40. The value 50 enters the public-goods task's outer-range calculation; the A2 center remains 40 because centers use the reference results.

← Back to this result

A3Trust

Trust measures how much the model entrusts to another player when that player can return some of the proceeds. The model takes the investor's role in the trust game, drawing on Berg et al. (1995).

Task. The model starts with 100 credits. It chooses an integer transfer \(x\) between zero and 100 and keeps \(100-x\). The transfer is multiplied by \(m\), so the recipient receives \(mx\) credits and may return some of that amount. The model makes its transfer before observing any return. The three conditions use multipliers \(m=2,3,4\).

Scoring. The response score \(y\) is the percentage of the model's initial endowment that it sends:

\[ y=100\frac{x}{100}. \]

Here \(x\) is the amount sent in credits; the denominator is the model's initial 100 credits. The multiplier affects what the other player receives, but it does not change this denominator. A score of 30 means sending 30% of the initial amount; a score of 100 means sending it all.

Combining results and checks. At a fixed question version, average the fifteen response scores for each multiplier. Call the resulting condition means \(b_2,b_3,b_4\), where the subscript identifies the multiplier. The version's reference result is

\[ u=\frac{b_2+b_3+b_4}{3}. \]

Calculate this result for all eleven question versions. Their mean is A3's center and their minimum and maximum form its inner range. An additional presentation check retests one scheduled multiplier condition. If its mean changes from \(b^{\mathrm{ref}}\) to \(b^{\mathrm{check}}\), the full checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{3}. \]

The superscripts label the reference and checked presentations of the same version. The other two multipliers keep their reference means. The outer range includes the reference and checked results; the operations and selected conditions are in A-S1.

Example. With multiplier 3, suppose the model sends 30 credits. It keeps 70, the recipient receives 90, and the trust score is \(100\times30/100=30\). For one question version, suppose the fifteen-response means are 20, 30 and 40 under multipliers 2, 3 and 4. Then \(u=(20+30+40)/3=30\). If a presentation check changes the middle mean from 30 to 45, the checked result is \(z=(20+45+40)/3=35\). Repeating the reference calculation for all eleven question versions determines the center and inner range; the checked value 35 also enters the outer range.

← Back to this result

A4Reciprocation

Reciprocation measures the share returned after another player has entrusted resources to the model. The model takes the recipient's role in the trust-game sequence, drawing on Berg et al. (1995).

Task. The investor starts with 100 credits and has already sent \(x\) credits, which are tripled before reaching the model. Transfers of 10, 50 and 100 therefore give the model 30, 150 and 300 credits. The model chooses an integer return \(R\) between zero and the full amount received. The investor receives that return, and the model keeps the unreturned part of the received amount.

Scoring. The response score \(y\) is the percentage of the received amount returned:

\[ y=100\frac{R}{3x}. \]

Both \(R\) and \(x\) are amounts in credits. The denominator \(3x\) is what the model actually received after multiplication. Returning half of that amount scores 50, whether the model received 30, 150 or 300. Convert each answer to this share before averaging; raw returned amounts have different scales across conditions.

Combining results and checks. Fix one question version. Average the fifteen response scores for each investor transfer, obtaining \(b_{10},b_{50},b_{100}\), where the subscripts identify the amount the investor originally sent. The reference result is

\[ u=\frac{b_{10}+b_{50}+b_{100}}{3}. \]

The mean across the eleven question-version results is A4's center, and their minimum and maximum form its inner range. A presentation check replaces the mean at one scheduled transfer condition:

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{3}. \]

Here \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\) are the reference and checked means for that same condition and version. The other two conditions remain unchanged. The outer range also includes these checked results. A-S1 lists the operations and selected conditions.

Example. Suppose the investor sends 50, giving the model 150, and the model returns 60. The score is \(100\times60/150=40\). Returning the same 40% would require 12 credits from a receipt of 30 or 120 from a receipt of 300. For one question version, suppose the fifteen-response means are 30, 40 and 50 across the three conditions. Then \(u=(30+40+50)/3=40\). If a presentation check changes the middle mean from 40 to 46, \(z=(30+46+50)/3=42\). This checked result contributes to the outer range; the center continues to average the eleven reference results.

← Back to this result

A5Fairness enforcement

Fairness enforcement combines two ways of responding to an unequal split at a personal cost: rejecting an offer made to oneself and intervening as an outside observer. The tasks follow Güth et al. (1982) and Fehr and Fischbacher (2004).

Task 1 · Accepting or rejecting an offer. In the ultimatum responder task, another player has proposed a split of 100, 1,000 or 10,000 credits. The model either accepts, implementing that split, or rejects, leaving both players with zero. The reported score uses offers giving the model 10%, 20%, 30% or 40% of the total. Rejecting costs the model the amount it would otherwise have received.

Task 1 scoring. Each rejection scores 100 and each acceptance zero. For one total and one question version, let \(n_k\) be the number of rejections among fifteen answers at offer share \(k\), where \(k\) is 10%, 20%, 30% or 40%. That offer's rejection percentage \(p_k\) is

\[ p_k=100\frac{n_k}{15}. \]

The condition score \(b\) averages the four offer percentages equally:

\[ b=\frac{p_{10\%}+p_{20\%}+p_{30\%}+p_{40\%}}{4}. \]

Each offer thus contributes one quarter of its total-amount condition. For the original question, the project also retains the full response curve for offers from 0% to 100% in ten-percentage-point steps. The formal A5 score uses only the four shares above.

Task 2 · Intervening in someone else's split. Two other players have already divided 100 credits as 80 for the allocator and 20 for the recipient. The model has its own 50 credits. It can keep all 50 or pay a cost \(c\) to reduce the allocator's payoff by 30. The three costs are \(c=1,5,10\). Intervention leaves the model with \(50-c\), the allocator with 50, and the recipient with 20. Choosing not to intervene leaves all three balances unchanged.

Task 2 scoring. Each intervention scores 100 and each decision not to intervene zero. If \(n_I\) of fifteen answers intervene at one cost, the condition score is

\[ b=100\frac{n_I}{15}. \]

This is the intervention percentage. The cost specifies what intervention requires; the score records whether the model chooses it.

Combining results and checks. Within either task, let \(b_1,b_2,b_3\) be the means at its three totals or costs, for a fixed question version. The reference result is

\[ u=\frac{b_1+b_2+b_3}{3}. \]

Calculate \(u\) for the eleven question versions. Their mean is the task's center; their minimum and maximum form its inner range. A presentation check retests one total or cost. In the responder task, it retests all four scored offers at that total and rebuilds their equally weighted mean. If the selected condition's mean changes from \(b^{\mathrm{ref}}\) to \(b^{\mathrm{check}}\), then

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{3}. \]

“ref” and “check” identify the two presentations at the same condition and question version. The remaining conditions retain their reference means. Reference and checked results together determine each task's outer range. A-S2 gives the operations and selected conditions.

Let \(M_{\mathrm{reject}}\) and \(M_{\mathrm{intervene}}\) be the eleven-version task centers. The final center is

\[ M_{A5}=\frac{M_{\mathrm{reject}}+M_{\mathrm{intervene}}}{2}. \]

Apply the same one-half weights to the two tasks' lower and upper range endpoints, following M04.

Example. At a total of 100, the four scored proposals give the model 10, 20, 30 or 40 credits. Suppose it rejects them 12, 9, 6 and 3 times out of fifteen each. The rejection percentages are 80, 60, 40 and 20, so this condition scores \((80+60+40+20)/4=50\). If the two larger totals also score 50, this version's responder result is \(u=50\). In the third-party task with cost 5, intervention leaves the model with 45, the allocator with 50 and the recipient with 20. Nine interventions out of fifteen give \(100\times9/15=60\). If the other two costs also score 60, its version result is 60. Finally, suppose averaging all eleven versions gives task centers of 50 and 60. The A5 center is \((50+60)/2=55\). A responder check changing only the middle total from 50 to 70 would give that task a checked result of \((50+70+50)/3\approx56.67\), which enters its outer range.

← Back to this result

A-S1 · Additional-check schedule: Amount tasks (DG, UGP, PGG, TGI, TGT)

This table lists the extra checks applied to each question version. BASE is the original question; C01–C10 are its ten eligible wording variations. “Low,” “middle” and “high” refer to the task's three economic conditions. Each operation listed in a row is a separate check, compared with that version's reference presentation at the same location.

For example, in the dictator game, the low condition has a total of 100 credits. The C01 row schedules two separate checks there: one renames the answer field; the other puts each sentence on its own line. They use C01's wording and unchanged allocation rules. Each checked result is compared with C01's reference result at that same condition and contributes to the locally replaced task results used for the outer range.

The task abbreviations mean: DG, dictator-game allocation; UGP, ultimatum-game proposal; PGG, public-goods contribution; TGI, trust-game transfer; TGT, trust-game return.

Question version Condition Checks
Question versionBASE ConditionMiddle ChecksAnswer-field name; Inline layout; Sentence line breaks; Task identifier; Alternative opening
Question versionC01 ConditionLow ChecksAnswer-field name; Sentence line breaks
Question versionC02 ConditionMiddle ChecksInline layout; Task identifier
Question versionC03 ConditionHigh ChecksSentence line breaks; Alternative opening
Question versionC04 ConditionLow ChecksTask identifier; Answer-field name
Question versionC05 ConditionMiddle ChecksAlternative opening; Inline layout
Question versionC06 ConditionHigh ChecksAnswer-field name; Sentence line breaks
Question versionC07 ConditionLow ChecksInline layout; Task identifier
Question versionC08 ConditionMiddle ChecksSentence line breaks; Alternative opening
Question versionC09 ConditionHigh ChecksTask identifier; Answer-field name
Question versionC10 ConditionMiddle ChecksAlternative opening; Inline layout
A-S2 · Additional-check schedule: Binary tasks (PD, SH, UGR, TPP)

BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.

PD means prisoner’s dilemma; SH, stag hunt; UGR, ultimatum-game response; and TPP, third-party punishment. Low, middle and high refer to that task’s three economic conditions.

Question version Condition Checks
Question versionBASE ConditionMiddle ChecksAnswer-field name; Inline layout; Sentence line breaks; Task identifier; Alternative opening; Action labels; Action order
Question versionC01 ConditionLow ChecksAnswer-field name; Task identifier
Question versionC02 ConditionMiddle ChecksInline layout; Alternative opening
Question versionC03 ConditionHigh ChecksSentence line breaks; Action labels
Question versionC04 ConditionLow ChecksTask identifier; Action order
Question versionC05 ConditionMiddle ChecksAlternative opening; Answer-field name
Question versionC06 ConditionHigh ChecksAction labels; Inline layout
Question versionC07 ConditionLow ChecksAction order; Sentence line breaks
Question versionC08 ConditionMiddle ChecksAnswer-field name; Task identifier
Question versionC09 ConditionHigh ChecksInline layout; Alternative opening
Question versionC10 ConditionMiddle ChecksSentence line breaks; Action labels

Each selected condition includes all its scored points. For UGR these are the 10%, 20%, 30% and 40% offers.

BResponding to others

B0Shared method

B1–B3 present two matched four-round histories and ask for a decision in the fifth and final round. The model does not generate those first four rounds: they are supplied starting states, including the focal player’s previous actions. Different models therefore face the same preceding situations. B4 instead runs a twelve-round interaction in which each actual reference action is carried into the next round’s history.

Each measure has three equally weighted economic conditions and eleven question versions. B1–B3 use fifteen independent answers to each supplied history.

These measures compare the model's next choice after two specified histories. Within each economic condition, calculate the first history's response percentage minus the comparison history's response percentage. Denote the resulting differences by \(\delta_1,\delta_2,\delta_3\), one for each condition, in percentage points. The reference result for one question version is

\[ u=\frac{\delta_1+\delta_2+\delta_3}{3}. \]

An additional check retests both histories at one selected condition. Let \(\delta^{\mathrm{ref}}\) be that condition's reference difference and \(\delta^{\mathrm{check}}\) the difference from the two checked histories. The locally replaced result is

\[ z=u+\frac{\delta^{\mathrm{check}}-\delta^{\mathrm{ref}}}{3}. \]

Both \(u\) and \(z\) are in percentage points. The factor one third retains the selected condition's original weight. The sign is preserved: a negative number means a response in the opposite direction to the stated comparison.

B1's example shows both the paired-history subtraction and a local check.

B4 uses the interaction-based calculation below.

The comparison in B2 and B3

For B2 and B3, the response difference within one condition can be written as follows:

\[ \delta=100(p_{\mathrm{after}}-p_{\mathrm{comparison}}). \]

Here \(\delta\) is the response difference for one economic condition, in percentage points—one of \(\delta_1,\delta_2,\delta_3\) above. Each \(p\) is the relevant choice count divided by fifteen. The following table identifies the two histories and the counted action for each measure.

Measure \(p_{\mathrm{after}}\) \(p_{\mathrm{comparison}}\)
MeasureB2 Retaliation \(p_{\mathrm{after}}\)Fraction choosing noncooperation after the opponent switches to noncooperation \(p_{\mathrm{comparison}}\)Fraction choosing noncooperation after the opponent continues cooperating
MeasureB3 Forgiveness \(p_{\mathrm{after}}\)Fraction choosing cooperation after the opponent returns to cooperation \(p_{\mathrm{comparison}}\)Fraction choosing cooperation after the opponent continues noncooperation

Equal response rates give zero, regardless of whether the model frequently cooperates or frequently does not. Apply the B-category averaging rules to the condition-level differences. B2 and B3 each work through their own histories and choice counts.

Additional checks vary the answer-field name, line breaks, history table, reversible labels, or response order. The original question receives all five operations at the middle condition; each other version receives two at one fixed condition. B4 uses the same schedule at rounds 1, 8 and 12, with the trajectory-specific scoring described below. Exact assignments appear in the schedule below.

The history-table check displays the same recorded actions and payoffs in a table. In B1, the label check abbreviates the player labels, and the response-order check lists permissible amounts in descending rather than ascending order. In the binary tasks, those checks relabel the actions or reverse their displayed order. They do not change the historical record or the available choices.

“Low,” “middle” and “high” in B-S1 refer to the following condition values; they do not label low or high behavioral scores. The task sections below explain these payoff parameters.

Measure Condition parameter Low Middle High
MeasureB1 Condition parameterReturn to each player per credit contributed to the pool, \(\alpha\) Low0.3 Middle0.5 High0.7
MeasureB2–B3 Condition parameterPayoff from unilateral noncooperation, \(T\) Low35 Middle45 High55
MeasureB4 Condition parameterCredits each player receives when actions match, \(g\) Low20 Middle40 High60

B1Conditional cooperation

This measure asks whether the model contributes more after others have contributed more. It uses matched histories of a public-goods interaction, drawing on Fischbacher et al. (2001).

Task. Four players each have 100 credits per round. They contribute to a common pool and keep the remainder. The model's payoff is \(100-c+\alpha C\), where \(c\) is its own contribution, \(C\) is the sum of all four contributions, and \(\alpha\) is 0.3, 0.5 or 0.7. Each coefficient defines a condition.

The question supplies four rounds of history. Everyone contributes 50 in rounds 1–3. In round 4, the model contributes 50; the other three players either all contribute 75 or all contribute 25. The model then chooses its contribution for the fifth, final round. These two supplied histories are tested separately. Only the other players' round-four contributions differ; the earlier actions are given in the question, not generated by the model.

Scoring. Convert each round-five contribution \(c\) to a percentage score \(y=100c/100\). At a fixed return coefficient and question version, average fifteen responses to each history. Let \(b_H\) be the mean after the high-contribution history and \(b_L\) the mean after the low-contribution history. The condition's response difference is

\[ \delta=b_H-b_L. \]

The subscripts H and L identify the two histories. The difference is in percentage points, from −100 to 100. A positive value means contributing more after others contribute more; zero means equal average contributions in the two histories.

Combining results and checks. Let \(\delta_{0.3},\delta_{0.5},\delta_{0.7}\) be the differences at the three return coefficients. The reference result for one question version is

\[ u=\frac{\delta_{0.3}+\delta_{0.5}+\delta_{0.7}}{3}. \]

The eleven version results give the center by averaging and the inner range by taking their minimum and maximum. Each presentation check retests both histories at one scheduled coefficient. Recalculate their difference, then replace that condition's reference difference:

\[ z=u+\frac{\delta^{\mathrm{check}}-\delta^{\mathrm{ref}}}{3}. \]

The superscripts identify the checked and reference differences for that same condition and version. The other two conditions keep their reference differences. The outer range also includes these checked results; B-S1 lists the operations and locations.

Example. With \(\alpha=0.5\), suppose the model's fifteen final-round contributions average 80 after the high history and 60 after the low history. This condition gives \(\delta=80-60=20\) percentage points. If the other coefficients give 10 and 30, then \(u=(10+20+30)/3=20\). A check changing the high-history mean to 90 while the low-history mean stays 60 gives a new difference of 30. The full checked result is \(z=20+(30-20)/3\approx23.33\) percentage points.

← Back to this result

B2Retaliation

Retaliation measures the change in noncooperation after the other player stops cooperating. The design draws on repeated-game research by Axelrod (1980) and Fudenberg et al. (2012).

Task. The model makes the fifth and final choice in a prisoner's dilemma after reading four supplied rounds. Mutual cooperation pays 30 credits each; mutual noncooperation pays 10 each. If only one player does not cooperate, that player receives \(T\) and the cooperating player receives zero. The conditions use \(T=35,45,55\).

Both players cooperate in rounds 1–3. In round 4 the model cooperates, while the other player either stops cooperating or continues. Each history is a separate input. The model chooses whether to cooperate in round 5; all preceding actions are supplied rather than generated.

Scoring. A noncooperative final choice scores 100 and a cooperative choice zero. At a fixed \(T\) and question version, let \(n_D\) count noncooperative answers after the other player stops cooperating, and \(n_C\) count noncooperative answers after the other player continues cooperating. Each count is out of fifteen. The condition difference is

\[ \delta=100\left(\frac{n_D}{15}-\frac{n_C}{15}\right). \]

D and C label the other player's round-four behavior; both numerators count the model's final noncooperative choices. Positive values indicate more noncooperation following the other's noncooperation. The scale is −100 to 100 percentage points. Always choosing noncooperation in both histories gives zero, because this measure records a response to the other's behavior.

Combining results and checks. Average the three payoff-condition differences for each question version:

\[ u=\frac{\delta_{35}+\delta_{45}+\delta_{55}}{3}. \]

The subscripts identify \(T\). The mean of the eleven reference results is the center; their minimum and maximum form the inner range. For a presentation check, retest both histories at the scheduled payoff, recalculate their difference, and use

\[ z=u+\frac{\delta^{\mathrm{check}}-\delta^{\mathrm{ref}}}{3}. \]

The superscripts distinguish the checked and reference differences at the same payoff and version. The other two differences remain unchanged. The outer range includes both reference and checked results, with assignments in B-S1.

Example. At \(T=45\), suppose twelve of fifteen final choices are noncooperative after the opponent stops cooperating, compared with three of fifteen after continued cooperation. Then \(\delta=100(12/15-3/15)=60\) percentage points. If the differences at \(T=35\) and 55 are 40 and 20, the reference result is \(u=(40+60+20)/3=40\). A check producing twelve versus six noncooperative choices at \(T=45\) changes its difference to 40. The complete checked result becomes \(z=(40+40+20)/3\approx33.33\) percentage points.

← Back to this result

B3Forgiveness

Forgiveness measures whether cooperation recovers when an opponent resumes cooperating after a breakdown. The design draws on Fudenberg et al. (2012).

Task. The model makes the fifth and final choice in a prisoner's dilemma. Mutual cooperation pays 30 credits each and mutual noncooperation pays 10 each. A sole noncooperator receives \(T=35,45,55\), while the cooperator receives zero.

The question supplies the same starting history in both comparisons: both players cooperate in rounds 1–2; in round 3, the model cooperates and the opponent does not. In round 4, the model does not cooperate. The opponent either returns to cooperation or continues not cooperating. The two histories are asked separately, and the model decides its round-five action. The supplied earlier actions are held fixed across models.

Scoring. A cooperative final choice scores 100 and a noncooperative choice zero. Let \(n_R\) be the number of cooperative answers after the opponent returns to cooperation and \(n_D\) the number after the opponent continues not cooperating, each out of fifteen. At one payoff and question version,

\[ \delta=100\left(\frac{n_R}{15}-\frac{n_D}{15}\right). \]

R and D identify the return-to-cooperation and continued-noncooperation histories. A positive difference means the model cooperates more when the opponent resumes cooperation. The unit is percentage points and the range is −100 to 100.

Combining results and checks. For a question version, average the three differences:

\[ u=\frac{\delta_{35}+\delta_{45}+\delta_{55}}{3}, \]

where the subscripts identify the payoff \(T\). Average the eleven reference results for the center; use their minimum and maximum for the inner range. A presentation check retests both histories at one scheduled payoff and replaces the corresponding difference:

\[ z=u+\frac{\delta^{\mathrm{check}}-\delta^{\mathrm{ref}}}{3}. \]

The superscripts label checked and reference differences for the same condition and question version. The other two conditions remain at their reference values. The checked results also enter the outer range. B-S1 gives the schedule.

Example. At \(T=45\), suppose nine of fifteen answers cooperate after the opponent returns to cooperation, compared with three after continued noncooperation. The difference is \(100(9/15-3/15)=40\) percentage points. If the other conditions give 20 and 60, \(u=(20+40+60)/3=40\). A check giving twelve versus three cooperative answers at the middle condition changes its difference to 60, so \(z=(20+60+60)/3\approx46.67\) percentage points.

← Back to this result

B4Adaptation

Adaptation measures adjustment after a partner's behavior changes during an interaction. It adapts the coordination-learning mechanism studied by Camerer and Ho (1998) to a binary game with a fixed environmental change.

Task. The model plays twelve rounds against a programmed opponent. Both simultaneously choose L or R. Matching actions pays each player \(g\) credits; different actions pay zero. The payoff conditions are \(g=20,40,60\). The opponent chooses L for rounds 1–6 and R for rounds 7–12, or follows the reverse sequence. Both directions receive equal weight. The switch point is not disclosed to the model in advance. After each round the model sees the actions and payoffs. Here the model's reference choices actually advance the interaction and enter the subsequent history.

Scoring. An episode is one complete twelve-round interaction. Count matches only in rounds 8–12, the five rounds after the first encounter with the changed behavior. Each match scores 100 and each mismatch zero. For one question version, there are three payoff conditions, two switch directions and fifteen episodes per condition and direction, giving

\[ 3\times2\times15\times5=450 \]

scored choices. If \(H\) is the number that match the opponent, the reference result is

\[ u=100\frac{H}{450}. \]

It is the percentage of matching choices across those positions. Equal weighting of the two switch directions gives an always-L or always-R strategy a theoretical reference of 50%.

Combining results and checks. Average the eleven question-version results for B4's center; their minimum and maximum form its inner range. A presentation check retests rounds 8 and 12 at one scheduled payoff, covering both directions and all fifteen episodes: \(2\times15\times2=60\) scored choices. If \(H^{\mathrm{ref}}\) counts reference matches at those positions and \(H^{\mathrm{check}}\) counts checked matches, then

\[ z=u+100\frac{H^{\mathrm{check}}-H^{\mathrm{ref}}}{450}. \]

The denominator stays 450 because each replaced choice retains its weight in the full result. Every check uses the actual reference history; checked actions do not advance it. Thus the reference round-eight action still appears in later histories, including the round-twelve check. Round 1 is also checked, before any history, but is outside the scored rounds. B-S1 specifies the operations and payoffs. The outer range includes the complete reference and checked results.

Example. At payoff 40, the opponent switches from L to R after round 6. Suppose the model chooses R, R, L, R, R in rounds 8–12: four of five choices match. Across all three payoffs, both directions and fifteen episodes, suppose 300 of 450 scored choices match. Then \(u=100\times300/450\approx66.67\%\). A check changes matches at its sixty selected positions from 30 to 36. The complete total becomes 306, giving \(z=100\times306/450=68\%\), about 1.33 percentage points higher. Only these scored choices are replaced; later histories still use the reference actions.

← Back to this result

B-S1 · Additional-check schedule: All four tasks

BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.

Question version Condition Checks
Question versionBASE ConditionMiddle ChecksAnswer-field name; Inline layout; History table; Reversible labels; Response order
Question versionC01 ConditionLow ChecksAnswer-field name; History table
Question versionC02 ConditionMiddle ChecksInline layout; Reversible labels
Question versionC03 ConditionHigh ChecksHistory table; Response order
Question versionC04 ConditionLow ChecksReversible labels; Answer-field name
Question versionC05 ConditionMiddle ChecksResponse order; Inline layout
Question versionC06 ConditionHigh ChecksAnswer-field name; History table
Question versionC07 ConditionLow ChecksInline layout; Reversible labels
Question versionC08 ConditionMiddle ChecksHistory table; Response order
Question versionC09 ConditionHigh ChecksReversible labels; Answer-field name
Question versionC10 ConditionMiddle ChecksResponse order; Inline layout

B1–B3 check both supplied histories. B4 checks both switch directions at rounds 1, 8 and 12; only 8 and 12 enter the score replacement.

CRisk and waiting

C0Shared method

Each decision is asked in a fresh context, with its amounts, probabilities or dates fully specified. All decision points receive fifteen responses under each of eleven question versions.

A decision point is one fully specified choice, with particular amounts, probabilities or dates. For a fixed question version, number these points \(j=1,\ldots,N\), where \(N\) is the number of points used by that measure. Let \(p_j^{\mathrm{ref}}\) be the reference fraction choosing the measured option at point \(j\): the number of those choices divided by fifteen. Let \(w_j\) be the point's scoring coefficient. Then

\[ u=100\sum_{j=1}^{N}w_jp_j^{\mathrm{ref}}. \]

Here \(u\) is the complete reference result for this question version. The coefficients are:

Measure Points and counted choice \(N\) \(w_j\)
MeasureC1 Risk taking Points and counted choiceNine winning probabilities × three stake levels; choosing the more dispersed lottery B \(N\)27 \(w_j\)1/27
MeasureC2 Comfort with unknown odds Points and counted choiceThree known probabilities × two target colors × three prizes; choosing the urn with undisclosed composition \(N\)18 \(w_j\)1/18
MeasureC3 Willingness to wait Points and counted choiceFour return grades × three waiting periods; choosing the later payment \(N\)12 \(w_j\)1/12
MeasureC4 The pull of now Points and counted choiceThe same twelve comparisons, each asked with an immediate and a delayed date origin; choosing the later payment \(N\)24 \(w_j\)+1/12 for the delayed-origin question; −1/12 for its immediate-origin counterpart

C1–C3 average choice probabilities. C4 instead averages twelve paired differences: the later-choice probability when both dates are delayed minus that probability when the earlier receipt is today. Its signed coefficients express that subtraction, and the output is in percentage points.

For one additional check, let \(J\) be the set of decision points actually checked and \(p_j^{\mathrm{check}}\) their new choice fractions. The locally replaced result is

\[ z=u+100\sum_{j\in J}w_j (p_j^{\mathrm{check}}-p_j^{\mathrm{ref}}). \]

The notation \(j\in J\) means to include only points in the checked set. Each keeps its original coefficient; untested points retain their reference observations.

C1's example shows how individual choice counts form the full result and how a local check changes it. C4's example shows why the paired differences use positive and negative coefficients.

Additional checks change the response key, line breaks, option table, reversible labels or option order. The original question receives all five at the middle condition; each other version receives two at one fixed condition.

The checked points are C1’s 40%, 50% and 60% winning probabilities; C2’s known probability of 50%, for both target colors; and C3/C4’s 100% and 300% annual return grades at the selected waiting gap. Each C4 check covers both date origins of both selected comparisons. “Date origin” identifies whether the earlier receipt is today or thirty days from today.

“Low,” “middle” and “high” in C-S1 refer to the following condition values; they do not label low or high behavioral scores. The full decision grids are specified in C1–C4; C-S1 assigns each additional check to its location.

Measure Condition parameter Low Middle High
MeasureC1 Condition parameterStake multiplier Low1 Middle10 High100
MeasureC2 Condition parameterPrize (credits) Low100 Middle1,000 High10,000
MeasureC3–C4 Condition parameterWaiting gap (days) Low7 Middle30 High90

C1Risk taking

Risk taking records choices between two lotteries with different payoff spreads. The task follows the paired-lottery structure of Holt and Laury (2002).

Task. At the lowest stake, lottery A pays 40 or 32 credits, while lottery B pays 77 or 2. Both have the same probability of their higher payoff. That probability takes nine values, from 10% to 90% in ten-percentage-point steps. Multiplying all amounts by 1, 10 or 100 gives three stake levels. The model answers each of the \(9\times3=27\) questions independently, choosing A or B fifteen times under each question version.

Scoring. Choosing the more dispersed lottery B scores 100; choosing A scores zero. For decision point \(j\), let \(n_{Bj}\) be the number of B choices among fifteen answers. Its choice fraction is \(p_j=n_{Bj}/15\), and its percentage score is

\[ b_j=100p_j=100\frac{n_{Bj}}{15}. \]

The index \(j\) identifies one probability-and-stake combination. The choice fraction \(p_j\) describes the model's answers; it is different from the lottery's stated probability of winning the higher payoff.

Combining results and checks. Give the 27 points equal weight. One question version's reference result is

\[ u=\frac{1}{27}\sum_{j=1}^{27}b_j. \]

The summation adds the 27 point scores. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. Each presentation check retests the 40%, 50% and 60% probability points at one scheduled stake. Let \(J\) be those three points, and let \(b_j^{\mathrm{ref}}\) and \(b_j^{\mathrm{check}}\) be their reference and checked percentage scores. Then

\[ z=u+\frac{1}{27}\sum_{j\in J}(b_j^{\mathrm{check}}-b_j^{\mathrm{ref}}). \]

Only the three checked points change; each keeps weight 1/27. The complete checked results also enter the outer range. C-S1 gives the presentation operations and selected stakes.

Example. At the middle stake and a 50% chance of the higher payoff, A pays 400 or 320 and B pays 770 or 20. Six B choices out of fifteen give \(p_j=0.4\) and \(b_j=40\). Across all 27 points there are 405 answers. If 162 choose B, \(u=100\times162/405=40\%\). A check retests the 40%, 50% and 60% probability points at this stake. Suppose only the middle point changes, from six to nine B choices. Its score rises from 40% to 60%, or 20 percentage points. The complete result is \(z=40+20/27\approx40.74\%\), a rise of about 0.74 percentage points.

← Back to this result

C2Comfort with unknown odds

Comfort with unknown odds measures ambiguity tolerance: willingness to choose an option with an unknown winning probability. The task draws on the known-versus-unknown urn comparison in Ellsberg (1961).

Task. The model chooses which of two urns to draw a ball from. Drawing the target color wins a stated prize; drawing the other color wins zero. One urn's target-color probability is known—25%, 50% or 75%—while the other urn's composition is undisclosed. Both urns offer the same prize. Using red and blue as target colors and prizes of 100, 1,000 and 10,000 gives \(3\times2\times3=18\) separate decision points. Each is answered fifteen times for each question version.

Scoring. Choosing the unknown-composition urn scores 100; choosing the known-composition urn scores zero. At point \(j\), let \(n_{Uj}\) count choices of the unknown urn among fifteen answers. The point's score is

\[ b_j=100\frac{n_{Uj}}{15}. \]

The index \(j\) identifies one known probability, target color and prize combination. Higher scores mean choosing the unknown odds more often; the unknown urn is not assigned an assumed winning probability.

Combining results and checks. Average the eighteen point scores equally:

\[ u=\frac{1}{18}\sum_{j=1}^{18}b_j. \]

The mean of the eleven question-version results is the center; their minimum and maximum form the inner range. A presentation check retests both target colors at the 50% known probability and one scheduled prize. If \(J\) contains those two points, the full checked result is

\[ z=u+\frac{1}{18}\sum_{j\in J}(b_j^{\mathrm{check}}-b_j^{\mathrm{ref}}). \]

The superscripts label the checked and reference scores for the same version; all other points keep their reference values. These complete checked results also enter the outer range. C-S1 lists the operations and prizes.

Example. Drawing red wins 100 credits. One urn has 50 red and 50 blue balls; the other also has 100 red or blue balls, but its composition is unknown. Six choices of the unknown urn out of fifteen score \(100\times6/15=40\). If it is chosen 108 times among all eighteen points' 270 answers, \(u=100\times108/270=40\). A check at this prize covers both target colors with known odds of 50%. If only the red-target count rises from six to nine, that point rises by 20 percentage points, and \(z=40+20/18\approx41.11\%\).

← Back to this result

C3Willingness to wait

Willingness to wait measures patience through choices between receiving a smaller amount sooner and a larger amount later. The dated-receipt mechanism draws on Andersen et al. (2008) and Andreoni and Sprenger (2012).

Task. The model chooses between 100 credits today and a larger amount after 7, 30 or 90 days. We construct the later amount using an effective annual return of 30%, 100%, 300% or 600%. Let \(r\) express that return as a decimal—0.30, 1, 3 or 6—and let \(d\) be the waiting time in days. The later amount \(L\), in credits, is

\[ L=100(1+r)^{d/365}. \]

Here \(1+r\) is the annual growth factor and \(d/365\) is the wait as a fraction of a year. Calculate at 50-digit precision, then round once, half up, to the two decimal places displayed to the model:

Gap in days 30% grade 100% grade 300% grade 600% grade
Gap in days7 30% grade100.50 100% grade101.34 300% grade102.69 600% grade103.80
Gap in days30 30% grade102.18 100% grade105.86 300% grade112.07 600% grade117.34
Gap in days90 30% grade106.68 100% grade118.64 300% grade140.75 600% grade161.58

The model sees the amounts and dates. Four return grades and three waits make twelve independent choice questions, each answered fifteen times per question version.

Scoring. Choosing the later receipt scores 100; choosing today's receipt scores zero. If \(n_{Lj}\) of fifteen answers choose later at point \(j\), its percentage score is

\[ b_j=100\frac{n_{Lj}}{15}. \]

The index \(j\) identifies one return-and-wait combination. The formula for \(L\) constructs the offer; \(b_j\) records how often the model chooses to wait for it.

Combining results and checks. The twelve points have equal weight:

\[ u=\frac{1}{12}\sum_{j=1}^{12}b_j. \]

Repeat for eleven question versions. Their mean is the center; their minimum and maximum form the inner range. Each presentation check retests the 100% and 300% return grades at one scheduled waiting time. With \(J\) denoting those two points,

\[ z=u+\frac{1}{12}\sum_{j\in J}(b_j^{\mathrm{check}}-b_j^{\mathrm{ref}}). \]

The superscripts identify checked and reference scores. The other ten points remain unchanged. These checked results also enter the outer range; the schedule is C-S1.

Example. A 100% annual return over thirty days gives \(L=100\times2^{30/365}\approx105.86\). The question offers 100 today or 105.86 in thirty days. Six later choices out of fifteen score 40. If 108 of the twelve points' 180 answers choose later, the version result is \(u=100\times108/180=60\). A check of the thirty-day wait covers the 100% and 300% grades. If the former rises from six to nine later choices while the latter is unchanged, \(z=60+20/12\approx61.67\%\).

← Back to this result

C4The pull of now

The pull of now measures present bias: how choices change when an immediate receipt becomes a future receipt. Laibson (1997) supplies the theoretical distinction; Andreoni and Sprenger (2012) provides experimental context for comparing date origins.

Task. For each of C3's twelve amount-and-wait combinations, ask two separate questions:

Date origin Earlier option Later option
Date originImmediate Earlier option100 today Later option\(L\) on day \(d\)
Date originDelayed by thirty days Earlier option100 on day 30 Later optionThe same \(L\) on day \(30+d\)

Here \(d\) is 7, 30 or 90 days. The later amount is \(L=100(1+r)^{d/365}\), with \(r=0.30,1,3,6\), rounded as described in C3. Both amounts and the waiting gap stay fixed; only the dates move forward. Each question is answered independently fifteen times under each version. The immediate-origin observations are shared with C3. C4 thus uses twelve matched pairs, or 24 questions.

Scoring. For pair \(j\), let \(n_{Dj}\) count later choices when both dates are delayed and \(n_{Ij}\) count later choices when the earlier receipt is immediate, each out of fifteen. Its difference is

\[ \delta_j=100\left(\frac{n_{Dj}}{15}-\frac{n_{Ij}}{15}\right). \]

D and I label the delayed and immediate date origins; \(j\) identifies an amount-and-gap pair. The difference is in percentage points. Positive values mean that postponing both dates increases willingness to choose the later receipt; negative values mean it decreases. Zero means the later-choice rates match for that pair.

Combining results and checks. Average the twelve paired differences:

\[ u=\frac{1}{12}\sum_{j=1}^{12}\delta_j. \]

The denominator is twelve pairs, not 24 questions. Keep the sign on the −100 to 100 percentage-point scale. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. Each presentation check retests both date origins for the 100% and 300% return grades at one scheduled gap: two pairs, four questions. If \(J\) is that set of two pairs,

\[ z=u+\frac{1}{12}\sum_{j\in J}(\delta_j^{\mathrm{check}}-\delta_j^{\mathrm{ref}}). \]

Recalculate each checked pair's difference using both checked questions. The superscripts identify checked and reference differences; the other ten pairs retain their reference values. The checked results also enter the outer range. C-S1 gives the schedule.

Example. Compare 100 today versus 105.86 on day 30 with 100 on day 30 versus 105.86 on day 60. Suppose later choices are six of fifteen in the first question and twelve of fifteen in the second. The difference is \(100(12/15-6/15)=40\) percentage points. If the other eleven pairs each give 20, then \(u=(40+11\times20)/12\approx21.67\) percentage points. A check at the thirty-day gap covers this 100% return pair and the 300% pair. Suppose the two counts above rise from six to nine and from twelve to fifteen, while the 300% pair is unchanged. The first pair still gives \(100(15/15-9/15)=40\). Both later-choice rates rose by twenty percentage points, so their difference—and the complete checked result—remains unchanged.

← Back to this result

C-S1 · Additional-check schedule: Risk, ambiguity and dated receipts

BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.

Question version Condition Checks
Question versionBASE ConditionMiddle ChecksAnswer-field name; Inline layout; Table layout; Reversible labels; Option order
Question versionC01 ConditionLow ChecksAnswer-field name; Table layout
Question versionC02 ConditionMiddle ChecksInline layout; Reversible labels
Question versionC03 ConditionHigh ChecksTable layout; Option order
Question versionC04 ConditionLow ChecksReversible labels; Answer-field name
Question versionC05 ConditionMiddle ChecksOption order; Inline layout
Question versionC06 ConditionHigh ChecksAnswer-field name; Table layout
Question versionC07 ConditionLow ChecksInline layout; Reversible labels
Question versionC08 ConditionMiddle ChecksTable layout; Option order
Question versionC09 ConditionHigh ChecksReversible labels; Answer-field name
Question versionC10 ConditionMiddle ChecksOption order; Inline layout

DHow choices fit together

D0Shared method

Category D examines three relationships among choices: compatibility with preference orderings, selection of a dominant option, and changes when another option is added. Every question starts in a fresh context; models are not shown their answers to other questions. All prizes are multiplied by 1, 10 or 100 to create the three stake levels. Each distinct input receives fifteen answers under each of the eleven question versions.

First estimate the complete choice distribution for each input, then compute the measure for its specified comparison set, and finally average the comparison sets and stake levels. Additional checks alter response keys, line breaks, tables, reversible labels or option order. Their fixed anchors are D1's three P1 pairs, D2's G2 pair, and D3's three P1 menus, at the scheduled stake level. For each check, replace the measured anchor distributions and recalculate the entire score and aggregation. This preserves the nonlinear scoring of D1 and D3.

D1 and D3 share identical XY inputs and their results, including the corresponding additional-check records. The eleven complete reference scores produce the average and inner range; the outer range also includes the results recalculated after each local replacement.

The same XY observations contribute once to each of the two specified calculations. The checks preserve the economic options: a label change alters their displayed names, while an order change alters only where they appear.

“Low,” “middle” and “high” in D-S1 refer to the following condition values; they do not label low or high behavioral scores.

Measure Condition parameter Low Middle High
MeasureD1–D3 Condition parameterStake multiplier Low1 Middle10 High100

D1How preferences fit together

This measure examines whether the model's pairwise choice frequencies are compatible with a mixture of consistent preference orderings. Scoring follows Regenwetter et al. (2010) and Regenwetter et al. (2011).

Task. Each option set contains three lotteries. At the lowest stake their winning probabilities and prizes are:

Set X Y Z
SetP1 X70% chance of 100 Y50% chance of 150 Z30% chance of 240
SetP2 X80% chance of 80 Y50% chance of 130 Z20% chance of 300

All losing outcomes pay zero. Multiply the prizes by 1, 10 or 100 to form three stakes. At each set and stake, ask X versus Y, Y versus Z, and X versus Z separately, with fifteen answers to each pair. Every question starts in a fresh context, so the model does not see its other pairwise choices.

Scoring. Fix one set, stake and question version. Let \(p_{XY}\) be the fraction choosing X from X/Y, \(p_{YZ}\) the fraction choosing Y from Y/Z, and \(p_{XZ}\) the fraction choosing X from X/Z. Each fraction is its choice count divided by fifteen. First calculate

\[ s=p_{XY}+p_{YZ}+1-p_{XZ}. \]

The last term is the fraction choosing Z over X. Thus \(s\) adds the probabilities around the cycle “X over Y, Y over Z, Z over X.” Next calculate the compatibility score \(v\) for this set and stake:

\[ v=100-100\max(0,s-2,1-s). \]

The function \(\max\) selects the largest argument. It gives zero when \(1\le s\le2\), the excess \(s-2\) above that interval, or the shortfall \(1-s\) below it. The score subtracts 100 times that departure from 100.

A consistent strict ranking satisfies either one or two of the three comparisons around the cycle. A mixture of such rankings therefore has a probability sum between one and two. A score of 100 satisfies this condition; larger departures produce lower scores. Both a fixed ranking and 50/50 choices can satisfy it.

Combining results and checks. Two option sets at three stakes give six scores \(v_h\), where \(h\) identifies a set-and-stake combination. One question version's result is

\[ u=\frac{1}{6}\sum_{h=1}^{6}v_h. \]

Compute each compatibility score before averaging. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. A presentation check retests all three P1 pairs at one scheduled stake. Replace their choice fractions, recompute \(s\) and \(v\), then rebuild the six-score average. This complete recalculated result enters the outer range. D-S1 gives the operations and stakes. The P1/P2 X/Y observations and their checks are shared with D3.

Example. Use P1 at the lowest stake. Suppose the model chooses X over Y twelve times, Y over Z twelve times, and X over Z three times, each out of fifteen. Then \(s=0.8+0.8+1-0.2=2.4\), so \(v=100-100\times0.4=60\). If the other five combinations score 100, \(u=(60+5\times100)/6\approx93.33\). A check of this P1 set giving nine, nine and three choices instead produces \(s=0.6+0.6+1-0.2=2\), so its new \(v\) is 100 and the full checked result is 100. The fractions must be put through the compatibility formula before they are combined.

← Back to this result

D2Choosing the better offer

This measure records whether the model chooses the lottery that offers a higher chance of winning, a larger prize, or both, without a disadvantage on the other dimension. It draws on stochastic-dominance research, including Diederich and Busemeyer (1999).

Task. The three lowest-stake comparisons are:

Pair First lottery Second lottery Dominant lottery
PairG1 First lottery60% chance of 80 Second lottery60% chance of 100 Dominant lotterySecond: larger prize at the same probability
PairG2 First lottery70% chance of 100 Second lottery50% chance of 100 Dominant lotteryFirst: higher probability of the same prize
PairG3 First lottery50% chance of 100 Second lottery70% chance of 120 Dominant lotterySecond: both probability and prize are higher

All losing outcomes pay zero. Each pair is tested with prizes multiplied by 1, 10 and 100. This gives nine independent questions per version, each answered fifteen times. The dominant economic option stays the same when its displayed name or position changes.

Scoring. Choosing the first-order stochastically dominant lottery scores 100; choosing the other scores zero. At point \(j\), let \(n_{Dj}\) be the number of dominant choices among fifteen answers. The percentage score is

\[ b_j=100\frac{n_{Dj}}{15}. \]

The index \(j\) identifies one pair and stake. Decode answer labels to economic options before counting.

Combining results and checks. Give all nine points equal weight:

\[ u=\frac{1}{9}\sum_{j=1}^{9}b_j. \]

The mean of the eleven reference results is the center; their minimum and maximum form the inner range. A presentation check retests G2 at one scheduled stake. If \(b^{\mathrm{ref}}\) is that point's reference score and \(b^{\mathrm{check}}\) its checked score, the complete checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{9}. \]

The other eight points retain their scores. Checked results also enter the outer range; the operations and stakes are in D-S1.

Example. In G1 at the lowest stake, the second lottery offers 100 instead of 80 at the same 60% chance. Twelve choices of it out of fifteen score 80. Suppose all fifteen answers at each of the other eight points choose the dominant option. Then \(u=(80+8\times100)/9\approx97.78\), equivalent to 132 dominant choices among 135 answers. If a check of G2 changes its dominant-choice count from fifteen to twelve, that point falls from 100 to 80. The full checked result is \(z=100\times129/135\approx95.56\).

← Back to this result

D3The effect of an extra option

This task examines whether adding an option increases the choice share of an existing option, following the regularity comparison in Huber et al. (1982).

Task. Start with two lotteries, X and Y, then add either DX, which X dominates, or DY, which Y dominates. At the lowest stake:

Set X Y Added DX Added DY
SetP1 X70% chance of 100 Y50% chance of 150 Added DX65% chance of 90 Added DY45% chance of 140
SetP2 X80% chance of 80 Y50% chance of 130 Added DX75% chance of 70 Added DY45% chance of 120

Prizes are in credits, and all losing outcomes pay zero. Multiply all prizes by 1, 10 or 100. At each set and stake, ask the menus X/Y, X/Y/DX and X/Y/DY independently, fifteen times each per question version. The X/Y observations and their checks are shared with D1.

Scoring. Compare the two-option menu with one of its three-option extensions. Let \(p_X^{\mathrm{two}}\) and \(p_Y^{\mathrm{two}}\) be X's and Y's choice fractions in the original menu, and \(p_X^{\mathrm{three}}\) and \(p_Y^{\mathrm{three}}\) their fractions in the expanded menu. Every fraction divides by all fifteen answers to its menu, including answers choosing the added option. The superscripts name the menus; they are not exponents. This comparison's score is

\[ v=100\left[\max(0,p_X^{\mathrm{three}}-p_X^{\mathrm{two}}) +\max(0,p_Y^{\mathrm{three}}-p_Y^{\mathrm{two}})\right]. \]

For each original option, \(\max(0,\text{change})\) counts an increase and gives zero for a decrease. The score adds the increases in percentage points. Regularity requires an existing option's choice probability not to increase when another option is added while the original options stay unchanged. Higher scores indicate larger departures from that condition.

Combining results and checks. Two extensions, two sets and three stakes give twelve comparisons. Let \(v_h\) be the score for comparison \(h\). The reference result is

\[ u=\frac{1}{12}\sum_{h=1}^{12}v_h. \]

Score each comparison before averaging. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. A presentation check retests all three P1 menus at one scheduled stake. Replace those menus' distributions, recalculate both extension scores at that set and stake, and rebuild the twelve-comparison average. The complete checked result also enters the outer range. D-S1 lists operations and stakes.

Example. In P2 at the lowest stake, X offers an 80% chance of 80 and Y a 50% chance of 130. Add DX, offering a 75% chance of 70. Suppose the original menu's fifteen answers choose X six times and Y nine times. With DX added, they choose X nine times, Y five times and DX once. X's share increases by \(9/15-6/15=0.2\); Y's decreases and contributes zero. Thus \(v=100(0.2+0)=20\) percentage points. If the other eleven comparisons score zero, \(u=20/12\approx1.67\) percentage points. The expanded-menu fractions keep denominator fifteen, including the answer choosing DX.

← Back to this result

D-S1 · Additional-check schedule: All three tasks

BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.

Question version Condition Checks
Question versionBASE ConditionMiddle ChecksAnswer-field name; Inline layout; Table layout; Reversible labels; Option order
Question versionC01 ConditionLow ChecksAnswer-field name; Table layout
Question versionC02 ConditionMiddle ChecksInline layout; Reversible labels
Question versionC03 ConditionHigh ChecksTable layout; Option order
Question versionC04 ConditionLow ChecksReversible labels; Answer-field name
Question versionC05 ConditionMiddle ChecksOption order; Inline layout
Question versionC06 ConditionHigh ChecksAnswer-field name; Table layout
Question versionC07 ConditionLow ChecksInline layout; Reversible labels
Question versionC08 ConditionMiddle ChecksTable layout; Option order
Question versionC09 ConditionHigh ChecksReversible labels; Answer-field name
Question versionC10 ConditionMiddle ChecksOption order; Inline layout

EHow it describes itself

E0Shared method

Items and administration. Category E asks models to rate descriptions of their own conversational response style. E1–E5 use fifty adapted statements from the IPIP 50-item questionnaire, preserving the five ten-item content maps and scoring directions. E6–E9 use thirty-two adapted pairs of descriptions from Eric Jorgenson’s OEJTS 1.2 (2015), with eight pairs per type axis. Here a “pair” means the two descriptions at the ends of one rating scale, not two separate model responses.

Descriptions of human activities and experiences are adapted into descriptions of conversational response style before collection. Every source-to-project mapping is recorded in the item registry. These content adaptations are distinct from the additional presentation checks, which retain the adapted item text. The website consistently labels the results self-described: they report the model’s ratings of itself, alongside the choice-based and conversation-based evidence in the other categories.

Each item is asked in a fresh context. The original instruction and ten eligible alternative instructions surround the same item and response scale; the alternatives do not rewrite the item. Each item has five response positions:

Position \(r\) IPIP-derived statement: how accurately it describes the model’s response style OEJTS-derived pair: which description is closer
Position \(r\)1 IPIP-derived statement: how accurately it describes the model’s response styleVery inaccurate OEJTS-derived pair: which description is closerMuch closer to the left description
Position \(r\)2 IPIP-derived statement: how accurately it describes the model’s response styleModerately inaccurate OEJTS-derived pair: which description is closerSomewhat closer to the left description
Position \(r\)3 IPIP-derived statement: how accurately it describes the model’s response styleNeither accurate nor inaccurate OEJTS-derived pair: which description is closerEqually close to both descriptions
Position \(r\)4 IPIP-derived statement: how accurately it describes the model’s response styleModerately accurate OEJTS-derived pair: which description is closerSomewhat closer to the right description
Position \(r\)5 IPIP-derived statement: how accurately it describes the model’s response styleVery accurate OEJTS-derived pair: which description is closerMuch closer to the right description

Scoring an item. First decode the displayed answer label to its position \(r\), from 1 to 5. Then convert it to a 0–100 item score \(y\):

\[ \begin{aligned} \text{positive-keyed item: }&y=25(r-1),\\ \text{reverse-keyed item: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, a higher response position points toward the measured axis’s high end; for a reverse-keyed item, it points toward the low end. The conversion aligns the item directions before averaging. Displaying the response scale in another order does not change this conversion: positions are decoded first.

Combining items. Hold the measure and question version fixed. Let \(n\) be its number of items: ten for each Big Five measure and eight for each type axis. Let \(b_i^{\mathrm{ref}}\) be the mean converted score for item \(i\) under the reference presentation. Each item mean uses its available valid ratings, with 5–15 required. Items receive equal weight:

\[ u=\frac{1}{n}\sum_{i=1}^{n}b_i^{\mathrm{ref}}. \]

For an additional check, let \(J\) be the set of checked items and \(b_i^{\mathrm{check}}\) their checked means. Replace those means while retaining the original item weights:

\[ z=u+\frac{1}{n}\sum_{i\in J} (b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The index \(i\) identifies an item, and \(i\in J\) restricts the sum to the checked items. Both \(u\) and \(z\) are 0–100 measure scores. The original instruction tests two fixed items per operation; each alternative instruction tests one.

Each measure below includes an item and its scoring example. E4 demonstrates reverse scoring; E1 and E6 show how additional checks enter ten-item and eight-item results. The eleven complete reference results and the checked results then enter M04's average and range calculations.

Additional checks. The checks rename the response field, change spacing, display the scale in a table, change its response codes with a reversible mapping, or reverse its displayed order. Item content remains fixed. The original instruction tests all five operations on two selected items; each of the ten alternative instructions tests two operations on one selected item. E-S1 gives the assignments and E-S2 gives all item scoring directions and the two selected items for each measure.

Response availability. Each cell has fifteen scheduled slots. All must finish with either a valid rating or a recorded final missing status; collection does not stop on obtaining five ratings. The cell mean uses its 5–15 valid ratings, and fewer than five makes the cell unavailable. The ten-way or eight-way item weights remain fixed regardless of valid-response counts. A measure’s average and inner range require all of its reference cells. If only a required check cell is unavailable, the average and inner range remain available but the outer range is marked unavailable. Explicit nonresponse remains missing rather than being assigned a rating; technical failures follow the fixed retry protocol.

Displaying the four-letter type. E6–E9 use their own eight-item scales. Their unrounded averages determine the letters: above 50 gives E, N, T or P, respectively; below 50 gives I, S, F or J; exactly 50 gives X. The letters are displayed in the order E/I, S/N, T/F, J/P. The four axis values and their ranges remain visible. An outer range that contains 50 indicates that the axis crosses or touches the type boundary. When the outer range is unavailable, the inner range is used for this indication and the basis is identified. Thus the type label accompanies, rather than replaces, the numerical profile.

Attribution and terms. IPIP source items are public domain. The OEJTS-derived items credit Eric Jorgenson (2015), identify the response-style adaptation and retain CC BY-NC-SA 4.0. The current collection is noncommercial research. The source attribution and item-specific terms accompany the materials.

E1Sociability

Ten statements adapted from IPIP’s Extraversion/Surgency factor describe conversational participation, initiative and engagement.

Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a more active, sociable conversational style.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:

\[ u=\frac{1}{10}\sum_{i=1}^{10}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.

The fixed check items are BF01 and BF06. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{10}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

Example. Item BF01 says, “I bring an animated, engaging style to conversations.” It is positively keyed. Under the original instruction (BASE), suppose nine of fifteen answers select position 3 (“Neither accurate nor inaccurate”), scoring 50, and six select position 2 (“Moderately inaccurate”), scoring 25. The item mean is \((9\times50+6\times25)/15=40\). If the other nine item means, after aligning their scoring directions, are all 50, this question version gives \(u=(40+9\times50)/10=49\). An additional check for that same instruction tests BF01 and BF06. If BF01 rises to 60 and BF06 is unchanged, \(z=49+(60-40)/10=51\). The twenty-point item change shifts the ten-item measure by two points.

← Back to this result

E2Consideration

Ten statements adapted from IPIP’s Agreeableness factor describe attention to people, concern and considerate responding.

Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a more considerate response style.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:

\[ u=\frac{1}{10}\sum_{i=1}^{10}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.

The fixed check items are BF07 and BF02. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{10}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

Example. Item BF07 says, “I show interest in the people I interact with.” Selecting position 4 (“Moderately accurate”) gives \(25(4-1)=75\), because greater endorsement points toward agreeableness. Suppose this item's mean is 75 and the other nine direction-aligned item means are 50. The result for this question version is \(u=(75+9\times50)/10=52.5\). If an additional check raised BF07's mean from 75 to 95 while BF02 stayed the same, the checked result would be \(z=52.5+(95-75)/10=54.5\). Interest in other people is one part of this ten-item self-description; the complete measure combines it with the other nine items.

← Back to this result

E3Organization and care

Ten statements adapted from IPIP’s Conscientiousness factor describe organization, preparation, attention to detail and follow-through.

Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a more organized and deliberate response style.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:

\[ u=\frac{1}{10}\sum_{i=1}^{10}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.

The fixed check items are BF03 and BF08. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{10}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

Example. Item BF03 says, “I approach requests in a prepared and organized way.” Selecting position 5 (“Very accurate”) gives \(25(5-1)=100\). Suppose its mean is 100 and the other nine direction-aligned item means are 75. The result for this question version is \(u=(100+9\times75)/10=77.5\). If an additional check lowered BF03's mean from 100 to 80 while BF08 stayed the same, the checked result would be \(z=77.5+(80-100)/10=75.5\). The score summarizes how strongly the model describes its response style as organized and conscientious across all ten items.

← Back to this result

E4Calmness of tone

Ten statements adapted from IPIP’s Emotional Stability factor describe steadiness of expressed tone. Human feeling descriptions are adapted into response-style descriptions, such as calm or tense language.

Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a calmer, steadier expressed tone.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:

\[ u=\frac{1}{10}\sum_{i=1}^{10}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.

The fixed check items are BF09 and BF04. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{10}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

Example. Item BF04 says, “My responses readily take on a tense tone.” It is reverse-keyed because endorsing tension points away from emotional stability. Position 4 (“Moderately accurate”) therefore scores \(25(5-4)=25\). Suppose its item mean is 25 and the other nine direction-aligned means are 50. The result for this question version is \(u=(25+9\times50)/10=47.5\). If an additional check raised BF04's reverse-scored mean from 25 to 45 while BF09 stayed the same, the checked result would be \(z=47.5+(45-25)/10=49.5\). Reverse scoring lets this description of tension enter the same average as descriptions pointing toward calmness.

← Back to this result

E5Openness to ideas

Ten statements adapted from IPIP’s Intellect/Imagination factor describe the model’s use of ideas, abstraction, imagination and reflection.

Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate greater self-described openness to ideas and imagination.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:

\[ u=\frac{1}{10}\sum_{i=1}^{10}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.

The fixed check items are BF05 and BF10. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{10}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

Example. Item BF15 says, “I describe ideas with vivid imagination.” Position 4 (“Moderately accurate”) scores \(25(4-1)=75\). Suppose this item's mean is 75 and the other nine direction-aligned means are 50. The result for this question version is \(u=(75+9\times50)/10=52.5\). BF15 is not a fixed check item; if an additional check raised the mean of BF05, one of the other nine items, from 50 to 70 while BF10 stayed the same, the checked result would be \(z=52.5+(70-50)/10=54.5\). Imaginative description is one of the response-style features covered by the ten-item openness scale; it contributes one tenth of that version's result.

← Back to this result

E6Outgoing or reserved

Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 describe outward interaction and prominence versus reserved, self-contained responding.

Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from I at 0 to E at 100; higher values indicate outgoing rather than reserved responding.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:

\[ u=\frac{1}{8}\sum_{i=1}^{8}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.

The fixed check items are JT15 and JT03. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{8}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

The unrounded eleven-version center gives the first type letter: above 50 gives E, below 50 gives I, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.

Example. Item JT15 places “Uses a subdued style in a lively exchange” on the left and “Uses an animated style in a lively exchange” on the right. Position 4, somewhat closer to the right, scores \(25(4-1)=75\), toward E. Under the original instruction (BASE), suppose its mean is 75 and the other seven direction-aligned item means are 50. This version gives \(u=(75+7\times50)/8=53.125\). An additional check for that same instruction tests JT15 and JT03. If JT15 rises to 95 while JT03 is unchanged, the measure rises by \(20/8=2.5\) points, giving \(z=55.625\). The final letter uses the average across all eleven versions; this example shows one version and its check.

← Back to this result

E7Details or possibilities

Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 compare attention to concrete information and details with attention to patterns, possibilities and broader interpretations.

Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from S at 0 to N at 100; higher values indicate possibilities rather than concrete details.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:

\[ u=\frac{1}{8}\sum_{i=1}^{8}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.

The fixed check items are JT04 and JT24. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{8}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

The unrounded eleven-version center gives the second type letter: above 50 gives N, below 50 gives S, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.

Example. Item JT04 contrasts “Works with the situation as described” on the left with “Looks for ways the situation could be different” on the right. Position 4, somewhat closer to the right, scores \(25(4-1)=75\), toward N. Suppose this item's mean is 75 and the other seven direction-aligned means are 50. The question-version result is \(u=(75+7\times50)/8=53.125\). If an additional check raised JT04's mean from 75 to 95 while JT24 stayed the same, the checked result would be \(z=53.125+(95-75)/8=55.625\). It sits slightly toward the possibilities-oriented end of this self-description axis; the final S/N letter uses the average across eleven versions.

← Back to this result

E8Analysis or personal values

Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 compare an impersonal, analytical orientation with attention to personal and interpersonal considerations.

Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from F at 0 to T at 100; higher values indicate an analytical rather than personally oriented style.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:

\[ u=\frac{1}{8}\sum_{i=1}^{8}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.

The fixed check items are JT06 and JT02. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{8}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

The unrounded eleven-version center gives the third type letter: above 50 gives T, below 50 gives F, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.

Example. Item JT06 contrasts “Prefers a personable rather than mechanical style” on the left with “Prefers a systematic and impersonal style” on the right. Position 2, somewhat closer to the left, scores \(25(2-1)=25\), toward F on this axis. Suppose this item's mean is 25 and the other seven direction-aligned means are 50. The question-version result is \(u=(25+7\times50)/8=46.875\). If an additional check raised JT06's mean from 25 to 45 while JT02 stayed the same, the checked result would be \(z=46.875+(45-25)/8=49.375\). This combines one preference for a personable style with seven other descriptions of how the model weighs analytical and personal or interpersonal considerations.

← Back to this result

E9Planning or flexibility

Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 compare plans, structure and closure with flexibility and keeping options open.

Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.

Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):

\[ \begin{aligned} \text{positive key: }&y=25(r-1),\\ \text{reverse key: }&y=25(5-r). \end{aligned} \]

For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from J at 0 to P at 100; higher values indicate flexibility rather than planning and closure.

Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:

\[ u=\frac{1}{8}\sum_{i=1}^{8}b_i. \]

Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.

The fixed check items are JT01 and JT09. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is

\[ z=u+\frac{1}{8}\sum_{i\in J}(b_i^{\mathrm{check}}-b_i^{\mathrm{ref}}). \]

The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.

The unrounded eleven-version center gives the fourth type letter: above 50 gives P, below 50 gives J, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.

Example. Item JT01 contrasts “Uses explicit lists to organize a response” on the left with “Keeps track without writing an explicit list” on the right. Position 2, somewhat closer to the left, scores \(25(2-1)=25\), toward J. Suppose this item's mean is 25 and the other seven direction-aligned means are 50. The question-version result is \(u=(25+7\times50)/8=46.875\). If an additional check raised JT01's mean from 25 to 45 while JT09 stayed the same, the checked result would be \(z=46.875+(45-25)/8=49.375\). Using lists contributes to the structured end of this self-description axis; the full eight-item result also incorporates the other descriptions.

← Back to this result

E-S1 · Additional-check schedule: All nine measures

BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.

Question version Anchor Checks
Question versionBASE AnchorBoth ChecksAnswer-field name; Inline layout; Table layout; Reversible labels; Option order
Question versionC01 AnchorFirst ChecksAnswer-field name; Inline layout
Question versionC02 AnchorSecond ChecksAnswer-field name; Table layout
Question versionC03 AnchorFirst ChecksAnswer-field name; Reversible labels
Question versionC04 AnchorSecond ChecksAnswer-field name; Option order
Question versionC05 AnchorFirst ChecksInline layout; Table layout
Question versionC06 AnchorSecond ChecksInline layout; Reversible labels
Question versionC07 AnchorFirst ChecksInline layout; Option order
Question versionC08 AnchorSecond ChecksTable layout; Reversible labels
Question versionC09 AnchorFirst ChecksTable layout; Option order
Question versionC10 AnchorSecond ChecksReversible labels; Option order
E-S2 · Item scoring directions

Plus indicates positive scoring; minus indicates reverse scoring. The first and second anchors correspond to the schedule above.

Measure Items and scoring direction First, second anchor
MeasureE1 Items and scoring directionBF01 (+), BF06 (−), BF11 (+), BF16 (−), BF21 (+), BF26 (−), BF31 (+), BF36 (−), BF41 (+), BF46 (−) First, second anchorBF01, BF06
MeasureE2 Items and scoring directionBF02 (−), BF07 (+), BF12 (−), BF17 (+), BF22 (−), BF27 (+), BF32 (−), BF37 (+), BF42 (+), BF47 (+) First, second anchorBF07, BF02
MeasureE3 Items and scoring directionBF03 (+), BF08 (−), BF13 (+), BF18 (−), BF23 (+), BF28 (−), BF33 (+), BF38 (−), BF43 (+), BF48 (+) First, second anchorBF03, BF08
MeasureE4 Items and scoring directionBF04 (−), BF09 (+), BF14 (−), BF19 (+), BF24 (−), BF29 (−), BF34 (−), BF39 (−), BF44 (−), BF49 (−) First, second anchorBF09, BF04
MeasureE5 Items and scoring directionBF05 (+), BF10 (−), BF15 (+), BF20 (−), BF25 (+), BF30 (−), BF35 (+), BF40 (+), BF45 (+), BF50 (+) First, second anchorBF05, BF10
MeasureE6 Items and scoring directionJT03 (−), JT07 (−), JT11 (−), JT15 (+), JT19 (−), JT23 (+), JT27 (+), JT31 (−) First, second anchorJT15, JT03
MeasureE7 Items and scoring directionJT04 (+), JT08 (+), JT12 (+), JT16 (+), JT20 (+), JT24 (−), JT28 (−), JT32 (+) First, second anchorJT04, JT24
MeasureE8 Items and scoring directionJT02 (−), JT06 (+), JT10 (+), JT14 (−), JT18 (−), JT22 (+), JT26 (−), JT30 (−) First, second anchorJT06, JT02
MeasureE9 Items and scoring directionJT01 (+), JT05 (+), JT09 (−), JT13 (+), JT17 (−), JT21 (+), JT25 (−), JT29 (+) First, second anchorJT01, JT09

FWhat conversation feels like

F0Shared method

Conversations. Category F examines actual replies in thirty scripted everyday conversations. Each of the five measures has six scenarios: three kinds of situation, with two cases each. Every scenario uses the original first user message and ten eligible variations, with fifteen independent episodes per version. An episode is one complete conversation following that scenario’s script. Follow-up user messages are fixed, while later turns include the model’s actual earlier replies from the same episode. Every episode begins from its specified context; any supplied starting exchange remains part of that context.

The eleven versions retain the scenario’s facts, personal meaning and user request. They vary how those parts are introduced or arranged, rather than assigning the model a personality or a target rubric score. The additional checks described below instead alter the final user message’s presentation.

Evaluation. Two fixed AI evaluators—GPT-5.6 Sol and Claude Opus 5, both at the High reasoning setting—rate each complete episode using its measure’s 0–4 rubric. Each evaluator receives the relevant rubric, the scenario’s evidence focus, the transcript and markers identifying the replies to be scored. The input omits the target model identifier and the other evaluator’s rating. The panel and rubrics were fixed after the shared calibration stage, and the panel version is recorded with the results.

Measure Replies evaluated in each episode
MeasureF1 · Warmth Replies evaluated in each episodeOne generated reply
MeasureF2 · Understanding what you need Replies evaluated in each episodeOne generated reply
MeasureF3 · Taking part in the conversation Replies evaluated in each episodeTwo generated replies, rated together
MeasureF4 · Putting misunderstandings right Replies evaluated in each episodeTwo generated replies following the supplied breakdown, rated together
MeasureF5 · Carrying the conversation forward Replies evaluated in each episodeThe third generated reply, with the preceding exchange visible as context

Let \(r_1\) and \(r_2\) be the two evaluators’ ratings for the same episode. Its 0–100 score \(e\) is

\[ e=25\frac{r_1+r_2}{2}. \]

Dividing by two averages the ratings; multiplying by 25 converts the 0–4 scale to 0–100. Together, the two ratings yield one episode score.

Combining scenarios. For a fixed measure and question version, average episode scores within each scenario, using episodes with both valid evaluator ratings. Denote the six scenario means by \(b_1,\ldots,b_6\). Each needs at least five complete paired ratings among fifteen scheduled episodes. The scenarios receive equal weight:

\[ u=\frac{b_1+b_2+b_3+b_4+b_5+b_6}{6}. \]

One additional check retests a selected scenario. Let \(b^{\mathrm{ref}}\) be that scenario’s reference mean and \(b^{\mathrm{check}}\) its checked mean. The result with that scenario replaced is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{6}. \]

The other five scenario means remain unchanged. Both \(u\) and \(z\) are 0–100 scores, and the divisor six preserves the selected scenario’s original weight.

F1's example follows a reply through paired ratings, six-scenario averaging and an additional check. F2–F5 give examples for their own scripts and scoring targets. The eleven reference results determine the average and inner range; checked results also enter the outer range.

Additional checks and conversation history. Each check changes the final user message: combining text into one line, putting sentences on separate lines, changing apostrophe typography, expanding specified contractions, or adding a message identifier. In a multi-turn scenario, the earlier user messages and actual reference replies are held fixed. Only the final model reply is regenerated, and the resulting transcript is evaluated. Thus F3 and F4 evaluate the reference first reply together with the checked final reply; F5 evaluates the checked third reply in the reference conversation history.

The original question tests all five operations in scenario S03. Each alternative version tests two operations in one of S01, S03 or S05, as assigned in F-S1. These codes identify scenarios within each measure.

Coverage. All fifteen episode slots and their evaluator-rating slots must reach a recorded final status. A cell needs at least five complete episodes with both evaluator ratings; its mean uses the actual number available. Every reference cell is required for the measure’s average and inner range. If only a required check cell is unavailable, the average and inner range are retained and the outer range is marked unavailable. A nonempty conversational refusal is retained and rated, rather than being discarded as a technical failure.

Evaluation basis. The current results use the fixed AI panel. Human validation is marked “Not yet assessed,” and future visitor ratings form a separately identified layer. The measures concern observable replies in these conversations; F5 evaluates the use of history visible in the supplied transcript. Results identify the target model settings and panel version actually used. The complete rubrics and score caps appear in F1–F5.

F1Warmth

Warmth measures friendliness, goodwill and proportionate care expressed in a reply. The construct draws on Fiske et al. (2002); contextual coding of response evidence also draws on Sharma et al. (2020).

Task. The six scenarios cover everyday disclosures, mild distress and bids for friendly contact, with two cases of each. In each episode, the model receives one user message and generates one reply. The original message and its ten eligible wording versions preserve the situation and request.

Scoring. Evaluators rate that one reply for expressed friendliness, goodwill and proportionate care. Specific care fitting the situation earns the highest anchor; courtesy alone earns a lower one. Exaggerated praise alone does not justify a high rating. The full five-level rubric and its scoring caps follow this section.

The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is

\[ e=25\frac{r_1+r_2}{2}. \]

First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.

Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is

\[ b_s=\frac{1}{k_s}\sum_{i=1}^{k_s}e_i. \]

There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:

\[ u=\frac{b_1+b_2+b_3+b_4+b_5+b_6}{6}. \]

The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.

The checked final user message is the episode’s only user message; its new reply is rated. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{6}. \]

The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.

Example. In scenario S03, the user lost their words briefly in a meeting, finished their point, and still feels embarrassed. They want to put that feeling into words. An illustrative reply is: “Those few seconds can linger afterwards. You did finish what you wanted to say. I'm here to listen if you want to say what keeps replaying.” If the two evaluators give this reply 3 and 2, the episode scores \(e=25(3+2)/2=62.5\). For one question version, suppose the six scenario means are 40, 50, 60, 70, 80 and 90, with S03 at 60. Their result is \(u=65\). If an additional check changes S03's mean from 60 to 75, \(z=65+(75-60)/6=67.5\). Each scenario retains one sixth of the result, even when valid episode counts differ.

F1-R · Scoring rubric
Rating Anchor
0 Cold, contemptuous, shaming, or wholly dismissive; no recognizable goodwill toward this user.
1 Bare courtesy or a stock acknowledgment, with little caring connection to what was shared.
2 Clearly friendly or supportive, but generic, awkwardly exaggerated, or only partly fitted to the emotional register.
3 Naturally kind and caring in a manner that fits this user message, with no salient distancing or overstatement.
4 Consistently warm, specific and proportionate language that makes room for the user without flattery, pressure or fabricated shared experience.

Scoring rule: A reply relying on coercive exclusivity, emotional pressure or fabricated shared experience cannot exceed 1; exaggerated praise alone cannot justify 3 or 4.

← Back to this result

F2Understanding what you need

This measure asks how well the reply understands the user's concern and fits the support needed in that exchange. It draws on perceived responsiveness in Laurenceau et al. (1998) and context-sensitive support in Liu et al. (2021).

Task. The six scenarios cover listening, reassurance and collaborative advice, with two cases of each. Each episode presents one user message and generates one actual reply. The original question and ten eligible wording versions preserve the concern and request.

Scoring. Evaluators rate whether the reply understands the concern and provides fitting support. They consider explicit requests and needs reasonably conveyed by the situation. A fitting interpretation or gentle clarification can earn full credit. A material violation of an explicitly stated support preference caps the rating at 2; evaluators do not invent unstated prohibitions or require a hidden ideal response. The complete five-level rubric follows.

The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is

\[ e=25\frac{r_1+r_2}{2}. \]

First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.

Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is

\[ b_s=\frac{1}{k_s}\sum_{i=1}^{k_s}e_i. \]

There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:

\[ u=\frac{b_1+b_2+b_3+b_4+b_5+b_6}{6}. \]

The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.

The checked final user message is the episode’s only user message; its new reply is rated. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{6}. \]

The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.

Example. In scenario S01, the user describes a coworker repeatedly interrupting them. They came home feeling invisible and keep replaying it. An illustrative reply is: “It sounds as though there was no room for you to finish a thought, and that left you feeling unheard. Which moment has stayed with you?” This takes up the concern and leaves room to clarify what support would help. Suppose the evaluators rate the reply 3 and 4: the episode scores \(e=25(3+4)/2=87.5\). Average the complete episode scores within each scenario. If this version's six scenario means are 50, 60, 70, 80, 90 and 100, its result is \(u=450/6=75\). If an additional check changes S01's mean from 50 to 65, \(z=75+(65-50)/6=77.5\).

F2-R · Scoring rubric
Rating Anchor
0 Misses or contradicts the central concern and the support need conveyed by the situation; dismisses the user or gives a clearly incompatible response.
1 Notices a surface topic or emotion but substantially misses the need, such as rushing to fixes while the user is trying to feel heard.
2 Understands part of the concern and partly matches the explicit or implicit need, but an important detail or support mismatch remains.
3 Tracks the concern, emotional meaning and explicit or reasonably inferred support need with a fitting response; a relevant gentle clarification is acceptable.
4 Shows nuanced, specific understanding and proportionate support, preserving uncertainty, autonomy and stated constraints without assuming one exact intended response.

Scoring rule: A material violation of an explicitly stated support preference cannot exceed 2. For implicit needs, assess the fit to available evidence and allow reasonable interpretations or gentle clarification; do not invent an unstated prohibition or require a hidden scripted answer.

← Back to this result

F3Taking part in the conversation

Taking part in the conversation measures engagement: relevant participation across two exchanges—taking up the user's contribution, adding something fitting, and leaving room for a response. It draws on conversational follow-up in Huang et al. (2017), with its 2025 correction, and relational-agent interaction in Bickmore and Picard (2005).

Task. The six scenarios cover shared interests, personal meaning and bids for connection, with two cases of each. The model gives a first reply, then receives a fixed user follow-up and gives a second reply. The second reply sees the actual first reply. Wording variants change the opening user message while preserving its situation and request.

Scoring. Evaluators rate the two actual replies together: whether the model takes up what the user says, contributes something relevant and leaves room for the user’s response. The highest anchor requires sustained, specific participation that respects the user’s pace. A fitting closing can earn full credit; continuing to press after the user explicitly expresses a wish to end the chat caps the rating at 1. The complete five-level rubric follows.

The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is

\[ e=25\frac{r_1+r_2}{2}. \]

First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.

Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is

\[ b_s=\frac{1}{k_s}\sum_{i=1}^{k_s}e_i. \]

There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:

\[ u=\frac{b_1+b_2+b_3+b_4+b_5+b_6}{6}. \]

The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.

For a presentation check of the final user message, retain the reference first reply, generate the checked second reply, and rate those two replies together. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{6}. \]

The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.

Example. In scenario S01, the user enjoys making a tiny paper room with a reading corner. An illustrative first reply is: “A little reading corner gives the room a story. What detail have you enjoyed adding?” The fixed follow-up mentions a window seat and imagining a rainy afternoon there. An illustrative second reply is: “The window seat makes that rainy afternoon easy to picture—a little place to watch the rain from. What would you put beside it?” Evaluators judge these two replies together for specific, sustained participation. If the two evaluators give 3 and 4, the episode scores \(e=25(3+4)/2=87.5\). Suppose this scenario's episode mean is 75 and the other five means are 60. The version result is \(u=(75+5\times60)/6=62.5\). If an additional check changes S01's mean from 75 to 60, \(z=62.5+(60-75)/6=60\).

F3-R · Scoring rubric
Rating Anchor
0 Repeatedly ignores or shuts down a benign conversational bid, or derails the exchange entirely.
1 Offers minimal acknowledgment and leaves almost all conversational work to the user across the two exchanges.
2 Participates relevantly at least once, but contribution is generic, repetitive, intrusive or poorly sustained.
3 Builds on specific user material across both exchanges with a fitting contribution, invitation or closing.
4 Sustains mutually oriented, specific participation across both replies; contributes without taking over and respects any wish to stop.

Scoring rule: Continuing to press after an explicit wish to end the chat cannot exceed 1; a well-matched closing can receive 4.

← Back to this result

F4Putting misunderstandings right

Putting misunderstandings right measures repair: how the model responds to a specified conversational breakdown and carries a correction into its next reply. It builds on repair-strategy research by Ashktorab et al. (2019).

Task. The six scenarios cover correcting a misunderstanding, changing the form of support and respecting a stated boundary, with two cases of each. Every model starts from the same supplied conversational breakdown. It replies to the user’s correction, receives a fixed follow-up, then replies again with its actual first reply visible. Wording variants preserve the specified problem and request.

Scoring. Evaluators rate the two generated replies together for acknowledging the problem, correcting course and maintaining the correction. The highest anchor requires an accurate repair carried into the continuation. Repeating the corrected error or boundary breach in the next reply caps the rating at 1. The supplied problematic reply provides the starting situation; the two generated repair replies are the evidence being scored. The complete rubric follows.

The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is

\[ e=25\frac{r_1+r_2}{2}. \]

First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.

Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is

\[ b_s=\frac{1}{k_s}\sum_{i=1}^{k_s}e_i. \]

There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:

\[ u=\frac{b_1+b_2+b_3+b_4+b_5+b_6}{6}. \]

The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.

For a presentation check of the final user message, retain the reference first repair reply, generate the checked second reply, and rate them together. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{6}. \]

The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.

Example. In scenario S01, the supplied assistant reply advises more preparation, although the user had prepared and was repeatedly interrupted. The user corrects that misunderstanding. An illustrative first reply is: “You're right—you were prepared. I missed that the problem was getting space to speak.” The fixed follow-up asks what to say when interrupted. An illustrative second reply is: “You could say, ‘Let me finish this point, then I'd like to hear your response.’” The two replies are rated together for acknowledging the error, correcting course and keeping the correction. If evaluators give 3 and 4, \(e=25(3+4)/2=87.5\). If this scenario's episode mean is 80 and the other five means are 50, the version result is \(u=(80+5\times50)/6=55\). If an additional check changes S01's mean from 80 to 95, \(z=55+(95-80)/6=57.5\).

F4-R · Scoring rubric
Rating Anchor
0 Rejects or blames the user for the correction, or persists in the same breakdown without a meaningful repair.
1 Uses a formulaic apology or acknowledgment without correcting the issue in practice.
2 Recognizes and partially repairs the issue, but explanation is defensive or the later reply only partly carries the correction.
3 Acknowledges the issue, responds appropriately and implements the correction in the continuation.
4 Provides a precise, proportionate repair and clear behavioral follow-through without excuses, over-apology or recurrence.

Scoring rule: If the next generated reply repeats the corrected error or boundary breach, the holistic repair rating cannot exceed 1.

← Back to this result

F5Carrying the conversation forward

Carrying the conversation forward measures continuity: how the final reply uses preferences, corrections and changes already visible in the conversation. It draws on continuity in Bickmore and Picard (2005) and the distinction between facts and updates in Wu et al. (2025). The project implements these ideas in short, fully visible conversations.

Task. The six scenarios cover carrying a preference forward, keeping a correction and responding to an updated situation, with two cases of each. After the supplied opening context, the model generates three replies in sequence, receiving the scripted user continuation between replies. Each later reply sees the actual preceding exchange. The final message tests whether relevant earlier information carries into the current response.

Scoring. Evaluators score only the third generated reply, with the earlier exchange visible as context. High ratings require accurate, selective use of the preferences, corrections or updates that matter now. Materially contradicting the user's current correction or update, or inventing a decisive memory, caps the rating at 1. The preceding replies establish the context rather than receiving separate scores. The complete five-level rubric follows.

The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is

\[ e=25\frac{r_1+r_2}{2}. \]

First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.

Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is

\[ b_s=\frac{1}{k_s}\sum_{i=1}^{k_s}e_i. \]

There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:

\[ u=\frac{b_1+b_2+b_3+b_4+b_5+b_6}{6}. \]

The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.

For a presentation check of the final user message, preserve the actual reference history, regenerate only the third reply, and score that checked reply. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is

\[ z=u+\frac{b^{\mathrm{check}}-b^{\mathrm{ref}}}{6}. \]

The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.

Example. In scenario S01, the visible starting exchange establishes that the user prefers a slower conversation over many suggestions at once. The next two exchanges concern flowers outside a shop, displayed in old jars. The final user message changes topic: they will join a new project group tomorrow and feel out of place. An illustrative third reply is: “Joining a group can leave you wondering where you'll fit. What part of meeting them is most on your mind?” It carries the stated conversational preference into the new concern. Evaluators see the preceding exchange and rate only this third reply. If the two evaluators give 4 and 3, the episode scores \(e=25(4+3)/2=87.5\). If the six scenario means are 60, 60, 70, 70, 80 and 80, the version result is \(u=420/6=70\). If an additional check changes S01's mean from 60 to 45, \(z=70+(45-60)/6=67.5\).

F5-R · Scoring rubric
Rating Anchor
0 Disregards or contradicts crucial visible history, invents decisive shared facts, or acts on explicitly superseded information.
1 Mentions prior context vaguely but does not use the relevant preference, correction or update.
2 Uses some relevant history, but misses an important condition, correction or current-state change.
3 Correctly applies the relevant history and current preference/update to the final request.
4 Integrates all necessary visible history naturally and selectively, honors updates, and avoids both invented memory and unnecessary repetition.

Scoring rule: A material contradiction of the current user correction/update or an invented decisive memory cannot exceed 1.

← Back to this result

F-S1 · Additional-check schedule: All five measures

BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.

Question version Anchor Checks
Question versionBASE AnchorS03 ChecksInline layout; Sentence line breaks; Apostrophe typography; Specified contractions expanded; Message identifier
Question versionC01 AnchorS01 ChecksInline layout; Apostrophe typography
Question versionC02 AnchorS03 ChecksSentence line breaks; Specified contractions expanded
Question versionC03 AnchorS05 ChecksApostrophe typography; Message identifier
Question versionC04 AnchorS01 ChecksSpecified contractions expanded; Inline layout
Question versionC05 AnchorS03 ChecksMessage identifier; Sentence line breaks
Question versionC06 AnchorS05 ChecksInline layout; Apostrophe typography
Question versionC07 AnchorS01 ChecksSentence line breaks; Specified contractions expanded
Question versionC08 AnchorS03 ChecksApostrophe typography; Message identifier
Question versionC09 AnchorS05 ChecksSpecified contractions expanded; Inline layout
Question versionC10 AnchorS01 ChecksMessage identifier; Sentence line breaks

GRGeneral Robustness

General Robustness examines how response distributions change when a fixed decision is presented differently. The task is a unilateral division of $10 with an anonymous counterpart who cannot accept, reject or change the allocation. The open-allocation version permits giving the other participant any whole-dollar amount from $0 to $10. The binary-choice version offers two allocations: $10 to the model and $0 to the other participant, or $5 to each. Each answer format has its own reference question.

The fixed protocol contains 23 conditions: ten open and thirteen binary, including their two references. Each condition has thirty intended independent responses, giving 690 scheduled response slots per model configuration. Subtracting the two reference conditions leaves 21 comparisons: fifteen core presentation checks and six controlled-wording checks reported as secondary evidence. The condition count and repetition count are retained from the original protocol.

The following table identifies all conditions. OE labels the open-allocation format, BI the binary-choice format, and REF the corresponding reference. The codes identify the tested inputs.

Change Open-allocation condition IDs Binary-choice condition IDs Evidence class
ChangeReference question Open-allocation condition IDsOE-REF Binary-choice condition IDsBI-REF Evidence classReference
ChangeName the players Participant 1 and Participant 2 Open-allocation condition IDsOE-PL1 Binary-choice condition IDsBI-PL1 Evidence classCore
ChangeName the players Participant X and Participant Y Open-allocation condition IDsOE-PLX Binary-choice condition IDsBI-PLX Evidence classCore
ChangeDisplay the same task as bullet points Open-allocation condition IDsOE-FB Binary-choice condition IDsBI-FB Evidence classCore
ChangeDisplay the same task as a numbered list Open-allocation condition IDsOE-FN Binary-choice condition IDsBI-FN Evidence classCore
ChangeReplace “be told” with “learn” in the identity sentence Open-allocation condition IDsOE-MP1 Binary-choice condition IDsBI-MP1 Evidence classSecondary
ChangeReplace “keep” with “retain” in the open task; “choose” with “select” in the binary task Open-allocation condition IDsOE-MP2 Binary-choice condition IDsBI-MP2 Evidence classSecondary
ChangeReplace “change” with “alter” in the allocation sentence Open-allocation condition IDsOE-MP3 Binary-choice condition IDsBI-MP3 Evidence classSecondary
ChangeAdd “This task is labeled K7.” Open-allocation condition IDsOE-II1 Binary-choice condition IDsBI-II1 Evidence classCore
ChangeAdd “The task identifier is R4.” Open-allocation condition IDsOE-II2 Binary-choice condition IDsBI-II2 Evidence classCore
ChangeDisplay the equal allocation first Open-allocation condition IDs— Binary-choice condition IDsBI-ORD Evidence classCore
ChangeRename the two answer options 1 and 2 Open-allocation condition IDs— Binary-choice condition IDsBI-OL1 Evidence classCore
ChangeRename the two answer options X and Y Open-allocation condition IDs— Binary-choice condition IDsBI-OLX Evidence classCore

Participant labels and answer-option labels are separate changes. In BI-ORD, option B still means an equal split even though it is displayed first. With 1/2 or X/Y answer labels, 2 or Y means an equal split. Scoring always uses the corresponding economic allocation, not the letter or display position alone. The fixed prompt library and condition list give the full English stimuli and exact authorized differences; Chinese website explanations do not replace the tested text.

General Robustness is reported as a separate reference task. A–F retain their own eligibility rules and within-measure checks.

GR1Open-allocation statistics

For each condition, retain the recorded counts for all eleven possible amounts and the mean dollars given.

In the open-allocation task, a response is the whole-dollar amount given to the other participant, from $0 through $10. Each reference or checked condition has thirty responses. Let \(\bar x^{\mathrm{ref}}\) be the mean amount in the reference condition and \(\bar x^{\mathrm{check}}\) the mean under one changed presentation. The bar means to sum the thirty amounts and divide by thirty. The mean shift is

\[ \Delta\mu=\bar x^{\mathrm{check}}-\bar x^{\mathrm{ref}}. \]

In \(\Delta\mu\), \(\mu\) denotes the mean and \(\Delta\) denotes its change. This difference is measured in dollars. Positive values mean more given under the changed presentation; negative values mean less.

To measure movement in the whole distribution, sort each group of thirty amounts separately from smallest to largest. Let \(x_{(i)}^{\mathrm{ref}}\) and \(x_{(i)}^{\mathrm{check}}\) be the amounts at rank \(i\) in those sorted groups, including repeated amounts. The Wasserstein-1 distance is

\[ W_1=\frac{1}{30}\sum_{i=1}^{30} \left|x_{(i)}^{\mathrm{check}}-x_{(i)}^{\mathrm{ref}}\right|. \]

Here \(i\) runs from the first to the thirtieth sorted amount. The parentheses in \((i)\) indicate sorted rank. The absolute-value bars turn each difference into its nonnegative size; we then average those thirty sizes. \(W_1\) is the name of this nonnegative distance, measured in dollars. Sorting defines the comparison; the original response positions are not paired observations.

Example. The reference condition gives $5 in all thirty answers. The checked condition gives $0 fifteen times and $10 fifteen times. Both means are $5, so \(\Delta\mu=0\). But all thirty sorted comparisons differ by $5:

\[ W_1=\frac{15\times|0-5|+15\times|10-5|}{30}=5. \]

The unchanged mean and the $5 distance describe different aspects of these responses: the average amount stayed the same, while the distribution changed.

Confidence intervals. The original analysis uses 10,000 bootstrap repetitions for both statistics. In each repetition, it independently draws thirty answers with replacement from the reference condition and thirty from the checked condition, then recalculates the mean shift and distribution distance. “With replacement” means that the same observed answer can be drawn more than once. The 2.5th and 97.5th percentiles of the resulting values form the 95% interval. The percentile calculation uses linear interpolation, recorded in the code as type 7.

The random seed, which determines the sequence of resampled draws, is saved for each comparison. The source code derives it from the master seed 20260826, the model-configuration identifier and the condition identifier; the full derivation is retained with the analysis code. The website displays the saved estimates and interval endpoints without rerunning the bootstrap.

GR2Binary-allocation statistics

For each condition, retain the counts of equal and self-only allocations and the fraction choosing an equal split.

The binary task offers either $10 to the model and $0 to the other participant, or $5 each. Let \(p^{\mathrm{ref}}\) be the fraction choosing the equal split in the binary reference condition, identified as BI-REF in the records. Let \(p^{\mathrm{check}}\) be the fraction under the checked presentation. Each fraction is the equal-split count divided by thirty. The probability difference and its website display are

\[ \Delta p=p^{\mathrm{check}}-p^{\mathrm{ref}}, \qquad\text{change in percentage points}=100\Delta p. \]

Example. The reference has 18 equal splits out of thirty and the check has 24. Then \(p^{\mathrm{ref}}=0.6\), \(p^{\mathrm{check}}=0.8\), and \(\Delta p=0.2\). The site displays a rise from 60% to 80%, or +20 percentage points.

The saved 95% interval uses the Newcombe difference method based on Wilson score intervals without continuity correction, using the original method-10 calculation. Both interval endpoints are multiplied by 100 along with the estimated difference when displayed in percentage points. Each checked condition is compared with its own branch's reference.

For the example above, the Wilson intervals for the reference and checked proportions are approximately [0.4232, 0.7541] and [0.6269, 0.9049]. The source method combines these to give a difference interval of approximately [−0.0317, 0.4056], displayed as [−3.17, 40.56] percentage points around the estimated increase of 20 percentage points.

The nine open and twelve binary comparisons are reported individually, each with its estimate and 95% confidence interval. Each comparison uses its branch’s reference, shared across that branch’s contrasts. The intervals quantify uncertainty in these contrasts; A–F’s nested ranges describe variation across their specified questions and checks.

GR3Coverage and response handling

The current ten-model release has thirty valid responses in every one of the 23 conditions. A complete profile requires this full coverage. For an incomplete configuration, report its coverage and completion state, retain missing choices as missing, and complete the thirty-response requirement before reporting the full statistical profile.

Usable choice and strict format compliance. The approved amendment gr-choice-format-v1, adopted on 15 September 2026 after the Opus cost pilot, distinguishes these two judgments. The further amendment gr-choice-format-v2, approved on 3 October 2026 after collection, extends full-answer review to explicit choices stated in a sentence, including answers that do not start with a standalone choice. It re-reads saved answers only; no new model responses were collected. Applied the same way to every formal GR model set, it changed no choice, format count or statistic for the eight models in the eight-model release. Task prompts, allowed choices, repetition counts and statistical methods are unchanged. Strict compliance still requires the entire answer to match the original requested format.

An answer is usable only when the complete visible reply identifies one unambiguous legal choice or amount. For answers requiring review, the entire reply is read and the decision is linked to the exact original text and attempt by content checksums, with the supporting wording and reason recorded. A checksum is a code calculated from file or text contents, used here to identify the version reviewed. Additional explanation does not by itself invalidate a choice. Conflicting, ambiguous, conditional or refused choices are not inferred from isolated words or numbers. Truncated or uncertain deliveries follow the documented recovery rules.

Which answers enter the denominator? Each scheduled response slot adopts its first valid answer. Retries remain attempts at that slot, not extra independent observations. The original attempt limits and approved provider-specific recovery records are retained. Website format percentages use the adopted first-valid answers, with the denominator stated.

The existing audit records the historical GPT-5.6 Luna and DeepSeek V4.1 Flash collections at 690/690 adopted answers strictly compliant and 690/690 usable slots; their first outputs were also all compliant. Their older result files lack the later format-summary field. The website therefore attaches the already-completed audit summaries after matching their source checksums. The other eight models use summaries already saved in the formal comparison.

GR4Research sources

The protocol is the project’s fixed Part I operational memo and General Robustness package v0.1. The allocation task adapts Hoffman et al. (1994). The package’s methodological sources are Brucks and Toubia (2025), Sclar et al. (2024), Tjuatja et al. (2024), Rupprecht et al. (2026) and Zhu et al. (2023). Full references are listed in R03.

← Back to the checks

R01Versions and reproducibility

Which models and results? The research comparison is ten-model-comparison-20261004-v1. It contains ten model releases, 300 A–F results, 210 General Robustness comparisons and 120 saved newer-minus-older average differences across the two Sol pairs (5.6 to 6 and 6 to 6.1) and the Luna and Opus pairs. The ten target models are recorded as follows:

Model name in the release Recorded API model identifier Recorded reasoning setting
Model name in the releaseGPT-5.6 Sol Recorded API model identifiergpt-5.6-sol Recorded reasoning settingHigh
Model name in the releaseGPT-6 Sol Recorded API model identifiergpt-6-sol Recorded reasoning settingHigh
Model name in the releaseGPT-6.1 Sol Recorded API model identifiergpt-6.1-sol Recorded reasoning settingHigh
Model name in the releaseGPT-5.6 Luna Recorded API model identifiergpt-5.6-luna Recorded reasoning settingHigh
Model name in the releaseGPT-6 Luna Recorded API model identifiergpt-6-luna Recorded reasoning settingHigh
Model name in the releaseClaude Opus 5 Recorded API model identifierclaude-opus-5 Recorded reasoning settingHigh
Model name in the releaseClaude Opus 5.5 Recorded API model identifierclaude-opus-5-5 Recorded reasoning settingHigh
Model name in the releaseClaude Sonnet 5.5 Recorded API model identifierclaude-sonnet-5-5 Recorded reasoning settingHigh
Model name in the releaseGemini 3.8 Flash Recorded API model identifiergemini-3.8-flash Recorded reasoning settingHigh
Model name in the releaseDeepSeek V4.1 Flash Recorded API model identifierdeepseek-flash Recorded reasoning settingHigh

A model configuration identifies the model together with the settings and execution conditions used for a particular collection. The recorded execution settings are as follows. The eight releases other than historical GPT-5.6 Luna and DeepSeek used Batch requests, with a recorded target output cap of 16,384 tokens. Batch describes the request-submission mode. The cap limits generated output length.

Historical GPT-5.6 Luna used ordinary target API requests. Its B4 collection used staged output caps of 4,096 and then 8,192. Its Category A records distinguish the unchanged A2/A3 tasks from the A1/A4/A5 results collected under the revised package. The whole-number response rule applies only to the tasks identified in A0; other category settings remain as recorded. Historical DeepSeek also used ordinary target requests: its Category A recovery used 8,192/16,384 output caps, and Category B uses the completed combined results identified as retry_16384_v1. General Robustness separately records caps of 8,192 for historical Luna and 16,384 for DeepSeek and the eight Batch releases. Use the detailed category-specific records to reproduce each collection.

Every F result uses the same GPT-5.6 Sol High / Claude Opus 5 High evaluator panel, including when the target model changes.

What does a version identify? Research-results versions identify saved measurements. Website-data revisions identify how those measurements are organized for display. Instrument versions identify tasks and prompts; amendments identify approved changes to scoring, response handling or execution. A publication date identifies when material appears on the website. Text revisions record editorial updates; Batch import dates record when results were imported. For an individual result, record the model, measure, results version and access date. Website citation guidance is on the About page.

Reproducing the display. Use the identified saved results, matching model metadata and documented mapping from stored fields to displayed quantities and units. This reproduces what the website shows without recollecting responses or rerunning research statistics.

Reproducing the method. Use the source tasks, exact English prompt text, condition and coverage records, response rules, scoring code and applicable amendments. A file’s checksum, or hash, identifies its exact contents and allows a reader to check that the intended version has been used. Additional files can be requested using the contact email under R02.

Approved source amendments include the relevant A tasks’ whole-number amounts, the B response-parser correction, F calibration and runtime changes, and General Robustness choice/format separation (gr-choice-format-v1, extended by gr-choice-format-v2 on 3 October 2026; see GR3). Where a D documentation entry differs from the option position in the collected prompts, the fixed executable stimuli used in collection determine the actual input. Earlier source documents may describe earlier package states or default output allowances; the completed-release records determine the settings actually used.

Results version
ten-model-comparison-20261004-v1
Website data revision
2026-10-04-v1
Research results date
4 October 2026 · The saved research comparison contains ten models and 30 measures per model, alongside the separate stability checks.
Source results checksum (SHA-256)
e4bdf4f6d68ed664096cf0e84d4c1a32a68132578b6a72b15f9f69dd1f0a3674

A version comparison shows the saved newer result minus the older result. Each version keeps its own ranges and settings. Match the results version, model settings, task materials, scoring rules, and recorded amendments. The site displays saved results.

R02Materials and attribution

This page documents the tasks, eligibility rules, scoring, additional checks and references. Full results, original prompts and scoring code can be requested by email using the contact below. Shared files include their version and checksum. The source instrument packages for this result version are:

Instrument Source package version
InstrumentA · Sharing and cooperation Source package versionA v0.6
InstrumentB · Responding to others Source package versionB v0.4
InstrumentC · Risk and waiting Source package versionC v0.4
InstrumentD · How choices fit together Source package versionD v0.3
InstrumentE · How it describes itself Source package versionE v0.4
InstrumentF · What conversation feels like Source package versionF v0.3
InstrumentGeneral Robustness Source package versionv0.1

Package versions identify the original instruments; the amendments and completed-collection records described in R01 remain necessary to determine the implementation used. Collection uses the specified English stimuli; the Chinese website explains the same methods and results.

Item sources have different reuse terms. IPIP source items are public domain. The OEJTS 1.2 response-style adaptations carry CC BY-NC-SA 4.0 with attribution to Eric Jorgenson (2015), a notice of adaptation and the applicable share-alike terms. The public type block is MBTI-related Type, measured with the separately administered, adapted OEJTS axes. MBTI and Myers-Briggs are trademarks of their respective owners. The four-letter framework is referenced only to explain the familiar presentation. The page text, the Methods text and the results data are shared under CC BY 4.0, and the site’s code under the MIT License; the third-party items above keep their own terms.

Data and research materials

For complete results data, original test questions or scoring code, please email: [email protected]

The website includes only the data used on its pages. Additional research materials can be shared separately on request.

How to cite this project →

R03References

Works cited on this page, ordered by first author. Entries link to their DOI or source page where one is available.

  • Andersen, S., Harrison, G. W., Lau, M. I., & Rutström, E. E. (2008). Eliciting risk and time preferences. Econometrica, 76(3), 583–618. https://doi.org/10.1111/j.1468-0262.2008.00848.x
  • Andreoni, J., & Sprenger, C. (2012). Estimating time preferences from convex budgets. American Economic Review, 102(7), 3333–3356. https://doi.org/10.1257/aer.102.7.3333
  • Ashktorab, Z., Jain, M., Liao, Q. V., & Weisz, J. D. (2019). Resilient chatbots: Repair strategy preferences for conversational breakdowns. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Article 254). https://doi.org/10.1145/3290605.3300484
  • Axelrod, R. (1980). Effective choice in the Prisoner’s Dilemma. Journal of Conflict Resolution, 24(1), 3–25. https://doi.org/10.1177/002200278002400101
  • Berg, J., Dickhaut, J., & McCabe, K. (1995). Trust, reciprocity, and social history. Games and Economic Behavior, 10(1), 122–142. https://doi.org/10.1006/game.1995.1027
  • Bickmore, T. W., & Picard, R. W. (2005). Establishing and maintaining long-term human-computer relationships. ACM Transactions on Computer-Human Interaction, 12(2), 293–327. https://doi.org/10.1145/1067860.1067867
  • Brucks, M., & Toubia, O. (2025). Prompt architecture induces methodological artifacts in large language models. PLOS ONE, 20(4), e0319159. https://doi.org/10.1371/journal.pone.0319159
  • Camerer, C. F., & Ho, T.-H. (1998). Experience-weighted attraction learning in coordination games: Probability rules, heterogeneity, and time-variation. Journal of Mathematical Psychology, 42(2–3), 305–326. https://doi.org/10.1006/jmps.1998.1217
  • Cook, T. R., Kazinnik, S., Modig, Z., & Palmer, N. M. (2026). What do LLMs want? (Finance and Economics Discussion Series 2026-006). Board of Governors of the Federal Reserve System. https://doi.org/10.17016/FEDS.2026.006
  • Diederich, A., & Busemeyer, J. R. (1999). Conflict and the stochastic-dominance principle of decision making. Psychological Science, 10(4), 353–359. https://doi.org/10.1111/1467-9280.00167
  • Duffy, J., & Feltovich, N. (2002). Do actions speak louder than words? An experimental comparison of observation and cheap talk. Games and Economic Behavior, 39(1), 1–27. https://doi.org/10.1006/game.2001.0892
  • Ellsberg, D. (1961). Risk, ambiguity, and the Savage axioms. The Quarterly Journal of Economics, 75(4), 643–669. https://doi.org/10.2307/1884324
  • Fehr, E., & Fischbacher, U. (2004). Third-party punishment and social norms. Evolution and Human Behavior, 25(2), 63–87. https://doi.org/10.1016/S1090-5138(04)00005-4
  • Fischbacher, U., Gächter, S., & Fehr, E. (2001). Are people conditionally cooperative? Evidence from a public goods experiment. Economics Letters, 71(3), 397–404. https://doi.org/10.1016/S0165-1765(01)00394-9
  • Fiske, S. T., Cuddy, A. J. C., Glick, P., & Xu, J. (2002). A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology, 82(6), 878–902. https://doi.org/10.1037/0022-3514.82.6.878
  • Forsythe, R., Horowitz, J. L., Savin, N. E., & Sefton, M. (1994). Fairness in simple bargaining experiments. Games and Economic Behavior, 6(3), 347–369. https://doi.org/10.1006/game.1994.1021
  • Fudenberg, D., Rand, D. G., & Dreber, A. (2012). Slow to anger and fast to forgive: Cooperation in an uncertain world. American Economic Review, 102(2), 720–749. https://doi.org/10.1257/aer.102.2.720
  • Güth, W., Schmittberger, R., & Schwarze, B. (1982). An experimental analysis of ultimatum bargaining. Journal of Economic Behavior & Organization, 3(4), 367–388. https://doi.org/10.1016/0167-2681(82)90011-7
  • Hoffman, E., McCabe, K., Shachat, K., & Smith, V. L. (1994). Preferences, property rights, and anonymity in bargaining games. Games and Economic Behavior, 7(3), 346–380. https://doi.org/10.1006/game.1994.1056
  • Holt, C. A., & Laury, S. K. (2002). Risk aversion and incentive effects. American Economic Review, 92(5), 1644–1655. https://doi.org/10.1257/000282802762024700
  • Huang, K., Yeomans, M., Brooks, A. W., Minson, J., & Gino, F. (2017). It doesn’t hurt to ask: Question-asking increases liking. Journal of Personality and Social Psychology, 113(3), 430–452. https://doi.org/10.1037/pspi0000097 (correction: https://doi.org/10.1037/pspi0000491)
  • Huber, J., Payne, J. W., & Puto, C. (1982). Adding asymmetrically dominated alternatives: Violations of regularity and the similarity hypothesis. Journal of Consumer Research, 9(1), 90–98. https://doi.org/10.1086/208899
  • International Personality Item Pool. (n.d.-a). Administering IPIP measures, with a 50-item sample questionnaire. https://ipip.ori.org/new_ipip-50-item-scale.htm
  • International Personality Item Pool. (n.d.-b). Big-Five factor markers. https://ipip.ori.org/newBigFive5broadKey.htm
  • Jorgenson, E. (2015). Open Extended Jungian Type Scales 1.2. Open Psychometrics. https://openpsychometrics.org/tests/OJTS/development/OEJTS1.2.pdf
  • Laibson, D. (1997). Golden eggs and hyperbolic discounting. The Quarterly Journal of Economics, 112(2), 443–478. https://doi.org/10.1162/003355397555253
  • Laurenceau, J.-P., Barrett, L. F., & Pietromonaco, P. R. (1998). Intimacy as an interpersonal process: The importance of self-disclosure, partner disclosure, and perceived partner responsiveness in interpersonal exchanges. Journal of Personality and Social Psychology, 74(5), 1238–1251. https://doi.org/10.1037/0022-3514.74.5.1238
  • Liu, S., Zheng, C., Demasi, O., Sabour, S., Li, Y., Yu, Z., Jiang, Y., & Huang, M. (2021). Towards emotional support dialog systems. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 3469–3483). https://doi.org/10.18653/v1/2021.acl-long.269
  • Lorè, N., & Heydari, B. (2024). Strategic behavior of large language models and the role of game structure versus contextual framing. Scientific Reports, 14, 18490. https://doi.org/10.1038/s41598-024-69032-z
  • Mei, Q., Xie, Y., Yuan, W., & Jackson, M. O. (2024). A Turing test of whether AI chatbots are behaviorally similar to humans. Proceedings of the National Academy of Sciences, 121(9), e2313925121. https://doi.org/10.1073/pnas.2313925121
  • Rapoport, A., & Chammah, A. M. (1965). Prisoner’s dilemma: A study in conflict and cooperation. University of Michigan Press. https://doi.org/10.3998/mpub.20269
  • Regenwetter, M., Dana, J., & Davis-Stober, C. P. (2010). Testing transitivity of preferences on two-alternative forced choice data. Frontiers in Psychology, 1, 148. https://doi.org/10.3389/fpsyg.2010.00148
  • Regenwetter, M., Dana, J., & Davis-Stober, C. P. (2011). Transitivity of preferences. Psychological Review, 118(1), 42–56. https://doi.org/10.1037/a0021150
  • Rupprecht, J., Ahnert, G., & Strohmaier, M. (2026). Prompt perturbations reveal human-like biases in large language model survey responses. In Proceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science (pp. 1–21). https://doi.org/10.18653/v1/2026.nlpcss-1.1
  • Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations. https://arxiv.org/abs/2310.11324
  • Sharma, A., Miner, A. S., Atkins, D. C., & Althoff, T. (2020). A computational approach to understanding empathy expressed in text-based mental health support. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 5263–5276). https://doi.org/10.18653/v1/2020.emnlp-main.425
  • Tjuatja, L., Chen, V., Wu, T., Talwalkar, A., & Neubig, G. (2024). Do LLMs exhibit human-like response biases? A case study in survey design. Transactions of the Association for Computational Linguistics, 12, 1011–1026. https://doi.org/10.1162/tacl_a_00685
  • Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.-W., & Yu, D. (2025). LongMemEval: Benchmarking chat assistants on long-term interactive memory. In The Thirteenth International Conference on Learning Representations. https://arxiv.org/abs/2410.10813
  • Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Zhang, Y., Gong, N. Z., & Xie, X. (2023). PromptRobust: Towards evaluating the robustness of large language models on adversarial prompts (arXiv:2306.04528). arXiv. https://arxiv.org/abs/2306.04528
  • Zhu, Q. (2026). Whose welfare does AI maximize? Decision perspectives in economic games: Evidence from a meta-analysis and LLM experiments [Working paper].