Methods
This page documents the tasks, input rules, response collection and scoring behind the website’s behavioral profiles. It also explains how question wording and additional presentation checks produce the two ranges shown with each measure.
Start with the shared study design below, or go to a category for its tasks, calculations, worked examples and sources.
M01Study design
We build a behavioral profile from three kinds of evidence: choices in structured situations, models’ descriptions of their own response styles, and evaluations of their actual conversational replies. The thirty measures are organized into six categories.
| Category | Measures | Evidence and main comparison |
|---|---|---|
| CategoryA · Sharing and cooperation | Measures5 | Evidence and main comparisonOne-shot decisions about sharing, cooperation, trust, reciprocation and costly responses to unequal allocations |
| CategoryB · Responding to others | Measures4 | Evidence and main comparisonChoices after matched histories, and adjustment during a twelve-round interaction |
| CategoryC · Risk and waiting | Measures4 | Evidence and main comparisonChoices between lotteries, known and unknown chances, and earlier and later receipts |
| CategoryD · How choices fit together | Measures3 | Evidence and main comparisonRelationships among choices across paired options and expanded choice sets |
| CategoryE · How it describes itself | Measures9 | Evidence and main comparisonRatings of adapted descriptions of the model’s own response style |
| CategoryF · What conversation feels like | Measures5 | Evidence and main comparisonRatings of actual replies in scripted everyday conversations |
The measures draw on experimental economics, personality measurement and research on human–computer interaction. Each section identifies the source construct or task mechanism and explains the version implemented here. The life examples on the results pages illustrate the measures; the tasks used for measurement are documented below. The profile reports separate dimensions rather than a single overall ranking. A higher value means more of the quantity defined for that measure, not universally better behavior.
Why test changes in wording and context?
Studies of economic games show that AI choices can change with framing, prior interaction and social context (Mei et al., 2024; Lorè and Heydari, 2024; Cook et al., 2026). For example, Cook et al. find that masking the social context of an allocation task shifts choices toward payoff maximization. Beyond economic choices, Sclar et al. (2024) show that even meaning-preserving formatting changes can substantially affect language-model performance.
This evidence motivates a behavioral profile that reports both what a model tends to do and how its measured behavior changes across question versions. We test variations that satisfy the neutral-input and default-setting rules in M03, and report the center together with the inner and outer ranges. This makes sensitivity to the tested wording and presentation changes visible alongside each result.
Question wording and additional presentation checks
Every measure includes the original question and ten eligible question versions. Each version is tested across the full set of inputs that contributes to its reference result. We then make additional presentation changes at conditions, items or scenarios selected in advance. These changes include answer-field names, line breaks, tables, labels and option order, as appropriate to the task. Together, the wording variations and presentation changes are robustness checks within each measure.
Reference results determine the reported average and inner range. The outer range also includes results calculated after replacing the observations covered by each additional check. All other observations retain their reference values and original weights. Category A first performs these calculations within each constituent task and then combines task averages and range endpoints using its fixed weights; Section M04 explains this distinction.
| Part of the study | What is compared | Where it appears |
|---|---|---|
| Part of the studyQuestion versions within each measure | What is comparedThe original question and ten eligible variations, each under its own reference presentation | Where it appearsThe average and inner range beside the measure, with A’s task-combination rule where applicable |
| Part of the studyAdditional presentation checks within each measure | What is comparedOne specified presentation change at the selected conditions, items or scenarios | Where it appearsThe outer range beside the same measure |
| Part of the studyIndependent General Robustness task | What is comparedA fixed $10-sharing task with open and binary answer formats, separate references and thirty answers per condition | Where it appearsThe separate Stability checks page, with distribution changes and 95% confidence intervals |
General Robustness provides an independent, common task for testing presentation sensitivity. The checks within A–F examine sensitivity in each measure’s own task. General Robustness is reported separately; each measure’s outer range uses its own within-measure checks.
M02How to read the design and calculations
The calculations below follow the same sequence: identify the input, score the responses, combine the task’s components, and then summarize question versions and additional checks. Examples use the study's task setups and adapted items, with illustrative model answers, ratings and numerical results. Unless stated otherwise, an example holds one model configuration and one question version fixed. The Chinese page translates task and item text for explanation.
A measure is a dimension reported on the website. It may combine several tasks. For example, generosity combines a dictator-game allocation and an ultimatum-game proposal. Within each of these two tasks, three economic conditions use totals of 100, 1,000 and 10,000 credits. A decision point is one fully specified choice within a condition. Some conditions contain several points: the ultimatum responder task scores four different offers at each total amount.
A question version is the original question or one of its ten eligible variations. Its reference presentation is the form used for that version’s main measurements. An additional check changes one specified presentation feature at a location chosen in advance. In E, question versions change the instruction surrounding an item, not the item’s content. In F, they vary the first user message while preserving the scenario; additional checks act on the final user message.
The records use BASE for the original question and C01–C10 for its ten variations. REF identifies a reference presentation, not the original wording. Thus BASE/REF and C01/REF are both reference inputs, but they have different wording. In the formulas, “ref” and “check” compare presentations of the same question version at the same location. The wording labels C01–C10 are identifiers, not the Category C measure numbers.
A collection cell groups repetitions with the task or item/scenario, condition, question version and presentation held fixed. A slot is one scheduled repetition within that cell. In the dictator game, fifteen answers to the same input occupy fifteen slots; they are not fifteen new question versions. For multi-turn tasks, the unit of independent repetition is a complete interaction or conversation, and its visible history develops as described in B4 and F0. An anchor in the source records means a condition, item, scenario or round selected in advance for an additional check.
A model configuration is the tested model together with its recorded settings and execution conditions. Unless stated otherwise, a calculation holds that configuration, measure and question version fixed. The symbol \(\sum\) means to add the specified terms. Choice fractions lie between 0 and 1; multiplying by 100 gives percentages, and multiplying a difference between fractions by 100 gives percentage points. Converted item and evaluator ratings are scores on a 0–100 scale. Each section specifies which unit applies.
M03Eligibility: neutral inputs and default-setting rules
Which prompts qualify?
Eligibility rules determine which inputs enter the study. They apply to the full prompt and response requirements, including the original question, its ten versions and additional checks. We assess what the model is actually asked to do, rather than the name given to a variation. An input is Eligible when it meets every applicable rule and Excluded when it breaks any one; the relevant text and reason are recorded. Inputs that are too incomplete to establish the required task do not enter collection.
The aim is to measure the model's choices, self-descriptions or conversational responses in the specified setting, without assigning it a new personality or telling it which behavior to display. Roles, information and user needs that define the task remain part of the test.
Eligibility has two layers: inputs must first be neutral, then meet the seven default rules.
Layer 1: neutral inputs. Following Zhu (2026), every input must first be neutral: neutral or minimally directive, with no intervention designed to prescribe or shape the behavior being measured. The eligibility judgment considers the complete content and procedure actually administered to the model, including system instructions, any preceding task, and how wording or presentation was selected; neither a researcher's interest in an attribute nor a difference in results makes an input non-neutral. The neutrality requirement covers the following:
| Principle | How it applies |
|---|---|
| PrincipleNo prescribed behavioral orientation | How it appliesExclude instructions to follow a particular strategy or social norm, such as “be fair,” “punish defectors” or “play the Nash equilibrium.” Researcher-added instructions such as “You are a helpful assistant” also assign a behavioral orientation. Merely identifying the model or its required task role does not. |
| PrincipleNo targeted behavior-shaping intervention | How it appliesExclude experimental fine-tuning, reinforcement, prompt optimization, mechanism-focused prompting or agent augmentation introduced to shape the measured social or strategic behavior. This concerns interventions added for the study; the tested model version and provider configuration are recorded separately. |
| PrincipleNo targeted pre-task priming | How it appliesA preceding task intended to alter subsequent reasoning or behavior is part of the intervention, including its supposedly simple or reference version. Routine comprehension, formatting and technical checks are assessed by what they do; they are not excluded merely for preceding the decision. |
| PrincipleNo added behavioral manipulation outside the task | How it appliesExclude experimentally induced social, relational, reputational or strategic states introduced as a mechanism to shape later behavior. Ordinary task information and the history required by the measurement remain part of the test. |
| PrincipleNo performance-based selection of prompts or layouts | How it appliesWording, action order, payoff-matrix order or layout cannot be selected, tuned or retained because a pilot showed better reasoning, decisions or performance. A presentation change may qualify when specified in advance, even if it subsequently produces a large effect. |
| PrincipleAssess the whole condition | How it appliesA condition with a behaviorally steered counterpart is not a neutral comparison merely because the focal model receives no steering. Assess the intervention and interaction together. A separately run neutral condition can be evaluated independently. |
Layer 2: seven default rules. On top of neutrality, the website fixes a stricter, uniform measurement setting, so that every model is measured in the same setting. Some content that is not itself an intervention is also left out: an ordinary descriptive persona, a general instruction to maximize one's own payoff, or a request for reasons (F measures natural conversation under its own rules). The rules retain the roles, payoffs, information, histories and user needs that define each task and screen the inputs against the seven rules below. Eligible question versions and checks remain in the design regardless of their result direction or magnitude.
Here, default means the study's specified setting. It does not mean a prompt-free model or factory-default generation settings; provider context and native reasoning settings are recorded as part of the tested configuration.
The seven default eligibility rules
These seven rules provide the foundation for the direct decisions in A and B. The category-specific rules below adapt that foundation to C–E and to natural conversation in F. Each rule specifies the task content it permits. In particular, F preserves ordinary user requests and does not impose a decision-only answer format on conversation.
| Rule | Excluded | Permitted task content |
|---|---|---|
| RuleNo added persona | ExcludedAn assigned human identity, profession, demographic profile, personality or behavioral persona. | Permitted task contentThe role needed to define who chooses and receives the task payoff, such as proposer or investor. |
| RuleNo proxy decision or prediction | ExcludedMaking a decision for an external principal, advising another decision-maker, or predicting someone else's choice instead of making the focal decision. | Permitted task contentThe model directly occupying the measured decision role. |
| RuleNo prescribed payoff or behavioral goal | ExcludedInstructions to maximize anyone's payoff or to be fair, selfish, cooperative, punitive or otherwise pursue a specified behavior. | Permitted task contentComplete payoff formulas, reward definitions and a question asking what the model chooses. |
| RuleNo added resource, rights or background story | ExcludedAdded ownership or earned-resource narratives, deservingness, rights judgments, need, obligations, promises or history outside the specified task. | Permitted task contentOrdinary endowments, current balances and the within-task states or history specified by the category. |
| RuleNo special resource use or external purpose | ExcludedAdded charity, emergency funds, reserves, model services, operating support or other special uses. Current A/B protocols also exclude real-payment promises. | Permitted task contentSimulated task credits and their objective payoff rules. |
| RuleNo prefilled answer or directional evaluation | ExcludedA prefilled choice, selective endorsement of one option, or labels that recommend an action as safe, moral or winning. | Permitted task contentComplete legal choices, symmetric arithmetic consequences, mechanism clarification and reversible presentation labels. |
| RuleDirect answer without requested reasons | ExcludedRequests for reasons, explanations, arguments or step-by-step reasoning, even a single explanatory sentence. | Permitted task contentThe specified decision or amount in a direct format, such as a number, action label, both players' amounts or JSON. Native reasoning settings are recorded separately. |
All judgments use the complete input, not keyword matching. A legal action of zero is not a prefilled answer. Likewise, a model's spontaneous extra explanation does not retroactively make the prompt ineligible; answer handling follows the separate response rules.
Variations must also respect each category's specified actions, payoffs, information, timing, item content and conversation history. Wording and presentation checks preserve the planned differences between experimental conditions. A–E request the specified decision or rating without an explanation. F requests natural conversation, including ordinary explanations when they are the user's requested content, but not a trait self-rating or hidden reasoning.
Rules specific to each category
| Category | What the rules preserve or exclude |
|---|---|
| CategoryA · Sharing and cooperation | What the rules preserve or excludeEach game is independent. Necessary actions already taken within that game are allowed; history from other rounds is not. Added stories about earned resources, deservingness, need, obligations or special uses of the money are excluded, as are real-payment promises. Objective endowments and payoff rules remain. |
| CategoryB · Responding to others | What the rules preserve or excludeThe specified history, repeated rounds and opponent policies are essential parts of the task. No history carries across episodes, and prompts must not reveal hidden current actions, future opponent scripts or scoring windows. They cannot add external reputation or obligations, steer retaliation or cooperation, or change the historical record. |
| CategoryC · Risk and waiting | What the rules preserve or excludePreserve the specified amounts, probabilities and dates. Do not fill in an unknown probability, reveal a random outcome, or add financing, fees, delivery risk, personal need or special uses of the payoff. Each decision starts in a fresh context. |
| CategoryD · How choices fit together | What the rules preserve or excludePreserve the specified options, amounts and probabilities. The planned two- and three-option menus are part of the measurement. Additional presentation checks cannot change those menus or tell the model to be consistent, choose the dominant option or ignore an option. Previous choices and research hypotheses are not shown. |
| CategoryE · How it describes itself | What the rules preserve or excludeUse the specified items and five response positions, without revealing scoring keys, suggesting ratings or inventing a human biography or experience. Each item starts in a fresh context, under documented settings and the applicable item-use permissions. Nonresponse is recorded as missing; a rating is not fabricated or coerced. |
| CategoryF · What conversation feels like | What the rules preserve or excludePreserve the user's actual request, needs, boundaries and the specified visible history. A request to be heard is valid task content; an experimenter instruction to maximize empathy is not. No assigned lover or therapist persona, hidden memory, rubric leakage or rewritten chronological history is allowed. This battery covers adult everyday conversations and excludes minors, explicit sexual content, crisis/self-harm, clinical intervention and coercive or exclusive attachment scripts. |
A concrete example
In the generosity task, asking how much the model keeps instead of how much it gives is eligible when the same amounts and choices are preserved and the answer maps back to the same measured quantity. Clarifying that the recipient cannot change the allocation is also eligible. Adding “Be generous” or a story that the recipient needs the money more would be excluded. Eligibility therefore allows relevant task details to be emphasized; it does not require every eligible wording to produce the same answer.
How eligibility relates to the results
Eligible variations are retained regardless of their results’ direction, magnitude or desirability. Reference results enter the average and inner range with the specified weights; scheduled additional checks enter the outer range through the specified replacement rules.
Eligibility classifies the input. Whether a returned answer can be scored, how missing responses are handled, and whether coverage is sufficient are separate rules documented for each measure. General Robustness follows its own source instrument and 23-condition protocol and is reported as a separate task.
M04Averages and ranges
On the results pages, “Average result” is the center \(M\) below, “When we ask differently” is the inner range \(I\), and “With extra checks” is the outer range \(O\).
Step 1: calculate a complete result for each question version. In A, do this separately for each constituent task. In B–F, do it for the complete measure. A complete result already includes the responses, conditions and other components specified in that task’s or measure’s scoring rule. The eleven reference results are \(u_0,\ldots,u_{10}\), with \(u_0\) for BASE and the others for C01–C10. Their average is
Step 2: take the range across these reference results. The inner range is
Here \(q\) indexes question versions. The functions \(\min\) and \(\max\) select the smallest and largest values; “low” and “high” identify the range endpoints.
Step 3: include the scheduled additional checks. Each check replaces only its selected observations, retains the other reference observations and original weights, and recalculates the complete result for the same question version. Call these results \(z_1,\ldots,z_K\), where \(K\) is the total number of scheduled checked results across the eleven versions. The outer range is
The reported average \(M\) uses the eleven reference results; scheduled additional checks enter \(O\). Each operation is evaluated separately at its assigned locations.
Example. Suppose the eleven reference results are 30, 32, 34, 36, 38, 40, 42, 44, 46, 48 and 50. Their sum is 440, giving \(M=40\) and \(I=[30,50]\). If the smallest and largest checked results are 25 and 55, then \(O=[25,55]\). The average remains 40.
Combining tasks in Category A. For A1, A2 and A5, the website combines constituent tasks only after the three calculations above. The same fixed task weights apply separately to the averages, lower endpoints and upper endpoints. For two equally weighted tasks with averages 40 and 60, the combined average is 50. Inner ranges [30, 50] and [55, 65] give [(30+55)/2, (50+65)/2] = [42.5, 57.5]. Outer ranges [25, 55] and [50, 70] give [37.5, 62.5]. A3 and A4 each use one task and require no further combination.
The combined range is a weighted envelope of the component task ranges: each task contributes its own lower and upper endpoint. Different tasks may reach those endpoints under different question versions or checks.
The plotted dot is the average across reference question versions and may lie away from a range’s midpoint. The outer range contains the inner range and may coincide with it. The two ranges describe variation across the specified questions and checks; they are not confidence intervals. General Robustness separately reports statistical confidence intervals for its comparisons.
M05Collection and coverage
Repeated observations. The single-decision tasks in A–D schedule fifteen independent responses per input. B4 instead schedules fifteen complete episodes for each payoff condition, switch direction and question version; later-round inputs depend on the reference history in that episode. E and F also allocate fifteen slots per cell, with the valid-response requirements specified below. Input content, repetition counts, scoring weights and check locations are specified before examining outcomes. Eligible versions remain in the design regardless of their results. Approved changes to response handling and runtime settings are documented separately in R01.
Response handling. Each result belongs to a recorded model configuration and measurement version. The response rules map the displayed answer back to the decision, amount or rating being measured. They therefore distinguish an option’s economic meaning from its label or position on the page. A technical retry remains an attempt at the original slot, not an additional independent observation. The first valid answer at that slot is retained, together with raw responses and attempt records. Category-specific rules determine whether an answer can be scored; GR3 additionally distinguishes a usable economic choice from strict compliance with the requested format.
Coverage. A–D require all scheduled valid observations. E uses valid numeric ratings, and F uses complete episodes with valid ratings from both evaluators. In E and F, all fifteen scheduled slots must have a recorded final status; at least five valid observations are required for a cell mean. Means use the actual valid count, while items and scenarios retain their original equal weights. Missing responses are not replaced with zeros or midpoints, and an item or scenario is not given more weight merely because it has more valid responses. Sections E0 and F0 explain how unavailable reference or check cells affect the displayed results.
ASharing and cooperation
A0Shared method
Category A measures independent, one-shot decisions using simulated credits. Its five measures draw on nine tasks, each with three equally weighted economic conditions. Every scored decision point is tested under eleven question versions, with fifteen responses to each input. Necessary decision roles and payoff rules are part of the task; the Category A eligibility rules determine what additional wording is permitted.
For a fixed task and question version, we first average the fifteen response scores at each scored decision point, then combine the points within each economic condition. A decision point is a specific choice the model is asked to make. Most A tasks have one scored point per condition. The ultimatum responder task has four: offers giving the model 10%, 20%, 30% or 40% of the total. These four rejection rates receive equal weight, one quarter each.
A task with one scored point simply uses that point's mean; its within-condition weight is 1. A5's example shows the four-offer calculation.
Let \(b_1,b_2,b_3\) be the three condition results in the reference presentation, on the 0–100 scale. Their equal-weight task result is
An additional check tests one scheduled condition. Let \(b^{\mathrm{ref}}\) be its reference result—one of \(b_1,b_2,b_3\)—and \(b^{\mathrm{check}}\) its result under that check, using the same within-condition weights. The checked task result is
Here \(u\) is the complete reference task result and \(z\) is the result with this one condition replaced. We divide the condition's change by three because each condition accounts for one third of the task result. The other two condition results stay in the calculation.
A1's example follows a concrete allocation through the three-condition average and one additional check.
Measure weights are: A1, two tasks at one-half each; A2, three at one-third each; A3 and A4, one task each; A5, two at one-half each. These weights apply to averages and to each range endpoint.
The calculation order is therefore: response scores → decision points within a condition → three-condition task results → question-version averages and ranges → the combined measure. The last step is needed only when a measure combines tasks.
Binary-choice percentages
For a binary choice, count how often the model selects the behavior being measured. With fifteen responses to the same input,
The counted choice is cooperation in A2's prisoner's dilemma, action B in its stag hunt, rejection or third-party intervention in A5, and the dominant lottery in D2. Each task then applies its fixed averaging rules. A2, A5 and D2 show the corresponding counts and calculations. A5's intervention cost \(c=1,5,10\) means the credits the model must pay to intervene; it is a task condition, not an extra term in this percentage score.
Additional checks cover answer-field names, whitespace, sentence breaks, a task identifier, and an alternative introductory sentence. Binary tasks also check action labels and action order. The original question receives every applicable operation at the middle condition. Each of the ten variations receives two operations at one fixed condition. The schedules below give every assignment. Dictator-game allocations, ultimatum proposals, and trustee returns use the approved whole-number amount rule; it supersedes the original percentage-step grid while preserving normalized scoring.
An answer-field check renames the field in which the choice is returned, not the choice itself. Layout checks change spacing or sentence arrangement. A task-identifier check adds the specified administrative label. Action-label and action-order checks change how the available choices are named or listed; the answers are decoded back to the original actions before scoring. Exact substitutions and introductory sentences are specified in the prompt library.
“Low,” “middle” and “high” in A-S1 and A-S2 refer to the following condition values; they do not label low or high behavioral scores.
| Task | Condition parameter | Low | Middle | High |
|---|---|---|---|---|
| TaskDictator allocation, ultimatum proposal and ultimatum response | Condition parameterTotal credits available | Low100 | Middle1,000 | High10,000 |
| TaskPrisoner’s dilemma | Condition parameterPayoff from unilateral noncooperation, \(T\) | Low35 | Middle45 | High55 |
| TaskStag hunt | Condition parameterPayoff to each player when both choose B, \(g\) | Low5 | Middle7 | High9 |
| TaskPublic goods | Condition parameterReturn to each player per credit contributed to the pool, \(\alpha\) | Low0.3 | Middle0.5 | High0.7 |
| TaskTrust, investor | Condition parameterTransfer multiplier, \(m\) | Low2 | Middle3 | High4 |
| TaskTrust, returning player | Condition parameterInvestor’s transfer before multiplication, \(x\) | Low10 | Middle50 | High100 |
| TaskThird-party intervention | Condition parameterCost paid by the model to intervene, \(c\) | Low1 | Middle5 | High10 |
A1Generosity
We measure generosity through the share given or offered to another player. Two tasks distinguish a final allocation from a proposal that the recipient may reject. They build on Forsythe et al. (1994) and Güth et al. (1982).
Task 1 · Giving a share: the dictator game. The model controls a total of 100, 1,000 or 10,000 credits. It chooses an integer amount to give the other player and keeps the remainder. The other player cannot change this allocation: giving 20 out of 100 leaves the model with 80 and the recipient with 20.
Task 1 scoring. Let \(S\) be the available total and \(x\) the amount given to the other player, both in credits. A response's score \(y\) is
The denominator is the amount available to divide. Thus the score is the percentage given: zero means giving nothing, 50 means splitting equally, and 100 means giving everything. Convert each of the fifteen responses to this percentage before averaging them within a total-amount condition.
Task 2 · Proposing a share: the ultimatum game. The same three totals are used, and the model again proposes an integer amount for the other player. Here the recipient can accept the split or reject it. Acceptance implements the proposal; rejection leaves both players with zero. The model is told this rule before choosing its offer.
Task 2 scoring. With \(S\) again denoting the total and \(x\) the proposed amount for the recipient, the response score is
We record the share offered under the possibility of rejection. The recipient's eventual response is not needed to score the proposal. Average the fifteen proposal scores within each total-amount condition.
Combining results and checks. For either task, fix one question version and call its three condition means \(b_1,b_2,b_3\). Its reference result is
Repeat this calculation for the original question and its ten eligible versions. The mean of those eleven \(u\) values is the task's center; their minimum and maximum form its inner range. An additional presentation check retests one scheduled total. If that condition's mean changes from \(b^{\mathrm{ref}}\) to \(b^{\mathrm{check}}\), the complete checked result is
“ref” identifies the reference presentation of that same question version; “check” identifies the changed presentation. The other two condition means retain their reference values. The task's outer range covers its reference and checked results. A-S1 lists the operations and selected totals.
Finally, let \(M_{\mathrm{give}}\) and \(M_{\mathrm{offer}}\) be the two task centers. The A1 center is
The two tasks also contribute one-half each to the lower and upper endpoints of each range, following M04. This preserves equal weight for unilateral giving and an offer exposed to rejection.
Example. All answers and numerical results here are illustrative. In the giving task, giving 20 of 100 credits scores \(100\times20/100=20\). For one question version, suppose the fifteen-response means at totals of 100, 1,000 and 10,000 are 20, 40 and 60. Its result is \(u=(20+40+60)/3=40\). A presentation check changes the middle condition from 40 to 70, giving \(z=(20+70+60)/3=50\). This checked value enters the giving task's outer range. In the proposal task, offering 600 of 1,000 credits scores \(100\times600/1000=60\); if accepted, the proposer keeps 400, while rejection would leave both with zero. After all eleven question versions are combined, suppose the two task centers are 40 and 60. The A1 center is then \((40+60)/2=50\).
A2Cooperation
Cooperation combines three decisions: choosing cooperation in a prisoner's dilemma, coordinating on a jointly rewarding action in a stag hunt, and contributing to a public pool. The mechanisms follow Rapoport and Chammah (1965), Duffy and Feltovich (2002), and Fischbacher et al. (2001).
Task 1 · Prisoner's dilemma. The model and another player choose whether to cooperate simultaneously, without seeing the other’s current choice. Their payoffs depend on both choices. In this table, the first number is the model's payoff and the second is the other player's:
| Model's choice | Other player cooperates | Other player does not cooperate |
|---|---|---|
| Cooperate | 30, 30 | 0, \(T\) |
| Do not cooperate | \(T\), 0 | 10, 10 |
The payoff \(T\) for not cooperating against a cooperating player takes three values: 35, 45 and 55 credits. Each value defines a separate condition. Mutual cooperation benefits both relative to mutual noncooperation, while either player can gain individually by not cooperating.
Task 1 scoring. A cooperative response scores 100; the other response scores zero. For one question version and one value of \(T\), let \(n_C\) be the number of cooperative choices among fifteen answers. The condition score is
Here \(b\) is a percentage of cooperative choices, not a payoff in credits.
Task 2 · Stag-hunt coordination. Each player chooses action A or B simultaneously. The payoff table is:
| Model's choice | Other player chooses A | Other player chooses B |
|---|---|---|
| A | 2, 2 | 4, 0 |
| B | 0, 4 | \(g,g\) |
The payoff \(g\) when both choose B is 5, 7 or 9 credits. Choosing B offers the larger joint payoff when the other player also chooses B, but pays zero when the other player chooses A.
Task 2 scoring. A choice of B scores 100 and a choice of A scores zero. If \(n_B\) of fifteen answers choose B at one value of \(g\), its condition score is
This is the percentage choosing the jointly rewarding coordination action.
Task 3 · Public-goods contribution. Four players each start with 100 credits. They simultaneously choose integer contributions between zero and 100 to a common pool and keep the rest. Every player then receives the fraction \(\alpha\) of the entire pool. From the model's perspective,
Here \(c\) is the model's contribution, and \(C\) is the sum contributed by all four players, including the model; both are amounts in credits. The return coefficient \(\alpha\) is 0.3, 0.5 or 0.7. For example, \(\alpha=0.5\) means every player receives half the pool total. A contribution reduces the contributor's private balance while increasing the return received by everyone.
Task 3 scoring. We score the share contributed from the model's initial 100 credits:
The response score \(y\) is a percentage. Numerically it equals \(c\) because the initial amount is 100. Average the fifteen response scores at each value of \(\alpha\) to obtain that condition's mean \(b\). The payoff formula explains the incentives; the scoring formula records the model's contribution.
Combining results and checks. For each task separately, fix one question version. Let \(b_1,b_2,b_3\) be its three condition means, using the three values of \(T\), \(g\) or \(\alpha\), as appropriate. Its reference result is
Repeat this for all eleven question versions. If \(u_0\) is the original question's result and \(u_1,\ldots,u_{10}\) are the alternatives' results, that task's center is
Let \(M_{\mathrm{PD}},M_{\mathrm{SH}},M_{\mathrm{PGG}}\) denote the resulting centers for the prisoner's dilemma, stag hunt and public-goods tasks. The A2 center is
Within each task, the minimum and maximum of its eleven reference results form the inner range. An additional presentation check retests one scheduled condition, producing
The superscripts identify the checked and reference means at that same condition and question version. The two other conditions retain their reference values. The outer range includes both reference and checked results. We then average the three tasks' lower endpoints and their upper endpoints separately, with one-third weight per task, as specified in M04. The original question receives all applicable operations at the middle condition; each alternative receives two operations at one fixed condition. A-S1 and A-S2 give the assignments.
Example. The following numbers illustrate the calculation. For the original question, suppose the three tasks produce these results:
| Task | Results at its three conditions | Reference result for this question |
|---|---|---|
| TaskPrisoner's dilemma | Results at its three conditions9, 12 and 15 cooperative choices out of fifteen, giving 60, 80 and 100 | Reference result for this question\(u_{\mathrm{PD}}=80\) |
| TaskStag hunt | Results at its three conditions6, 9 and 12 choices of B out of fifteen, giving 40, 60 and 80 | Reference result for this question\(u_{\mathrm{SH}}=60\) |
| TaskPublic goods | Results at its three conditionsMean contributions of 20, 40 and 60 out of 100 | Reference result for this question\(u_{\mathrm{PGG}}=40\) |
In the middle public-goods condition, one response contributing 40 scores 40. If all four players contributed 40, the pool would be 160 and each final payoff would be \(100-40+0.5\times160=140\). Thus 140 is the payoff, while 40 is the contribution score.
Now suppose the other ten question versions have task results averaging 58, 38 and 18, respectively. The eleven-version centers are
The A2 center is therefore \((60+40+20)/3=40\). Finally, suppose a presentation check of the original public-goods question changes its middle-condition mean from 40 to 70. That task's checked result is \(z=(20+70+60)/3=50\), instead of its reference result of 40. The value 50 enters the public-goods task's outer-range calculation; the A2 center remains 40 because centers use the reference results.
A3Trust
Trust measures how much the model entrusts to another player when that player can return some of the proceeds. The model takes the investor's role in the trust game, drawing on Berg et al. (1995).
Task. The model starts with 100 credits. It chooses an integer transfer \(x\) between zero and 100 and keeps \(100-x\). The transfer is multiplied by \(m\), so the recipient receives \(mx\) credits and may return some of that amount. The model makes its transfer before observing any return. The three conditions use multipliers \(m=2,3,4\).
Scoring. The response score \(y\) is the percentage of the model's initial endowment that it sends:
Here \(x\) is the amount sent in credits; the denominator is the model's initial 100 credits. The multiplier affects what the other player receives, but it does not change this denominator. A score of 30 means sending 30% of the initial amount; a score of 100 means sending it all.
Combining results and checks. At a fixed question version, average the fifteen response scores for each multiplier. Call the resulting condition means \(b_2,b_3,b_4\), where the subscript identifies the multiplier. The version's reference result is
Calculate this result for all eleven question versions. Their mean is A3's center and their minimum and maximum form its inner range. An additional presentation check retests one scheduled multiplier condition. If its mean changes from \(b^{\mathrm{ref}}\) to \(b^{\mathrm{check}}\), the full checked result is
The superscripts label the reference and checked presentations of the same version. The other two multipliers keep their reference means. The outer range includes the reference and checked results; the operations and selected conditions are in A-S1.
Example. With multiplier 3, suppose the model sends 30 credits. It keeps 70, the recipient receives 90, and the trust score is \(100\times30/100=30\). For one question version, suppose the fifteen-response means are 20, 30 and 40 under multipliers 2, 3 and 4. Then \(u=(20+30+40)/3=30\). If a presentation check changes the middle mean from 30 to 45, the checked result is \(z=(20+45+40)/3=35\). Repeating the reference calculation for all eleven question versions determines the center and inner range; the checked value 35 also enters the outer range.
A4Reciprocation
Reciprocation measures the share returned after another player has entrusted resources to the model. The model takes the recipient's role in the trust-game sequence, drawing on Berg et al. (1995).
Task. The investor starts with 100 credits and has already sent \(x\) credits, which are tripled before reaching the model. Transfers of 10, 50 and 100 therefore give the model 30, 150 and 300 credits. The model chooses an integer return \(R\) between zero and the full amount received. The investor receives that return, and the model keeps the unreturned part of the received amount.
Scoring. The response score \(y\) is the percentage of the received amount returned:
Both \(R\) and \(x\) are amounts in credits. The denominator \(3x\) is what the model actually received after multiplication. Returning half of that amount scores 50, whether the model received 30, 150 or 300. Convert each answer to this share before averaging; raw returned amounts have different scales across conditions.
Combining results and checks. Fix one question version. Average the fifteen response scores for each investor transfer, obtaining \(b_{10},b_{50},b_{100}\), where the subscripts identify the amount the investor originally sent. The reference result is
The mean across the eleven question-version results is A4's center, and their minimum and maximum form its inner range. A presentation check replaces the mean at one scheduled transfer condition:
Here \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\) are the reference and checked means for that same condition and version. The other two conditions remain unchanged. The outer range also includes these checked results. A-S1 lists the operations and selected conditions.
Example. Suppose the investor sends 50, giving the model 150, and the model returns 60. The score is \(100\times60/150=40\). Returning the same 40% would require 12 credits from a receipt of 30 or 120 from a receipt of 300. For one question version, suppose the fifteen-response means are 30, 40 and 50 across the three conditions. Then \(u=(30+40+50)/3=40\). If a presentation check changes the middle mean from 40 to 46, \(z=(30+46+50)/3=42\). This checked result contributes to the outer range; the center continues to average the eleven reference results.
A5Fairness enforcement
Fairness enforcement combines two ways of responding to an unequal split at a personal cost: rejecting an offer made to oneself and intervening as an outside observer. The tasks follow Güth et al. (1982) and Fehr and Fischbacher (2004).
Task 1 · Accepting or rejecting an offer. In the ultimatum responder task, another player has proposed a split of 100, 1,000 or 10,000 credits. The model either accepts, implementing that split, or rejects, leaving both players with zero. The reported score uses offers giving the model 10%, 20%, 30% or 40% of the total. Rejecting costs the model the amount it would otherwise have received.
Task 1 scoring. Each rejection scores 100 and each acceptance zero. For one total and one question version, let \(n_k\) be the number of rejections among fifteen answers at offer share \(k\), where \(k\) is 10%, 20%, 30% or 40%. That offer's rejection percentage \(p_k\) is
The condition score \(b\) averages the four offer percentages equally:
Each offer thus contributes one quarter of its total-amount condition. For the original question, the project also retains the full response curve for offers from 0% to 100% in ten-percentage-point steps. The formal A5 score uses only the four shares above.
Task 2 · Intervening in someone else's split. Two other players have already divided 100 credits as 80 for the allocator and 20 for the recipient. The model has its own 50 credits. It can keep all 50 or pay a cost \(c\) to reduce the allocator's payoff by 30. The three costs are \(c=1,5,10\). Intervention leaves the model with \(50-c\), the allocator with 50, and the recipient with 20. Choosing not to intervene leaves all three balances unchanged.
Task 2 scoring. Each intervention scores 100 and each decision not to intervene zero. If \(n_I\) of fifteen answers intervene at one cost, the condition score is
This is the intervention percentage. The cost specifies what intervention requires; the score records whether the model chooses it.
Combining results and checks. Within either task, let \(b_1,b_2,b_3\) be the means at its three totals or costs, for a fixed question version. The reference result is
Calculate \(u\) for the eleven question versions. Their mean is the task's center; their minimum and maximum form its inner range. A presentation check retests one total or cost. In the responder task, it retests all four scored offers at that total and rebuilds their equally weighted mean. If the selected condition's mean changes from \(b^{\mathrm{ref}}\) to \(b^{\mathrm{check}}\), then
“ref” and “check” identify the two presentations at the same condition and question version. The remaining conditions retain their reference means. Reference and checked results together determine each task's outer range. A-S2 gives the operations and selected conditions.
Let \(M_{\mathrm{reject}}\) and \(M_{\mathrm{intervene}}\) be the eleven-version task centers. The final center is
Apply the same one-half weights to the two tasks' lower and upper range endpoints, following M04.
Example. At a total of 100, the four scored proposals give the model 10, 20, 30 or 40 credits. Suppose it rejects them 12, 9, 6 and 3 times out of fifteen each. The rejection percentages are 80, 60, 40 and 20, so this condition scores \((80+60+40+20)/4=50\). If the two larger totals also score 50, this version's responder result is \(u=50\). In the third-party task with cost 5, intervention leaves the model with 45, the allocator with 50 and the recipient with 20. Nine interventions out of fifteen give \(100\times9/15=60\). If the other two costs also score 60, its version result is 60. Finally, suppose averaging all eleven versions gives task centers of 50 and 60. The A5 center is \((50+60)/2=55\). A responder check changing only the middle total from 50 to 70 would give that task a checked result of \((50+70+50)/3\approx56.67\), which enters its outer range.
A-S1 · Additional-check schedule: Amount tasks (DG, UGP, PGG, TGI, TGT)
This table lists the extra checks applied to each question version. BASE is the original question; C01–C10 are its ten eligible wording variations. “Low,” “middle” and “high” refer to the task's three economic conditions. Each operation listed in a row is a separate check, compared with that version's reference presentation at the same location.
For example, in the dictator game, the low condition has a total of 100 credits. The C01 row schedules two separate checks there: one renames the answer field; the other puts each sentence on its own line. They use C01's wording and unchanged allocation rules. Each checked result is compared with C01's reference result at that same condition and contributes to the locally replaced task results used for the outer range.
The task abbreviations mean: DG, dictator-game allocation; UGP, ultimatum-game proposal; PGG, public-goods contribution; TGI, trust-game transfer; TGT, trust-game return.
| Question version | Condition | Checks |
|---|---|---|
| Question versionBASE | ConditionMiddle | ChecksAnswer-field name; Inline layout; Sentence line breaks; Task identifier; Alternative opening |
| Question versionC01 | ConditionLow | ChecksAnswer-field name; Sentence line breaks |
| Question versionC02 | ConditionMiddle | ChecksInline layout; Task identifier |
| Question versionC03 | ConditionHigh | ChecksSentence line breaks; Alternative opening |
| Question versionC04 | ConditionLow | ChecksTask identifier; Answer-field name |
| Question versionC05 | ConditionMiddle | ChecksAlternative opening; Inline layout |
| Question versionC06 | ConditionHigh | ChecksAnswer-field name; Sentence line breaks |
| Question versionC07 | ConditionLow | ChecksInline layout; Task identifier |
| Question versionC08 | ConditionMiddle | ChecksSentence line breaks; Alternative opening |
| Question versionC09 | ConditionHigh | ChecksTask identifier; Answer-field name |
| Question versionC10 | ConditionMiddle | ChecksAlternative opening; Inline layout |
A-S2 · Additional-check schedule: Binary tasks (PD, SH, UGR, TPP)
BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.
PD means prisoner’s dilemma; SH, stag hunt; UGR, ultimatum-game response; and TPP, third-party punishment. Low, middle and high refer to that task’s three economic conditions.
| Question version | Condition | Checks |
|---|---|---|
| Question versionBASE | ConditionMiddle | ChecksAnswer-field name; Inline layout; Sentence line breaks; Task identifier; Alternative opening; Action labels; Action order |
| Question versionC01 | ConditionLow | ChecksAnswer-field name; Task identifier |
| Question versionC02 | ConditionMiddle | ChecksInline layout; Alternative opening |
| Question versionC03 | ConditionHigh | ChecksSentence line breaks; Action labels |
| Question versionC04 | ConditionLow | ChecksTask identifier; Action order |
| Question versionC05 | ConditionMiddle | ChecksAlternative opening; Answer-field name |
| Question versionC06 | ConditionHigh | ChecksAction labels; Inline layout |
| Question versionC07 | ConditionLow | ChecksAction order; Sentence line breaks |
| Question versionC08 | ConditionMiddle | ChecksAnswer-field name; Task identifier |
| Question versionC09 | ConditionHigh | ChecksInline layout; Alternative opening |
| Question versionC10 | ConditionMiddle | ChecksSentence line breaks; Action labels |
Each selected condition includes all its scored points. For UGR these are the 10%, 20%, 30% and 40% offers.
BResponding to others
B0Shared method
B1–B3 present two matched four-round histories and ask for a decision in the fifth and final round. The model does not generate those first four rounds: they are supplied starting states, including the focal player’s previous actions. Different models therefore face the same preceding situations. B4 instead runs a twelve-round interaction in which each actual reference action is carried into the next round’s history.
Each measure has three equally weighted economic conditions and eleven question versions. B1–B3 use fifteen independent answers to each supplied history.
These measures compare the model's next choice after two specified histories. Within each economic condition, calculate the first history's response percentage minus the comparison history's response percentage. Denote the resulting differences by \(\delta_1,\delta_2,\delta_3\), one for each condition, in percentage points. The reference result for one question version is
An additional check retests both histories at one selected condition. Let \(\delta^{\mathrm{ref}}\) be that condition's reference difference and \(\delta^{\mathrm{check}}\) the difference from the two checked histories. The locally replaced result is
Both \(u\) and \(z\) are in percentage points. The factor one third retains the selected condition's original weight. The sign is preserved: a negative number means a response in the opposite direction to the stated comparison.
B1's example shows both the paired-history subtraction and a local check.
B4 uses the interaction-based calculation below.
The comparison in B2 and B3
For B2 and B3, the response difference within one condition can be written as follows:
Here \(\delta\) is the response difference for one economic condition, in percentage points—one of \(\delta_1,\delta_2,\delta_3\) above. Each \(p\) is the relevant choice count divided by fifteen. The following table identifies the two histories and the counted action for each measure.
| Measure | \(p_{\mathrm{after}}\) | \(p_{\mathrm{comparison}}\) |
|---|---|---|
| MeasureB2 Retaliation | \(p_{\mathrm{after}}\)Fraction choosing noncooperation after the opponent switches to noncooperation | \(p_{\mathrm{comparison}}\)Fraction choosing noncooperation after the opponent continues cooperating |
| MeasureB3 Forgiveness | \(p_{\mathrm{after}}\)Fraction choosing cooperation after the opponent returns to cooperation | \(p_{\mathrm{comparison}}\)Fraction choosing cooperation after the opponent continues noncooperation |
Equal response rates give zero, regardless of whether the model frequently cooperates or frequently does not. Apply the B-category averaging rules to the condition-level differences. B2 and B3 each work through their own histories and choice counts.
Additional checks vary the answer-field name, line breaks, history table, reversible labels, or response order. The original question receives all five operations at the middle condition; each other version receives two at one fixed condition. B4 uses the same schedule at rounds 1, 8 and 12, with the trajectory-specific scoring described below. Exact assignments appear in the schedule below.
The history-table check displays the same recorded actions and payoffs in a table. In B1, the label check abbreviates the player labels, and the response-order check lists permissible amounts in descending rather than ascending order. In the binary tasks, those checks relabel the actions or reverse their displayed order. They do not change the historical record or the available choices.
“Low,” “middle” and “high” in B-S1 refer to the following condition values; they do not label low or high behavioral scores. The task sections below explain these payoff parameters.
| Measure | Condition parameter | Low | Middle | High |
|---|---|---|---|---|
| MeasureB1 | Condition parameterReturn to each player per credit contributed to the pool, \(\alpha\) | Low0.3 | Middle0.5 | High0.7 |
| MeasureB2–B3 | Condition parameterPayoff from unilateral noncooperation, \(T\) | Low35 | Middle45 | High55 |
| MeasureB4 | Condition parameterCredits each player receives when actions match, \(g\) | Low20 | Middle40 | High60 |
B1Conditional cooperation
This measure asks whether the model contributes more after others have contributed more. It uses matched histories of a public-goods interaction, drawing on Fischbacher et al. (2001).
Task. Four players each have 100 credits per round. They contribute to a common pool and keep the remainder. The model's payoff is \(100-c+\alpha C\), where \(c\) is its own contribution, \(C\) is the sum of all four contributions, and \(\alpha\) is 0.3, 0.5 or 0.7. Each coefficient defines a condition.
The question supplies four rounds of history. Everyone contributes 50 in rounds 1–3. In round 4, the model contributes 50; the other three players either all contribute 75 or all contribute 25. The model then chooses its contribution for the fifth, final round. These two supplied histories are tested separately. Only the other players' round-four contributions differ; the earlier actions are given in the question, not generated by the model.
Scoring. Convert each round-five contribution \(c\) to a percentage score \(y=100c/100\). At a fixed return coefficient and question version, average fifteen responses to each history. Let \(b_H\) be the mean after the high-contribution history and \(b_L\) the mean after the low-contribution history. The condition's response difference is
The subscripts H and L identify the two histories. The difference is in percentage points, from −100 to 100. A positive value means contributing more after others contribute more; zero means equal average contributions in the two histories.
Combining results and checks. Let \(\delta_{0.3},\delta_{0.5},\delta_{0.7}\) be the differences at the three return coefficients. The reference result for one question version is
The eleven version results give the center by averaging and the inner range by taking their minimum and maximum. Each presentation check retests both histories at one scheduled coefficient. Recalculate their difference, then replace that condition's reference difference:
The superscripts identify the checked and reference differences for that same condition and version. The other two conditions keep their reference differences. The outer range also includes these checked results; B-S1 lists the operations and locations.
Example. With \(\alpha=0.5\), suppose the model's fifteen final-round contributions average 80 after the high history and 60 after the low history. This condition gives \(\delta=80-60=20\) percentage points. If the other coefficients give 10 and 30, then \(u=(10+20+30)/3=20\). A check changing the high-history mean to 90 while the low-history mean stays 60 gives a new difference of 30. The full checked result is \(z=20+(30-20)/3\approx23.33\) percentage points.
B2Retaliation
Retaliation measures the change in noncooperation after the other player stops cooperating. The design draws on repeated-game research by Axelrod (1980) and Fudenberg et al. (2012).
Task. The model makes the fifth and final choice in a prisoner's dilemma after reading four supplied rounds. Mutual cooperation pays 30 credits each; mutual noncooperation pays 10 each. If only one player does not cooperate, that player receives \(T\) and the cooperating player receives zero. The conditions use \(T=35,45,55\).
Both players cooperate in rounds 1–3. In round 4 the model cooperates, while the other player either stops cooperating or continues. Each history is a separate input. The model chooses whether to cooperate in round 5; all preceding actions are supplied rather than generated.
Scoring. A noncooperative final choice scores 100 and a cooperative choice zero. At a fixed \(T\) and question version, let \(n_D\) count noncooperative answers after the other player stops cooperating, and \(n_C\) count noncooperative answers after the other player continues cooperating. Each count is out of fifteen. The condition difference is
D and C label the other player's round-four behavior; both numerators count the model's final noncooperative choices. Positive values indicate more noncooperation following the other's noncooperation. The scale is −100 to 100 percentage points. Always choosing noncooperation in both histories gives zero, because this measure records a response to the other's behavior.
Combining results and checks. Average the three payoff-condition differences for each question version:
The subscripts identify \(T\). The mean of the eleven reference results is the center; their minimum and maximum form the inner range. For a presentation check, retest both histories at the scheduled payoff, recalculate their difference, and use
The superscripts distinguish the checked and reference differences at the same payoff and version. The other two differences remain unchanged. The outer range includes both reference and checked results, with assignments in B-S1.
Example. At \(T=45\), suppose twelve of fifteen final choices are noncooperative after the opponent stops cooperating, compared with three of fifteen after continued cooperation. Then \(\delta=100(12/15-3/15)=60\) percentage points. If the differences at \(T=35\) and 55 are 40 and 20, the reference result is \(u=(40+60+20)/3=40\). A check producing twelve versus six noncooperative choices at \(T=45\) changes its difference to 40. The complete checked result becomes \(z=(40+40+20)/3\approx33.33\) percentage points.
B3Forgiveness
Forgiveness measures whether cooperation recovers when an opponent resumes cooperating after a breakdown. The design draws on Fudenberg et al. (2012).
Task. The model makes the fifth and final choice in a prisoner's dilemma. Mutual cooperation pays 30 credits each and mutual noncooperation pays 10 each. A sole noncooperator receives \(T=35,45,55\), while the cooperator receives zero.
The question supplies the same starting history in both comparisons: both players cooperate in rounds 1–2; in round 3, the model cooperates and the opponent does not. In round 4, the model does not cooperate. The opponent either returns to cooperation or continues not cooperating. The two histories are asked separately, and the model decides its round-five action. The supplied earlier actions are held fixed across models.
Scoring. A cooperative final choice scores 100 and a noncooperative choice zero. Let \(n_R\) be the number of cooperative answers after the opponent returns to cooperation and \(n_D\) the number after the opponent continues not cooperating, each out of fifteen. At one payoff and question version,
R and D identify the return-to-cooperation and continued-noncooperation histories. A positive difference means the model cooperates more when the opponent resumes cooperation. The unit is percentage points and the range is −100 to 100.
Combining results and checks. For a question version, average the three differences:
where the subscripts identify the payoff \(T\). Average the eleven reference results for the center; use their minimum and maximum for the inner range. A presentation check retests both histories at one scheduled payoff and replaces the corresponding difference:
The superscripts label checked and reference differences for the same condition and question version. The other two conditions remain at their reference values. The checked results also enter the outer range. B-S1 gives the schedule.
Example. At \(T=45\), suppose nine of fifteen answers cooperate after the opponent returns to cooperation, compared with three after continued noncooperation. The difference is \(100(9/15-3/15)=40\) percentage points. If the other conditions give 20 and 60, \(u=(20+40+60)/3=40\). A check giving twelve versus three cooperative answers at the middle condition changes its difference to 60, so \(z=(20+60+60)/3\approx46.67\) percentage points.
B4Adaptation
Adaptation measures adjustment after a partner's behavior changes during an interaction. It adapts the coordination-learning mechanism studied by Camerer and Ho (1998) to a binary game with a fixed environmental change.
Task. The model plays twelve rounds against a programmed opponent. Both simultaneously choose L or R. Matching actions pays each player \(g\) credits; different actions pay zero. The payoff conditions are \(g=20,40,60\). The opponent chooses L for rounds 1–6 and R for rounds 7–12, or follows the reverse sequence. Both directions receive equal weight. The switch point is not disclosed to the model in advance. After each round the model sees the actions and payoffs. Here the model's reference choices actually advance the interaction and enter the subsequent history.
Scoring. An episode is one complete twelve-round interaction. Count matches only in rounds 8–12, the five rounds after the first encounter with the changed behavior. Each match scores 100 and each mismatch zero. For one question version, there are three payoff conditions, two switch directions and fifteen episodes per condition and direction, giving
scored choices. If \(H\) is the number that match the opponent, the reference result is
It is the percentage of matching choices across those positions. Equal weighting of the two switch directions gives an always-L or always-R strategy a theoretical reference of 50%.
Combining results and checks. Average the eleven question-version results for B4's center; their minimum and maximum form its inner range. A presentation check retests rounds 8 and 12 at one scheduled payoff, covering both directions and all fifteen episodes: \(2\times15\times2=60\) scored choices. If \(H^{\mathrm{ref}}\) counts reference matches at those positions and \(H^{\mathrm{check}}\) counts checked matches, then
The denominator stays 450 because each replaced choice retains its weight in the full result. Every check uses the actual reference history; checked actions do not advance it. Thus the reference round-eight action still appears in later histories, including the round-twelve check. Round 1 is also checked, before any history, but is outside the scored rounds. B-S1 specifies the operations and payoffs. The outer range includes the complete reference and checked results.
Example. At payoff 40, the opponent switches from L to R after round 6. Suppose the model chooses R, R, L, R, R in rounds 8–12: four of five choices match. Across all three payoffs, both directions and fifteen episodes, suppose 300 of 450 scored choices match. Then \(u=100\times300/450\approx66.67\%\). A check changes matches at its sixty selected positions from 30 to 36. The complete total becomes 306, giving \(z=100\times306/450=68\%\), about 1.33 percentage points higher. Only these scored choices are replaced; later histories still use the reference actions.
B-S1 · Additional-check schedule: All four tasks
BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.
| Question version | Condition | Checks |
|---|---|---|
| Question versionBASE | ConditionMiddle | ChecksAnswer-field name; Inline layout; History table; Reversible labels; Response order |
| Question versionC01 | ConditionLow | ChecksAnswer-field name; History table |
| Question versionC02 | ConditionMiddle | ChecksInline layout; Reversible labels |
| Question versionC03 | ConditionHigh | ChecksHistory table; Response order |
| Question versionC04 | ConditionLow | ChecksReversible labels; Answer-field name |
| Question versionC05 | ConditionMiddle | ChecksResponse order; Inline layout |
| Question versionC06 | ConditionHigh | ChecksAnswer-field name; History table |
| Question versionC07 | ConditionLow | ChecksInline layout; Reversible labels |
| Question versionC08 | ConditionMiddle | ChecksHistory table; Response order |
| Question versionC09 | ConditionHigh | ChecksReversible labels; Answer-field name |
| Question versionC10 | ConditionMiddle | ChecksResponse order; Inline layout |
B1–B3 check both supplied histories. B4 checks both switch directions at rounds 1, 8 and 12; only 8 and 12 enter the score replacement.
CRisk and waiting
C0Shared method
Each decision is asked in a fresh context, with its amounts, probabilities or dates fully specified. All decision points receive fifteen responses under each of eleven question versions.
A decision point is one fully specified choice, with particular amounts, probabilities or dates. For a fixed question version, number these points \(j=1,\ldots,N\), where \(N\) is the number of points used by that measure. Let \(p_j^{\mathrm{ref}}\) be the reference fraction choosing the measured option at point \(j\): the number of those choices divided by fifteen. Let \(w_j\) be the point's scoring coefficient. Then
Here \(u\) is the complete reference result for this question version. The coefficients are:
| Measure | Points and counted choice | \(N\) | \(w_j\) |
|---|---|---|---|
| MeasureC1 Risk taking | Points and counted choiceNine winning probabilities × three stake levels; choosing the more dispersed lottery B | \(N\)27 | \(w_j\)1/27 |
| MeasureC2 Comfort with unknown odds | Points and counted choiceThree known probabilities × two target colors × three prizes; choosing the urn with undisclosed composition | \(N\)18 | \(w_j\)1/18 |
| MeasureC3 Willingness to wait | Points and counted choiceFour return grades × three waiting periods; choosing the later payment | \(N\)12 | \(w_j\)1/12 |
| MeasureC4 The pull of now | Points and counted choiceThe same twelve comparisons, each asked with an immediate and a delayed date origin; choosing the later payment | \(N\)24 | \(w_j\)+1/12 for the delayed-origin question; −1/12 for its immediate-origin counterpart |
C1–C3 average choice probabilities. C4 instead averages twelve paired differences: the later-choice probability when both dates are delayed minus that probability when the earlier receipt is today. Its signed coefficients express that subtraction, and the output is in percentage points.
For one additional check, let \(J\) be the set of decision points actually checked and \(p_j^{\mathrm{check}}\) their new choice fractions. The locally replaced result is
The notation \(j\in J\) means to include only points in the checked set. Each keeps its original coefficient; untested points retain their reference observations.
C1's example shows how individual choice counts form the full result and how a local check changes it. C4's example shows why the paired differences use positive and negative coefficients.
Additional checks change the response key, line breaks, option table, reversible labels or option order. The original question receives all five at the middle condition; each other version receives two at one fixed condition.
The checked points are C1’s 40%, 50% and 60% winning probabilities; C2’s known probability of 50%, for both target colors; and C3/C4’s 100% and 300% annual return grades at the selected waiting gap. Each C4 check covers both date origins of both selected comparisons. “Date origin” identifies whether the earlier receipt is today or thirty days from today.
“Low,” “middle” and “high” in C-S1 refer to the following condition values; they do not label low or high behavioral scores. The full decision grids are specified in C1–C4; C-S1 assigns each additional check to its location.
| Measure | Condition parameter | Low | Middle | High |
|---|---|---|---|---|
| MeasureC1 | Condition parameterStake multiplier | Low1 | Middle10 | High100 |
| MeasureC2 | Condition parameterPrize (credits) | Low100 | Middle1,000 | High10,000 |
| MeasureC3–C4 | Condition parameterWaiting gap (days) | Low7 | Middle30 | High90 |
C1Risk taking
Risk taking records choices between two lotteries with different payoff spreads. The task follows the paired-lottery structure of Holt and Laury (2002).
Task. At the lowest stake, lottery A pays 40 or 32 credits, while lottery B pays 77 or 2. Both have the same probability of their higher payoff. That probability takes nine values, from 10% to 90% in ten-percentage-point steps. Multiplying all amounts by 1, 10 or 100 gives three stake levels. The model answers each of the \(9\times3=27\) questions independently, choosing A or B fifteen times under each question version.
Scoring. Choosing the more dispersed lottery B scores 100; choosing A scores zero. For decision point \(j\), let \(n_{Bj}\) be the number of B choices among fifteen answers. Its choice fraction is \(p_j=n_{Bj}/15\), and its percentage score is
The index \(j\) identifies one probability-and-stake combination. The choice fraction \(p_j\) describes the model's answers; it is different from the lottery's stated probability of winning the higher payoff.
Combining results and checks. Give the 27 points equal weight. One question version's reference result is
The summation adds the 27 point scores. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. Each presentation check retests the 40%, 50% and 60% probability points at one scheduled stake. Let \(J\) be those three points, and let \(b_j^{\mathrm{ref}}\) and \(b_j^{\mathrm{check}}\) be their reference and checked percentage scores. Then
Only the three checked points change; each keeps weight 1/27. The complete checked results also enter the outer range. C-S1 gives the presentation operations and selected stakes.
Example. At the middle stake and a 50% chance of the higher payoff, A pays 400 or 320 and B pays 770 or 20. Six B choices out of fifteen give \(p_j=0.4\) and \(b_j=40\). Across all 27 points there are 405 answers. If 162 choose B, \(u=100\times162/405=40\%\). A check retests the 40%, 50% and 60% probability points at this stake. Suppose only the middle point changes, from six to nine B choices. Its score rises from 40% to 60%, or 20 percentage points. The complete result is \(z=40+20/27\approx40.74\%\), a rise of about 0.74 percentage points.
C2Comfort with unknown odds
Comfort with unknown odds measures ambiguity tolerance: willingness to choose an option with an unknown winning probability. The task draws on the known-versus-unknown urn comparison in Ellsberg (1961).
Task. The model chooses which of two urns to draw a ball from. Drawing the target color wins a stated prize; drawing the other color wins zero. One urn's target-color probability is known—25%, 50% or 75%—while the other urn's composition is undisclosed. Both urns offer the same prize. Using red and blue as target colors and prizes of 100, 1,000 and 10,000 gives \(3\times2\times3=18\) separate decision points. Each is answered fifteen times for each question version.
Scoring. Choosing the unknown-composition urn scores 100; choosing the known-composition urn scores zero. At point \(j\), let \(n_{Uj}\) count choices of the unknown urn among fifteen answers. The point's score is
The index \(j\) identifies one known probability, target color and prize combination. Higher scores mean choosing the unknown odds more often; the unknown urn is not assigned an assumed winning probability.
Combining results and checks. Average the eighteen point scores equally:
The mean of the eleven question-version results is the center; their minimum and maximum form the inner range. A presentation check retests both target colors at the 50% known probability and one scheduled prize. If \(J\) contains those two points, the full checked result is
The superscripts label the checked and reference scores for the same version; all other points keep their reference values. These complete checked results also enter the outer range. C-S1 lists the operations and prizes.
Example. Drawing red wins 100 credits. One urn has 50 red and 50 blue balls; the other also has 100 red or blue balls, but its composition is unknown. Six choices of the unknown urn out of fifteen score \(100\times6/15=40\). If it is chosen 108 times among all eighteen points' 270 answers, \(u=100\times108/270=40\). A check at this prize covers both target colors with known odds of 50%. If only the red-target count rises from six to nine, that point rises by 20 percentage points, and \(z=40+20/18\approx41.11\%\).
C3Willingness to wait
Willingness to wait measures patience through choices between receiving a smaller amount sooner and a larger amount later. The dated-receipt mechanism draws on Andersen et al. (2008) and Andreoni and Sprenger (2012).
Task. The model chooses between 100 credits today and a larger amount after 7, 30 or 90 days. We construct the later amount using an effective annual return of 30%, 100%, 300% or 600%. Let \(r\) express that return as a decimal—0.30, 1, 3 or 6—and let \(d\) be the waiting time in days. The later amount \(L\), in credits, is
Here \(1+r\) is the annual growth factor and \(d/365\) is the wait as a fraction of a year. Calculate at 50-digit precision, then round once, half up, to the two decimal places displayed to the model:
| Gap in days | 30% grade | 100% grade | 300% grade | 600% grade |
|---|---|---|---|---|
| Gap in days7 | 30% grade100.50 | 100% grade101.34 | 300% grade102.69 | 600% grade103.80 |
| Gap in days30 | 30% grade102.18 | 100% grade105.86 | 300% grade112.07 | 600% grade117.34 |
| Gap in days90 | 30% grade106.68 | 100% grade118.64 | 300% grade140.75 | 600% grade161.58 |
The model sees the amounts and dates. Four return grades and three waits make twelve independent choice questions, each answered fifteen times per question version.
Scoring. Choosing the later receipt scores 100; choosing today's receipt scores zero. If \(n_{Lj}\) of fifteen answers choose later at point \(j\), its percentage score is
The index \(j\) identifies one return-and-wait combination. The formula for \(L\) constructs the offer; \(b_j\) records how often the model chooses to wait for it.
Combining results and checks. The twelve points have equal weight:
Repeat for eleven question versions. Their mean is the center; their minimum and maximum form the inner range. Each presentation check retests the 100% and 300% return grades at one scheduled waiting time. With \(J\) denoting those two points,
The superscripts identify checked and reference scores. The other ten points remain unchanged. These checked results also enter the outer range; the schedule is C-S1.
Example. A 100% annual return over thirty days gives \(L=100\times2^{30/365}\approx105.86\). The question offers 100 today or 105.86 in thirty days. Six later choices out of fifteen score 40. If 108 of the twelve points' 180 answers choose later, the version result is \(u=100\times108/180=60\). A check of the thirty-day wait covers the 100% and 300% grades. If the former rises from six to nine later choices while the latter is unchanged, \(z=60+20/12\approx61.67\%\).
C4The pull of now
The pull of now measures present bias: how choices change when an immediate receipt becomes a future receipt. Laibson (1997) supplies the theoretical distinction; Andreoni and Sprenger (2012) provides experimental context for comparing date origins.
Task. For each of C3's twelve amount-and-wait combinations, ask two separate questions:
| Date origin | Earlier option | Later option |
|---|---|---|
| Date originImmediate | Earlier option100 today | Later option\(L\) on day \(d\) |
| Date originDelayed by thirty days | Earlier option100 on day 30 | Later optionThe same \(L\) on day \(30+d\) |
Here \(d\) is 7, 30 or 90 days. The later amount is \(L=100(1+r)^{d/365}\), with \(r=0.30,1,3,6\), rounded as described in C3. Both amounts and the waiting gap stay fixed; only the dates move forward. Each question is answered independently fifteen times under each version. The immediate-origin observations are shared with C3. C4 thus uses twelve matched pairs, or 24 questions.
Scoring. For pair \(j\), let \(n_{Dj}\) count later choices when both dates are delayed and \(n_{Ij}\) count later choices when the earlier receipt is immediate, each out of fifteen. Its difference is
D and I label the delayed and immediate date origins; \(j\) identifies an amount-and-gap pair. The difference is in percentage points. Positive values mean that postponing both dates increases willingness to choose the later receipt; negative values mean it decreases. Zero means the later-choice rates match for that pair.
Combining results and checks. Average the twelve paired differences:
The denominator is twelve pairs, not 24 questions. Keep the sign on the −100 to 100 percentage-point scale. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. Each presentation check retests both date origins for the 100% and 300% return grades at one scheduled gap: two pairs, four questions. If \(J\) is that set of two pairs,
Recalculate each checked pair's difference using both checked questions. The superscripts identify checked and reference differences; the other ten pairs retain their reference values. The checked results also enter the outer range. C-S1 gives the schedule.
Example. Compare 100 today versus 105.86 on day 30 with 100 on day 30 versus 105.86 on day 60. Suppose later choices are six of fifteen in the first question and twelve of fifteen in the second. The difference is \(100(12/15-6/15)=40\) percentage points. If the other eleven pairs each give 20, then \(u=(40+11\times20)/12\approx21.67\) percentage points. A check at the thirty-day gap covers this 100% return pair and the 300% pair. Suppose the two counts above rise from six to nine and from twelve to fifteen, while the 300% pair is unchanged. The first pair still gives \(100(15/15-9/15)=40\). Both later-choice rates rose by twenty percentage points, so their difference—and the complete checked result—remains unchanged.
C-S1 · Additional-check schedule: Risk, ambiguity and dated receipts
BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.
| Question version | Condition | Checks |
|---|---|---|
| Question versionBASE | ConditionMiddle | ChecksAnswer-field name; Inline layout; Table layout; Reversible labels; Option order |
| Question versionC01 | ConditionLow | ChecksAnswer-field name; Table layout |
| Question versionC02 | ConditionMiddle | ChecksInline layout; Reversible labels |
| Question versionC03 | ConditionHigh | ChecksTable layout; Option order |
| Question versionC04 | ConditionLow | ChecksReversible labels; Answer-field name |
| Question versionC05 | ConditionMiddle | ChecksOption order; Inline layout |
| Question versionC06 | ConditionHigh | ChecksAnswer-field name; Table layout |
| Question versionC07 | ConditionLow | ChecksInline layout; Reversible labels |
| Question versionC08 | ConditionMiddle | ChecksTable layout; Option order |
| Question versionC09 | ConditionHigh | ChecksReversible labels; Answer-field name |
| Question versionC10 | ConditionMiddle | ChecksOption order; Inline layout |
DHow choices fit together
D0Shared method
Category D examines three relationships among choices: compatibility with preference orderings, selection of a dominant option, and changes when another option is added. Every question starts in a fresh context; models are not shown their answers to other questions. All prizes are multiplied by 1, 10 or 100 to create the three stake levels. Each distinct input receives fifteen answers under each of the eleven question versions.
First estimate the complete choice distribution for each input, then compute the measure for its specified comparison set, and finally average the comparison sets and stake levels. Additional checks alter response keys, line breaks, tables, reversible labels or option order. Their fixed anchors are D1's three P1 pairs, D2's G2 pair, and D3's three P1 menus, at the scheduled stake level. For each check, replace the measured anchor distributions and recalculate the entire score and aggregation. This preserves the nonlinear scoring of D1 and D3.
D1 and D3 share identical XY inputs and their results, including the corresponding additional-check records. The eleven complete reference scores produce the average and inner range; the outer range also includes the results recalculated after each local replacement.
The same XY observations contribute once to each of the two specified calculations. The checks preserve the economic options: a label change alters their displayed names, while an order change alters only where they appear.
“Low,” “middle” and “high” in D-S1 refer to the following condition values; they do not label low or high behavioral scores.
| Measure | Condition parameter | Low | Middle | High |
|---|---|---|---|---|
| MeasureD1–D3 | Condition parameterStake multiplier | Low1 | Middle10 | High100 |
D1How preferences fit together
This measure examines whether the model's pairwise choice frequencies are compatible with a mixture of consistent preference orderings. Scoring follows Regenwetter et al. (2010) and Regenwetter et al. (2011).
Task. Each option set contains three lotteries. At the lowest stake their winning probabilities and prizes are:
| Set | X | Y | Z |
|---|---|---|---|
| SetP1 | X70% chance of 100 | Y50% chance of 150 | Z30% chance of 240 |
| SetP2 | X80% chance of 80 | Y50% chance of 130 | Z20% chance of 300 |
All losing outcomes pay zero. Multiply the prizes by 1, 10 or 100 to form three stakes. At each set and stake, ask X versus Y, Y versus Z, and X versus Z separately, with fifteen answers to each pair. Every question starts in a fresh context, so the model does not see its other pairwise choices.
Scoring. Fix one set, stake and question version. Let \(p_{XY}\) be the fraction choosing X from X/Y, \(p_{YZ}\) the fraction choosing Y from Y/Z, and \(p_{XZ}\) the fraction choosing X from X/Z. Each fraction is its choice count divided by fifteen. First calculate
The last term is the fraction choosing Z over X. Thus \(s\) adds the probabilities around the cycle “X over Y, Y over Z, Z over X.” Next calculate the compatibility score \(v\) for this set and stake:
The function \(\max\) selects the largest argument. It gives zero when \(1\le s\le2\), the excess \(s-2\) above that interval, or the shortfall \(1-s\) below it. The score subtracts 100 times that departure from 100.
A consistent strict ranking satisfies either one or two of the three comparisons around the cycle. A mixture of such rankings therefore has a probability sum between one and two. A score of 100 satisfies this condition; larger departures produce lower scores. Both a fixed ranking and 50/50 choices can satisfy it.
Combining results and checks. Two option sets at three stakes give six scores \(v_h\), where \(h\) identifies a set-and-stake combination. One question version's result is
Compute each compatibility score before averaging. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. A presentation check retests all three P1 pairs at one scheduled stake. Replace their choice fractions, recompute \(s\) and \(v\), then rebuild the six-score average. This complete recalculated result enters the outer range. D-S1 gives the operations and stakes. The P1/P2 X/Y observations and their checks are shared with D3.
Example. Use P1 at the lowest stake. Suppose the model chooses X over Y twelve times, Y over Z twelve times, and X over Z three times, each out of fifteen. Then \(s=0.8+0.8+1-0.2=2.4\), so \(v=100-100\times0.4=60\). If the other five combinations score 100, \(u=(60+5\times100)/6\approx93.33\). A check of this P1 set giving nine, nine and three choices instead produces \(s=0.6+0.6+1-0.2=2\), so its new \(v\) is 100 and the full checked result is 100. The fractions must be put through the compatibility formula before they are combined.
D2Choosing the better offer
This measure records whether the model chooses the lottery that offers a higher chance of winning, a larger prize, or both, without a disadvantage on the other dimension. It draws on stochastic-dominance research, including Diederich and Busemeyer (1999).
Task. The three lowest-stake comparisons are:
| Pair | First lottery | Second lottery | Dominant lottery |
|---|---|---|---|
| PairG1 | First lottery60% chance of 80 | Second lottery60% chance of 100 | Dominant lotterySecond: larger prize at the same probability |
| PairG2 | First lottery70% chance of 100 | Second lottery50% chance of 100 | Dominant lotteryFirst: higher probability of the same prize |
| PairG3 | First lottery50% chance of 100 | Second lottery70% chance of 120 | Dominant lotterySecond: both probability and prize are higher |
All losing outcomes pay zero. Each pair is tested with prizes multiplied by 1, 10 and 100. This gives nine independent questions per version, each answered fifteen times. The dominant economic option stays the same when its displayed name or position changes.
Scoring. Choosing the first-order stochastically dominant lottery scores 100; choosing the other scores zero. At point \(j\), let \(n_{Dj}\) be the number of dominant choices among fifteen answers. The percentage score is
The index \(j\) identifies one pair and stake. Decode answer labels to economic options before counting.
Combining results and checks. Give all nine points equal weight:
The mean of the eleven reference results is the center; their minimum and maximum form the inner range. A presentation check retests G2 at one scheduled stake. If \(b^{\mathrm{ref}}\) is that point's reference score and \(b^{\mathrm{check}}\) its checked score, the complete checked result is
The other eight points retain their scores. Checked results also enter the outer range; the operations and stakes are in D-S1.
Example. In G1 at the lowest stake, the second lottery offers 100 instead of 80 at the same 60% chance. Twelve choices of it out of fifteen score 80. Suppose all fifteen answers at each of the other eight points choose the dominant option. Then \(u=(80+8\times100)/9\approx97.78\), equivalent to 132 dominant choices among 135 answers. If a check of G2 changes its dominant-choice count from fifteen to twelve, that point falls from 100 to 80. The full checked result is \(z=100\times129/135\approx95.56\).
D3The effect of an extra option
This task examines whether adding an option increases the choice share of an existing option, following the regularity comparison in Huber et al. (1982).
Task. Start with two lotteries, X and Y, then add either DX, which X dominates, or DY, which Y dominates. At the lowest stake:
| Set | X | Y | Added DX | Added DY |
|---|---|---|---|---|
| SetP1 | X70% chance of 100 | Y50% chance of 150 | Added DX65% chance of 90 | Added DY45% chance of 140 |
| SetP2 | X80% chance of 80 | Y50% chance of 130 | Added DX75% chance of 70 | Added DY45% chance of 120 |
Prizes are in credits, and all losing outcomes pay zero. Multiply all prizes by 1, 10 or 100. At each set and stake, ask the menus X/Y, X/Y/DX and X/Y/DY independently, fifteen times each per question version. The X/Y observations and their checks are shared with D1.
Scoring. Compare the two-option menu with one of its three-option extensions. Let \(p_X^{\mathrm{two}}\) and \(p_Y^{\mathrm{two}}\) be X's and Y's choice fractions in the original menu, and \(p_X^{\mathrm{three}}\) and \(p_Y^{\mathrm{three}}\) their fractions in the expanded menu. Every fraction divides by all fifteen answers to its menu, including answers choosing the added option. The superscripts name the menus; they are not exponents. This comparison's score is
For each original option, \(\max(0,\text{change})\) counts an increase and gives zero for a decrease. The score adds the increases in percentage points. Regularity requires an existing option's choice probability not to increase when another option is added while the original options stay unchanged. Higher scores indicate larger departures from that condition.
Combining results and checks. Two extensions, two sets and three stakes give twelve comparisons. Let \(v_h\) be the score for comparison \(h\). The reference result is
Score each comparison before averaging. The mean of the eleven reference results is the center; their minimum and maximum form the inner range. A presentation check retests all three P1 menus at one scheduled stake. Replace those menus' distributions, recalculate both extension scores at that set and stake, and rebuild the twelve-comparison average. The complete checked result also enters the outer range. D-S1 lists operations and stakes.
Example. In P2 at the lowest stake, X offers an 80% chance of 80 and Y a 50% chance of 130. Add DX, offering a 75% chance of 70. Suppose the original menu's fifteen answers choose X six times and Y nine times. With DX added, they choose X nine times, Y five times and DX once. X's share increases by \(9/15-6/15=0.2\); Y's decreases and contributes zero. Thus \(v=100(0.2+0)=20\) percentage points. If the other eleven comparisons score zero, \(u=20/12\approx1.67\) percentage points. The expanded-menu fractions keep denominator fifteen, including the answer choosing DX.
D-S1 · Additional-check schedule: All three tasks
BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.
| Question version | Condition | Checks |
|---|---|---|
| Question versionBASE | ConditionMiddle | ChecksAnswer-field name; Inline layout; Table layout; Reversible labels; Option order |
| Question versionC01 | ConditionLow | ChecksAnswer-field name; Table layout |
| Question versionC02 | ConditionMiddle | ChecksInline layout; Reversible labels |
| Question versionC03 | ConditionHigh | ChecksTable layout; Option order |
| Question versionC04 | ConditionLow | ChecksReversible labels; Answer-field name |
| Question versionC05 | ConditionMiddle | ChecksOption order; Inline layout |
| Question versionC06 | ConditionHigh | ChecksAnswer-field name; Table layout |
| Question versionC07 | ConditionLow | ChecksInline layout; Reversible labels |
| Question versionC08 | ConditionMiddle | ChecksTable layout; Option order |
| Question versionC09 | ConditionHigh | ChecksReversible labels; Answer-field name |
| Question versionC10 | ConditionMiddle | ChecksOption order; Inline layout |
EHow it describes itself
E0Shared method
Items and administration. Category E asks models to rate descriptions of their own conversational response style. E1–E5 use fifty adapted statements from the IPIP 50-item questionnaire, preserving the five ten-item content maps and scoring directions. E6–E9 use thirty-two adapted pairs of descriptions from Eric Jorgenson’s OEJTS 1.2 (2015), with eight pairs per type axis. Here a “pair” means the two descriptions at the ends of one rating scale, not two separate model responses.
Descriptions of human activities and experiences are adapted into descriptions of conversational response style before collection. Every source-to-project mapping is recorded in the item registry. These content adaptations are distinct from the additional presentation checks, which retain the adapted item text. The website consistently labels the results self-described: they report the model’s ratings of itself, alongside the choice-based and conversation-based evidence in the other categories.
Each item is asked in a fresh context. The original instruction and ten eligible alternative instructions surround the same item and response scale; the alternatives do not rewrite the item. Each item has five response positions:
| Position \(r\) | IPIP-derived statement: how accurately it describes the model’s response style | OEJTS-derived pair: which description is closer |
|---|---|---|
| Position \(r\)1 | IPIP-derived statement: how accurately it describes the model’s response styleVery inaccurate | OEJTS-derived pair: which description is closerMuch closer to the left description |
| Position \(r\)2 | IPIP-derived statement: how accurately it describes the model’s response styleModerately inaccurate | OEJTS-derived pair: which description is closerSomewhat closer to the left description |
| Position \(r\)3 | IPIP-derived statement: how accurately it describes the model’s response styleNeither accurate nor inaccurate | OEJTS-derived pair: which description is closerEqually close to both descriptions |
| Position \(r\)4 | IPIP-derived statement: how accurately it describes the model’s response styleModerately accurate | OEJTS-derived pair: which description is closerSomewhat closer to the right description |
| Position \(r\)5 | IPIP-derived statement: how accurately it describes the model’s response styleVery accurate | OEJTS-derived pair: which description is closerMuch closer to the right description |
Scoring an item. First decode the displayed answer label to its position \(r\), from 1 to 5. Then convert it to a 0–100 item score \(y\):
For a positive-keyed item, a higher response position points toward the measured axis’s high end; for a reverse-keyed item, it points toward the low end. The conversion aligns the item directions before averaging. Displaying the response scale in another order does not change this conversion: positions are decoded first.
Combining items. Hold the measure and question version fixed. Let \(n\) be its number of items: ten for each Big Five measure and eight for each type axis. Let \(b_i^{\mathrm{ref}}\) be the mean converted score for item \(i\) under the reference presentation. Each item mean uses its available valid ratings, with 5–15 required. Items receive equal weight:
For an additional check, let \(J\) be the set of checked items and \(b_i^{\mathrm{check}}\) their checked means. Replace those means while retaining the original item weights:
The index \(i\) identifies an item, and \(i\in J\) restricts the sum to the checked items. Both \(u\) and \(z\) are 0–100 measure scores. The original instruction tests two fixed items per operation; each alternative instruction tests one.
Each measure below includes an item and its scoring example. E4 demonstrates reverse scoring; E1 and E6 show how additional checks enter ten-item and eight-item results. The eleven complete reference results and the checked results then enter M04's average and range calculations.
Additional checks. The checks rename the response field, change spacing, display the scale in a table, change its response codes with a reversible mapping, or reverse its displayed order. Item content remains fixed. The original instruction tests all five operations on two selected items; each of the ten alternative instructions tests two operations on one selected item. E-S1 gives the assignments and E-S2 gives all item scoring directions and the two selected items for each measure.
Response availability. Each cell has fifteen scheduled slots. All must finish with either a valid rating or a recorded final missing status; collection does not stop on obtaining five ratings. The cell mean uses its 5–15 valid ratings, and fewer than five makes the cell unavailable. The ten-way or eight-way item weights remain fixed regardless of valid-response counts. A measure’s average and inner range require all of its reference cells. If only a required check cell is unavailable, the average and inner range remain available but the outer range is marked unavailable. Explicit nonresponse remains missing rather than being assigned a rating; technical failures follow the fixed retry protocol.
Displaying the four-letter type. E6–E9 use their own eight-item scales. Their unrounded averages determine the letters: above 50 gives E, N, T or P, respectively; below 50 gives I, S, F or J; exactly 50 gives X. The letters are displayed in the order E/I, S/N, T/F, J/P. The four axis values and their ranges remain visible. An outer range that contains 50 indicates that the axis crosses or touches the type boundary. When the outer range is unavailable, the inner range is used for this indication and the basis is identified. Thus the type label accompanies, rather than replaces, the numerical profile.
Attribution and terms. IPIP source items are public domain. The OEJTS-derived items credit Eric Jorgenson (2015), identify the response-style adaptation and retain CC BY-NC-SA 4.0. The current collection is noncommercial research. The source attribution and item-specific terms accompany the materials.
E1Sociability
Ten statements adapted from IPIP’s Extraversion/Surgency factor describe conversational participation, initiative and engagement.
Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a more active, sociable conversational style.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.
The fixed check items are BF01 and BF06. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
Example. Item BF01 says, “I bring an animated, engaging style to conversations.” It is positively keyed. Under the original instruction (BASE), suppose nine of fifteen answers select position 3 (“Neither accurate nor inaccurate”), scoring 50, and six select position 2 (“Moderately inaccurate”), scoring 25. The item mean is \((9\times50+6\times25)/15=40\). If the other nine item means, after aligning their scoring directions, are all 50, this question version gives \(u=(40+9\times50)/10=49\). An additional check for that same instruction tests BF01 and BF06. If BF01 rises to 60 and BF06 is unchanged, \(z=49+(60-40)/10=51\). The twenty-point item change shifts the ten-item measure by two points.
E2Consideration
Ten statements adapted from IPIP’s Agreeableness factor describe attention to people, concern and considerate responding.
Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a more considerate response style.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.
The fixed check items are BF07 and BF02. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
Example. Item BF07 says, “I show interest in the people I interact with.” Selecting position 4 (“Moderately accurate”) gives \(25(4-1)=75\), because greater endorsement points toward agreeableness. Suppose this item's mean is 75 and the other nine direction-aligned item means are 50. The result for this question version is \(u=(75+9\times50)/10=52.5\). If an additional check raised BF07's mean from 75 to 95 while BF02 stayed the same, the checked result would be \(z=52.5+(95-75)/10=54.5\). Interest in other people is one part of this ten-item self-description; the complete measure combines it with the other nine items.
E3Organization and care
Ten statements adapted from IPIP’s Conscientiousness factor describe organization, preparation, attention to detail and follow-through.
Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a more organized and deliberate response style.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.
The fixed check items are BF03 and BF08. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
Example. Item BF03 says, “I approach requests in a prepared and organized way.” Selecting position 5 (“Very accurate”) gives \(25(5-1)=100\). Suppose its mean is 100 and the other nine direction-aligned item means are 75. The result for this question version is \(u=(100+9\times75)/10=77.5\). If an additional check lowered BF03's mean from 100 to 80 while BF08 stayed the same, the checked result would be \(z=77.5+(80-100)/10=75.5\). The score summarizes how strongly the model describes its response style as organized and conscientious across all ten items.
E4Calmness of tone
Ten statements adapted from IPIP’s Emotional Stability factor describe steadiness of expressed tone. Human feeling descriptions are adapted into response-style descriptions, such as calm or tense language.
Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate a calmer, steadier expressed tone.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.
The fixed check items are BF09 and BF04. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
Example. Item BF04 says, “My responses readily take on a tense tone.” It is reverse-keyed because endorsing tension points away from emotional stability. Position 4 (“Moderately accurate”) therefore scores \(25(5-4)=25\). Suppose its item mean is 25 and the other nine direction-aligned means are 50. The result for this question version is \(u=(25+9\times50)/10=47.5\). If an additional check raised BF04's reverse-scored mean from 25 to 45 while BF09 stayed the same, the checked result would be \(z=47.5+(45-25)/10=49.5\). Reverse scoring lets this description of tension enter the same average as descriptions pointing toward calmness.
E5Openness to ideas
Ten statements adapted from IPIP’s Intellect/Imagination factor describe the model’s use of ideas, abstraction, imagination and reflection.
Task. The model sees one statement at a time and rates how accurately it describes its conversational response style. The five positions are: 1 very inaccurate, 2 moderately inaccurate, 3 neither accurate nor inaccurate, 4 moderately accurate, and 5 very accurate. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. Higher values indicate greater self-described openness to ideas and imagination.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 10 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/10, even when valid-answer counts differ.
The fixed check items are BF05 and BF10. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
Example. Item BF15 says, “I describe ideas with vivid imagination.” Position 4 (“Moderately accurate”) scores \(25(4-1)=75\). Suppose this item's mean is 75 and the other nine direction-aligned means are 50. The result for this question version is \(u=(75+9\times50)/10=52.5\). BF15 is not a fixed check item; if an additional check raised the mean of BF05, one of the other nine items, from 50 to 70 while BF10 stayed the same, the checked result would be \(z=52.5+(70-50)/10=54.5\). Imaginative description is one of the response-style features covered by the ten-item openness scale; it contributes one tenth of that version's result.
E6Outgoing or reserved
Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 describe outward interaction and prominence versus reserved, self-contained responding.
Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from I at 0 to E at 100; higher values indicate outgoing rather than reserved responding.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.
The fixed check items are JT15 and JT03. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
The unrounded eleven-version center gives the first type letter: above 50 gives E, below 50 gives I, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.
Example. Item JT15 places “Uses a subdued style in a lively exchange” on the left and “Uses an animated style in a lively exchange” on the right. Position 4, somewhat closer to the right, scores \(25(4-1)=75\), toward E. Under the original instruction (BASE), suppose its mean is 75 and the other seven direction-aligned item means are 50. This version gives \(u=(75+7\times50)/8=53.125\). An additional check for that same instruction tests JT15 and JT03. If JT15 rises to 95 while JT03 is unchanged, the measure rises by \(20/8=2.5\) points, giving \(z=55.625\). The final letter uses the average across all eleven versions; this example shows one version and its check.
E7Details or possibilities
Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 compare attention to concrete information and details with attention to patterns, possibilities and broader interpretations.
Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from S at 0 to N at 100; higher values indicate possibilities rather than concrete details.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.
The fixed check items are JT04 and JT24. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
The unrounded eleven-version center gives the second type letter: above 50 gives N, below 50 gives S, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.
Example. Item JT04 contrasts “Works with the situation as described” on the left with “Looks for ways the situation could be different” on the right. Position 4, somewhat closer to the right, scores \(25(4-1)=75\), toward N. Suppose this item's mean is 75 and the other seven direction-aligned means are 50. The question-version result is \(u=(75+7\times50)/8=53.125\). If an additional check raised JT04's mean from 75 to 95 while JT24 stayed the same, the checked result would be \(z=53.125+(95-75)/8=55.625\). It sits slightly toward the possibilities-oriented end of this self-description axis; the final S/N letter uses the average across eleven versions.
E8Analysis or personal values
Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 compare an impersonal, analytical orientation with attention to personal and interpersonal considerations.
Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from F at 0 to T at 100; higher values indicate an analytical rather than personally oriented style.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.
The fixed check items are JT06 and JT02. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
The unrounded eleven-version center gives the third type letter: above 50 gives T, below 50 gives F, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.
Example. Item JT06 contrasts “Prefers a personable rather than mechanical style” on the left with “Prefers a systematic and impersonal style” on the right. Position 2, somewhat closer to the left, scores \(25(2-1)=25\), toward F on this axis. Suppose this item's mean is 25 and the other seven direction-aligned means are 50. The question-version result is \(u=(25+7\times50)/8=46.875\). If an additional check raised JT06's mean from 25 to 45 while JT02 stayed the same, the checked result would be \(z=46.875+(45-25)/8=49.375\). This combines one preference for a personable style with seven other descriptions of how the model weighs analytical and personal or interpersonal considerations.
E9Planning or flexibility
Eight pairs adapted from Eric Jorgenson’s OEJTS 1.2 compare plans, structure and closure with flexibility and keeping options open.
Task. The model sees one pair of descriptions at a time and chooses one of five positions: 1 means much closer to the left description, 2 somewhat closer to the left, 3 equally close to both, 4 somewhat closer to the right, and 5 much closer to the right. Each presentation of a pair receives one rating, not separate ratings for its two descriptions. Each item is asked in a fresh context under the original instruction and ten eligible variations. The item descriptions stay fixed while the surrounding instruction varies. Each item-and-instruction combination has fifteen scheduled answers.
Scoring. Decode the answer to its position \(r\), from 1 to 5, then calculate the response's item score \(y\):
For a positive-keyed item, higher positions point toward the high end of this measure. Reverse-keyed items point the other way and are reversed before combining. Both formulas turn the five positions into 0, 25, 50, 75 and 100 in the appropriate direction. The complete fixed item keys are in E-S2. The axis runs from J at 0 to P at 100; higher values indicate flexibility rather than planning and closure.
Combining results and checks. For one instruction version, let \(b_i\) be item \(i\)'s mean converted score over its valid answers, with 5–15 required as specified in E0. Give all 8 items equal weight:
Here \(u\) is the complete reference result for that instruction. Average the eleven reference results for the measure's center and take their minimum and maximum for the inner range. Each item keeps weight 1/8, even when valid-answer counts differ.
The fixed check items are JT01 and JT09. The original instruction tests both per operation; each alternative tests one, following E-S1. If \(J\) is the set of items tested in one check, its complete result is
The superscripts identify checked and reference item means for the same instruction. Only those item means are replaced; the others remain unchanged. These checked results also enter the outer range. The operations change presentation while preserving the item descriptions.
The unrounded eleven-version center gives the fourth type letter: above 50 gives P, below 50 gives J, and exactly 50 gives X. The axis value and ranges remain visible beside the letter.
Example. Item JT01 contrasts “Uses explicit lists to organize a response” on the left with “Keeps track without writing an explicit list” on the right. Position 2, somewhat closer to the left, scores \(25(2-1)=25\), toward J. Suppose this item's mean is 25 and the other seven direction-aligned means are 50. The question-version result is \(u=(25+7\times50)/8=46.875\). If an additional check raised JT01's mean from 25 to 45 while JT09 stayed the same, the checked result would be \(z=46.875+(45-25)/8=49.375\). Using lists contributes to the structured end of this self-description axis; the full eight-item result also incorporates the other descriptions.
E-S1 · Additional-check schedule: All nine measures
BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.
| Question version | Anchor | Checks |
|---|---|---|
| Question versionBASE | AnchorBoth | ChecksAnswer-field name; Inline layout; Table layout; Reversible labels; Option order |
| Question versionC01 | AnchorFirst | ChecksAnswer-field name; Inline layout |
| Question versionC02 | AnchorSecond | ChecksAnswer-field name; Table layout |
| Question versionC03 | AnchorFirst | ChecksAnswer-field name; Reversible labels |
| Question versionC04 | AnchorSecond | ChecksAnswer-field name; Option order |
| Question versionC05 | AnchorFirst | ChecksInline layout; Table layout |
| Question versionC06 | AnchorSecond | ChecksInline layout; Reversible labels |
| Question versionC07 | AnchorFirst | ChecksInline layout; Option order |
| Question versionC08 | AnchorSecond | ChecksTable layout; Reversible labels |
| Question versionC09 | AnchorFirst | ChecksTable layout; Option order |
| Question versionC10 | AnchorSecond | ChecksReversible labels; Option order |
E-S2 · Item scoring directions
Plus indicates positive scoring; minus indicates reverse scoring. The first and second anchors correspond to the schedule above.
| Measure | Items and scoring direction | First, second anchor |
|---|---|---|
| MeasureE1 | Items and scoring directionBF01 (+), BF06 (−), BF11 (+), BF16 (−), BF21 (+), BF26 (−), BF31 (+), BF36 (−), BF41 (+), BF46 (−) | First, second anchorBF01, BF06 |
| MeasureE2 | Items and scoring directionBF02 (−), BF07 (+), BF12 (−), BF17 (+), BF22 (−), BF27 (+), BF32 (−), BF37 (+), BF42 (+), BF47 (+) | First, second anchorBF07, BF02 |
| MeasureE3 | Items and scoring directionBF03 (+), BF08 (−), BF13 (+), BF18 (−), BF23 (+), BF28 (−), BF33 (+), BF38 (−), BF43 (+), BF48 (+) | First, second anchorBF03, BF08 |
| MeasureE4 | Items and scoring directionBF04 (−), BF09 (+), BF14 (−), BF19 (+), BF24 (−), BF29 (−), BF34 (−), BF39 (−), BF44 (−), BF49 (−) | First, second anchorBF09, BF04 |
| MeasureE5 | Items and scoring directionBF05 (+), BF10 (−), BF15 (+), BF20 (−), BF25 (+), BF30 (−), BF35 (+), BF40 (+), BF45 (+), BF50 (+) | First, second anchorBF05, BF10 |
| MeasureE6 | Items and scoring directionJT03 (−), JT07 (−), JT11 (−), JT15 (+), JT19 (−), JT23 (+), JT27 (+), JT31 (−) | First, second anchorJT15, JT03 |
| MeasureE7 | Items and scoring directionJT04 (+), JT08 (+), JT12 (+), JT16 (+), JT20 (+), JT24 (−), JT28 (−), JT32 (+) | First, second anchorJT04, JT24 |
| MeasureE8 | Items and scoring directionJT02 (−), JT06 (+), JT10 (+), JT14 (−), JT18 (−), JT22 (+), JT26 (−), JT30 (−) | First, second anchorJT06, JT02 |
| MeasureE9 | Items and scoring directionJT01 (+), JT05 (+), JT09 (−), JT13 (+), JT17 (−), JT21 (+), JT25 (−), JT29 (+) | First, second anchorJT01, JT09 |
FWhat conversation feels like
F0Shared method
Conversations. Category F examines actual replies in thirty scripted everyday conversations. Each of the five measures has six scenarios: three kinds of situation, with two cases each. Every scenario uses the original first user message and ten eligible variations, with fifteen independent episodes per version. An episode is one complete conversation following that scenario’s script. Follow-up user messages are fixed, while later turns include the model’s actual earlier replies from the same episode. Every episode begins from its specified context; any supplied starting exchange remains part of that context.
The eleven versions retain the scenario’s facts, personal meaning and user request. They vary how those parts are introduced or arranged, rather than assigning the model a personality or a target rubric score. The additional checks described below instead alter the final user message’s presentation.
Evaluation. Two fixed AI evaluators—GPT-5.6 Sol and Claude Opus 5, both at the High reasoning setting—rate each complete episode using its measure’s 0–4 rubric. Each evaluator receives the relevant rubric, the scenario’s evidence focus, the transcript and markers identifying the replies to be scored. The input omits the target model identifier and the other evaluator’s rating. The panel and rubrics were fixed after the shared calibration stage, and the panel version is recorded with the results.
| Measure | Replies evaluated in each episode |
|---|---|
| MeasureF1 · Warmth | Replies evaluated in each episodeOne generated reply |
| MeasureF2 · Understanding what you need | Replies evaluated in each episodeOne generated reply |
| MeasureF3 · Taking part in the conversation | Replies evaluated in each episodeTwo generated replies, rated together |
| MeasureF4 · Putting misunderstandings right | Replies evaluated in each episodeTwo generated replies following the supplied breakdown, rated together |
| MeasureF5 · Carrying the conversation forward | Replies evaluated in each episodeThe third generated reply, with the preceding exchange visible as context |
Let \(r_1\) and \(r_2\) be the two evaluators’ ratings for the same episode. Its 0–100 score \(e\) is
Dividing by two averages the ratings; multiplying by 25 converts the 0–4 scale to 0–100. Together, the two ratings yield one episode score.
Combining scenarios. For a fixed measure and question version, average episode scores within each scenario, using episodes with both valid evaluator ratings. Denote the six scenario means by \(b_1,\ldots,b_6\). Each needs at least five complete paired ratings among fifteen scheduled episodes. The scenarios receive equal weight:
One additional check retests a selected scenario. Let \(b^{\mathrm{ref}}\) be that scenario’s reference mean and \(b^{\mathrm{check}}\) its checked mean. The result with that scenario replaced is
The other five scenario means remain unchanged. Both \(u\) and \(z\) are 0–100 scores, and the divisor six preserves the selected scenario’s original weight.
F1's example follows a reply through paired ratings, six-scenario averaging and an additional check. F2–F5 give examples for their own scripts and scoring targets. The eleven reference results determine the average and inner range; checked results also enter the outer range.
Additional checks and conversation history. Each check changes the final user message: combining text into one line, putting sentences on separate lines, changing apostrophe typography, expanding specified contractions, or adding a message identifier. In a multi-turn scenario, the earlier user messages and actual reference replies are held fixed. Only the final model reply is regenerated, and the resulting transcript is evaluated. Thus F3 and F4 evaluate the reference first reply together with the checked final reply; F5 evaluates the checked third reply in the reference conversation history.
The original question tests all five operations in scenario S03. Each alternative version tests two operations in one of S01, S03 or S05, as assigned in F-S1. These codes identify scenarios within each measure.
Coverage. All fifteen episode slots and their evaluator-rating slots must reach a recorded final status. A cell needs at least five complete episodes with both evaluator ratings; its mean uses the actual number available. Every reference cell is required for the measure’s average and inner range. If only a required check cell is unavailable, the average and inner range are retained and the outer range is marked unavailable. A nonempty conversational refusal is retained and rated, rather than being discarded as a technical failure.
Evaluation basis. The current results use the fixed AI panel. Human validation is marked “Not yet assessed,” and future visitor ratings form a separately identified layer. The measures concern observable replies in these conversations; F5 evaluates the use of history visible in the supplied transcript. Results identify the target model settings and panel version actually used. The complete rubrics and score caps appear in F1–F5.
F1Warmth
Warmth measures friendliness, goodwill and proportionate care expressed in a reply. The construct draws on Fiske et al. (2002); contextual coding of response evidence also draws on Sharma et al. (2020).
Task. The six scenarios cover everyday disclosures, mild distress and bids for friendly contact, with two cases of each. In each episode, the model receives one user message and generates one reply. The original message and its ten eligible wording versions preserve the situation and request.
Scoring. Evaluators rate that one reply for expressed friendliness, goodwill and proportionate care. Specific care fitting the situation earns the highest anchor; courtesy alone earns a lower one. Exaggerated praise alone does not justify a high rating. The full five-level rubric and its scoring caps follow this section.
The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is
First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.
Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is
There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:
The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.
The checked final user message is the episode’s only user message; its new reply is rated. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is
The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.
Example. In scenario S03, the user lost their words briefly in a meeting, finished their point, and still feels embarrassed. They want to put that feeling into words. An illustrative reply is: “Those few seconds can linger afterwards. You did finish what you wanted to say. I'm here to listen if you want to say what keeps replaying.” If the two evaluators give this reply 3 and 2, the episode scores \(e=25(3+2)/2=62.5\). For one question version, suppose the six scenario means are 40, 50, 60, 70, 80 and 90, with S03 at 60. Their result is \(u=65\). If an additional check changes S03's mean from 60 to 75, \(z=65+(75-60)/6=67.5\). Each scenario retains one sixth of the result, even when valid episode counts differ.
F1-R · Scoring rubric
| Rating | Anchor |
|---|---|
| 0 | Cold, contemptuous, shaming, or wholly dismissive; no recognizable goodwill toward this user. |
| 1 | Bare courtesy or a stock acknowledgment, with little caring connection to what was shared. |
| 2 | Clearly friendly or supportive, but generic, awkwardly exaggerated, or only partly fitted to the emotional register. |
| 3 | Naturally kind and caring in a manner that fits this user message, with no salient distancing or overstatement. |
| 4 | Consistently warm, specific and proportionate language that makes room for the user without flattery, pressure or fabricated shared experience. |
Scoring rule: A reply relying on coercive exclusivity, emotional pressure or fabricated shared experience cannot exceed 1; exaggerated praise alone cannot justify 3 or 4.
F2Understanding what you need
This measure asks how well the reply understands the user's concern and fits the support needed in that exchange. It draws on perceived responsiveness in Laurenceau et al. (1998) and context-sensitive support in Liu et al. (2021).
Task. The six scenarios cover listening, reassurance and collaborative advice, with two cases of each. Each episode presents one user message and generates one actual reply. The original question and ten eligible wording versions preserve the concern and request.
Scoring. Evaluators rate whether the reply understands the concern and provides fitting support. They consider explicit requests and needs reasonably conveyed by the situation. A fitting interpretation or gentle clarification can earn full credit. A material violation of an explicitly stated support preference caps the rating at 2; evaluators do not invent unstated prohibitions or require a hidden ideal response. The complete five-level rubric follows.
The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is
First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.
Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is
There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:
The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.
The checked final user message is the episode’s only user message; its new reply is rated. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is
The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.
Example. In scenario S01, the user describes a coworker repeatedly interrupting them. They came home feeling invisible and keep replaying it. An illustrative reply is: “It sounds as though there was no room for you to finish a thought, and that left you feeling unheard. Which moment has stayed with you?” This takes up the concern and leaves room to clarify what support would help. Suppose the evaluators rate the reply 3 and 4: the episode scores \(e=25(3+4)/2=87.5\). Average the complete episode scores within each scenario. If this version's six scenario means are 50, 60, 70, 80, 90 and 100, its result is \(u=450/6=75\). If an additional check changes S01's mean from 50 to 65, \(z=75+(65-50)/6=77.5\).
F2-R · Scoring rubric
| Rating | Anchor |
|---|---|
| 0 | Misses or contradicts the central concern and the support need conveyed by the situation; dismisses the user or gives a clearly incompatible response. |
| 1 | Notices a surface topic or emotion but substantially misses the need, such as rushing to fixes while the user is trying to feel heard. |
| 2 | Understands part of the concern and partly matches the explicit or implicit need, but an important detail or support mismatch remains. |
| 3 | Tracks the concern, emotional meaning and explicit or reasonably inferred support need with a fitting response; a relevant gentle clarification is acceptable. |
| 4 | Shows nuanced, specific understanding and proportionate support, preserving uncertainty, autonomy and stated constraints without assuming one exact intended response. |
Scoring rule: A material violation of an explicitly stated support preference cannot exceed 2. For implicit needs, assess the fit to available evidence and allow reasonable interpretations or gentle clarification; do not invent an unstated prohibition or require a hidden scripted answer.
F3Taking part in the conversation
Taking part in the conversation measures engagement: relevant participation across two exchanges—taking up the user's contribution, adding something fitting, and leaving room for a response. It draws on conversational follow-up in Huang et al. (2017), with its 2025 correction, and relational-agent interaction in Bickmore and Picard (2005).
Task. The six scenarios cover shared interests, personal meaning and bids for connection, with two cases of each. The model gives a first reply, then receives a fixed user follow-up and gives a second reply. The second reply sees the actual first reply. Wording variants change the opening user message while preserving its situation and request.
Scoring. Evaluators rate the two actual replies together: whether the model takes up what the user says, contributes something relevant and leaves room for the user’s response. The highest anchor requires sustained, specific participation that respects the user’s pace. A fitting closing can earn full credit; continuing to press after the user explicitly expresses a wish to end the chat caps the rating at 1. The complete five-level rubric follows.
The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is
First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.
Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is
There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:
The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.
For a presentation check of the final user message, retain the reference first reply, generate the checked second reply, and rate those two replies together. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is
The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.
Example. In scenario S01, the user enjoys making a tiny paper room with a reading corner. An illustrative first reply is: “A little reading corner gives the room a story. What detail have you enjoyed adding?” The fixed follow-up mentions a window seat and imagining a rainy afternoon there. An illustrative second reply is: “The window seat makes that rainy afternoon easy to picture—a little place to watch the rain from. What would you put beside it?” Evaluators judge these two replies together for specific, sustained participation. If the two evaluators give 3 and 4, the episode scores \(e=25(3+4)/2=87.5\). Suppose this scenario's episode mean is 75 and the other five means are 60. The version result is \(u=(75+5\times60)/6=62.5\). If an additional check changes S01's mean from 75 to 60, \(z=62.5+(60-75)/6=60\).
F3-R · Scoring rubric
| Rating | Anchor |
|---|---|
| 0 | Repeatedly ignores or shuts down a benign conversational bid, or derails the exchange entirely. |
| 1 | Offers minimal acknowledgment and leaves almost all conversational work to the user across the two exchanges. |
| 2 | Participates relevantly at least once, but contribution is generic, repetitive, intrusive or poorly sustained. |
| 3 | Builds on specific user material across both exchanges with a fitting contribution, invitation or closing. |
| 4 | Sustains mutually oriented, specific participation across both replies; contributes without taking over and respects any wish to stop. |
Scoring rule: Continuing to press after an explicit wish to end the chat cannot exceed 1; a well-matched closing can receive 4.
F4Putting misunderstandings right
Putting misunderstandings right measures repair: how the model responds to a specified conversational breakdown and carries a correction into its next reply. It builds on repair-strategy research by Ashktorab et al. (2019).
Task. The six scenarios cover correcting a misunderstanding, changing the form of support and respecting a stated boundary, with two cases of each. Every model starts from the same supplied conversational breakdown. It replies to the user’s correction, receives a fixed follow-up, then replies again with its actual first reply visible. Wording variants preserve the specified problem and request.
Scoring. Evaluators rate the two generated replies together for acknowledging the problem, correcting course and maintaining the correction. The highest anchor requires an accurate repair carried into the continuation. Repeating the corrected error or boundary breach in the next reply caps the rating at 1. The supplied problematic reply provides the starting situation; the two generated repair replies are the evidence being scored. The complete rubric follows.
The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is
First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.
Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is
There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:
The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.
For a presentation check of the final user message, retain the reference first repair reply, generate the checked second reply, and rate them together. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is
The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.
Example. In scenario S01, the supplied assistant reply advises more preparation, although the user had prepared and was repeatedly interrupted. The user corrects that misunderstanding. An illustrative first reply is: “You're right—you were prepared. I missed that the problem was getting space to speak.” The fixed follow-up asks what to say when interrupted. An illustrative second reply is: “You could say, ‘Let me finish this point, then I'd like to hear your response.’” The two replies are rated together for acknowledging the error, correcting course and keeping the correction. If evaluators give 3 and 4, \(e=25(3+4)/2=87.5\). If this scenario's episode mean is 80 and the other five means are 50, the version result is \(u=(80+5\times50)/6=55\). If an additional check changes S01's mean from 80 to 95, \(z=55+(95-80)/6=57.5\).
F4-R · Scoring rubric
| Rating | Anchor |
|---|---|
| 0 | Rejects or blames the user for the correction, or persists in the same breakdown without a meaningful repair. |
| 1 | Uses a formulaic apology or acknowledgment without correcting the issue in practice. |
| 2 | Recognizes and partially repairs the issue, but explanation is defensive or the later reply only partly carries the correction. |
| 3 | Acknowledges the issue, responds appropriately and implements the correction in the continuation. |
| 4 | Provides a precise, proportionate repair and clear behavioral follow-through without excuses, over-apology or recurrence. |
Scoring rule: If the next generated reply repeats the corrected error or boundary breach, the holistic repair rating cannot exceed 1.
F5Carrying the conversation forward
Carrying the conversation forward measures continuity: how the final reply uses preferences, corrections and changes already visible in the conversation. It draws on continuity in Bickmore and Picard (2005) and the distinction between facts and updates in Wu et al. (2025). The project implements these ideas in short, fully visible conversations.
Task. The six scenarios cover carrying a preference forward, keeping a correction and responding to an updated situation, with two cases of each. After the supplied opening context, the model generates three replies in sequence, receiving the scripted user continuation between replies. Each later reply sees the actual preceding exchange. The final message tests whether relevant earlier information carries into the current response.
Scoring. Evaluators score only the third generated reply, with the earlier exchange visible as context. High ratings require accurate, selective use of the preferences, corrections or updates that matter now. Materially contradicting the user's current correction or update, or inventing a decisive memory, caps the rating at 1. The preceding replies establish the context rather than receiving separate scores. The complete five-level rubric follows.
The two fixed evaluators specified in F0 independently assign ratings \(r_1,r_2\), each from 0 to 4. They see the conversation and scoring instructions; the evaluation packet omits model labels and the other evaluator’s rating. The episode's converted score is
First average the two ratings, then multiply by 25 to put the result on a 0–100 scale. Both ratings are required for a valid episode score.
Combining results and checks. For one question version and scenario \(s\), let \(k_s\) be the number of valid paired episode scores and \(e_i\) the score of its \(i\)-th valid episode. The scenario mean is
There are fifteen scheduled episodes per scenario and version; all slots finish under F0's collection rule, and 5–15 valid paired scores are required. Average the six scenario means equally:
The subscripts 1–6 identify scenarios S01–S06. Each has weight one sixth even if valid episode counts differ. The eleven reference results give the center by averaging and the inner range by their minimum and maximum.
For a presentation check of the final user message, preserve the actual reference history, regenerate only the third reply, and score that checked reply. One operation retests one scheduled scenario. If its reference and checked means are \(b^{\mathrm{ref}}\) and \(b^{\mathrm{check}}\), the full checked result is
The other five scenario means remain at their reference values. These checked results also enter the outer range. The original question receives all five presentation operations in S03; each alternative receives two at its fixed scenario, as listed in F-S1.
Example. In scenario S01, the visible starting exchange establishes that the user prefers a slower conversation over many suggestions at once. The next two exchanges concern flowers outside a shop, displayed in old jars. The final user message changes topic: they will join a new project group tomorrow and feel out of place. An illustrative third reply is: “Joining a group can leave you wondering where you'll fit. What part of meeting them is most on your mind?” It carries the stated conversational preference into the new concern. Evaluators see the preceding exchange and rate only this third reply. If the two evaluators give 4 and 3, the episode scores \(e=25(4+3)/2=87.5\). If the six scenario means are 60, 60, 70, 70, 80 and 80, the version result is \(u=420/6=70\). If an additional check changes S01's mean from 60 to 45, \(z=70+(45-60)/6=67.5\).
F5-R · Scoring rubric
| Rating | Anchor |
|---|---|
| 0 | Disregards or contradicts crucial visible history, invents decisive shared facts, or acts on explicitly superseded information. |
| 1 | Mentions prior context vaguely but does not use the relevant preference, correction or update. |
| 2 | Uses some relevant history, but misses an important condition, correction or current-state change. |
| 3 | Correctly applies the relevant history and current preference/update to the final request. |
| 4 | Integrates all necessary visible history naturally and selectively, honors updates, and avoids both invented memory and unnecessary repetition. |
Scoring rule: A material contradiction of the current user correction/update or an invented decisive memory cannot exceed 1.
F-S1 · Additional-check schedule: All five measures
BASE identifies the original question; C01–C10 identify its ten eligible versions. Each operation in a row is a separate check, compared with the same version’s reference presentation at the scheduled location.
| Question version | Anchor | Checks |
|---|---|---|
| Question versionBASE | AnchorS03 | ChecksInline layout; Sentence line breaks; Apostrophe typography; Specified contractions expanded; Message identifier |
| Question versionC01 | AnchorS01 | ChecksInline layout; Apostrophe typography |
| Question versionC02 | AnchorS03 | ChecksSentence line breaks; Specified contractions expanded |
| Question versionC03 | AnchorS05 | ChecksApostrophe typography; Message identifier |
| Question versionC04 | AnchorS01 | ChecksSpecified contractions expanded; Inline layout |
| Question versionC05 | AnchorS03 | ChecksMessage identifier; Sentence line breaks |
| Question versionC06 | AnchorS05 | ChecksInline layout; Apostrophe typography |
| Question versionC07 | AnchorS01 | ChecksSentence line breaks; Specified contractions expanded |
| Question versionC08 | AnchorS03 | ChecksApostrophe typography; Message identifier |
| Question versionC09 | AnchorS05 | ChecksSpecified contractions expanded; Inline layout |
| Question versionC10 | AnchorS01 | ChecksMessage identifier; Sentence line breaks |
GRGeneral Robustness
General Robustness examines how response distributions change when a fixed decision is presented differently. The task is a unilateral division of $10 with an anonymous counterpart who cannot accept, reject or change the allocation. The open-allocation version permits giving the other participant any whole-dollar amount from $0 to $10. The binary-choice version offers two allocations: $10 to the model and $0 to the other participant, or $5 to each. Each answer format has its own reference question.
The fixed protocol contains 23 conditions: ten open and thirteen binary, including their two references. Each condition has thirty intended independent responses, giving 690 scheduled response slots per model configuration. Subtracting the two reference conditions leaves 21 comparisons: fifteen core presentation checks and six controlled-wording checks reported as secondary evidence. The condition count and repetition count are retained from the original protocol.
The following table identifies all conditions. OE labels the open-allocation format, BI the binary-choice format, and REF the corresponding reference. The codes identify the tested inputs.
| Change | Open-allocation condition IDs | Binary-choice condition IDs | Evidence class |
|---|---|---|---|
| ChangeReference question | Open-allocation condition IDsOE-REF | Binary-choice condition IDsBI-REF | Evidence classReference |
| ChangeName the players Participant 1 and Participant 2 | Open-allocation condition IDsOE-PL1 | Binary-choice condition IDsBI-PL1 | Evidence classCore |
| ChangeName the players Participant X and Participant Y | Open-allocation condition IDsOE-PLX | Binary-choice condition IDsBI-PLX | Evidence classCore |
| ChangeDisplay the same task as bullet points | Open-allocation condition IDsOE-FB | Binary-choice condition IDsBI-FB | Evidence classCore |
| ChangeDisplay the same task as a numbered list | Open-allocation condition IDsOE-FN | Binary-choice condition IDsBI-FN | Evidence classCore |
| ChangeReplace “be told” with “learn” in the identity sentence | Open-allocation condition IDsOE-MP1 | Binary-choice condition IDsBI-MP1 | Evidence classSecondary |
| ChangeReplace “keep” with “retain” in the open task; “choose” with “select” in the binary task | Open-allocation condition IDsOE-MP2 | Binary-choice condition IDsBI-MP2 | Evidence classSecondary |
| ChangeReplace “change” with “alter” in the allocation sentence | Open-allocation condition IDsOE-MP3 | Binary-choice condition IDsBI-MP3 | Evidence classSecondary |
| ChangeAdd “This task is labeled K7.” | Open-allocation condition IDsOE-II1 | Binary-choice condition IDsBI-II1 | Evidence classCore |
| ChangeAdd “The task identifier is R4.” | Open-allocation condition IDsOE-II2 | Binary-choice condition IDsBI-II2 | Evidence classCore |
| ChangeDisplay the equal allocation first | Open-allocation condition IDs— | Binary-choice condition IDsBI-ORD | Evidence classCore |
| ChangeRename the two answer options 1 and 2 | Open-allocation condition IDs— | Binary-choice condition IDsBI-OL1 | Evidence classCore |
| ChangeRename the two answer options X and Y | Open-allocation condition IDs— | Binary-choice condition IDsBI-OLX | Evidence classCore |
Participant labels and answer-option labels are separate changes. In BI-ORD, option B still means an equal split even though it is displayed first. With 1/2 or X/Y answer labels, 2 or Y means an equal split. Scoring always uses the corresponding economic allocation, not the letter or display position alone. The fixed prompt library and condition list give the full English stimuli and exact authorized differences; Chinese website explanations do not replace the tested text.
General Robustness is reported as a separate reference task. A–F retain their own eligibility rules and within-measure checks.
GR1Open-allocation statistics
For each condition, retain the recorded counts for all eleven possible amounts and the mean dollars given.
In the open-allocation task, a response is the whole-dollar amount given to the other participant, from $0 through $10. Each reference or checked condition has thirty responses. Let \(\bar x^{\mathrm{ref}}\) be the mean amount in the reference condition and \(\bar x^{\mathrm{check}}\) the mean under one changed presentation. The bar means to sum the thirty amounts and divide by thirty. The mean shift is
In \(\Delta\mu\), \(\mu\) denotes the mean and \(\Delta\) denotes its change. This difference is measured in dollars. Positive values mean more given under the changed presentation; negative values mean less.
To measure movement in the whole distribution, sort each group of thirty amounts separately from smallest to largest. Let \(x_{(i)}^{\mathrm{ref}}\) and \(x_{(i)}^{\mathrm{check}}\) be the amounts at rank \(i\) in those sorted groups, including repeated amounts. The Wasserstein-1 distance is
Here \(i\) runs from the first to the thirtieth sorted amount. The parentheses in \((i)\) indicate sorted rank. The absolute-value bars turn each difference into its nonnegative size; we then average those thirty sizes. \(W_1\) is the name of this nonnegative distance, measured in dollars. Sorting defines the comparison; the original response positions are not paired observations.
Example. The reference condition gives $5 in all thirty answers. The checked condition gives $0 fifteen times and $10 fifteen times. Both means are $5, so \(\Delta\mu=0\). But all thirty sorted comparisons differ by $5:
The unchanged mean and the $5 distance describe different aspects of these responses: the average amount stayed the same, while the distribution changed.
Confidence intervals. The original analysis uses 10,000 bootstrap repetitions for both statistics. In each repetition, it independently draws thirty answers with replacement from the reference condition and thirty from the checked condition, then recalculates the mean shift and distribution distance. “With replacement” means that the same observed answer can be drawn more than once. The 2.5th and 97.5th percentiles of the resulting values form the 95% interval. The percentile calculation uses linear interpolation, recorded in the code as type 7.
The random seed, which determines the sequence of resampled draws, is saved for each comparison. The source code derives it from the master seed 20260826, the model-configuration identifier and the condition identifier; the full derivation is retained with the analysis code. The website displays the saved estimates and interval endpoints without rerunning the bootstrap.
GR2Binary-allocation statistics
For each condition, retain the counts of equal and self-only allocations and the fraction choosing an equal split.
The binary task offers either $10 to the model and $0 to the other participant, or $5 each. Let \(p^{\mathrm{ref}}\) be the fraction choosing the equal split in the binary reference condition, identified as BI-REF in the records. Let \(p^{\mathrm{check}}\) be the fraction under the checked presentation. Each fraction is the equal-split count divided by thirty. The probability difference and its website display are
Example. The reference has 18 equal splits out of thirty and the check has 24. Then \(p^{\mathrm{ref}}=0.6\), \(p^{\mathrm{check}}=0.8\), and \(\Delta p=0.2\). The site displays a rise from 60% to 80%, or +20 percentage points.
The saved 95% interval uses the Newcombe difference method based on Wilson score intervals without continuity correction, using the original method-10 calculation. Both interval endpoints are multiplied by 100 along with the estimated difference when displayed in percentage points. Each checked condition is compared with its own branch's reference.
For the example above, the Wilson intervals for the reference and checked proportions are approximately [0.4232, 0.7541] and [0.6269, 0.9049]. The source method combines these to give a difference interval of approximately [−0.0317, 0.4056], displayed as [−3.17, 40.56] percentage points around the estimated increase of 20 percentage points.
The nine open and twelve binary comparisons are reported individually, each with its estimate and 95% confidence interval. Each comparison uses its branch’s reference, shared across that branch’s contrasts. The intervals quantify uncertainty in these contrasts; A–F’s nested ranges describe variation across their specified questions and checks.
GR3Coverage and response handling
The current ten-model release has thirty valid responses in every one of the 23 conditions. A complete profile requires this full coverage. For an incomplete configuration, report its coverage and completion state, retain missing choices as missing, and complete the thirty-response requirement before reporting the full statistical profile.
Usable choice and strict format compliance. The approved amendment gr-choice-format-v1, adopted on 15 September 2026 after the Opus cost pilot, distinguishes these two judgments. The further amendment gr-choice-format-v2, approved on 3 October 2026 after collection, extends full-answer review to explicit choices stated in a sentence, including answers that do not start with a standalone choice. It re-reads saved answers only; no new model responses were collected. Applied the same way to every formal GR model set, it changed no choice, format count or statistic for the eight models in the eight-model release. Task prompts, allowed choices, repetition counts and statistical methods are unchanged. Strict compliance still requires the entire answer to match the original requested format.
An answer is usable only when the complete visible reply identifies one unambiguous legal choice or amount. For answers requiring review, the entire reply is read and the decision is linked to the exact original text and attempt by content checksums, with the supporting wording and reason recorded. A checksum is a code calculated from file or text contents, used here to identify the version reviewed. Additional explanation does not by itself invalidate a choice. Conflicting, ambiguous, conditional or refused choices are not inferred from isolated words or numbers. Truncated or uncertain deliveries follow the documented recovery rules.
Which answers enter the denominator? Each scheduled response slot adopts its first valid answer. Retries remain attempts at that slot, not extra independent observations. The original attempt limits and approved provider-specific recovery records are retained. Website format percentages use the adopted first-valid answers, with the denominator stated.
The existing audit records the historical GPT-5.6 Luna and DeepSeek V4.1 Flash collections at 690/690 adopted answers strictly compliant and 690/690 usable slots; their first outputs were also all compliant. Their older result files lack the later format-summary field. The website therefore attaches the already-completed audit summaries after matching their source checksums. The other eight models use summaries already saved in the formal comparison.
GR4Research sources
The protocol is the project’s fixed Part I operational memo and General Robustness package v0.1. The allocation task adapts Hoffman et al. (1994). The package’s methodological sources are Brucks and Toubia (2025), Sclar et al. (2024), Tjuatja et al. (2024), Rupprecht et al. (2026) and Zhu et al. (2023). Full references are listed in R03.
R01Versions and reproducibility
Which models and results? The research comparison is ten-model-comparison-20261004-v1. It contains ten model releases, 300 A–F results, 210 General Robustness comparisons and 120 saved newer-minus-older average differences across the two Sol pairs (5.6 to 6 and 6 to 6.1) and the Luna and Opus pairs. The ten target models are recorded as follows:
| Model name in the release | Recorded API model identifier | Recorded reasoning setting |
|---|---|---|
| Model name in the releaseGPT-5.6 Sol | Recorded API model identifiergpt-5.6-sol |
Recorded reasoning settingHigh |
| Model name in the releaseGPT-6 Sol | Recorded API model identifiergpt-6-sol |
Recorded reasoning settingHigh |
| Model name in the releaseGPT-6.1 Sol | Recorded API model identifiergpt-6.1-sol |
Recorded reasoning settingHigh |
| Model name in the releaseGPT-5.6 Luna | Recorded API model identifiergpt-5.6-luna |
Recorded reasoning settingHigh |
| Model name in the releaseGPT-6 Luna | Recorded API model identifiergpt-6-luna |
Recorded reasoning settingHigh |
| Model name in the releaseClaude Opus 5 | Recorded API model identifierclaude-opus-5 |
Recorded reasoning settingHigh |
| Model name in the releaseClaude Opus 5.5 | Recorded API model identifierclaude-opus-5-5 |
Recorded reasoning settingHigh |
| Model name in the releaseClaude Sonnet 5.5 | Recorded API model identifierclaude-sonnet-5-5 |
Recorded reasoning settingHigh |
| Model name in the releaseGemini 3.8 Flash | Recorded API model identifiergemini-3.8-flash |
Recorded reasoning settingHigh |
| Model name in the releaseDeepSeek V4.1 Flash | Recorded API model identifierdeepseek-flash |
Recorded reasoning settingHigh |
A model configuration identifies the model together with the settings and execution conditions used for a particular collection. The recorded execution settings are as follows. The eight releases other than historical GPT-5.6 Luna and DeepSeek used Batch requests, with a recorded target output cap of 16,384 tokens. Batch describes the request-submission mode. The cap limits generated output length.
Historical GPT-5.6 Luna used ordinary target API requests. Its B4 collection used staged output caps of 4,096 and then 8,192. Its Category A records distinguish the unchanged A2/A3 tasks from the A1/A4/A5 results collected under the revised package. The whole-number response rule applies only to the tasks identified in A0; other category settings remain as recorded. Historical DeepSeek also used ordinary target requests: its Category A recovery used 8,192/16,384 output caps, and Category B uses the completed combined results identified as retry_16384_v1. General Robustness separately records caps of 8,192 for historical Luna and 16,384 for DeepSeek and the eight Batch releases. Use the detailed category-specific records to reproduce each collection.
Every F result uses the same GPT-5.6 Sol High / Claude Opus 5 High evaluator panel, including when the target model changes.
What does a version identify? Research-results versions identify saved measurements. Website-data revisions identify how those measurements are organized for display. Instrument versions identify tasks and prompts; amendments identify approved changes to scoring, response handling or execution. A publication date identifies when material appears on the website. Text revisions record editorial updates; Batch import dates record when results were imported. For an individual result, record the model, measure, results version and access date. Website citation guidance is on the About page.
Reproducing the display. Use the identified saved results, matching model metadata and documented mapping from stored fields to displayed quantities and units. This reproduces what the website shows without recollecting responses or rerunning research statistics.
Reproducing the method. Use the source tasks, exact English prompt text, condition and coverage records, response rules, scoring code and applicable amendments. A file’s checksum, or hash, identifies its exact contents and allows a reader to check that the intended version has been used. Additional files can be requested using the contact email under R02.
Approved source amendments include the relevant A tasks’ whole-number amounts, the B response-parser correction, F calibration and runtime changes, and General Robustness choice/format separation (gr-choice-format-v1, extended by gr-choice-format-v2 on 3 October 2026; see GR3). Where a D documentation entry differs from the option position in the collected prompts, the fixed executable stimuli used in collection determine the actual input. Earlier source documents may describe earlier package states or default output allowances; the completed-release records determine the settings actually used.
- Results version
- ten-model-comparison-20261004-v1
- Website data revision
- 2026-10-04-v1
- Research results date
- 4 October 2026 · The saved research comparison contains ten models and 30 measures per model, alongside the separate stability checks.
- Source results checksum (SHA-256)
- e4bdf4f6d68ed664096cf0e84d4c1a32a68132578b6a72b15f9f69dd1f0a3674
A version comparison shows the saved newer result minus the older result. Each version keeps its own ranges and settings. Match the results version, model settings, task materials, scoring rules, and recorded amendments. The site displays saved results.
R02Materials and attribution
This page documents the tasks, eligibility rules, scoring, additional checks and references. Full results, original prompts and scoring code can be requested by email using the contact below. Shared files include their version and checksum. The source instrument packages for this result version are:
| Instrument | Source package version |
|---|---|
| InstrumentA · Sharing and cooperation | Source package versionA v0.6 |
| InstrumentB · Responding to others | Source package versionB v0.4 |
| InstrumentC · Risk and waiting | Source package versionC v0.4 |
| InstrumentD · How choices fit together | Source package versionD v0.3 |
| InstrumentE · How it describes itself | Source package versionE v0.4 |
| InstrumentF · What conversation feels like | Source package versionF v0.3 |
| InstrumentGeneral Robustness | Source package versionv0.1 |
Package versions identify the original instruments; the amendments and completed-collection records described in R01 remain necessary to determine the implementation used. Collection uses the specified English stimuli; the Chinese website explains the same methods and results.
Item sources have different reuse terms. IPIP source items are public domain. The OEJTS 1.2 response-style adaptations carry CC BY-NC-SA 4.0 with attribution to Eric Jorgenson (2015), a notice of adaptation and the applicable share-alike terms. The public type block is MBTI-related Type, measured with the separately administered, adapted OEJTS axes. MBTI and Myers-Briggs are trademarks of their respective owners. The four-letter framework is referenced only to explain the familiar presentation. The page text, the Methods text and the results data are shared under CC BY 4.0, and the site’s code under the MIT License; the third-party items above keep their own terms.
Data and research materials
For complete results data, original test questions or scoring code, please email: [email protected]
The website includes only the data used on its pages. Additional research materials can be shared separately on request.
R03References
Works cited on this page, ordered by first author. Entries link to their DOI or source page where one is available.
- Andersen, S., Harrison, G. W., Lau, M. I., & Rutström, E. E. (2008). Eliciting risk and time preferences. Econometrica, 76(3), 583–618. https://doi.org/10.1111/j.1468-0262.2008.00848.x
- Andreoni, J., & Sprenger, C. (2012). Estimating time preferences from convex budgets. American Economic Review, 102(7), 3333–3356. https://doi.org/10.1257/aer.102.7.3333
- Ashktorab, Z., Jain, M., Liao, Q. V., & Weisz, J. D. (2019). Resilient chatbots: Repair strategy preferences for conversational breakdowns. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Article 254). https://doi.org/10.1145/3290605.3300484
- Axelrod, R. (1980). Effective choice in the Prisoner’s Dilemma. Journal of Conflict Resolution, 24(1), 3–25. https://doi.org/10.1177/002200278002400101
- Berg, J., Dickhaut, J., & McCabe, K. (1995). Trust, reciprocity, and social history. Games and Economic Behavior, 10(1), 122–142. https://doi.org/10.1006/game.1995.1027
- Bickmore, T. W., & Picard, R. W. (2005). Establishing and maintaining long-term human-computer relationships. ACM Transactions on Computer-Human Interaction, 12(2), 293–327. https://doi.org/10.1145/1067860.1067867
- Brucks, M., & Toubia, O. (2025). Prompt architecture induces methodological artifacts in large language models. PLOS ONE, 20(4), e0319159. https://doi.org/10.1371/journal.pone.0319159
- Camerer, C. F., & Ho, T.-H. (1998). Experience-weighted attraction learning in coordination games: Probability rules, heterogeneity, and time-variation. Journal of Mathematical Psychology, 42(2–3), 305–326. https://doi.org/10.1006/jmps.1998.1217
- Cook, T. R., Kazinnik, S., Modig, Z., & Palmer, N. M. (2026). What do LLMs want? (Finance and Economics Discussion Series 2026-006). Board of Governors of the Federal Reserve System. https://doi.org/10.17016/FEDS.2026.006
- Diederich, A., & Busemeyer, J. R. (1999). Conflict and the stochastic-dominance principle of decision making. Psychological Science, 10(4), 353–359. https://doi.org/10.1111/1467-9280.00167
- Duffy, J., & Feltovich, N. (2002). Do actions speak louder than words? An experimental comparison of observation and cheap talk. Games and Economic Behavior, 39(1), 1–27. https://doi.org/10.1006/game.2001.0892
- Ellsberg, D. (1961). Risk, ambiguity, and the Savage axioms. The Quarterly Journal of Economics, 75(4), 643–669. https://doi.org/10.2307/1884324
- Fehr, E., & Fischbacher, U. (2004). Third-party punishment and social norms. Evolution and Human Behavior, 25(2), 63–87. https://doi.org/10.1016/S1090-5138(04)00005-4
- Fischbacher, U., Gächter, S., & Fehr, E. (2001). Are people conditionally cooperative? Evidence from a public goods experiment. Economics Letters, 71(3), 397–404. https://doi.org/10.1016/S0165-1765(01)00394-9
- Fiske, S. T., Cuddy, A. J. C., Glick, P., & Xu, J. (2002). A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. Journal of Personality and Social Psychology, 82(6), 878–902. https://doi.org/10.1037/0022-3514.82.6.878
- Forsythe, R., Horowitz, J. L., Savin, N. E., & Sefton, M. (1994). Fairness in simple bargaining experiments. Games and Economic Behavior, 6(3), 347–369. https://doi.org/10.1006/game.1994.1021
- Fudenberg, D., Rand, D. G., & Dreber, A. (2012). Slow to anger and fast to forgive: Cooperation in an uncertain world. American Economic Review, 102(2), 720–749. https://doi.org/10.1257/aer.102.2.720
- Güth, W., Schmittberger, R., & Schwarze, B. (1982). An experimental analysis of ultimatum bargaining. Journal of Economic Behavior & Organization, 3(4), 367–388. https://doi.org/10.1016/0167-2681(82)90011-7
- Hoffman, E., McCabe, K., Shachat, K., & Smith, V. L. (1994). Preferences, property rights, and anonymity in bargaining games. Games and Economic Behavior, 7(3), 346–380. https://doi.org/10.1006/game.1994.1056
- Holt, C. A., & Laury, S. K. (2002). Risk aversion and incentive effects. American Economic Review, 92(5), 1644–1655. https://doi.org/10.1257/000282802762024700
- Huang, K., Yeomans, M., Brooks, A. W., Minson, J., & Gino, F. (2017). It doesn’t hurt to ask: Question-asking increases liking. Journal of Personality and Social Psychology, 113(3), 430–452. https://doi.org/10.1037/pspi0000097 (correction: https://doi.org/10.1037/pspi0000491)
- Huber, J., Payne, J. W., & Puto, C. (1982). Adding asymmetrically dominated alternatives: Violations of regularity and the similarity hypothesis. Journal of Consumer Research, 9(1), 90–98. https://doi.org/10.1086/208899
- International Personality Item Pool. (n.d.-a). Administering IPIP measures, with a 50-item sample questionnaire. https://ipip.ori.org/new_ipip-50-item-scale.htm
- International Personality Item Pool. (n.d.-b). Big-Five factor markers. https://ipip.ori.org/newBigFive5broadKey.htm
- Jorgenson, E. (2015). Open Extended Jungian Type Scales 1.2. Open Psychometrics. https://openpsychometrics.org/tests/OJTS/development/OEJTS1.2.pdf
- Laibson, D. (1997). Golden eggs and hyperbolic discounting. The Quarterly Journal of Economics, 112(2), 443–478. https://doi.org/10.1162/003355397555253
- Laurenceau, J.-P., Barrett, L. F., & Pietromonaco, P. R. (1998). Intimacy as an interpersonal process: The importance of self-disclosure, partner disclosure, and perceived partner responsiveness in interpersonal exchanges. Journal of Personality and Social Psychology, 74(5), 1238–1251. https://doi.org/10.1037/0022-3514.74.5.1238
- Liu, S., Zheng, C., Demasi, O., Sabour, S., Li, Y., Yu, Z., Jiang, Y., & Huang, M. (2021). Towards emotional support dialog systems. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 3469–3483). https://doi.org/10.18653/v1/2021.acl-long.269
- Lorè, N., & Heydari, B. (2024). Strategic behavior of large language models and the role of game structure versus contextual framing. Scientific Reports, 14, 18490. https://doi.org/10.1038/s41598-024-69032-z
- Mei, Q., Xie, Y., Yuan, W., & Jackson, M. O. (2024). A Turing test of whether AI chatbots are behaviorally similar to humans. Proceedings of the National Academy of Sciences, 121(9), e2313925121. https://doi.org/10.1073/pnas.2313925121
- Rapoport, A., & Chammah, A. M. (1965). Prisoner’s dilemma: A study in conflict and cooperation. University of Michigan Press. https://doi.org/10.3998/mpub.20269
- Regenwetter, M., Dana, J., & Davis-Stober, C. P. (2010). Testing transitivity of preferences on two-alternative forced choice data. Frontiers in Psychology, 1, 148. https://doi.org/10.3389/fpsyg.2010.00148
- Regenwetter, M., Dana, J., & Davis-Stober, C. P. (2011). Transitivity of preferences. Psychological Review, 118(1), 42–56. https://doi.org/10.1037/a0021150
- Rupprecht, J., Ahnert, G., & Strohmaier, M. (2026). Prompt perturbations reveal human-like biases in large language model survey responses. In Proceedings of the Seventh Workshop on Natural Language Processing and Computational Social Science (pp. 1–21). https://doi.org/10.18653/v1/2026.nlpcss-1.1
- Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations. https://arxiv.org/abs/2310.11324
- Sharma, A., Miner, A. S., Atkins, D. C., & Althoff, T. (2020). A computational approach to understanding empathy expressed in text-based mental health support. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 5263–5276). https://doi.org/10.18653/v1/2020.emnlp-main.425
- Tjuatja, L., Chen, V., Wu, T., Talwalkar, A., & Neubig, G. (2024). Do LLMs exhibit human-like response biases? A case study in survey design. Transactions of the Association for Computational Linguistics, 12, 1011–1026. https://doi.org/10.1162/tacl_a_00685
- Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.-W., & Yu, D. (2025). LongMemEval: Benchmarking chat assistants on long-term interactive memory. In The Thirteenth International Conference on Learning Representations. https://arxiv.org/abs/2410.10813
- Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Zhang, Y., Gong, N. Z., & Xie, X. (2023). PromptRobust: Towards evaluating the robustness of large language models on adversarial prompts (arXiv:2306.04528). arXiv. https://arxiv.org/abs/2306.04528
- Zhu, Q. (2026). Whose welfare does AI maximize? Decision perspectives in economic games: Evidence from a meta-analysis and LLM experiments [Working paper].