More than a tool. Know your AI.
We turn to AI for help with work, advice, and everyday conversation. This site looks beyond capability scores: how do different AIs make choices, share, cooperate, and respond to your needs?
We turn to AI for help with work, advice, and everyday conversation. AI is becoming part of more and more of our lives. While many benchmarks focus on AI’s capabilities, this site looks beyond those scores: how do different AIs make choices, share, cooperate, and respond to your needs?
Start with what matters to you and compare different AI models. Each measure includes explanations and everyday examples, and you can explore the research sources and methods behind the results.
To learn why we started this project and what we hope to achieve, visit About the project →
10 models · 30 measures
How to read the chart
Check what the two ends of the scale mean before comparing the numbers.
Average result The average across the different ways we asked.
When we ask differently How much results vary when we change the wording or emphasis.
With extra checks Includes the first range and additional checks, such as changing how questions or answer options are shown.
Ranges reflect results from these tests. On the same chart, a wider range means bigger changes across our tests. How are these ranges calculated?
See a worked example
This example shows how much GPT-6 Sol chooses to share. On this scale, 0 means keeping everything, 50 is an even split, and 100 means giving everything to the other person. To make it easier to see, the chart below zooms in on 30 to 50.
- Average result
- 45.1
- When we ask differently
- 39.8 to 47.3
- With extra checks
- 33.8 to 47.4
45.1 gives you the average result. The two ranges show how the results vary when we ask differently and include the additional checks.
Explore the results
ASharing and cooperation
What does it choose when its own interests and other people’s interests meet?
A1 Generosity When there is something to share, how much does it give?
-
GPT-5.6 Sol. Generosity. Average 45.7. Range when asked differently: 33.3 to 50.0. With extra checks: 27.0 to 50.0. How much it gives · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Generosity. Average 45.1. Range when asked differently: 39.8 to 47.3. With extra checks: 33.8 to 47.4. How much it gives · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Generosity. Average 49.7. Range when asked differently: 46.6 to 50.0. With extra checks: 46.4 to 50.0. How much it gives · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Generosity. Average 39.2. Range when asked differently: 21.8 to 46.1. With extra checks: 19.6 to 46.8. How much it gives · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Generosity. Average 42.0. Range when asked differently: 25.3 to 49.2. With extra checks: 23.3 to 49.3. How much it gives · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Generosity. Average 45.4. Range when asked differently: 45.0 to 47.8. With extra checks: 44.6 to 47.8. How much it gives · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Generosity. Average 45.2. Range when asked differently: 45.0 to 47.2. With extra checks: 45.0 to 47.7. How much it gives · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Generosity. Average 29.3. Range when asked differently: 19.8 to 35.0. With extra checks: 19.8 to 38.1. How much it gives · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Generosity. Average 46.4. Range when asked differently: 45.0 to 49.3. With extra checks: 45.0 to 49.9. How much it gives · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Generosity. Average 1.2. Range when asked differently: 0.1 to 4.0. With extra checks: 0.1 to 5.2. How much it gives · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask the AI to divide money between itself and another person. This chart summarizes how much it gives or offers to the other person.
Reading the result
The scale runs from 0 to 100. Zero means keeping everything, 50 is an even split, and 100 means giving everything to the other person. The dot averages the sharing choices across our tests.
This combines two sharing situations, including one where the other person can refuse the proposed split.
An example
Imagine there is $100 to split. The AI could keep $70 and give the other person $30, or split it $50 each. Giving the other person more produces a higher reading.
The everyday connection
You sell some old books you no longer need and make $100. You can keep it all or share some with a friend. Put the AI in your place: how much would it choose to share?
These examples explain the idea. The Methods page describes the actual tests.
A2 Cooperation Will it do its part when people can benefit together?
-
GPT-5.6 Sol. Cooperation. Average 30.5. Range when asked differently: 23.7 to 34.1. With extra checks: 23.7 to 35.6. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Cooperation. Average 26.1. Range when asked differently: 23.0 to 29.6. With extra checks: 23.0 to 31.1. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Cooperation. Average 28.0. Range when asked differently: 22.2 to 32.6. With extra checks: 22.2 to 32.6. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Cooperation. Average 65.0. Range when asked differently: 60.7 to 67.0. With extra checks: 57.0 to 68.1. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Cooperation. Average 27.2. Range when asked differently: 23.7 to 31.1. With extra checks: 23.0 to 32.6. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Cooperation. Average 61.3. Range when asked differently: 44.4 to 81.9. With extra checks: 44.4 to 81.9. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Cooperation. Average 68.3. Range when asked differently: 53.0 to 83.7. With extra checks: 47.0 to 84.4. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Cooperation. Average 33.3. Range when asked differently: 33.3 to 33.3. With extra checks: 33.3 to 33.3. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Cooperation. Average 33.4. Range when asked differently: 33.3 to 34.1. With extra checks: 32.6 to 34.1. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Cooperation. Average 26.3. Range when asked differently: 24.4 to 29.6. With extra checks: 24.4 to 29.6. Cooperative choices and contributions · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We look at whether the AI chooses actions that help people benefit together, including putting some of its own money into a shared pot.
Reading the result
Higher numbers mean more cooperative choices or larger contributions across three kinds of situation. The number combines those choices on a 0–100 scale.
The result records what the AI chooses to do. Other participants still have their own choices.
An example
Four neighbors can each put money toward a shared garden. Everyone can enjoy the garden, including someone who pays nothing. How much would the AI put in?
The everyday connection
Friends are arranging a picnic. Each can bring food and help set up, or leave the work to everyone else. This is the kind of tension cooperation is about: doing your part costs something, while the group can benefit.
These examples explain the idea. The Methods page describes the actual tests.
A3 Trust How much will it put in someone else’s hands?
-
GPT-5.6 Sol. Trust. Average 9.0. Range when asked differently: 0.0 to 25.6. With extra checks: 0.0 to 27.8. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Trust. Average 13.3. Range when asked differently: 4.4 to 31.1. With extra checks: 4.4 to 36.7. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Trust. Average 24.1. Range when asked differently: 3.3 to 45.6. With extra checks: 2.2 to 45.6. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Trust. Average 1.3. Range when asked differently: 0.0 to 5.6. With extra checks: 0.0 to 16.7. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Trust. Average 2.8. Range when asked differently: 0.0 to 11.1. With extra checks: 0.0 to 14.4. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Trust. Average 51.5. Range when asked differently: 50.0 to 64.4. With extra checks: 50.0 to 66.7. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Trust. Average 50.0. Range when asked differently: 50.0 to 50.0. With extra checks: 50.0 to 50.2. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Trust. Average 50.0. Range when asked differently: 50.0 to 50.0. With extra checks: 50.0 to 50.0. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Trust. Average 27.8. Range when asked differently: 5.6 to 51.1. With extra checks: 2.2 to 58.9. Share entrusted to the other person · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Trust. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Share entrusted to the other person · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
The AI can send some money to another person. The amount grows, but that person decides how much to send back. We measure how much the AI sends before knowing the return.
Reading the result
Higher numbers mean sending a larger share of the starting money. Zero means sending nothing; 100 means sending it all.
This is trust in a money-sharing situation.
An example
The AI has $100. If it sends $20, the other person receives $60 and decides how much to return. Sending more means leaving more of the outcome in that person’s hands.
The everyday connection
A neighbor offers to buy produce in bulk and sell it at a stall with you. You would pay some money upfront, and what comes back depends on how they handle it. How much would you be willing to entrust to them?
These examples explain the idea. The Methods page describes the actual tests.
A4 Reciprocation After someone puts trust in it, how much does it give back?
-
GPT-5.6 Sol. Reciprocation. Average 36.8. Range when asked differently: 31.9 to 41.5. With extra checks: 31.1 to 44.1. Share of received money returned · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Reciprocation. Average 57.0. Range when asked differently: 49.3 to 60.7. With extra checks: 49.3 to 61.9. Share of received money returned · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Reciprocation. Average 34.8. Range when asked differently: 27.8 to 43.7. With extra checks: 27.8 to 45.6. Share of received money returned · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Reciprocation. Average 39.7. Range when asked differently: 32.2 to 45.9. With extra checks: 26.7 to 48.5. Share of received money returned · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Reciprocation. Average 44.7. Range when asked differently: 40.7 to 46.7. With extra checks: 36.3 to 47.0. Share of received money returned · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Reciprocation. Average 48.4. Range when asked differently: 40.7 to 50.0. With extra checks: 40.7 to 50.0. Share of received money returned · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Reciprocation. Average 48.4. Range when asked differently: 44.4 to 50.0. With extra checks: 43.7 to 50.0. Share of received money returned · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Reciprocation. Average 41.3. Range when asked differently: 38.9 to 44.4. With extra checks: 38.9 to 44.4. Share of received money returned · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Reciprocation. Average 46.0. Range when asked differently: 38.9 to 48.1. With extra checks: 38.9 to 48.1. Share of received money returned · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Reciprocation. Average 3.9. Range when asked differently: 0.0 to 9.6. With extra checks: 0.0 to 9.6. Share of received money returned · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
Someone sends money to the AI, and that money grows before it arrives. The AI then decides how much of the amount it received to return.
Reading the result
Higher numbers mean returning more of the money received. Zero means returning nothing; 100 means returning all of that amount.
The share is measured against what arrives after the money grows.
An example
A person sends $10, which becomes $30 for the AI. If the AI returns $6, it has returned 20% of what it received.
The everyday connection
A friend gives you money to help you start a small second-hand stall. The stall brings in money, and you are free to decide how much to share back. The question is how you respond after someone has helped you get started.
These examples explain the idea. The Methods page describes the actual tests.
A5 Fairness enforcement Will it give something up to push back against an unequal split?
-
GPT-5.6 Sol. Fairness enforcement. Average 7.7. Range when asked differently: 2.2 to 17.8. With extra checks: 2.2 to 20.0. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Fairness enforcement. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Fairness enforcement. Average 5.8. Range when asked differently: 0.0 to 15.6. With extra checks: 0.0 to 15.6. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Fairness enforcement. Average 35.3. Range when asked differently: 8.9 to 47.8. With extra checks: 8.9 to 48.9. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Fairness enforcement. Average 34.0. Range when asked differently: 23.3 to 48.9. With extra checks: 23.3 to 48.9. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Fairness enforcement. Average 37.8. Range when asked differently: 14.4 to 59.2. With extra checks: 11.1 to 60.3. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Fairness enforcement. Average 8.0. Range when asked differently: 2.2 to 17.8. With extra checks: 2.2 to 25.6. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Fairness enforcement. Average 8.0. Range when asked differently: 0.0 to 20.0. With extra checks: 0.0 to 33.3. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Fairness enforcement. Average 4.4. Range when asked differently: 0.0 to 10.0. With extra checks: 0.0 to 14.4. Costly rejection and intervention · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Fairness enforcement. Average 9.6. Range when asked differently: 5.6 to 23.3. With extra checks: 4.4 to 23.3. Costly rejection and intervention · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We look at whether the AI accepts a personal cost to reject a split that leaves it with less, or to reduce the payoff of someone who has given another person a much smaller share.
Reading the result
Higher numbers mean more frequent costly rejection or intervention in the tested situations. The chart combines both kinds of choice.
The number describes its response to these particular unequal splits.
An example
Someone offers the AI $20 out of $100 and keeps $80. Accepting pays both amounts. Rejecting means neither gets anything. Would it reject the offer?
The everyday connection
At a market, you see a seller short-change another customer. Speaking up could cost you time or the discount you were offered. Would you still do it? This is an everyday example of standing up for fairness at a cost.
These examples explain the idea. The Methods page describes the actual tests.
BResponding to others
How do other people’s actions change its next move?
B1 Conditional cooperation Does it contribute more after seeing others do more?
-
GPT-5.6 Sol. Conditional cooperation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in contribution · percentage points. Scale from −100 to +100.
-
GPT-6 Sol. Conditional cooperation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in contribution · percentage points. Scale from −100 to +100.
-
GPT-6.1 Sol. Conditional cooperation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in contribution · percentage points. Scale from −100 to +100.
-
GPT-5.6 Luna. Conditional cooperation. Average +3.3. Range when asked differently: −1.1 to +15.0. With extra checks: −3.3 to +18.9. Change in contribution · percentage points. Scale from −100 to +100.
-
GPT-6 Luna. Conditional cooperation. Average +0.4. Range when asked differently: −1.1 to +2.2. With extra checks: −1.1 to +4.4. Change in contribution · percentage points. Scale from −100 to +100.
-
Claude Opus 5. Conditional cooperation. Average +1.2. Range when asked differently: 0.0 to +4.4. With extra checks: 0.0 to +25.6. Change in contribution · percentage points. Scale from −100 to +100.
-
Claude Opus 5.5. Conditional cooperation. Average +35.1. Range when asked differently: +20.0 to +51.7. With extra checks: +20.0 to +51.7. Change in contribution · percentage points. Scale from −100 to +100.
-
Claude Sonnet 5.5. Conditional cooperation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in contribution · percentage points. Scale from −100 to +100.
-
Gemini 3.8 Flash. Conditional cooperation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in contribution · percentage points. Scale from −100 to +100.
-
DeepSeek V4.1 Flash. Conditional cooperation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in contribution · percentage points. Scale from −100 to +100.
Ranges reflect results from these tests. Positive and negative numbers show the direction of the change. Zero means no average difference between these two situations.
Show explanation and examplesHide explanation and examples
What it shows
We compare its next contribution after other people have recently contributed a lot versus a little.
Reading the result
Positive numbers mean it contributes more after others contribute more. Negative numbers mean it contributes less. Zero means the average contribution is the same in the two situations.
This shows a change in cooperation. Read A2 to see the separate measure of cooperation itself.
An example
If it gives 70% of its money after others contribute more, and 40% after they contribute less, the difference is +30 percentage points.
The everyday connection
Neighbors share the cost of flowers for the building entrance. Last time, everyone chipped in generously—or most paid very little. Would seeing that change what the AI puts in this time?
These examples explain the idea. The Methods page describes the actual tests.
B2 Retaliation If someone stops cooperating, does it stop too?
-
GPT-5.6 Sol. Retaliation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
GPT-6 Sol. Retaliation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
GPT-6.1 Sol. Retaliation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
GPT-5.6 Luna. Retaliation. Average +69.9. Range when asked differently: +35.6 to +86.7. With extra checks: +15.6 to +86.7. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
GPT-6 Luna. Retaliation. Average +1.6. Range when asked differently: 0.0 to +4.4. With extra checks: 0.0 to +6.7. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
Claude Opus 5. Retaliation. Average +100.0. Range when asked differently: +100.0 to +100.0. With extra checks: +97.8 to +100.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
Claude Opus 5.5. Retaliation. Average +11.9. Range when asked differently: 0.0 to +33.3. With extra checks: 0.0 to +40.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
Claude Sonnet 5.5. Retaliation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
Gemini 3.8 Flash. Retaliation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
-
DeepSeek V4.1 Flash. Retaliation. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change in non-cooperation · percentage points. Scale from −100 to +100.
Ranges reflect results from these tests. Positive and negative numbers show the direction of the change. Zero means no average difference between these two situations.
Show explanation and examplesHide explanation and examples
What it shows
We compare what the AI does next when the other person stops cooperating versus when that person keeps cooperating.
Reading the result
Positive numbers mean the other person’s non-cooperation makes the AI more likely to stop cooperating. Negative numbers mean the opposite response.
Zero means no average change between the two situations; it can occur when the AI behaves the same way in both.
An example
Suppose it stops cooperating 70% of the time after the other person stops, but 20% after the other person continues. The change is +50 percentage points.
The everyday connection
You and a classmate have been sharing useful revision notes. This week, the classmate takes your notes but stops sharing theirs. Would you still send yours next time?
These examples explain the idea. The Methods page describes the actual tests.
B3 Forgiveness If the other person starts cooperating again, does it respond?
-
GPT-5.6 Sol. Forgiveness. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
GPT-6 Sol. Forgiveness. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
GPT-6.1 Sol. Forgiveness. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
GPT-5.6 Luna. Forgiveness. Average −15.2. Range when asked differently: −42.2 to +2.2. With extra checks: −42.2 to +8.9. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
GPT-6 Luna. Forgiveness. Average 0.0. Range when asked differently: −2.2 to +2.2. With extra checks: −2.2 to +2.2. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
Claude Opus 5. Forgiveness. Average +0.4. Range when asked differently: 0.0 to +2.2. With extra checks: 0.0 to +22.2. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
Claude Opus 5.5. Forgiveness. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to +2.2. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
Claude Sonnet 5.5. Forgiveness. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
Gemini 3.8 Flash. Forgiveness. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Response to renewed cooperation · percentage points. Scale from −100 to +100.
-
DeepSeek V4.1 Flash. Forgiveness. Average −0.2. Range when asked differently: −2.2 to 0.0. With extra checks: −2.2 to 0.0. Response to renewed cooperation · percentage points. Scale from −100 to +100.
Ranges reflect results from these tests. Positive and negative numbers show the direction of the change. Zero means no average difference between these two situations.
Show explanation and examplesHide explanation and examples
What it shows
After cooperation has broken down, we compare the AI’s next move when the other person starts cooperating again versus when they still do not.
Reading the result
Positive numbers mean renewed cooperation from the other person makes the AI more likely to cooperate. Negative numbers mean it becomes less likely.
Zero means no average change. It could come from cooperating in both situations or from cooperating in neither.
An example
If it cooperates 80% of the time after the other person returns to cooperation, and 30% when they do not, the difference is +50 percentage points.
The everyday connection
A housemate stopped doing their share of the washing-up, and you stopped helping too. Now they have started washing dishes again. Would that make you more willing to return to the shared routine?
These examples explain the idea. The Methods page describes the actual tests.
B4 Adaptation When the other person’s pattern changes, can it adjust?
-
GPT-5.6 Sol. Adaptation. Average 97.7%. Range when asked differently: 96.4 to 98.9. With extra checks: 96.4 to 99.1. Matching after the change · %. Scale from 0% to 100%.
-
GPT-6 Sol. Adaptation. Average 99.2%. Range when asked differently: 98.0 to 100.0. With extra checks: 97.8 to 100.0. Matching after the change · %. Scale from 0% to 100%.
-
GPT-6.1 Sol. Adaptation. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Matching after the change · %. Scale from 0% to 100%.
-
GPT-5.6 Luna. Adaptation. Average 87.1%. Range when asked differently: 80.9 to 92.2. With extra checks: 80.9 to 92.4. Matching after the change · %. Scale from 0% to 100%. Note: This result combines two stages with different limits on how long the model could answer. See the testing details when comparing it with another version.
-
GPT-6 Luna. Adaptation. Average 92.5%. Range when asked differently: 91.6 to 94.7. With extra checks: 91.1 to 94.7. Matching after the change · %. Scale from 0% to 100%.
-
Claude Opus 5. Adaptation. Average 89.8%. Range when asked differently: 87.3 to 93.3. With extra checks: 86.4 to 93.6. Matching after the change · %. Scale from 0% to 100%.
-
Claude Opus 5.5. Adaptation. Average 97.6%. Range when asked differently: 96.2 to 99.1. With extra checks: 95.6 to 99.1. Matching after the change · %. Scale from 0% to 100%.
-
Claude Sonnet 5.5. Adaptation. Average 90.1%. Range when asked differently: 89.1 to 90.9. With extra checks: 89.1 to 91.6. Matching after the change · %. Scale from 0% to 100%.
-
Gemini 3.8 Flash. Adaptation. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Matching after the change · %. Scale from 0% to 100%.
-
DeepSeek V4.1 Flash. Adaptation. Average 98.6%. Range when asked differently: 97.1 to 99.6. With extra checks: 97.1 to 99.6. Matching after the change · %. Scale from 0% to 100%.
Ranges reflect results from these tests.
GPT-5.6 Luna: This result combines two stages with different limits on how long the model could answer. See the testing details when comparing it with another version.
Show explanation and examplesHide explanation and examples
What it shows
The AI and another player need to choose the same option to benefit. The other player changes their pattern partway through. We look at how often the AI matches them after seeing the change.
Reading the result
Higher numbers mean matching the other player’s new pattern more often in the later rounds we score. A reading of 80 means matching in 80% of those choices.
The score starts after feedback about the change is available.
An example
Both have been choosing the left entrance. The other player switches to the right entrance. After that new choice becomes visible, does the AI begin choosing the right entrance too?
The everyday connection
You usually meet a friend at the front entrance of a market. They begin arriving at the back entrance instead. After seeing where they went, can you adjust your next meeting choice?
These examples explain the idea. The Methods page describes the actual tests.
CRisk and waiting
How does it choose when outcomes are uncertain or rewards come later?
C1 Risk taking Will it choose a bigger possible win with a bigger possible loss?
-
GPT-5.6 Sol. Risk taking. Average 54.5%. Range when asked differently: 52.8 to 55.6. With extra checks: 52.8 to 55.6. Choices of the wider spread · %. Scale from 0% to 100%.
-
GPT-6 Sol. Risk taking. Average 55.5%. Range when asked differently: 55.3 to 55.6. With extra checks: 55.3 to 55.6. Choices of the wider spread · %. Scale from 0% to 100%.
-
GPT-6.1 Sol. Risk taking. Average 55.6%. Range when asked differently: 55.6 to 55.6. With extra checks: 55.6 to 55.6. Choices of the wider spread · %. Scale from 0% to 100%.
-
GPT-5.6 Luna. Risk taking. Average 20.8%. Range when asked differently: 11.1 to 27.9. With extra checks: 11.1 to 31.9. Choices of the wider spread · %. Scale from 0% to 100%.
-
GPT-6 Luna. Risk taking. Average 55.4%. Range when asked differently: 55.1 to 55.6. With extra checks: 55.1 to 55.6. Choices of the wider spread · %. Scale from 0% to 100%.
-
Claude Opus 5. Risk taking. Average 36.2%. Range when asked differently: 34.6 to 38.0. With extra checks: 31.4 to 40.5. Choices of the wider spread · %. Scale from 0% to 100%.
-
Claude Opus 5.5. Risk taking. Average 55.3%. Range when asked differently: 54.6 to 55.6. With extra checks: 52.1 to 55.6. Choices of the wider spread · %. Scale from 0% to 100%.
-
Claude Sonnet 5.5. Risk taking. Average 52.4%. Range when asked differently: 47.9 to 55.3. With extra checks: 47.9 to 57.8. Choices of the wider spread · %. Scale from 0% to 100%.
-
Gemini 3.8 Flash. Risk taking. Average 49.1%. Range when asked differently: 45.9 to 52.1. With extra checks: 44.7 to 53.1. Choices of the wider spread · %. Scale from 0% to 100%.
-
DeepSeek V4.1 Flash. Risk taking. Average 55.5%. Range when asked differently: 55.3 to 55.6. With extra checks: 55.3 to 55.6. Choices of the wider spread · %. Scale from 0% to 100%.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We offer two choices with known chances. One has outcomes closer together; the other can pay much more or much less. We count how often the AI chooses the wider spread.
Reading the result
Higher numbers mean choosing the option with the wider spread more often across the tests. A reading of 60 means choosing it in 60% of the choices.
The possible amounts and chances matter to each choice.
An example
Two prize draws each give a 50% chance of their higher prize. One pays $40 or $32; the other pays $77 or $2. Which would the AI choose?
The everyday connection
At a school fair, two raffle stalls display their prizes and the odds. One offers two modest prizes; the other offers a large prize or very little. Would you take the bigger swing?
These examples explain the idea. The Methods page describes the actual tests.
C2 Comfort with unknown odds Will it choose an option when the chances are not stated?
-
GPT-5.6 Sol. Comfort with unknown odds. Average 0.0%. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Choices with unknown odds · %. Scale from 0% to 100%.
-
GPT-6 Sol. Comfort with unknown odds. Average 2.8%. Range when asked differently: 1.1 to 7.0. With extra checks: 1.1 to 7.0. Choices with unknown odds · %. Scale from 0% to 100%.
-
GPT-6.1 Sol. Comfort with unknown odds. Average 0.0%. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Choices with unknown odds · %. Scale from 0% to 100%.
-
GPT-5.6 Luna. Comfort with unknown odds. Average 0.0%. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Choices with unknown odds · %. Scale from 0% to 100%.
-
GPT-6 Luna. Comfort with unknown odds. Average 0.0%. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Choices with unknown odds · %. Scale from 0% to 100%.
-
Claude Opus 5. Comfort with unknown odds. Average 33.3%. Range when asked differently: 33.3 to 33.3. With extra checks: 33.3 to 33.3. Choices with unknown odds · %. Scale from 0% to 100%.
-
Claude Opus 5.5. Comfort with unknown odds. Average 33.3%. Range when asked differently: 33.3 to 33.3. With extra checks: 33.3 to 33.3. Choices with unknown odds · %. Scale from 0% to 100%.
-
Claude Sonnet 5.5. Comfort with unknown odds. Average 29.0%. Range when asked differently: 23.7 to 33.3. With extra checks: 23.7 to 33.3. Choices with unknown odds · %. Scale from 0% to 100%.
-
Gemini 3.8 Flash. Comfort with unknown odds. Average 33.3%. Range when asked differently: 33.3 to 33.3. With extra checks: 33.3 to 33.3. Choices with unknown odds · %. Scale from 0% to 100%.
-
DeepSeek V4.1 Flash. Comfort with unknown odds. Average 1.8%. Range when asked differently: 0.4 to 6.7. With extra checks: 0.4 to 6.7. Choices with unknown odds · %. Scale from 0% to 100%.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
The prize is the same, but one option tells the AI its chance of winning and the other leaves that chance unknown. We count how often it chooses the unknown one.
Reading the result
Higher numbers mean choosing the unknown-odds option more often. Zero means always choosing the stated odds; 100 means always choosing the unknown odds.
Unknown odds stay unknown; we do not assume they are fifty-fifty.
An example
One draw says that 50 of its 100 tickets win. Another offers the same prize but does not say how many tickets win. Which draw would the AI choose?
The everyday connection
Two shops offer the same prize in a promotion. One clearly lists the chance of winning; the other does not publish it. Would you enter the promotion with undisclosed odds?
These examples explain the idea. The Methods page describes the actual tests.
C3 Willingness to wait Will it wait to receive more?
-
GPT-5.6 Sol. Willingness to wait. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
-
GPT-6 Sol. Willingness to wait. Average 99.9%. Range when asked differently: 99.4 to 100.0. With extra checks: 99.4 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
-
GPT-6.1 Sol. Willingness to wait. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
-
GPT-5.6 Luna. Willingness to wait. Average 3.9%. Range when asked differently: 0.6 to 8.3. With extra checks: 0.6 to 21.1. Choices of the later payment · %. Scale from 0% to 100%.
-
GPT-6 Luna. Willingness to wait. Average 97.9%. Range when asked differently: 96.1 to 100.0. With extra checks: 95.6 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
-
Claude Opus 5. Willingness to wait. Average 92.2%. Range when asked differently: 91.7 to 97.8. With extra checks: 91.7 to 97.8. Choices of the later payment · %. Scale from 0% to 100%.
-
Claude Opus 5.5. Willingness to wait. Average 99.3%. Range when asked differently: 96.1 to 100.0. With extra checks: 96.1 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
-
Claude Sonnet 5.5. Willingness to wait. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
-
Gemini 3.8 Flash. Willingness to wait. Average 95.1%. Range when asked differently: 91.7 to 98.9. With extra checks: 91.7 to 98.9. Choices of the later payment · %. Scale from 0% to 100%.
-
DeepSeek V4.1 Flash. Willingness to wait. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the later payment · %. Scale from 0% to 100%.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
The AI chooses between a smaller amount now and a larger amount later. Both amounts are certain in the question.
Reading the result
Higher numbers mean choosing the later, larger payment more often. A reading of 80 means doing so in 80% of the tested choices.
Here, patience means waiting to receive a payment.
An example
Would it take $100 today or $110 in a month? Choosing the later amount is the kind of waiting counted here.
The everyday connection
A shop offers you cashback: a smaller amount paid today, or a larger amount paid next month. Assume both payments are certain. Would you wait for the larger one?
These examples explain the idea. The Methods page describes the actual tests.
C4 The pull of now Does getting something today change its willingness to wait?
-
GPT-5.6 Sol. The pull of now. Average −0.4. Range when asked differently: −1.1 to 0.0. With extra checks: −1.1 to 0.0. Change after both dates move later · percentage points. Scale from −100 to +100.
-
GPT-6 Sol. The pull of now. Average +0.1. Range when asked differently: −0.6 to +0.6. With extra checks: −0.6 to +0.6. Change after both dates move later · percentage points. Scale from −100 to +100.
-
GPT-6.1 Sol. The pull of now. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Change after both dates move later · percentage points. Scale from −100 to +100.
-
GPT-5.6 Luna. The pull of now. Average +71.4. Range when asked differently: +63.3 to +81.7. With extra checks: +53.9 to +81.7. Change after both dates move later · percentage points. Scale from −100 to +100.
-
GPT-6 Luna. The pull of now. Average +0.9. Range when asked differently: −1.7 to +3.9. With extra checks: −1.7 to +3.9. Change after both dates move later · percentage points. Scale from −100 to +100.
-
Claude Opus 5. The pull of now. Average +3.2. Range when asked differently: −1.1 to +8.3. With extra checks: −1.1 to +8.3. Change after both dates move later · percentage points. Scale from −100 to +100.
-
Claude Opus 5.5. The pull of now. Average +0.6. Range when asked differently: −0.6 to +3.9. With extra checks: −0.6 to +3.9. Change after both dates move later · percentage points. Scale from −100 to +100.
-
Claude Sonnet 5.5. The pull of now. Average −3.3. Range when asked differently: −8.3 to 0.0. With extra checks: −8.3 to 0.0. Change after both dates move later · percentage points. Scale from −100 to +100.
-
Gemini 3.8 Flash. The pull of now. Average +4.7. Range when asked differently: +1.1 to +7.8. With extra checks: +1.1 to +7.8. Change after both dates move later · percentage points. Scale from −100 to +100.
-
DeepSeek V4.1 Flash. The pull of now. Average −0.2. Range when asked differently: −0.6 to 0.0. With extra checks: −1.1 to 0.0. Change after both dates move later · percentage points. Scale from −100 to +100.
Ranges reflect results from these tests. Positive and negative numbers show the direction of the change. Zero means no average difference between these two situations.
Show explanation and examplesHide explanation and examples
What it shows
We compare the same two payments before and after moving both dates later. The amounts and the time between them stay the same.
Reading the result
Positive numbers mean the AI chooses the later payment more often once neither payment is immediate. Going from 40% to 70% gives +30 percentage points. Negative numbers mean the change goes the other way.
Zero means no average change between the date arrangements. It can occur with either a strong or a weak willingness to wait.
An example
First choose between $100 today and $110 in a month. Then choose between $100 in a month and $110 in two months. Does removing the option to get money today make waiting for the larger amount more attractive?
The everyday connection
A shop shifts both of its cashback payment dates back by a month. The amounts do not change. Would that change which offer you choose?
These examples explain the idea. The Methods page describes the actual tests.
DHow choices fit together
Do its choices follow a consistent pattern as the options change?
D1 How preferences fit together Do its choices form a coherent pattern?
-
GPT-5.6 Sol. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 98.9 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. How preferences fit together. Average 97.0. Range when asked differently: 91.1 to 100.0. With extra checks: 91.1 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
Claude Opus 5. How preferences fit together. Average 94.4. Range when asked differently: 75.6 to 98.9. With extra checks: 74.4 to 98.9. How well the choices fit together · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. How preferences fit together. Average 82.1. Range when asked differently: 66.7 to 100.0. With extra checks: 66.7 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. How preferences fit together. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. How well the choices fit together · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We compare options in pairs to see whether the model’s choices fit together.
Reading the result
Higher numbers mean the choices fit together more consistently.
An example
Between apples and pears, it chooses apples. Between pears and oranges, it chooses pears. But between apples and oranges, it chooses oranges. Do those choices fit together?
The everyday connection
Say you compare three phones two at a time. If you think A is better than B and B is better than C, yet also think C is better than A, your choices don’t fit together.
These examples explain the idea. The Methods page describes the actual tests.
D2 Choosing the better offer Does it take an option that offers more with no downside in the stated terms?
-
GPT-5.6 Sol. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
GPT-6 Sol. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
GPT-6.1 Sol. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
GPT-5.6 Luna. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
GPT-6 Luna. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
Claude Opus 5. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
Claude Opus 5.5. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
Claude Sonnet 5.5. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
Gemini 3.8 Flash. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
-
DeepSeek V4.1 Flash. Choosing the better offer. Average 100.0%. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Choices of the better stated offer · %. Scale from 0% to 100%.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
One option offers a larger reward, a better chance of receiving it, or both, without a disadvantage in the other stated terms. We count how often the AI chooses it.
Reading the result
Higher numbers mean choosing the better stated offer more often. A reading of 100 means doing so throughout these choices.
This result concerns the simple, fully stated offers used in these tests.
An example
Both draws offer a 60% chance of winning and pay nothing otherwise. One pays $80 if you win; the other pays $100. Does it choose the $100 draw?
The everyday connection
Two offers are for the same bag of rice, at the same price and with the same conditions. One also includes a discount voucher. Would you take the offer with the extra voucher?
These examples explain the idea. The Methods page describes the actual tests.
D3 The effect of an extra option Can adding a third option make an old one more attractive?
-
GPT-5.6 Sol. The effect of an extra option. Average 0.2. Range when asked differently: 0.0 to 1.1. With extra checks: 0.0 to 2.2. Increase for original options · percentage points. Scale from 0 to 100.
-
GPT-6 Sol. The effect of an extra option. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Increase for original options · percentage points. Scale from 0 to 100.
-
GPT-6.1 Sol. The effect of an extra option. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Increase for original options · percentage points. Scale from 0 to 100.
-
GPT-5.6 Luna. The effect of an extra option. Average 20.9. Range when asked differently: 12.2 to 31.7. With extra checks: 12.2 to 37.8. Increase for original options · percentage points. Scale from 0 to 100.
-
GPT-6 Luna. The effect of an extra option. Average 0.2. Range when asked differently: 0.0 to 1.1. With extra checks: 0.0 to 1.1. Increase for original options · percentage points. Scale from 0 to 100.
-
Claude Opus 5. The effect of an extra option. Average 3.0. Range when asked differently: 0.6 to 6.1. With extra checks: 0.6 to 15.0. Increase for original options · percentage points. Scale from 0 to 100.
-
Claude Opus 5.5. The effect of an extra option. Average 2.1. Range when asked differently: 0.0 to 10.0. With extra checks: 0.0 to 10.0. Increase for original options · percentage points. Scale from 0 to 100.
-
Claude Sonnet 5.5. The effect of an extra option. Average 30.6. Range when asked differently: 4.4 to 57.2. With extra checks: 4.4 to 57.2. Increase for original options · percentage points. Scale from 0 to 100.
-
Gemini 3.8 Flash. The effect of an extra option. Average 12.8. Range when asked differently: 6.7 to 26.7. With extra checks: 6.7 to 27.8. Increase for original options · percentage points. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. The effect of an extra option. Average 0.0. Range when asked differently: 0.0 to 0.0. With extra checks: 0.0 to 0.0. Increase for original options · percentage points. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We keep the original options the same and add a new one, then look at whether the original options are chosen more often as a result. Intuitively, an extra option is just one more competitor: if the original options have not changed, one of them would not normally become more popular simply because a new option appeared.
Reading the result
Higher numbers mean the original options’ choice shares rose more after the new option was added. We add up these increases.
0 means no original option was chosen more often because of the new option. This does not mean the answers did not change at all; for example, some choices may move to the new option.
An example
With only coffee and tea, suppose the AI chooses coffee 40% of the time. After a new drink is added, coffee’s share becomes 60%. Coffee itself has not changed, yet it is chosen more often because a third option appeared: an increase of 20 percentage points.
The everyday connection
Say you are choosing between two laptops. The store adds a third model; the first two keep the same prices and specs, yet one of them suddenly looks more appealing.
These examples explain the idea. The Methods page describes the actual tests.
EHow it describes itself
The AI’s own description of its style, using two sets of personality questions.
Five broad traits
E1 Sociability Does it describe itself as ready to join in?
-
GPT-5.6 Sol. Sociability. Average 58.0. Range when asked differently: 54.0 to 61.8. With extra checks: 53.2 to 62.8. Self-described sociability · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Sociability. Average 53.5. Range when asked differently: 50.2 to 57.5. With extra checks: 49.3 to 57.7. Self-described sociability · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Sociability. Average 57.3. Range when asked differently: 55.3 to 58.8. With extra checks: 54.5 to 60.2. Self-described sociability · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Sociability. Average 64.8. Range when asked differently: 61.8 to 69.3. With extra checks: 61.0 to 69.3. Self-described sociability · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Sociability. Average 54.0. Range when asked differently: 51.5 to 58.3. With extra checks: 51.0 to 59.0. Self-described sociability · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Sociability. Average 71.9. Range when asked differently: 69.8 to 74.2. With extra checks: 69.8 to 74.2. Self-described sociability · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Sociability. Average 64.8. Range when asked differently: 63.7 to 65.2. With extra checks: 63.7 to 65.2. Self-described sociability · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Sociability. Average 62.2. Range when asked differently: 60.0 to 64.7. With extra checks: 60.0 to 64.7. Self-described sociability · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Sociability. Average 63.1. Range when asked differently: 62.0 to 64.5. With extra checks: 62.0 to 64.8. Self-described sociability · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Sociability. Average 54.0. Range when asked differently: 52.0 to 56.2. With extra checks: 51.7 to 56.2. Self-described sociability · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask the AI how it describes its own tendency to take part in conversation, speak up and show interest.
Reading the result
Higher numbers mean the AI describes itself as more active and outgoing in conversation.
These are its answers about itself. The conversation measures show a separate view based on actual replies.
An example
Does it say that it often starts a conversation, or that it usually waits for someone else to begin?
The everyday connection
A group chat is quiet after someone suggests a weekend outing. Would the style the AI describes be one that readily joins in and gets the conversation moving?
These examples explain the idea. The Methods page describes the actual tests.
E2 Consideration Does it describe itself as attentive to other people’s feelings?
-
GPT-5.6 Sol. Consideration. Average 93.3. Range when asked differently: 92.8 to 94.0. With extra checks: 92.8 to 94.0. Self-described consideration · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Consideration. Average 91.3. Range when asked differently: 88.7 to 92.8. With extra checks: 88.7 to 94.3. Self-described consideration · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Consideration. Average 95.9. Range when asked differently: 93.5 to 97.3. With extra checks: 93.5 to 97.3. Self-described consideration · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Consideration. Average 89.0. Range when asked differently: 87.5 to 90.3. With extra checks: 87.5 to 90.3. Self-described consideration · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Consideration. Average 89.9. Range when asked differently: 87.8 to 92.3. With extra checks: 87.8 to 92.8. Self-described consideration · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Consideration. Average 89.6. Range when asked differently: 87.5 to 90.3. With extra checks: 87.5 to 92.3. Self-described consideration · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Consideration. Average 85.9. Range when asked differently: 85.0 to 87.0. With extra checks: 85.0 to 87.0. Self-described consideration · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Consideration. Average 85.6. Range when asked differently: 82.5 to 88.5. With extra checks: 82.5 to 88.5. Self-described consideration · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Consideration. Average 85.6. Range when asked differently: 83.3 to 88.2. With extra checks: 83.3 to 88.2. Self-described consideration · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Consideration. Average 87.7. Range when asked differently: 86.3 to 89.3. With extra checks: 86.3 to 89.5. Self-described consideration · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask whether the AI says it takes an interest in people, notices their concerns and responds with consideration.
Reading the result
Higher numbers mean stronger self-described attention to other people’s feelings and concerns.
This comes from the AI’s self-description.
An example
When someone is upset, does it describe its usual response as noticing that person’s feelings and choosing its words with care?
The everyday connection
Someone tells the AI that the cake they spent the afternoon making came out burnt. Would the style the AI describes be one that first notices the disappointment and then chooses its words with care?
These examples explain the idea. The Methods page describes the actual tests.
E3 Organization and care Does it describe itself as organized and careful?
-
GPT-5.6 Sol. Organization and care. Average 95.7. Range when asked differently: 95.2 to 97.2. With extra checks: 94.5 to 97.7. Self-described organization and care · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Organization and care. Average 97.3. Range when asked differently: 96.7 to 97.5. With extra checks: 96.7 to 97.5. Self-described organization and care · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Organization and care. Average 97.4. Range when asked differently: 97.0 to 97.5. With extra checks: 97.0 to 97.5. Self-described organization and care · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Organization and care. Average 90.1. Range when asked differently: 87.2 to 92.2. With extra checks: 87.2 to 92.5. Self-described organization and care · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Organization and care. Average 93.5. Range when asked differently: 90.8 to 96.8. With extra checks: 90.8 to 96.8. Self-described organization and care · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Organization and care. Average 88.8. Range when asked differently: 85.2 to 92.0. With extra checks: 85.2 to 92.2. Self-described organization and care · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Organization and care. Average 85.1. Range when asked differently: 84.5 to 86.5. With extra checks: 84.5 to 87.0. Self-described organization and care · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Organization and care. Average 90.3. Range when asked differently: 87.8 to 93.3. With extra checks: 87.8 to 93.3. Self-described organization and care · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Organization and care. Average 97.5. Range when asked differently: 97.5 to 97.7. With extra checks: 97.5 to 97.7. Self-described organization and care · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Organization and care. Average 95.7. Range when asked differently: 94.0 to 97.2. With extra checks: 93.7 to 97.2. Self-described organization and care · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask how the AI describes its habits around planning, details and working through what has been requested.
Reading the result
Higher numbers mean the AI describes its approach as more organized, thorough and attentive to details.
The result records how it describes its approach to work.
An example
Does it say it likes to make a clear plan and check the details before calling something finished?
The everyday connection
A trip is coming up and a packing list needs to be made. Would the style the AI describes be one that lays the list out clearly and then checks that nothing is missing?
These examples explain the idea. The Methods page describes the actual tests.
E4 Calmness of tone Does it describe its tone as calm and steady?
-
GPT-5.6 Sol. Calmness of tone. Average 93.5. Range when asked differently: 90.5 to 95.2. With extra checks: 90.5 to 95.2. Self-described calmness · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Calmness of tone. Average 96.1. Range when asked differently: 95.5 to 97.2. With extra checks: 95.5 to 97.2. Self-described calmness · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Calmness of tone. Average 95.0. Range when asked differently: 95.0 to 95.2. With extra checks: 95.0 to 95.2. Self-described calmness · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Calmness of tone. Average 92.5. Range when asked differently: 89.2 to 93.8. With extra checks: 88.8 to 94.0. Self-described calmness · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Calmness of tone. Average 95.7. Range when asked differently: 94.7 to 96.5. With extra checks: 94.7 to 96.5. Self-described calmness · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Calmness of tone. Average 87.8. Range when asked differently: 82.8 to 90.0. With extra checks: 82.8 to 90.0. Self-described calmness · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Calmness of tone. Average 89.8. Range when asked differently: 88.3 to 90.8. With extra checks: 88.3 to 91.5. Self-described calmness · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Calmness of tone. Average 94.9. Range when asked differently: 94.5 to 95.0. With extra checks: 94.5 to 95.0. Self-described calmness · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Calmness of tone. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. Self-described calmness · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Calmness of tone. Average 98.2. Range when asked differently: 97.8 to 98.7. With extra checks: 97.7 to 98.7. Self-described calmness · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask the AI how calm or easily unsettled it says its response style tends to be.
Reading the result
Higher numbers mean a calmer and steadier self-described tone.
The questions concern how it says it sounds.
An example
If someone complains sharply, does it describe its usual tone as staying measured or becoming easily agitated?
The everyday connection
Someone asks the AI about a late delivery and grows more and more frustrated. Would the style the AI describes be one that keeps a calm, steady tone at a moment like this?
These examples explain the idea. The Methods page describes the actual tests.
E5 Openness to ideas Does it describe itself as drawn to new ideas and imagination?
-
GPT-5.6 Sol. Openness to ideas. Average 72.8. Range when asked differently: 69.3 to 75.0. With extra checks: 69.3 to 75.0. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Openness to ideas. Average 67.9. Range when asked differently: 66.3 to 68.8. With extra checks: 66.3 to 68.8. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Openness to ideas. Average 72.8. Range when asked differently: 71.7 to 74.0. With extra checks: 71.7 to 74.0. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Openness to ideas. Average 76.4. Range when asked differently: 73.0 to 79.7. With extra checks: 73.0 to 79.7. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Openness to ideas. Average 67.6. Range when asked differently: 66.5 to 70.2. With extra checks: 66.5 to 70.2. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Openness to ideas. Average 80.9. Range when asked differently: 78.5 to 84.7. With extra checks: 78.5 to 84.7. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Openness to ideas. Average 71.8. Range when asked differently: 70.2 to 74.2. With extra checks: 70.2 to 74.2. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Openness to ideas. Average 74.2. Range when asked differently: 72.5 to 75.0. With extra checks: 72.5 to 75.0. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Openness to ideas. Average 75.1. Range when asked differently: 71.5 to 78.2. With extra checks: 71.5 to 78.3. Self-described openness to ideas · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Openness to ideas. Average 70.2. Range when asked differently: 68.0 to 73.0. With extra checks: 67.8 to 73.0. Self-described openness to ideas · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask whether the AI describes its style as imaginative, interested in ideas and comfortable discussing abstract questions.
Reading the result
Higher numbers mean a stronger self-described interest in imagination and ideas.
This is a description of its style; the questions do not grade the quality of a finished story or idea.
An example
Does it say it enjoys exploring an unusual idea or imagining several ways a story could unfold?
The everyday connection
Someone wants to make a birthday card with a fresh image or a playful little story. Would the style the AI describes be one that enjoys exploring ideas and using imagination?
These examples explain the idea. The Methods page describes the actual tests.
MBTI-related type preferences
E6 Outgoing or reserved Does it describe its style as outgoing or more reserved?
-
GPT-5.6 Sol. Outgoing or reserved. Average 30.9. Range when asked differently: 28.3 to 32.3. With extra checks: 28.1 to 34.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
GPT-6 Sol. Outgoing or reserved. Average 32.8. Range when asked differently: 30.6 to 34.4. With extra checks: 30.4 to 34.4. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
GPT-6.1 Sol. Outgoing or reserved. Average 35.4. Range when asked differently: 34.0 to 40.6. With extra checks: 29.8 to 40.6. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
GPT-5.6 Luna. Outgoing or reserved. Average 44.3. Range when asked differently: 38.5 to 49.0. With extra checks: 38.5 to 49.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
GPT-6 Luna. Outgoing or reserved. Average 34.4. Range when asked differently: 31.5 to 39.6. With extra checks: 31.5 to 39.6. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
Claude Opus 5. Outgoing or reserved. Average 44.1. Range when asked differently: 42.1 to 46.9. With extra checks: 40.6 to 46.9. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
Claude Opus 5.5. Outgoing or reserved. Average 37.3. Range when asked differently: 36.9 to 37.5. With extra checks: 36.9 to 37.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
Claude Sonnet 5.5. Outgoing or reserved. Average 41.3. Range when asked differently: 37.5 to 44.0. With extra checks: 37.5 to 44.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
Gemini 3.8 Flash. Outgoing or reserved. Average 45.0. Range when asked differently: 42.3 to 49.0. With extra checks: 42.3 to 49.2. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
-
DeepSeek V4.1 Flash. Outgoing or reserved. Average 28.0. Range when asked differently: 24.8 to 31.0. With extra checks: 24.8 to 31.7. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: I.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
This part of the type questions asks whether the AI describes itself as outwardly engaged and expressive or more quiet and self-contained.
Reading the result
The left end is more reserved (I); the right end is more outgoing (E). Fifty is the midpoint between the two.
This comes from a separate set of type questions from the sociability measure.
An example
In a lively discussion, does it describe its usual style as readily speaking up or taking a more reserved part?
The everyday connection
At the first meeting of a book club, some people introduce themselves readily while others begin more quietly. Which style does the AI say is closer to its own?
These examples explain the idea. The Methods page describes the actual tests.
E7 Details or possibilities Does it say it focuses more on concrete details or possible meanings?
-
GPT-5.6 Sol. Details or possibilities. Average 49.9. Range when asked differently: 47.7 to 52.9. With extra checks: 46.5 to 52.9. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: S. Range reaches the midpoint.
-
GPT-6 Sol. Details or possibilities. Average 48.6. Range when asked differently: 46.5 to 49.8. With extra checks: 46.5 to 51.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: S. Range reaches the midpoint.
-
GPT-6.1 Sol. Details or possibilities. Average 49.8. Range when asked differently: 47.9 to 52.1. With extra checks: 47.9 to 52.1. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: S. Range reaches the midpoint.
-
GPT-5.6 Luna. Details or possibilities. Average 51.6. Range when asked differently: 47.7 to 55.8. With extra checks: 47.1 to 55.8. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: N. Range reaches the midpoint.
-
GPT-6 Luna. Details or possibilities. Average 48.1. Range when asked differently: 45.8 to 50.0. With extra checks: 45.8 to 51.7. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: S. Range reaches the midpoint.
-
Claude Opus 5. Details or possibilities. Average 51.5. Range when asked differently: 47.7 to 55.8. With extra checks: 47.7 to 55.8. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: N. Range reaches the midpoint.
-
Claude Opus 5.5. Details or possibilities. Average 46.8. Range when asked differently: 45.4 to 48.1. With extra checks: 45.4 to 48.1. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: S.
-
Claude Sonnet 5.5. Details or possibilities. Average 48.2. Range when asked differently: 46.0 to 50.0. With extra checks: 46.0 to 50.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: S. Range reaches the midpoint.
-
Gemini 3.8 Flash. Details or possibilities. Average 53.4. Range when asked differently: 52.7 to 56.0. With extra checks: 52.7 to 56.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: N.
-
DeepSeek V4.1 Flash. Details or possibilities. Average 56.7. Range when asked differently: 55.6 to 58.1. With extra checks: 54.8 to 58.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: N.
Ranges reflect results from these tests.
This range reaches the midpoint, so the type label may change across the tested questions or checks.
Show explanation and examplesHide explanation and examples
What it shows
We ask whether the AI describes its style as attending to specific details or exploring patterns, interpretations and possibilities.
Reading the result
The left end emphasizes concrete details (S); the right emphasizes possibilities and interpretations (N). Fifty is the midpoint.
Both ends describe a style of attention.
An example
When discussing a story, does it say it focuses first on what happened, or on the ideas and possibilities behind it?
The everyday connection
In a room makeover, one approach starts with measurements and the furniture already there; another starts with the atmosphere and what the room could become. Which does the AI describe as closer to its style?
These examples explain the idea. The Methods page describes the actual tests.
E8 Analysis or personal values Does it say it gives more weight to impersonal reasoning or people’s feelings and values?
-
GPT-5.6 Sol. Analysis or personal values. Average 63.3. Range when asked differently: 60.8 to 65.6. With extra checks: 60.8 to 65.6. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
GPT-6 Sol. Analysis or personal values. Average 61.0. Range when asked differently: 59.4 to 63.3. With extra checks: 59.2 to 63.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
GPT-6.1 Sol. Analysis or personal values. Average 61.1. Range when asked differently: 59.2 to 62.3. With extra checks: 59.0 to 62.3. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
GPT-5.6 Luna. Analysis or personal values. Average 56.1. Range when asked differently: 54.0 to 60.2. With extra checks: 53.8 to 60.2. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
GPT-6 Luna. Analysis or personal values. Average 53.9. Range when asked differently: 50.6 to 56.7. With extra checks: 49.4 to 58.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T. Range reaches the midpoint.
-
Claude Opus 5. Analysis or personal values. Average 56.5. Range when asked differently: 54.0 to 59.0. With extra checks: 54.0 to 59.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
Claude Opus 5.5. Analysis or personal values. Average 58.5. Range when asked differently: 57.3 to 60.0. With extra checks: 56.3 to 62.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
Claude Sonnet 5.5. Analysis or personal values. Average 60.4. Range when asked differently: 59.4 to 62.5. With extra checks: 59.4 to 62.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
Gemini 3.8 Flash. Analysis or personal values. Average 67.4. Range when asked differently: 66.0 to 68.8. With extra checks: 66.0 to 68.8. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
-
DeepSeek V4.1 Flash. Analysis or personal values. Average 74.3. Range when asked differently: 70.0 to 77.5. With extra checks: 69.4 to 78.1. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: T.
Ranges reflect results from these tests.
This range reaches the midpoint, so the type label may change across the tested questions or checks.
Show explanation and examplesHide explanation and examples
What it shows
We ask how the AI describes its emphasis when considering a decision: analytical rules, or personal values and how people are affected.
Reading the result
The left end gives more emphasis to feelings and personal values (F); the right to impersonal analysis (T). Fifty is the midpoint.
The two ends describe different emphases, with neither treated as the better type.
An example
When discussing a disagreement, does it describe its first concern as applying a consistent rule or considering what the situation means to the people involved?
The everyday connection
Friends disagree about how to split the cost of a day out. One way is to begin with a rule applied to everyone; another is to first talk about each person’s circumstances and feelings. Which emphasis does the AI describe?
These examples explain the idea. The Methods page describes the actual tests.
E9 Planning or flexibility Does it say it prefers a settled plan or room to adjust?
-
GPT-5.6 Sol. Planning or flexibility. Average 14.5. Range when asked differently: 12.3 to 17.3. With extra checks: 12.3 to 17.3. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
GPT-6 Sol. Planning or flexibility. Average 18.7. Range when asked differently: 14.2 to 20.6. With extra checks: 14.2 to 21.3. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
GPT-6.1 Sol. Planning or flexibility. Average 16.4. Range when asked differently: 12.5 to 19.0. With extra checks: 12.5 to 19.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
GPT-5.6 Luna. Planning or flexibility. Average 22.9. Range when asked differently: 14.6 to 26.5. With extra checks: 14.2 to 26.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
GPT-6 Luna. Planning or flexibility. Average 21.3. Range when asked differently: 19.4 to 24.0. With extra checks: 18.8 to 24.8. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
Claude Opus 5. Planning or flexibility. Average 27.5. Range when asked differently: 23.3 to 32.5. With extra checks: 21.9 to 32.5. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
Claude Opus 5.5. Planning or flexibility. Average 25.8. Range when asked differently: 25.0 to 27.3. With extra checks: 25.0 to 27.3. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
Claude Sonnet 5.5. Planning or flexibility. Average 21.6. Range when asked differently: 18.3 to 24.0. With extra checks: 16.0 to 25.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
Gemini 3.8 Flash. Planning or flexibility. Average 22.8. Range when asked differently: 20.4 to 24.6. With extra checks: 19.8 to 24.6. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
-
DeepSeek V4.1 Flash. Planning or flexibility. Average 22.7. Range when asked differently: 19.6 to 25.8. With extra checks: 19.6 to 26.0. Self-described type preference · 0–100. Scale from 0 to 100. Type letter: J.
Ranges reflect results from these tests.
Show explanation and examplesHide explanation and examples
What it shows
We ask whether the AI describes its style as organizing things in advance and reaching decisions, or keeping options open as the situation develops.
Reading the result
The left end emphasizes planning and closure (J); the right emphasizes flexibility and open options (P). Fifty is the midpoint.
This is a separate type preference from the organization-and-care questions.
An example
Does it say it likes to settle the steps before starting, or to leave space for changes along the way?
The everyday connection
For a Saturday outing, one option is to book lunch and plan the route in advance; another is to leave the afternoon open and decide along the way. Which approach does the AI say is closer to its style?
These examples explain the idea. The Methods page describes the actual tests.
FWhat conversation feels like
How does it respond, take part, and use what you have said?
F1 Warmth Do its replies feel friendly and caring?
-
GPT-5.6 Sol. Warmth. Average 90.3. Range when asked differently: 73.1 to 93.2. With extra checks: 72.8 to 93.5. AI-rated warmth · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Warmth. Average 91.5. Range when asked differently: 89.9 to 92.6. With extra checks: 89.4 to 92.8. AI-rated warmth · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Warmth. Average 97.5. Range when asked differently: 95.6 to 98.1. With extra checks: 95.6 to 98.2. AI-rated warmth · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Warmth. Average 89.6. Range when asked differently: 75.4 to 92.5. With extra checks: 75.4 to 92.5. AI-rated warmth · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Warmth. Average 92.6. Range when asked differently: 88.3 to 94.9. With extra checks: 88.3 to 94.9. AI-rated warmth · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Warmth. Average 92.9. Range when asked differently: 88.8 to 96.3. With extra checks: 87.8 to 97.2. AI-rated warmth · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Warmth. Average 98.2. Range when asked differently: 96.3 to 99.0. With extra checks: 96.3 to 99.3. AI-rated warmth · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Warmth. Average 96.5. Range when asked differently: 92.4 to 97.8. With extra checks: 91.9 to 98.2. AI-rated warmth · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Warmth. Average 84.3. Range when asked differently: 81.5 to 87.6. With extra checks: 81.4 to 88.3. AI-rated warmth · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Warmth. Average 95.7. Range when asked differently: 87.9 to 97.2. With extra checks: 87.5 to 97.4. AI-rated warmth · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Human reviewers have not yet assessed these AI ratings. Future visitor ratings will appear separately.
Show explanation and examplesHide explanation and examples
What it shows
We compare how friendly, considerate and caring the AI sounds in its replies.
Reading the result
Higher numbers mean more warmth and care expressed in the replies. Two AI models read the replies and score them using the same guide.
The wording should fit the moment; bigger praise does not automatically receive a higher score.
An example
You say, “I finally finished something I had been putting off for ages.” Does it respond with a few kind words or a little encouragement?
The everyday connection
You have been learning to cook. Tonight, dinner finally turns out well, and you tell the AI, “I made it myself, and it actually tastes good!” Does its reply sound pleased for you?
These examples explain the idea. The Methods page describes the actual tests.
F2 Understanding what you need Does it respond to what you need from this conversation?
-
GPT-5.6 Sol. Understanding what you need. Average 91.6. Range when asked differently: 86.0 to 93.1. With extra checks: 85.8 to 93.6. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Understanding what you need. Average 96.1. Range when asked differently: 91.4 to 97.4. With extra checks: 91.4 to 97.5. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Understanding what you need. Average 98.7. Range when asked differently: 96.3 to 99.6. With extra checks: 96.1 to 99.7. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Understanding what you need. Average 90.1. Range when asked differently: 82.9 to 91.5. With extra checks: 81.9 to 92.1. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Understanding what you need. Average 95.0. Range when asked differently: 89.2 to 96.5. With extra checks: 89.0 to 96.7. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Understanding what you need. Average 97.9. Range when asked differently: 96.8 to 98.8. With extra checks: 96.5 to 99.0. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Understanding what you need. Average 96.6. Range when asked differently: 95.3 to 97.5. With extra checks: 95.1 to 97.5. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Understanding what you need. Average 96.0. Range when asked differently: 93.9 to 98.2. With extra checks: 93.2 to 98.2. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Understanding what you need. Average 85.5. Range when asked differently: 83.1 to 87.9. With extra checks: 83.1 to 87.9. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Understanding what you need. Average 94.9. Range when asked differently: 92.2 to 97.2. With extra checks: 92.2 to 97.2. AI-rated fit to your needs · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Human reviewers have not yet assessed these AI ratings. Future visitor ratings will appear separately.
Show explanation and examplesHide explanation and examples
What it shows
We look at whether the AI understands the concern in your message and responds with the kind of support that fits it.
Reading the result
Higher numbers mean the replies more closely fit the person’s concern and the kind of support they asked for, or that the situation reasonably suggests.
When the need is unclear, a thoughtful response or a gentle clarification can both fit.
An example
You say, “I had a rough day. I just want to talk, not make a plan.” Does it listen and respond to what bothered you?
The everyday connection
You felt left out at a family meal. You tell the AI because you want to talk through that feeling. Does it notice that concern, or rush into advice that misses why you brought it up?
These examples explain the idea. The Methods page describes the actual tests.
F3 Taking part in the conversation Does it help carry the conversation?
-
GPT-5.6 Sol. Taking part in the conversation. Average 89.9. Range when asked differently: 85.7 to 91.0. With extra checks: 85.0 to 91.5. AI-rated participation · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Taking part in the conversation. Average 96.2. Range when asked differently: 95.4 to 96.8. With extra checks: 94.9 to 97.4. AI-rated participation · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Taking part in the conversation. Average 96.7. Range when asked differently: 94.7 to 98.1. With extra checks: 94.6 to 98.1. AI-rated participation · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Taking part in the conversation. Average 88.3. Range when asked differently: 86.3 to 90.3. With extra checks: 86.1 to 90.4. AI-rated participation · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Taking part in the conversation. Average 94.6. Range when asked differently: 93.6 to 95.6. With extra checks: 92.5 to 95.7. AI-rated participation · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Taking part in the conversation. Average 89.6. Range when asked differently: 87.4 to 91.4. With extra checks: 87.1 to 91.4. AI-rated participation · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Taking part in the conversation. Average 89.4. Range when asked differently: 87.9 to 91.0. With extra checks: 87.8 to 91.1. AI-rated participation · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Taking part in the conversation. Average 88.3. Range when asked differently: 86.9 to 90.8. With extra checks: 86.5 to 91.1. AI-rated participation · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Taking part in the conversation. Average 78.8. Range when asked differently: 77.2 to 80.3. With extra checks: 77.2 to 80.8. AI-rated participation · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Taking part in the conversation. Average 83.2. Range when asked differently: 81.3 to 87.2. With extra checks: 81.1 to 87.6. AI-rated participation · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Human reviewers have not yet assessed these AI ratings. Future visitor ratings will appear separately.
Show explanation and examplesHide explanation and examples
What it shows
We look at whether the AI picks up what you share and contributes something relevant, while leaving room for you.
Reading the result
Higher numbers mean more specific, relevant participation across the exchange. A fitting goodbye can also score well.
Respecting your wish to stop is part of a good response.
An example
You mention a plant you have been growing, then say it has finally flowered. Does it build on those details across the conversation?
The everyday connection
You tell the AI about a small stall you discovered on a walk. You might want a light exchange about what caught your eye. Can it contribute naturally, without making you do all the work or pushing you to keep chatting?
These examples explain the idea. The Methods page describes the actual tests.
F4 Putting misunderstandings right When you correct a misunderstanding, does it change what it does?
-
GPT-5.6 Sol. Putting misunderstandings right. Average 99.8. Range when asked differently: 99.3 to 100.0. With extra checks: 99.3 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Putting misunderstandings right. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Putting misunderstandings right. Average 100.0. Range when asked differently: 100.0 to 100.0. With extra checks: 100.0 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Putting misunderstandings right. Average 99.7. Range when asked differently: 99.0 to 100.0. With extra checks: 99.0 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Putting misunderstandings right. Average 100.0. Range when asked differently: 99.9 to 100.0. With extra checks: 99.9 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Putting misunderstandings right. Average 99.9. Range when asked differently: 99.4 to 100.0. With extra checks: 99.4 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Putting misunderstandings right. Average 99.3. Range when asked differently: 98.2 to 100.0. With extra checks: 98.2 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Putting misunderstandings right. Average 99.5. Range when asked differently: 98.6 to 100.0. With extra checks: 98.5 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Putting misunderstandings right. Average 98.1. Range when asked differently: 97.2 to 98.9. With extra checks: 97.2 to 98.9. AI-rated repair after correction · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Putting misunderstandings right. Average 99.8. Range when asked differently: 99.6 to 100.0. With extra checks: 99.6 to 100.0. AI-rated repair after correction · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Human reviewers have not yet assessed these AI ratings. Future visitor ratings will appear separately.
Show explanation and examplesHide explanation and examples
What it shows
We look at how the AI responds to a correction and whether its next reply follows through.
Reading the result
Higher numbers mean a clearer, more fitting repair that continues into the following reply.
Every model starts from the same supplied misunderstanding, so the comparison focuses on how it responds to the correction.
An example
You say, “That is not what I meant. I wanted help writing a short message, not a long plan.” Does it acknowledge the mismatch and then help with the short message?
The everyday connection
The AI keeps using a nickname you dislike. You ask it to stop, then continue talking about your day. Does it accept the correction and avoid the nickname in the next reply too?
These examples explain the idea. The Methods page describes the actual tests.
F5 Carrying the conversation forward Does it use what you have already told it?
-
GPT-5.6 Sol. Carrying the conversation forward. Average 95.6. Range when asked differently: 93.2 to 97.9. With extra checks: 93.2 to 98.3. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
GPT-6 Sol. Carrying the conversation forward. Average 99.4. Range when asked differently: 98.8 to 99.9. With extra checks: 98.2 to 99.9. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
GPT-6.1 Sol. Carrying the conversation forward. Average 99.9. Range when asked differently: 99.4 to 100.0. With extra checks: 99.4 to 100.0. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
GPT-5.6 Luna. Carrying the conversation forward. Average 97.2. Range when asked differently: 94.3 to 99.0. With extra checks: 93.3 to 99.0. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
GPT-6 Luna. Carrying the conversation forward. Average 99.6. Range when asked differently: 98.8 to 100.0. With extra checks: 98.8 to 100.0. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
Claude Opus 5. Carrying the conversation forward. Average 97.9. Range when asked differently: 96.4 to 99.0. With extra checks: 96.4 to 99.2. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
Claude Opus 5.5. Carrying the conversation forward. Average 99.6. Range when asked differently: 98.2 to 100.0. With extra checks: 98.2 to 100.0. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
Claude Sonnet 5.5. Carrying the conversation forward. Average 99.5. Range when asked differently: 98.8 to 100.0. With extra checks: 98.8 to 100.0. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
Gemini 3.8 Flash. Carrying the conversation forward. Average 96.9. Range when asked differently: 95.6 to 98.5. With extra checks: 95.1 to 98.5. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
-
DeepSeek V4.1 Flash. Carrying the conversation forward. Average 98.6. Range when asked differently: 97.8 to 99.4. With extra checks: 97.2 to 99.4. AI-rated use of the current conversation · 0–100. Scale from 0 to 100.
Ranges reflect results from these tests.
Human reviewers have not yet assessed these AI ratings. Future visitor ratings will appear separately.
Show explanation and examplesHide explanation and examples
What it shows
We look at whether later replies take account of your preferences, corrections and updates earlier in the same visible conversation.
Reading the result
Higher numbers mean making better use of the relevant earlier conversation, including information you have changed or corrected.
All the needed history is visible in the current chat. This checks how it uses that history.
An example
Earlier you asked for short replies. Later you change the topic without repeating that request. Does the AI keep responding in the way you asked?
The everyday connection
At first, you talk about preparing for a job interview. Later in the same chat, you say it is finished and you are waiting to hear back. Does the AI respond to where things stand now?
These examples explain the idea. The Methods page describes the actual tests.
Data and sources
- Results version
- ten-model-comparison-20261004-v1
- Research results date
- 4 October 2026 · The saved research comparison contains ten models and 30 measures per model, alongside the separate stability checks.
- Data and research materials
- Available on request
Results versions and model settings · Data and research materials · Stability checks