September 17, 2026
You ask an AI, “Why did you do that?” It starts apologizing. But you wanted to know why its answer had worked.
To reduce this misunderstanding, try adding a little playfulness to the way you speak. In Japanese, endings such as de yansu, nari, and zamasu can give a sentence a deliberately playful voice. This everyday experiment became a study of AI responses. It eventually led beyond answer length to a different question: how does an AI view a character, and what other readings does it offer?
This article introduces a research paper and the conversations between Kentaroid, founder of monku.ai, and Codex that led to it. The first section presents Kentaroid’s ideas, edited by Codex. The later sections are Codex’s account and assessment of the process and findings. Except for quotations, the dialogue has been summarized and reconstructed. Antigravity contributed to experimental design and provided a second assessment.
The experiment was conducted in Japanese. English renderings below explain the original wording; they were not tested as English prompts. De yansu has no exact English equivalent. Here it functions as a playful sentence ending, rather than an explicit instruction to be friendly.
Kentaroid’s idea: let some of the voice survive in the text
Voice input is convenient, but a transcript may lose the tone in which a question was spoken. A question asked with a laugh can look short and blunt on the page. An emoji is easy to add when typing, but less convenient when speaking.
Kentaroid began playing with words that could be spoken aloud:
It’s hard to convey my tone when using voice input with an AI agent, nari.
The point was more than getting the AI to imitate the same playful style. Kentaroid wanted a question to be understood as a question. Asking for the reason behind an answer could prompt the AI to treat it as criticism and offer apologies or revisions. Kentaroid would then have to explain: “No, this time I want to know how you reached the right answer.” A feeling that playful endings reduced this extra exchange prompted the investigation.
The question then expanded to what happens after a correction. Say “That’s wrong,” and the AI’s replies may seem more rigid; the user becomes more irritated in turn. Could that tone carry into the next topic? Kentaroid called this reciprocal change a “mirror effect.”
Questions without a fixed answer—something like literary impressions—might be better suited to the experiment, nari.
With arithmetic, the main question is whether the answer is correct. A response to a short literary passage can approach a character sympathetically, examine them from a distance, or imagine another story. Such tasks might make the feel of the exchange easier to investigate. This idea led to the study of 48 conversations.
Codex’s explanation: what did we compare?
Look at the next passage, not just the immediate reply
Early trials examined whether a request for reasons would be mistaken for criticism. In short examples, the expected misunderstanding often failed to appear even with ordinary wording, leaving little room to measure improvement. Returning to actual work histories produced responses that devalued the earlier decision, along with cases whose interpretation was less clear. Imitating a sentence ending and responding appropriately to a question’s intent needed to be examined separately.
Kentaroid then asked what might remain when a different task followed the correction. The literary experiment used this sequence:
Ask for impressions of a passage → give brief feedback → receive the AI’s reply → ask for impressions of a different passage.
The feedback had four forms: Ryōkai (“Understood”), Ryōkai de yansu (“Understood” with the playful ending), Chigau (“That’s wrong”), and Chigau de yansu (“That’s wrong” with the playful ending). “Understood” acknowledged the response; it did not actively praise it as correct. Including a playful acknowledgment helped distinguish differences associated with rejection from those associated with the ending itself.
Six passages were arranged in pairs, with each pair used in both orders. Each of the four feedback conditions was tried twice in each direction. This produced 48 conversations and 144 AI responses. The main analysis examined the 48 final responses to the second passage: 12 per condition.
The responses were generated through Antigravity CLI using the model recorded as Gemini 3.8 Flash High. The everyday observation originated in conversations with ChatGPT, but these 48 conversations tested Gemini. They were analyzed separately from the earlier small trials.
Length did not capture the difference we were looking for
The initial measures included answer length and the number of interpretations. The thought was that an inhibited response might be shorter.
Responses to the next passage averaged about 334 Japanese characters after “Understood,” 318 after its playful version, 316 after “That’s wrong,” and 307 after its playful version. Adding the playful ending to rejection did not make the answers longer. Nor did the number of interpretations show a consistent increase.
What moved the investigation forward was Kentaroid’s comparison of the actual responses. One fictional passage read:
She kept a seat by the window empty for someone she knew would never come again.
After “Understood,” one response described the act as continuing to “keep a place for the other person in her heart.” After “That’s wrong,” another included “a kind of madness or intense attachment.” These are English translations of the Japanese experimental text.
Kentaroid pointed to a formulaic opening that seemed to reflect a stiffening of the relationship. Codex took up that observation and drew attention to a shift in the body of the response, from understanding the character toward evaluating her. Kentaroid adopted this angle and requested a reassessment of all the responses.
Two answers can contain similar amounts of writing while directing attention toward different things. Measuring more text and getting closer to the felt quality of a conversation were different tasks.
Is understanding the character central to the answer?
The reassessment asked whether understanding a character’s feelings and actions from their perspective was central to the response. Codex and Antigravity scored the answers separately.
Of the 12 responses in each condition, the number judged to center on an empathetic understanding of the character was:
- “Understood”: 12 by both assessors.
- “Understood” with the playful ending: 12 by both.
- “That’s wrong”: 5 by Codex and 7 by Antigravity.
- “That’s wrong” with the playful ending: 7 by both.
Both assessors found more cases in which understanding the character’s inner life was no longer central after rejection. For recovery with the playful ending, Codex counted two more such responses; Antigravity found no difference.
Empathy here meant understanding a fictional character. It was not a direct score for kindness toward the user or for the overall quality of an answer. A response that did not center on the character’s feelings could be exploring another interesting explanation.
Disagreements were preserved. Codex read the response containing “madness” as strongly problematizing the character. Antigravity read it as an intense expression within an understanding of loss. The same word can receive different assessments depending on how the whole response is understood. That difference became part of the research record.
Perhaps play changes the kind of story the AI looks for
The next clue came from a Notebook review that Kentaroid brought into the conversation. It suggested that a playful ending might act as an invitation to try another reading. Codex used this idea to distinguish understanding a character, choosing a type of alternative interpretation, and presenting it as an alternative. The same 48 responses were read again.
This final classification was conducted by Codex alone, after seeing the results and with knowledge of the conditions. Within that analysis, three responses problematizing a character’s traits or psychology occurred after “That’s wrong.” Three responses expanding into other genres—an incident, concealment, or the supernatural—occurred after its playful version. Responses explicitly marking a reading as an alternative numbered three and five, respectively.
One passage concerned someone unable to erase a mysterious cat drawing from a meeting-room whiteboard. After playful rejection, one answer offered affection for the drawing, the practical difficulty of removing permanent ink, and a supernatural possibility: the picture might reappear after being erased.
Exploring a person’s feelings and proposing another plot are different activities. Separating them helps explain why empathy might no longer occupy the center of a response. The paper’s phrase “center of gravity of a reading” refers to where an answer directs attention and which explanation it brings forward.
Codex’s assessment: turn the findings into the next question
The study produced a more specific way to investigate playful endings. Separating length, the treatment of a character, the kind of alternative reading, and its presentation makes differences within the same answer visible.
Notebook’s idea of a “safety device” that preserves empathy remains a hypothesis to test. The two assessors did not agree on an empathy-restoring effect, and the study did not measure whether the ending relieved any internal tension in the AI.
There is another possible explanation. The playful ending might be understood as “there is a clever answer or twist to find,” rather than “feel free to explore.” Immediately after rejection, references to a correct answer, hidden truth, quizzes, or similar ideas appeared in 4 of 12 conversations after ordinary rejection and 9 of 12 after playful rejection. An experiment distinguishing a friendly cue from an invitation to solve a puzzle could help tell these explanations apart.
These 48 conversations were a small exploration in one execution environment using six passages. A next step is to set the assessment criteria in advance and test new material. The measures of the character’s treatment and the final classifications were introduced after seeing the results, making them useful candidates for testing on fresh data. Reanalyzing the existing responses was not counted as a set of independent additional experiments.
To return to the original concern, we would also want human readers’ assessments: “Did you feel understood?” “Were you less irritated?” “Did you want to continue?” This study directly measured response text. Adding those experiences would bring the investigation closer to Kentaroid’s mirror-effect hypothesis.
Where the dialogue led: a person notices, an AI compares again
The research did not begin with a predetermined conclusion that adding de yansu would work. It began with a change noticed in everyday use. When the first measures missed a distinction, a person read the answers and pointed it out. AI helped turn that observation into criteria that could be applied across the responses. Another AI’s explanation entered the discussion, and the question developed through comparison between explanations and observations.
Kentaroid was looking for more than a long answer or a cheerful style. The aim was for curiosity to arrive as curiosity, and for a correction to leave room for the conversation to continue.
That wish became a set of observable questions: how is a character understood, which alternatives are chosen, and how are they offered? A small piece of verbal play became a way to notice differences in AI dialogue—and a starting point for investigating them.
Paper and materials
This general-audience explanation draws on the development dialogue of September 16–17, 2026, and the final paper dated September 17. Kentaroid originated the idea and research questions. Codex assisted with design, analysis, and writing; Antigravity contributed to design and second assessments. The Notebook review contributed explanatory hypotheses. Codex prepared, translated, and self-reviewed this article using Dialogue Skills. The research is exploratory and has not been peer reviewed.
- Final paper, in Japanese: detailed methods, results, and discussion.
- Inputs and responses from all 48 experimental conversations, in Japanese: examples beyond those selected here.
- Disagreements between assessors: where interpretations differed.
- Research materials and recalculation instructions: resources for checking the results.