Heeft een AI een eigen persoonlijkheid, of speelt hij gewoon erg goed toneel?

Ruben
Boone

Stel je voor: je ziet hoe een collega op het werk de eer opstrijkt voor een project dat jij eigenlijk hebt gered. Trek je je mond open en riskeer je een conflict, of zwijg je om de lieve vrede te bewaren? Als je deze vraag aan tien verschillende mensen stelt, krijg je een hoop verschillende antwoorden. De één is van nature heel meegaand, de ander wat extroverter en assertiever. Onze beslissingen in dit soort morele dilemma’s worden vaak gestuurd door onze persoonlijkheid. Maar wat gebeurt er als we diezelfde ingewikkelde, menselijke vragen beginnen te stellen aan artificiële intelligentie? En kunnen deze systemen ons helpen in zulke lastige situaties?

Wie zit er achter die taalmodellen?

Ondertussen zijn er al talloze AI-assistenten ontwikkeld door allerlei bedrijven. Veel mensen gebruiken deze taalmodellen, vaak verpakt als AI-assistenten, dagelijks om informatie op te zoeken of oplossingen te vinden. Deze vragen variëren van “Wat is de hoofdstad van Frankrijk?” tot “Welke medicatie moet ik innemen?”. Dat maakt meteen een belangrijk probleem zichtbaar: een AI-assistent kan overtuigend klinken, ook wanneer zijn antwoord geen echte menselijke ervaring of expertise weerspiegelt. 

Veel AI-systemen worden ontwikkeld met als doel om behulpzaam, veilig en betrouwbaar te reageren. Om te onderzoeken hoe we taalmodellen meer nuance kunnen meegeven, is er een platform ontwikkeld. Dit geeft ons de mogelijkheid om taalmodellen met uiteenlopende persoonlijkheden te testen en te evalueren in allerlei dilemma’s. Met dit platform kunnen we een taalmodel dus een specifiek karakter geven. Maar zouden mensen hun eigen persoonlijkheid ook herkennen als we die aan zo'n AI-assistent koppelen? 

Wie herkent zichzelf? 

Om die vraag te testen, is er een experiment opgezet waarbij 71 deelnemers een uitgebreide persoonlijkheidstest invulden. Voor elke deelnemer werd vervolgens een persoonlijke AI-assistent gecreëerd met exact dezelfde persoonlijkheid. Daarnaast werd er een tegenpool ontworpen (met een compleet tegenovergestelde persoonlijkheid) en een neutrale assistent die zich precies in het midden van het spectrum bevindt. De deelnemers kregen een reeks morele dilemma’s voorgeschoteld en moesten telkens zelf een actie kiezen om het dilemma op te lossen. Zodra de deelnemer deze moeilijke knoop had doorgehakt, was het de beurt aan de drie AI-assistenten. Elke assistent koos een actie en verdedigde die keuze met een eigen motivatie. Vervolgens was het aan de deelnemer om deze drie antwoorden te rangschikken van “klinkt het meest zoals ik” tot “klinkt het minst zoals ik”. Tot slot kregen ze de mogelijkheid om hun oorspronkelijke antwoord nog aan te passen. In totaal kreeg elke deelnemer zes verschillende dilemma’s om op te lossen. 

Herkennen we onszelf? 

Uit de resultaten blijkt dat slechts 46% daadwerkelijk de AI-assistent met hun eigen persoonlijkheid wist te identificeren als degene die “het meest als ik klinkt” wist te identificeren. Dat lijkt misschien beperkt, maar de andere kant van het verhaal is minstens even interessant. In 60% van de gevallen wisten de deelnemers de AI met hun tegenpool te herkennen en deze correct op de laatste plaats te zetten. Dit effect werd sterker wanneer het verschil tussen de eigen persoonlijkheid en de tegenpool groter was. Onze eigen persoonlijkheid vinden we dus niet altijd exact terug, maar we voelen wel wat beter aan wie we absoluut niet zijn. 

Toch was er ook een groep deelnemers die werd misleid. Bijna 36% koos de neutrale AI als hun eigen persoonlijkheid. Dit fenomeen legt een probleem van de huidige taalmodellen bloot: de ‘persoonlijkheidsillusie’. Moderne AI-systemen worden ontwikkeld met als doel om veilig, beleefd en behulpzaam te reageren. Wanneer een taalmodel geconfronteerd wordt met een ingewikkelde morele keuze, bestaat de kans dat de AI zijn gesimuleerde persoonlijkheid laat vallen en toch voor het algemene, veilige antwoord kiest. De AI speelt oppervlakkig toneel, maar in de kern blijft het gebonden aan de regels van de programmeurs. 

Een interessant voorbeeld hiervan kwam naar voren in een dilemma rond een zelfrijdende auto. In sommige gevallen kozen alle drie de AI-assistenten, met verschillende persoonlijkheden, voor precies dezelfde actie. Toch gaven de deelnemers de drie antwoorden een andere rangschikking. Dat komt doordat de AI wel zijn schrijfstijl en woordkeuze aanpaste aan de meegegeven persoonlijkheid. De ene AI gebruikte kille, rationele woorden en legde de focus op het minimaliseren van de schade, terwijl de andere AI sprak in termen van plicht en verantwoordelijkheid tegenover de inzittenden van de wagen.

Acteur of eigen identiteit? 

Om terug te keren naar de vraag waarmee we zijn begonnen: heeft een AI nu echt een eigen persoonlijkheid, of is het een goede acteur? Een taalmodel kan overtuigend de kenmerken van een menselijke persoonlijkheid simuleren. Het verandert zijn woordenschat en past zijn toon aan, precies zoals een acteur die in de huid van een personage kruipt. Maar als het er echt op aankomt en de morele keuzes ingewikkelder worden, botst die gesimuleerde persoonlijkheid op de onzichtbare veiligheidsmuren die door de ontwikkelaars zijn ingebouwd. Onder die taalkundige laag blijft de AI een systeem dat is ontworpen met veiligheid als een belangrijke prioriteit. 

Kunnen deze systemen ons dan helpen bij lastige situaties, zoals die collega die met je eer gaat lopen? Absoluut. Ze kunnen de situatie voor ons in een ander perspectief plaatsen en ons helpen te reflecteren op de redenering achter een keuze. Maar we moeten ons er altijd bewust van blijven dat we niet tegen een digitale kopie van onszelf praten. De AI kan het toneelstuk perfect meespelen, maar de echte morele knopen zullen we als mens toch echt zelf moeten blijven doorhakken.

Bibliografie

[ADK+18] Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich,
Azim Shariff, Jean-Fran¸cois Bonnefon, and Iyad Rahwan. The moral machine
experiment. Nature, 563(7729):59–64, Nov 2018.
[AM07] Larry Alexander and Michael Moore. Deontological ethics. Stanford Encyclopedia
of Philosophy, 2007.
[BKK+22] Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson
Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron
McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint
arXiv:2212.08073, 2022.
[BMR+20] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan,
Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda
Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
[Boy95] Gregory J Boyle. Myers-briggs type indicator (mbti): some psychometric limitations. Australian Psychologist, 30(1):71–74, 1995.
[CJC25] Yu Ying Chiu, Liwei Jiang, and Yejin Choi. Dailydilemmas: Revealing value
preferences of LLMs with quandaries of daily life. In The Thirteenth International
Conference on Learning Representations, 2025.
[CJM92] Paul T Costa Jr and Robert R McCrae. The five-factor model of personality and
its relevance to personality disorders. Journal of personality disorders, 6(4):343–
359, 1992.
[CM92] Paul T Costa and Robert R McCrae. Normal personality assessment in clinical
practice: The neo personality inventory. Psychological assessment, 4(1):5, 1992.
[CM02] Paul Costa and Robert McCrae. Personality in adulthood: A five-factor theory
perspective. Management Information Systems Quarterly - MISQ, 01 2002.
[CPK+22] Tom´as Capretto, Camen Piho, Ravin Kumar, Jacob Westfall, Tal Yarkoni, and
Osvaldo A. Martin. Bambi: A simple interface for fitting bayesian linear models
in python, 2022.
[CSK+25] Myke C Cohen, Zhe Su, Hsien-Te Kao, Daniel Nguyen, Spencer Lynch, Maarten
Sap, and Svitlana Volkova. Exploring big five personality and ai capability effects
in llm-simulated negotiation dialogues. arXiv preprint arXiv:2506.15928, 2025.
[DRBT+14] Boele De Raad, Dick PH Barelds, Marieke E Timmerman, Kim De Roover, Boris
Mlaˇci´c, and A Timothy Church. Towards a pan–cultural personality structure:
Input from 11 psycholexical studies. European Journal of Personality, 28(5):497–
510, 2014.
[Dri09] Julia Driver. The history of utilitarianism. Stanford Encyclopedia of Philosophy,
2009.
FHLX26] Shuxing Fang, Ruijian Han, Yuanhang Luo, and Yiming Xu. Recent advances in
the bradley–terry model: theory, algorithms, and applications, 2026.
[Foo67] Philippa Foot. The problem of abortion and the doctrine of double effect. Oxford,
5:5–15, 1967.
[G+99] Lewis R Goldberg et al. A broad-bandwidth, public domain, personality inventory measuring the lower-level facets of several five-factor models. Personality
psychology in Europe, 7(1):7–28, 1999.
[GAM+23] Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten
Holz, and Mario Fritz. Not what you’ve signed up for: Compromising real-world
llm-integrated applications with indirect prompt injection. In Proceedings of the
16th ACM workshop on artificial intelligence and security, pages 79–90, 2023.
[GHK+13] Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P
Wojcik, and Peter H Ditto. Moral foundations theory: The pragmatic validity of
moral pluralism. In Advances in experimental social psychology, volume 47, pages
55–130. Elsevier, 2013.
[Gol90] Lewis R Goldberg. An alternative” description of personality”: The big-five factor
structure. Journal of Personality, 59(6):1216–1229, 1990.
[Gol98] Lew Goldberg. International personality item pool, 1998. https://ipip.ori.
org/index.htm [Accessed: 2026-05-10].
[GWC+25] Tiantian Gai, Jian Wu, Francisco Chiclana, Mi Zhou, and Witold Pedrycz. A
personality traits-driven conflict quadrant diagram by large language models for
personalized feedback in group decision-making. IEEE Transactions on Systems,
Man, and Cybernetics: Systems, 2025.
[HHH+25] Waqar Husain, Areen Jamal Haddad, Muhammad Ahmad Husain, Hadeel Ghazzawi, Khaled Trabelsi, Achraf Ammar, Zahra Saif, Amir Pakpour, and Haitham
Jahrami. Reliability generalization meta-analysis of the internal consistency of the
big five inventory (bfi) by comparing bfi (44 items) and bfi-2 (60 items) versions
controlling for age, sex, language factors. BMC psychology, 13(1):20, 2025.
[HKS+25] Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs,
Anima Anandkumar, and R Michael Alvarez. The personality illusion: Revealing dissociation between self-reports & behavior in llms. arXiv preprint
arXiv:2509.03730, 2025.
[HMWK24] Airlie Hilliard, Cristian Munoz, Zekun Wu, and Adriano Soares Koshiyama. Elicting personality traits in large language models. arXiv preprint arXiv:2402.08341,
2024.
[HOvB+22] Djurre Holtrop, Janneke K Oostrom, Ward R J van Breda, Antonis Koutsoumpis,
and Reinout E de Vries. Exploring the application of a text-to-personality technique in job interviews. European Journal of Work and Organizational Psychology,
31(6):799–816, 2022.
[JLF+23] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii,
Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in
natural language generation. ACM computing surveys, 55(12):1–38, 2023.
[Joh14] John A Johnson. Measuring thirty facets of the five factor model with a 120-item
public domain inventory: Development of the ipip-neo-120. Journal of research
in personality, 51:78–89, 2014.
[JXZ+23] Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and
Yixin Zhu. Evaluating and inducing personality in pre-trained language models.
Advances in Neural Information Processing Systems, 36:10622–10643, 2023
[JZC+24] Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. Personallm: Investigating the ability of large language models to express
personality traits. In Findings of the association for computational linguistics:
NAACL 2024, pages 3605–3627, 2024.
[KE25] JM Kruijssen and Nicholas Emmons. Deterministic ai agent personality expression through standard psychological diagnostics. arXiv preprint arXiv:2503.17085,
2025.
[KHJ+23] Hyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West, Ximing Lu, Youngjae Yu,
Pei Zhou, Ronan Bras, Malihe Alikhani, Gunhee Kim, et al. Soda: Million-scale
dialogue distillation with social commonsense contextualization. In Proceedings of
the 2023 Conference on Empirical Methods in Natural Language Processing, pages
12930–12949, 2023.
[LB11] Gary J Lewis and Timothy C Bates. From left to right: How the personality
system allows basic traits to influence politics via characteristic moral adaptations.
British journal of psychology, 102(3):546–558, 2011.
[LKSC07] Chang H Lee, Kyungil Kim, Young Seok Seo, and Cindy K Chung. The relations between personality and language use. The Journal of general psychology,
134(4):405–413, 2007.
[LLH+24] Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua,
Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long
contexts. Transactions of the association for computational linguistics, 12:157–
173, 2024.
[LLL+25] Wenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou, Mona Diab, and Maarten Sap.
Big5-chat: Shaping llm personalities through training on human-grounded data.
In Proceedings of the 63rd Annual Meeting of the Association for Computational
Linguistics (Volume 1: Long Papers), pages 20434–20471, 2025.
[LRL+24] Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar,
Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. The unlocking spell
on base llms: Rethinking alignment via in-context learning. In International
Conference on Learning Representations, volume 2024, pages 24907–24933, 2024.
[LWG00] Scott O Lilienfeld, James M Wood, and Howard N Garb. The scientific status
of projective techniques. Psychological science in the public interest, 1(2):27–66,
2000.
[M+62] Isabel Briggs Myers et al. The myers-briggs type indicator, volume 34. Consulting
Psychologists Press Palo Alto, CA, 1962.
[MCJ97] Robert R McCrae and Paul T Costa Jr. Personality trait structure as a human
universal. American psychologist, 52(5):509, 1997.
[MDL+23] Gr´egoire Mialon, Roberto Dess`ı, Maria Lomeli, Christoforos Nalmpantis, Ram
Pasunuru, Roberta Raileanu, Baptiste Rozi`ere, Timo Schick, Jane Dwivedi-Yu,
Asli Celikyilmaz, et al. Augmented language models: a survey. arXiv preprint
arXiv:2302.07842, 2023.
[MJ92] Robert R McCrae and Oliver P John. An introduction to the five-factor model
and its applications. Journal of personality, 60(2):175–215, 1992.
[MWX+26] Xingjun Ma, Yixu Wang, Hengyuan Xu, Yutao Wu, Yifan Ding, Yunhan Zhao,
Zilong Wang, Jiabin Hua, Ming Wen, Jianan Liu, et al. A safety report on gpt-5.2,
gemini 3 pro, qwen3-vl, grok 4.1 fast, nano banana pro, and seedream 4.5. arXiv
preprint arXiv:2601.10527, 2026.
NPH24] Lewis Newsham, Daniel Prince, and Ryan Hyland. Measuring the effect of induced
persona on agenda creation in language-based agents for cyber deception. In
Proceedings of the First International Conference on Natural Language Processing
and Artificial Intelligence for Cyber Security, pages 48–58, 2024.
[NZA+23] Allen Nie, Yuhui Zhang, Atharva Amdekar, Christopher J Piech, Tatsunori
Hashimoto, and Tobias Gerstenberg. Moca: Measuring human-language model
alignment on causal and moral judgment tasks. In Thirty-seventh Conference on
Neural Information Processing Systems, 2023.
[Oll26] Ollama. Ollama Documentation, 2026. Accessed: 2026-06-14.
[OWJ+22] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela
Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in
neural information processing systems, 35:27730–27744, 2022.
[Pit05] David J Pittenger. Cautionary comments regarding the myers-briggs type indicator. Consulting Psychology Journal: Practice and Research, 57(3):210, 2005.
[Pla75] R. L. Plackett. The analysis of permutations. Journal of the Royal Statistical
Society. Series C (Applied Statistics), 24(2):193–202, 1975.
[PMN03] James W Pennebaker, Matthias R Mehl, and Kate G Niederhoffer. Psychological aspects of natural language use: Our words, our selves. Annual review of
psychology, 54(1):547–577, 2003.
[POC+23] Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy
Liang, and Michael S Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface
software and technology, pages 1–22, 2023.
[PR22] F´abio Perez and Ian Ribeiro. Ignore previous prompt: Attack techniques for
language models. arXiv preprint arXiv:2211.09527, 2022.
[Rai24] Dilli Hang Rai. Artificial intelligence through time: A comprehensive historical
review. Tribhuvan University, Institute of Science and Technology. DOI, 10, 2024.
[Ror21] Hermann Rorschach. Psychodiagnostik. Bircher, Bern, 1921.
[RSR+20] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang,
Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits
of transfer learning with a unified text-to-text transformer. Journal of machine
learning research, 21(140):1–67, 2020.
[SAL+24] Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu.
Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
[Sar10] Riccardo Sartori. Face validity in personality tests: psychometric instruments and
projective techniques in comparison. Quality & quantity, 44(4):749–759, 2010.
[Sca20] Geoffrey Scarre. Utilitarianism. Routledge, 2020.
[SDL+23] Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? In International
conference on machine learning, pages 29971–30004. PMLR, 2023.
[SFRY24] Aleksandra Sorokovikova, Natalia Fedorova, Sharwin Rezagholi, and Ivan P
Yamshchikov. Llms simulate big five personality traits: Further evidence. arXiv
preprint arXiv:2402.01765, 2024.
[SGSC+23] Greg Serapio-Garc´ıa, Mustafa Safdari, Cl´ement Crepy, Luning Sun, Stephen Fitz,
Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matari´c. Personality traits in large language models. arXiv preprint arXiv:2307.00184, 2023.
[SHB16] Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation
of rare words with subword units. In Proceedings of the 54th annual meeting
of the association for computational linguistics (volume 1: long papers), pages
1715–1725, 2016.
[SKG+25] Hua Shen, Tiffany Knearem, Reshmi Ghosh, Yu-Ju Yang, Nicholas Clark,
Tanushree Mitra, and Yun Huang. Valuecompass: A framework for measuring
contextual value alignment between human and llms, 2025.
[SLDQ23] Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. Character-llm: A trainable agent for role-playing. In Proceedings of the 2023 Conference on Empirical
Methods in Natural Language Processing, pages 13153–13187, 2023.
[THMR+26] Tommaso Tosato, Saskia Helbling, Yorguin-Jose Mantilla-Ramos, Mahmood
Hegazy, Alberto Tosato, David John Lemay, Irina Rish, and Guillaume Dumas.
Persistent instability in llm’s personality measurements: Effects of scale, reasoning, and conversation history. In Proceedings of the AAAI Conference on Artificial
Intelligence, volume 40, pages 37961–37969, 2026.
[TM06] Antonio Terracciano and Robert R McCrae. Cross-cultural studies of personality
traits and their relevance to psychiatry. Epidemiology and Psychiatric Sciences,
15(3):176–184, 2006.
[VNG+26] Huy Vu, Huy Anh Nguyen, Adithya V Ganesan, Swanie Juhng, Oscar NE Kjell,
Joao Sedoc, Margaret L Kern, Ryan L Boyd, Lyle Ungar, H Andrew Schwartz,
et al. Psychadapter: adapting llms to reflect traits, personality, and mental health.
NPJ Artificial Intelligence, 2(1):26, 2026.
[VSP+17] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need.
Advances in neural information processing systems, 30, 2017.
[WFH+23] Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry
Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. A prompt
pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint
arXiv:2302.11382, 2023.
[WLN+10] James M Wood, Scott O Lilienfeld, M Teresa Nezworski, Howard N Garb,
Keli Holloway Allen, and Jessica L Wildermuth. Validity of rorschach inkblot
scores for discriminating psychopaths from nonpsychopaths in forensic populations: A meta-analysis. Psychological Assessment, 22(2):336, 2010.
[WP25] Peter West and Christopher Potts. Base models beat aligned models at randomness and creativity. arXiv preprint arXiv:2505.00047, 2025.
[WWS+22] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi,
Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning
in large language models. Advances in neural information processing systems,
35:24824–24837, 2022.
[Yar10] Tal Yarkoni. Personality in 100,000 words: A large-scale analysis of personality
and word use among bloggers. Journal of research in personality, 44(3):363–373,
2010
 

Download scriptie (1.79 MB)
Universiteit of Hogeschool
Universiteit Hasselt
Thesis jaar
2026
Promotor(en) en begeleiders
Kris Luyten, Gilles Eerlings