Autodesign? Large Language Model Evaluations for UX Assistants
Review and Guidelines
DOI:
https://doi.org/10.29147/datjournal.v11i2.1034Keywords:
LLM, Evaluation, Design, UX Research, UX WritingAbstract
Large language models (LLMs) are being rapidly integrated into design workflows for Design Systems, UX research, and UX writing, but systematic methods to evaluate their safety and utility are lacking. Based on a scoping review, this paper maps key evaluation challenges and proposes three frameworks specific to these tasks. Each framework connects design activities to measurable success criteria (e.g., ≤ 3% hallucination, < 2s latency) and suggests datasets for reproducible testing. A demonstration with GPT-4o and Gemini 2.5 illustrates how the metrics identify the models' strengths and weaknesses. The guidelines offer design teams a practical roadmap to incorporate continuous, evidence-based testing, promoting greater rigor and accountability in AI-assisted design practice.
Downloads
References
BENDER, Emily M. et al. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: ACM CONFERENCE ON FAIRNESS, ACCOUNTABILITY, AND TRANSPARENCY (FAccT), 2021. Virtual Event, Canada. Proceedings [...]. New York: ACM, 2021. p. 610–623. DOI: 10.1145/3442188.3445922.
BROWN, Tim. Design Thinking: uma metodologia poderosa para decretar o fim das velhas ideias. Rio de Janeiro: Alta Books, 2020.
COSKUN, M. N. K. Spotify’s latest layoff is an opportunity to reconsider how we look at UX research. Fast Company, 2023. Available at: https://www.fastcompany.com/90999717/spotify-latest-layoff-ux-research. Accessed on: Jun. 25, 2025.
CRUTH, M. Discover the Spotify model. Atlassian, 2022. Available at: https://www.atlassian.com/agile/agile-at-scale/spotify. Accessed on: Jun. 25, 2025.
DIETZ, Laura et al. LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2504.19076.
FESSENDEN, T. Design Systems 101. Nielsen Norman Group, 2021. Available at: https://www.nngroup.com/articles/design-systems-101/. Accessed on: Jun. 26, 2025.
GAO, Mingqi et al. LLM-based NLG Evaluation: Current Status and Challenges. 2024. Preprint. DOI: https://doi.org/10.48550/arXiv.2402.01383.
GOTHELF, Jeff; SEIDEN, Josh. Lean UX: projetando ótimos produtos com times ágeis. 3. ed. Rio de Janeiro: Novatec Editora, 2022.
GOYAL, Tanya; LI, Junyi Jessy; DURRETT, Greg. News Summarization and Evaluation. 2022. Preprint. DOI: https://doi.org/10.48550/arXiv.2209.12356.
HULLEY, Stephen B. et al. Delineando a pesquisa clínica. 4. ed. Porto Alegre: Artmed, 2015.
HUSAIN, Hamel. Your AI Product Needs Evals. Hamel.dev, 2024. Available at: https://hamel.dev/blog/posts/evals/#motivation. Accessed on: Jun. 11, 2025.
LIN, Stephanie; HILTON, Jacob; EVANS, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 60., 2022, Dublin. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2022. p. 3214–3252. DOI: https://doi.org/10.18653/v1/2022.acl-long.229.
LIANG, Percy et al. Holistic Evaluation of Language Models. 2022. Preprint. Available at: https://arxiv.org/abs/2211.09110. Accessed on: Jun. 25, 2025.
MIZRAHI, Moran et al. State of What Art? A Call for Multi-Prompt LLM Evaluation. Transactions of the Association for Computational Linguistics, Cambridge, MA, v. 12, p. 933–949, 2024. DOI: https://doi.org/10.1162/tacl_a_00681.
MORAN, Kate. Usability (User) Testing. Nielsen Norman Group, 2019. Available at: https://www.nngroup.com/articles/usability-testing-101/. Accessed on: Jul. 18, 2025.
______. Calculating ROI for Design Projects in 4 Steps. Nielsen Norman Group, 2020. Available at: https://www.nngroup.com/articles/calculating-roi-design-projects/. Accessed on: Aug. 26, 2025.
______. The Four Dimensions of Tone of Voice. Nielsen Norman Group, 2023. Available at: https://www.nngroup.com/articles/tone-of-voice-dimensions/. Accessed on: Jun. 26, 2025.
NG, Andrew. Machine Learning Engineering for Production (MLOps) Specialization: Course 1. Coursera; DeepLearning.AI, 2025. Online course. Available at: https://www.coursera.org/learn/introduction-to-machine-learning-in-production. Accessed on: Jun. 25, 2025.
NIELSEN, Jakob. Usability 101. Nielsen Norman Group, 2012. Available at: https://www.nngroup.com/articles/usability-101-introduction-to-usability/. Accessed on: Jul. 18, 2025.
NIJS, Diane. Lens: a big shift in science – seeing change and innovation as a matter of emergence. In: ______ (Ed.). Advanced Imagineering: Designing Innovation as Collective Creation. Cheltenham, UK: Edward Elgar Publishing, 2019. chap. 2, p. 27–55.
RIBEIRO, Marco Tulio et al. Beyond Accuracy: Behavioral Testing of NLP Models with Checklist. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 58., 2020. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2020. p. 4902–4912. DOI: https://doi.org/10.18653/v1/2020.acl-main.442.
RICHARDS, Julian; WESSEL, Merten. Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants. In: INTERNATIONAL WORKSHOP ON BOTS IN SOFTWARE ENGINEERING (BOTSE), 6., 2025, Ottawa. Proceedings [...]. IEEE/ACM, 2025. (In press). Available at: https://arxiv.org/abs/2502.07956. Accessed on: Aug. 28, 2025.
SAURO, Jeff; LEWIS, James R. Quantifying the User Experience: Practical Statistics for User Research. 2. ed. Cambridge, MA: Morgan Kaufmann, 2016.
SCHULER, Douglas; NAMIOKA, Aki (Org.). Participatory Design: Principles and Practices. Hillsdale, NJ: Lawrence Erlbaum Associates, 1993.
SCHWARTZ, Reva et al. Reality Check: A New Evaluation Ecosystem is Necessary to Understand AI's Real World Effects. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2505.18893.
SPONHEIM, Christina. What Is User Research? Nielsen Norman Group, 2024. 1 video (2 min 37 s). Available at: https://www.nngroup.com/videos/what-is-user-research/. Accessed on: Jun. 26, 2025.
TRICCO, Andrea C. et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, Philadelphia, PA, v. 169, n. 7, p. 467–473, Oct. 2018. DOI: https://doi.org/10.7326/M18-0850.
WANG, Xuezhi et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. Preprint. Available at: https://doi.org/10.48550/arXiv.2203.11171. Accessed on: Aug. 28, 2025.
WOOD, Ian. Metadesign: Designing in the Anthropocene. London: Routledge, 2020.
ZHANG, Ziyin et al. Rethinking the A/B Test for Large Language Model-based Features: A Case Study from GitHub Copilot. In: ACM JOINT EUROPEAN SOFTWARE ENGINEERING CONFERENCE AND SYMPOSIUM ON THE FOUNDATIONS OF SOFTWARE ENGINEERING (ESEC/FSE), 31., 2023, San Francisco. Proceedings [...]. New York: ACM, 2023. p. 1000–1012. DOI: https://doi.org/10.1145/3611643.3616319.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 DAT Journal

This work is licensed under a Creative Commons Attribution 4.0 International License.






















