Autodesign? Avaliações de Modelos de Linguagem de Grande Porte para Assistentes de UX
Revisão e Diretrizes
DOI:
https://doi.org/10.29147/datjournal.v11i2.1034Palavras-chave:
LLM, Avaliação, Design, Pesquisa em UX, Escrita em UXResumo
Modelos de linguagem de grande porte (LLMs) estão sendo rapidamente integrados aos fluxos de trabalho de Design Systems, pesquisa em UX e UX writing, mas faltam métodos sistemáticos para avaliar sua segurança e utilidade. Com base em uma revisão de escopo, este artigo mapeia os principais desafios de avaliação e propõe três frameworks específicos para essas tarefas. Cada framework conecta atividades de design a critérios de sucesso mensuráveis (por exemplo, ≤ 3% de alucinação, < 2s de latência) e sugere conjuntos de dados para testes reprodutíveis. Uma demonstração com o GPT-4o e o Gemini 2.5 ilustra como as métricas identificam pontos fortes e fracos dos modelos. As diretrizes oferecem às equipes de design um roteiro prático para incorporar testes contínuos e baseados em evidências, promovendo maior rigor e responsabilidade na prática de design assistido por IA.
Downloads
Referências
BENDER, Emily M. et al. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: ACM CONFERENCE ON FAIRNESS, ACCOUNTABILITY, AND TRANSPARENCY (FAccT), 2021. Virtual Event, Canada. Proceedings [...]. New York: ACM, 2021. p. 610–623. DOI: 10.1145/3442188.3445922.
BROWN, Tim. Design Thinking: uma metodologia poderosa para decretar o fim das velhas ideias. Rio de Janeiro: Alta Books, 2020.
COSKUN, M. N. K. Spotify’s latest layoff is an opportunity to reconsider how we look at UX research. Fast Company, 2023. Available at: https://www.fastcompany.com/90999717/spotify-latest-layoff-ux-research. Accessed on: Jun. 25, 2025.
CRUTH, M. Discover the Spotify model. Atlassian, 2022. Available at: https://www.atlassian.com/agile/agile-at-scale/spotify. Accessed on: Jun. 25, 2025.
DIETZ, Laura et al. LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2504.19076.
FESSENDEN, T. Design Systems 101. Nielsen Norman Group, 2021. Available at: https://www.nngroup.com/articles/design-systems-101/. Accessed on: Jun. 26, 2025.
GAO, Mingqi et al. LLM-based NLG Evaluation: Current Status and Challenges. 2024. Preprint. DOI: https://doi.org/10.48550/arXiv.2402.01383.
GOTHELF, Jeff; SEIDEN, Josh. Lean UX: projetando ótimos produtos com times ágeis. 3. ed. Rio de Janeiro: Novatec Editora, 2022.
GOYAL, Tanya; LI, Junyi Jessy; DURRETT, Greg. News Summarization and Evaluation. 2022. Preprint. DOI: https://doi.org/10.48550/arXiv.2209.12356.
HULLEY, Stephen B. et al. Delineando a pesquisa clínica. 4. ed. Porto Alegre: Artmed, 2015.
HUSAIN, Hamel. Your AI Product Needs Evals. Hamel.dev, 2024. Available at: https://hamel.dev/blog/posts/evals/#motivation. Accessed on: Jun. 11, 2025.
LIN, Stephanie; HILTON, Jacob; EVANS, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 60., 2022, Dublin. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2022. p. 3214–3252. DOI: https://doi.org/10.18653/v1/2022.acl-long.229.
LIANG, Percy et al. Holistic Evaluation of Language Models. 2022. Preprint. Available at: https://arxiv.org/abs/2211.09110. Accessed on: Jun. 25, 2025.
MIZRAHI, Moran et al. State of What Art? A Call for Multi-Prompt LLM Evaluation. Transactions of the Association for Computational Linguistics, Cambridge, MA, v. 12, p. 933–949, 2024. DOI: https://doi.org/10.1162/tacl_a_00681.
MORAN, Kate. Usability (User) Testing. Nielsen Norman Group, 2019. Available at: https://www.nngroup.com/articles/usability-testing-101/. Accessed on: Jul. 18, 2025.
______. Calculating ROI for Design Projects in 4 Steps. Nielsen Norman Group, 2020. Available at: https://www.nngroup.com/articles/calculating-roi-design-projects/. Accessed on: Aug. 26, 2025.
______. The Four Dimensions of Tone of Voice. Nielsen Norman Group, 2023. Available at: https://www.nngroup.com/articles/tone-of-voice-dimensions/. Accessed on: Jun. 26, 2025.
NG, Andrew. Machine Learning Engineering for Production (MLOps) Specialization: Course 1. Coursera; DeepLearning.AI, 2025. Online course. Available at: https://www.coursera.org/learn/introduction-to-machine-learning-in-production. Accessed on: Jun. 25, 2025.
NIELSEN, Jakob. Usability 101. Nielsen Norman Group, 2012. Available at: https://www.nngroup.com/articles/usability-101-introduction-to-usability/. Accessed on: Jul. 18, 2025.
NIJS, Diane. Lens: a big shift in science – seeing change and innovation as a matter of emergence. In: ______ (Ed.). Advanced Imagineering: Designing Innovation as Collective Creation. Cheltenham, UK: Edward Elgar Publishing, 2019. chap. 2, p. 27–55.
RIBEIRO, Marco Tulio et al. Beyond Accuracy: Behavioral Testing of NLP Models with Checklist. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 58., 2020. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2020. p. 4902–4912. DOI: https://doi.org/10.18653/v1/2020.acl-main.442.
RICHARDS, Julian; WESSEL, Merten. Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants. In: INTERNATIONAL WORKSHOP ON BOTS IN SOFTWARE ENGINEERING (BOTSE), 6., 2025, Ottawa. Proceedings [...]. IEEE/ACM, 2025. (In press). Available at: https://arxiv.org/abs/2502.07956. Accessed on: Aug. 28, 2025.
SAURO, Jeff; LEWIS, James R. Quantifying the User Experience: Practical Statistics for User Research. 2. ed. Cambridge, MA: Morgan Kaufmann, 2016.
SCHULER, Douglas; NAMIOKA, Aki (Org.). Participatory Design: Principles and Practices. Hillsdale, NJ: Lawrence Erlbaum Associates, 1993.
SCHWARTZ, Reva et al. Reality Check: A New Evaluation Ecosystem is Necessary to Understand AI's Real World Effects. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2505.18893.
SPONHEIM, Christina. What Is User Research? Nielsen Norman Group, 2024. 1 video (2 min 37 s). Available at: https://www.nngroup.com/videos/what-is-user-research/. Accessed on: Jun. 26, 2025.
TRICCO, Andrea C. et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, Philadelphia, PA, v. 169, n. 7, p. 467–473, Oct. 2018. DOI: https://doi.org/10.7326/M18-0850.
WANG, Xuezhi et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. Preprint. Available at: https://doi.org/10.48550/arXiv.2203.11171. Accessed on: Aug. 28, 2025.
WOOD, Ian. Metadesign: Designing in the Anthropocene. London: Routledge, 2020.
ZHANG, Ziyin et al. Rethinking the A/B Test for Large Language Model-based Features: A Case Study from GitHub Copilot. In: ACM JOINT EUROPEAN SOFTWARE ENGINEERING CONFERENCE AND SYMPOSIUM ON THE FOUNDATIONS OF SOFTWARE ENGINEERING (ESEC/FSE), 31., 2023, San Francisco. Proceedings [...]. New York: ACM, 2023. p. 1000–1012. DOI: https://doi.org/10.1145/3611643.3616319.
Downloads
Publicado
Como Citar
Edição
Seção
Licença
Copyright (c) 2026 DAT Journal

Este trabalho está licenciado sob uma licença Creative Commons Attribution 4.0 International License.





















