Autodesign? Avaliações de Modelos de Linguagem de Grande Porte para Assistentes de UX

Revisão e Diretrizes

Autores

DOI:

https://doi.org/10.29147/datjournal.v11i2.1034

Palavras-chave:

LLM, Avaliação, Design, Pesquisa em UX, Escrita em UX

Resumo

Modelos de linguagem de grande porte (LLMs) estão sendo rapidamente integrados aos fluxos de trabalho de Design Systems, pesquisa em UX e UX writing, mas faltam métodos sistemáticos para avaliar sua segurança e utilidade. Com base em uma revisão de escopo, este artigo mapeia os principais desafios de avaliação e propõe três frameworks específicos para essas tarefas. Cada framework conecta atividades de design a critérios de sucesso mensuráveis (por exemplo, ≤ 3% de alucinação, < 2s de latência) e sugere conjuntos de dados para testes reprodutíveis. Uma demonstração com o GPT-4o e o Gemini 2.5 ilustra como as métricas identificam pontos fortes e fracos dos modelos. As diretrizes oferecem às equipes de design um roteiro prático para incorporar testes contínuos e baseados em evidências, promovendo maior rigor e responsabilidade na prática de design assistido por IA.

Downloads

Não há dados estatísticos.

Biografia do Autor

Weynner Kenneth Bezerra Santos, UFPE

PhD candidate in FiGitAL Design (UFPE, 2028); MBA in Project Management (USP, 2022); Master’s in Digital Artifact Design (UFPE, 2019); and a Bachelor’s degree in Design (UFPE, 2016). Works as Senior Research Analyst at PicPay’s Design Center of Excellence, helping the company to apply AI for user research and other design contexts, leading the ResearchOps squad.

Referências

BENDER, Emily M. et al. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: ACM CONFERENCE ON FAIRNESS, ACCOUNTABILITY, AND TRANSPARENCY (FAccT), 2021. Virtual Event, Canada. Proceedings [...]. New York: ACM, 2021. p. 610–623. DOI: 10.1145/3442188.3445922.

BROWN, Tim. Design Thinking: uma metodologia poderosa para decretar o fim das velhas ideias. Rio de Janeiro: Alta Books, 2020.

COSKUN, M. N. K. Spotify’s latest layoff is an opportunity to reconsider how we look at UX research. Fast Company, 2023. Available at: https://www.fastcompany.com/90999717/spotify-latest-layoff-ux-research. Accessed on: Jun. 25, 2025.

CRUTH, M. Discover the Spotify model. Atlassian, 2022. Available at: https://www.atlassian.com/agile/agile-at-scale/spotify. Accessed on: Jun. 25, 2025.

DIETZ, Laura et al. LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2504.19076.

FESSENDEN, T. Design Systems 101. Nielsen Norman Group, 2021. Available at: https://www.nngroup.com/articles/design-systems-101/. Accessed on: Jun. 26, 2025.

GAO, Mingqi et al. LLM-based NLG Evaluation: Current Status and Challenges. 2024. Preprint. DOI: https://doi.org/10.48550/arXiv.2402.01383.

GOTHELF, Jeff; SEIDEN, Josh. Lean UX: projetando ótimos produtos com times ágeis. 3. ed. Rio de Janeiro: Novatec Editora, 2022.

GOYAL, Tanya; LI, Junyi Jessy; DURRETT, Greg. News Summarization and Evaluation. 2022. Preprint. DOI: https://doi.org/10.48550/arXiv.2209.12356.

HULLEY, Stephen B. et al. Delineando a pesquisa clínica. 4. ed. Porto Alegre: Artmed, 2015.

HUSAIN, Hamel. Your AI Product Needs Evals. Hamel.dev, 2024. Available at: https://hamel.dev/blog/posts/evals/#motivation. Accessed on: Jun. 11, 2025.

LIN, Stephanie; HILTON, Jacob; EVANS, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 60., 2022, Dublin. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2022. p. 3214–3252. DOI: https://doi.org/10.18653/v1/2022.acl-long.229.

LIANG, Percy et al. Holistic Evaluation of Language Models. 2022. Preprint. Available at: https://arxiv.org/abs/2211.09110. Accessed on: Jun. 25, 2025.

MIZRAHI, Moran et al. State of What Art? A Call for Multi-Prompt LLM Evaluation. Transactions of the Association for Computational Linguistics, Cambridge, MA, v. 12, p. 933–949, 2024. DOI: https://doi.org/10.1162/tacl_a_00681.

MORAN, Kate. Usability (User) Testing. Nielsen Norman Group, 2019. Available at: https://www.nngroup.com/articles/usability-testing-101/. Accessed on: Jul. 18, 2025.

______. Calculating ROI for Design Projects in 4 Steps. Nielsen Norman Group, 2020. Available at: https://www.nngroup.com/articles/calculating-roi-design-projects/. Accessed on: Aug. 26, 2025.

______. The Four Dimensions of Tone of Voice. Nielsen Norman Group, 2023. Available at: https://www.nngroup.com/articles/tone-of-voice-dimensions/. Accessed on: Jun. 26, 2025.

NG, Andrew. Machine Learning Engineering for Production (MLOps) Specialization: Course 1. Coursera; DeepLearning.AI, 2025. Online course. Available at: https://www.coursera.org/learn/introduction-to-machine-learning-in-production. Accessed on: Jun. 25, 2025.

NIELSEN, Jakob. Usability 101. Nielsen Norman Group, 2012. Available at: https://www.nngroup.com/articles/usability-101-introduction-to-usability/. Accessed on: Jul. 18, 2025.

NIJS, Diane. Lens: a big shift in science – seeing change and innovation as a matter of emergence. In: ______ (Ed.). Advanced Imagineering: Designing Innovation as Collective Creation. Cheltenham, UK: Edward Elgar Publishing, 2019. chap. 2, p. 27–55.

RIBEIRO, Marco Tulio et al. Beyond Accuracy: Behavioral Testing of NLP Models with Checklist. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 58., 2020. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2020. p. 4902–4912. DOI: https://doi.org/10.18653/v1/2020.acl-main.442.

RICHARDS, Julian; WESSEL, Merten. Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants. In: INTERNATIONAL WORKSHOP ON BOTS IN SOFTWARE ENGINEERING (BOTSE), 6., 2025, Ottawa. Proceedings [...]. IEEE/ACM, 2025. (In press). Available at: https://arxiv.org/abs/2502.07956. Accessed on: Aug. 28, 2025.

SAURO, Jeff; LEWIS, James R. Quantifying the User Experience: Practical Statistics for User Research. 2. ed. Cambridge, MA: Morgan Kaufmann, 2016.

SCHULER, Douglas; NAMIOKA, Aki (Org.). Participatory Design: Principles and Practices. Hillsdale, NJ: Lawrence Erlbaum Associates, 1993.

SCHWARTZ, Reva et al. Reality Check: A New Evaluation Ecosystem is Necessary to Understand AI's Real World Effects. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2505.18893.

SPONHEIM, Christina. What Is User Research? Nielsen Norman Group, 2024. 1 video (2 min 37 s). Available at: https://www.nngroup.com/videos/what-is-user-research/. Accessed on: Jun. 26, 2025.

TRICCO, Andrea C. et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, Philadelphia, PA, v. 169, n. 7, p. 467–473, Oct. 2018. DOI: https://doi.org/10.7326/M18-0850.

WANG, Xuezhi et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. Preprint. Available at: https://doi.org/10.48550/arXiv.2203.11171. Accessed on: Aug. 28, 2025.

WOOD, Ian. Metadesign: Designing in the Anthropocene. London: Routledge, 2020.

ZHANG, Ziyin et al. Rethinking the A/B Test for Large Language Model-based Features: A Case Study from GitHub Copilot. In: ACM JOINT EUROPEAN SOFTWARE ENGINEERING CONFERENCE AND SYMPOSIUM ON THE FOUNDATIONS OF SOFTWARE ENGINEERING (ESEC/FSE), 31., 2023, San Francisco. Proceedings [...]. New York: ACM, 2023. p. 1000–1012. DOI: https://doi.org/10.1145/3611643.3616319.

Downloads

Publicado

2026-08-28

Como Citar

Santos, W. K. B. (2026). Autodesign? Avaliações de Modelos de Linguagem de Grande Porte para Assistentes de UX: Revisão e Diretrizes. DAT Journal, 11(2), 194–211. https://doi.org/10.29147/datjournal.v11i2.1034