¿Autodiseño? Evaluaciones de Modelos de Lenguaje de Gran Escala para Asistentes de UX

Revisión y Directrices

Autores/as

DOI:

https://doi.org/10.29147/datjournal.v11i2.1034

Palabras clave:

LLM, Evaluación, Diseño, Investigación en UX, Redacción en UX

Resumen

Los modelos de lenguaje de gran escala (LLMs) están siendo rápidamente integrados en los flujos de trabajo de Design Systems, investigación en UX y redacción en UX, pero faltan métodos sistemáticos para evaluar su seguridad y utilidad. Basado en una revisión de alcance, este artículo mapea los principales desafíos de evaluación y propone tres marcos específicos para estas tareas. Cada marco conecta actividades de diseño con criterios de éxito medibles (p. ej., ≤ 3% de alucinación, < 2s de latencia) y sugiere conjuntos de datos para pruebas reproducibles. Una demostración con GPT-4o y Gemini 2.5 ilustra cómo las métricas identifican las fortalezas y debilidades de los modelos. Las directrices ofrecen a los equipos de diseño una hoja de ruta práctica para incorporar pruebas continuas y basadas en evidencia, promoviendo mayor rigor y responsabilidad en la práctica del diseño asistido por IA.

Descargas

Los datos de descargas todavía no están disponibles.

Biografía del autor/a

Weynner Kenneth Bezerra Santos, UFPE

PhD candidate in FiGitAL Design (UFPE, 2028); MBA in Project Management (USP, 2022); Master’s in Digital Artifact Design (UFPE, 2019); and a Bachelor’s degree in Design (UFPE, 2016). Works as Senior Research Analyst at PicPay’s Design Center of Excellence, helping the company to apply AI for user research and other design contexts, leading the ResearchOps squad.

Citas

BENDER, Emily M. et al. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In: ACM CONFERENCE ON FAIRNESS, ACCOUNTABILITY, AND TRANSPARENCY (FAccT), 2021. Virtual Event, Canada. Proceedings [...]. New York: ACM, 2021. p. 610–623. DOI: 10.1145/3442188.3445922.

BROWN, Tim. Design Thinking: uma metodologia poderosa para decretar o fim das velhas ideias. Rio de Janeiro: Alta Books, 2020.

COSKUN, M. N. K. Spotify’s latest layoff is an opportunity to reconsider how we look at UX research. Fast Company, 2023. Available at: https://www.fastcompany.com/90999717/spotify-latest-layoff-ux-research. Accessed on: Jun. 25, 2025.

CRUTH, M. Discover the Spotify model. Atlassian, 2022. Available at: https://www.atlassian.com/agile/agile-at-scale/spotify. Accessed on: Jun. 25, 2025.

DIETZ, Laura et al. LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2504.19076.

FESSENDEN, T. Design Systems 101. Nielsen Norman Group, 2021. Available at: https://www.nngroup.com/articles/design-systems-101/. Accessed on: Jun. 26, 2025.

GAO, Mingqi et al. LLM-based NLG Evaluation: Current Status and Challenges. 2024. Preprint. DOI: https://doi.org/10.48550/arXiv.2402.01383.

GOTHELF, Jeff; SEIDEN, Josh. Lean UX: projetando ótimos produtos com times ágeis. 3. ed. Rio de Janeiro: Novatec Editora, 2022.

GOYAL, Tanya; LI, Junyi Jessy; DURRETT, Greg. News Summarization and Evaluation. 2022. Preprint. DOI: https://doi.org/10.48550/arXiv.2209.12356.

HULLEY, Stephen B. et al. Delineando a pesquisa clínica. 4. ed. Porto Alegre: Artmed, 2015.

HUSAIN, Hamel. Your AI Product Needs Evals. Hamel.dev, 2024. Available at: https://hamel.dev/blog/posts/evals/#motivation. Accessed on: Jun. 11, 2025.

LIN, Stephanie; HILTON, Jacob; EVANS, Owain. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 60., 2022, Dublin. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2022. p. 3214–3252. DOI: https://doi.org/10.18653/v1/2022.acl-long.229.

LIANG, Percy et al. Holistic Evaluation of Language Models. 2022. Preprint. Available at: https://arxiv.org/abs/2211.09110. Accessed on: Jun. 25, 2025.

MIZRAHI, Moran et al. State of What Art? A Call for Multi-Prompt LLM Evaluation. Transactions of the Association for Computational Linguistics, Cambridge, MA, v. 12, p. 933–949, 2024. DOI: https://doi.org/10.1162/tacl_a_00681.

MORAN, Kate. Usability (User) Testing. Nielsen Norman Group, 2019. Available at: https://www.nngroup.com/articles/usability-testing-101/. Accessed on: Jul. 18, 2025.

______. Calculating ROI for Design Projects in 4 Steps. Nielsen Norman Group, 2020. Available at: https://www.nngroup.com/articles/calculating-roi-design-projects/. Accessed on: Aug. 26, 2025.

______. The Four Dimensions of Tone of Voice. Nielsen Norman Group, 2023. Available at: https://www.nngroup.com/articles/tone-of-voice-dimensions/. Accessed on: Jun. 26, 2025.

NG, Andrew. Machine Learning Engineering for Production (MLOps) Specialization: Course 1. Coursera; DeepLearning.AI, 2025. Online course. Available at: https://www.coursera.org/learn/introduction-to-machine-learning-in-production. Accessed on: Jun. 25, 2025.

NIELSEN, Jakob. Usability 101. Nielsen Norman Group, 2012. Available at: https://www.nngroup.com/articles/usability-101-introduction-to-usability/. Accessed on: Jul. 18, 2025.

NIJS, Diane. Lens: a big shift in science – seeing change and innovation as a matter of emergence. In: ______ (Ed.). Advanced Imagineering: Designing Innovation as Collective Creation. Cheltenham, UK: Edward Elgar Publishing, 2019. chap. 2, p. 27–55.

RIBEIRO, Marco Tulio et al. Beyond Accuracy: Behavioral Testing of NLP Models with Checklist. In: ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), 58., 2020. Proceedings [...]. Stroudsburg, PA: Association for Computational Linguistics, 2020. p. 4902–4912. DOI: https://doi.org/10.18653/v1/2020.acl-main.442.

RICHARDS, Julian; WESSEL, Merten. Bridging HCI and AI Research for the Evaluation of Conversational SE Assistants. In: INTERNATIONAL WORKSHOP ON BOTS IN SOFTWARE ENGINEERING (BOTSE), 6., 2025, Ottawa. Proceedings [...]. IEEE/ACM, 2025. (In press). Available at: https://arxiv.org/abs/2502.07956. Accessed on: Aug. 28, 2025.

SAURO, Jeff; LEWIS, James R. Quantifying the User Experience: Practical Statistics for User Research. 2. ed. Cambridge, MA: Morgan Kaufmann, 2016.

SCHULER, Douglas; NAMIOKA, Aki (Org.). Participatory Design: Principles and Practices. Hillsdale, NJ: Lawrence Erlbaum Associates, 1993.

SCHWARTZ, Reva et al. Reality Check: A New Evaluation Ecosystem is Necessary to Understand AI's Real World Effects. 2025. Preprint. DOI: https://doi.org/10.48550/arXiv.2505.18893.

SPONHEIM, Christina. What Is User Research? Nielsen Norman Group, 2024. 1 video (2 min 37 s). Available at: https://www.nngroup.com/videos/what-is-user-research/. Accessed on: Jun. 26, 2025.

TRICCO, Andrea C. et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine, Philadelphia, PA, v. 169, n. 7, p. 467–473, Oct. 2018. DOI: https://doi.org/10.7326/M18-0850.

WANG, Xuezhi et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. 2022. Preprint. Available at: https://doi.org/10.48550/arXiv.2203.11171. Accessed on: Aug. 28, 2025.

WOOD, Ian. Metadesign: Designing in the Anthropocene. London: Routledge, 2020.

ZHANG, Ziyin et al. Rethinking the A/B Test for Large Language Model-based Features: A Case Study from GitHub Copilot. In: ACM JOINT EUROPEAN SOFTWARE ENGINEERING CONFERENCE AND SYMPOSIUM ON THE FOUNDATIONS OF SOFTWARE ENGINEERING (ESEC/FSE), 31., 2023, San Francisco. Proceedings [...]. New York: ACM, 2023. p. 1000–1012. DOI: https://doi.org/10.1145/3611643.3616319.

Descargas

Publicado

2026-08-28

Cómo citar

Santos, W. K. B. (2026). ¿Autodiseño? Evaluaciones de Modelos de Lenguaje de Gran Escala para Asistentes de UX: Revisión y Directrices. DAT Journal, 11(2), 194–211. https://doi.org/10.29147/datjournal.v11i2.1034