UMD research study ($150): can a node-level view of LLM output spread beat trace-by-trace debugging? Final recruitment round for agent builders
热度趋势
百分比基于当前可用热度信号,而非评论数或独立用户人数。
马里兰大学的一名博士生正在进行一项经IRB批准的研究,旨在探讨开发者如何调试和迭代多智能体系统。该研究旨在比较LLM输出传播的节点级视图与逐迹调试的效率。参与者需完成一个75分钟的录音Zoom会话,并在其工作流程中使用调试工具约一周,随后进行30分钟的后续访谈。成功完成整个研究的参与者将获得150美元的礼品卡。目前正在进行最后一轮招募,持续到下周。
Hey folks — PhD student at the University of Maryland here, studying how developers debug and iterate on multi-agent systems. We're in the last stretch of recruitment, with sessions running now through next week.
The question we're testing: when you tweak a prompt in an agent workflow, you usually judge it by eyeballing a run or two. Our research tool shows the distribution of outputs each node produces across runs — does that actually beat clicking through traces one at a time, or is it just one more dashboard? "It doesn't help" is a publishable answer.
Participating: a 75-min Zoom session on structured debugging tasks (recorded, think-aloud), about a week using the tool in your own workflow, and a 30-min follow-up interview. $150 gift card on completing the full study.
If you've built with LangGraph/LangChain (or agent workflows generally), the screener takes ~2 min: https://forms.gle/Zwqvgd1h8DUnFRfC8
IRB-approved academic research, not a product pitch. Questions welcome — or zxu169@umd.edu.