RCreddit.com·
暂不在当前实时榜单
AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR
The AndroidLife benchmark tested the Qwen3.8-27b AI model's ability to perform 60 real-world tasks on a OnePlus phone, achieving a 56.7% success rate. The model struggled with multi-application tasks and instances requiring user interaction, failing 43% of the tasks. During the test, the phone's chip reached a peak temperature of 98.2 C, and 69% of the battery was consumed. This benchmark aims to evaluate AI agents in a realistic mobile environment, tracking performance, thermals, and battery usage.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月18日 07:00 UTC
- 收录
- 2026年9月18日 07:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →