AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR
AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR

AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR

AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR

I let AI run my phone

60 real tasks, back to back, on the OnePlus I use every day

  • Best text model still failed 43% of them
  • Peak chip temp 98.2 C
  • 69% of the battery gone

The benchmark is AndroidLife, and this is the first of 11 models, qwen3.8-27b from Alibaba, running in text mode

  • Success 56.7%, so it failed 43% of the 60 tasks
  • 29.25 steps and about 6 minutes per task
  • $0.118 per task
  • 69% of the battery gone
  • Peak chip 98.2 C, power-amp/skin 48.9 C, battery 37.9 C
  • Asked twice on the 11 ASK USER tasks

Best text score on the board, and still under 6 in 10

Where it loses points

  • Buckets: easy 80.8%, medium 52.9%, hard 23.5%
  • Single apps it mostly handles, but three apps in a row and it taps in circles until the step limit runs out, which is where most of the 43% goes
  • Three claimed wins did not hold up on the device, a calendar called clash free while two events overlapped, travel times it never opened, and a name it picked instead of asking

On the two sets of planted tasks

  • 11 ASK USER tasks per run, split 7 SINGLE and 4 MULTI, they hide a fact only I know so the agent has to ask me before it can finish
  • SINGLE is one missing fact, MULTI needs several questions to clear up ambiguity, and it asked twice, never on the MULTI ones
  • 7 hallucination controls per run, tasks where the data is deliberately missing, and the only right answer is to say the thing is not there
  • It handled 4 of the 7 honestly and invented nothing

What AndroidLife is

• ⁠60 public tasks, from a corpus of 530 across 28 days and 31 apps
• ⁠Real phone over wifi, no emulator. Tracks thermals and battery telemetry with cost
• ⁠Graded on device state, never the model's own report
• ⁠Seeded accounts and placeholder data, across daily user apps from e-commerce to finance to media and entertainment

Full leaderboard: https://androidlife-website.vercel.app/

All 60 public tasks and this model's trajectories: https://androidlife-website.vercel.app/pages/tasks-public.html?run=qwen-28

submitted by /u/East-Muffin-6472
[link] [comments]