I let AI run my phone
60 real tasks, back to back, on the OnePlus I use every day
- Best text model still failed 43% of them
- Peak chip temp 98.2 C
- 69% of the battery gone
The benchmark is AndroidLife, and this is the first of 11 models, qwen3.8-27b from Alibaba, running in text mode
- Success 56.7%, so it failed 43% of the 60 tasks
- 29.25 steps and about 6 minutes per task
- $0.118 per task
- 69% of the battery gone
- Peak chip 98.2 C, power-amp/skin 48.9 C, battery 37.9 C
- Asked twice on the 11 ASK USER tasks
Best text score on the board, and still under 6 in 10
Where it loses points
- Buckets: easy 80.8%, medium 52.9%, hard 23.5%
- Single apps it mostly handles, but three apps in a row and it taps in circles until the step limit runs out, which is where most of the 43% goes
- Three claimed wins did not hold up on the device, a calendar called clash free while two events overlapped, travel times it never opened, and a name it picked instead of asking
On the two sets of planted tasks
- 11 ASK USER tasks per run, split 7 SINGLE and 4 MULTI, they hide a fact only I know so the agent has to ask me before it can finish
- SINGLE is one missing fact, MULTI needs several questions to clear up ambiguity, and it asked twice, never on the MULTI ones
- 7 hallucination controls per run, tasks where the data is deliberately missing, and the only right answer is to say the thing is not there
- It handled 4 of the 7 honestly and invented nothing
What AndroidLife is
• 60 public tasks, from a corpus of 530 across 28 days and 31 apps
• Real phone over wifi, no emulator. Tracks thermals and battery telemetry with cost
• Graded on device state, never the model's own report
• Seeded accounts and placeholder data, across daily user apps from e-commerce to finance to media and entertainment
Full leaderboard: https://androidlife-website.vercel.app/
All 60 public tasks and this model's trajectories: https://androidlife-website.vercel.app/pages/tasks-public.html?run=qwen-28
submitted by