How should real-world AI-tool proficiency be measured without turning usage into a fake expertise score?
How should real-world AI-tool proficiency be measured without turning usage into a fake expertise score?

How should real-world AI-tool proficiency be measured without turning usage into a fake expertise score?

I’m exploring a measurement problem rather than proposing that token count equals skill.

I built a local-first technical alpha that records Claude Code and Codex activity, produces a signed privacy-sanitized snapshot, and separates activity telemetry from self-submitted identity, connected work, and outcomes. Prompts, responses, code, local paths, and credentials are excluded from the public payload.

The long-term question is whether a portable AI-work record could help researchers recruit genuine power users and help companies find people with sustained, demonstrable AI-tool experience.

Example implementation: https://ledger.imagineqira.com/#/u/bryan

Methodology and setup: https://ledger.imagineqira.com/#/join

Source: https://github.com/TheArtOfSound/TOKENS

Which measures would be defensible: active days, task completion, accepted changes, evaluations, independently confirmed outcomes, or something else?

submitted by /u/OGMYT
[link] [comments]