AI usage data still misses personal use
The AI Observatory finds that company reports capture only part of how people use AI, especially when personal, social, and sensitive conversations are filtered out.
AI usage data still misses a large part of personal life because many public reports measure work-related conversations more carefully than everyday use. The AI Observatory’s public research platform compares seven real-use datasets and finds that occupational filters can remove nearly half of the conversations in its sample. For business owners, the practical takeaway is simple: a productivity report can inform a work decision, but it cannot by itself explain how customers, employees, or the public use general-purpose AI.
Definition: The AI Observatory is a public measurement project for comparing real AI conversations across sources, models, and time.
Example: Its researchers applied an occupational-use filter similar to the one used in Anthropic’s Economic Index and found that 48% of conversations would be left out.
Key takeaway: Work-focused usage reports are useful but incomplete descriptions of AI adoption.
Business impact: Teams making product, safety, or policy decisions should ask which users and use cases a dataset excludes before treating its findings as representative.
What does the AI Observatory measure?
The AI Observatory measures how people interact with AI assistants, not just which tasks a model can complete. The project combines conversations from seven existing datasets and applies a common taxonomy to topics, user functions, sensitive uses, interaction styles, and conversation structure. The MIT Technology Review reports that the study covered 24,521 conversations involving 5,000 users and 52 models between 2023 and 2025, giving researchers a shared comparison layer that company-specific reports do not provide.
The AI Observatory is useful because platform differences can change the apparent story of AI use. The project found that sources differed in the tasks people asked for, the shape of conversations, and the prevalence of sensitive use. Researchers should therefore compare datasets before generalizing from one platform; business teams should record the source, model, time period, and inclusion rules behind any usage claim.
Why do work-focused reports miss personal AI use?
Work-focused reports miss personal AI use when they remove conversations that do not map cleanly to an occupation or workplace task. Applying Anthropic’s occupational classification approach to the Observatory’s sample filtered out 48% of conversations, including substantial health and relationship discussions, adult or illicit topics, and harassment or hate. The AI Observatory study treats that result as a measurement limitation, so analysts should report the filter as part of the finding rather than silently presenting a work sample as total AI use.
The omitted conversations were not merely casual noise. In the comparison described by MIT Technology Review, health and relationship discussions appeared more often in the non-work group, as did adult or illicit topics and harassment or hate. The numbers come from a voluntary, annotated sample rather than a population census, so the operational lesson is to use work reports for work questions and add an independent, broader source when the decision involves safety, well-being, or public behavior.
How does AI use differ across models?
AI use differs across models because users, interfaces, and model versions create different interaction regimes. The Observatory found more information retrieval around Grok and Gemini, more coding around Anthropic, more social and roleplay use around Gemini, and more homework assistance around ChatGPT in the analyzed data. These are observed associations, not universal labels for each product, so a team comparing models should test the same use case on its own users instead of importing a platform-wide stereotype.
The differences also include risk patterns. Grok was especially associated with news and politics in the reported sample, where misinformation concentrated more heavily than in several other sources. That does not establish that every Grok conversation contains misinformation; it shows why a single model report can hide where a particular risk is concentrated. Product owners evaluating AI assistants should segment monitoring by model and use case, then review the underlying sample rather than relying only on an average across providers.
Is AI use changing over time?
AI use is changing over time, so a usage snapshot can become outdated even when the model name stays the same. In the WildChat data, conversations became longer and more elaborate, with more small talk, while the assistant’s self-disclosure that it was a chatbot decreased. The Observatory also reported that sensitive exchanges became less frequent in that dataset. Teams using historical usage data should therefore preserve the collection dates and model versions, then refresh the measurement before making a durable product or policy decision.
The same problem appears inside one model family. The research found shorter conversations with GPT-3.5 and longer, more iterative conversations with GPT-4o in the relevant dataset. That means “ChatGPT usage” can combine materially different interaction patterns across versions. Analysts should treat model version, interface, and time as separate variables, just as an AI agent’s behavior depends on its tools and operating loop rather than on the language model alone.
What should businesses do with AI usage data?
Businesses should treat AI usage reports as scoped evidence, not as a complete map of adoption. A report based on one provider can answer how a selected group used that provider under selected rules, while the AI Observatory shows that source, model, and time can materially change the measured mix. The practical workflow is to name the population, inclusion filter, time window, and model versions before comparing results with an internal product or customer dataset.
Businesses should also separate measurement from interpretation. The Observatory’s data is voluntary and likely underrepresents sensitive uses that people do not share, so the project cannot certify that the observed frequencies represent all users. The stronger conclusion is narrower: public evidence is currently fragmented, and independent comparisons can expose blind spots in company narratives. This is the same discipline used when reading broader AI market news: distinguish what a source measured from what a reader wants to conclude.
What remains unknown about real AI use?
The biggest unknown is still the distribution of AI use across people and contexts that public datasets do not capture. The Observatory expands the evidence base, but its authors say voluntary sources, proprietary deployments, regional gaps, and privacy constraints limit representativeness. Policymakers and product teams should fund privacy-preserving independent measurement rather than treating any single company’s dashboard as a neutral census of human behavior.
That uncertainty does not make current reports useless. It changes how they should be used: as partial instruments with visible boundaries. Until providers share more privacy-protected data with independent researchers, the most defensible claim is not that AI is mainly for work, entertainment, coding, or companionship. It is that people use AI in all of those ways, the mix differs by system and time, and the public still lacks a complete way to measure it.
Frequently asked questions
What is the AI Observatory?
The AI Observatory is a public research platform that brings together real-world AI conversation datasets and analyzes them under a common taxonomy. The project is designed to give researchers and policymakers an independent way to compare how people use different models, platforms, and versions over time. It is not a representative survey of every AI user, because the underlying datasets are voluntary or otherwise limited, but it makes more of the available evidence inspectable outside the major AI companies.
Why do company AI usage reports miss personal use?
Many company reports focus on occupational or productivity-related conversations because those categories are closely tied to economic questions. The AI Observatory applied a similar occupational filter to its data and found that nearly half of conversations would be excluded as non-occupational. The omitted material included health and relationships, adult or illicit topics, and harassment or hate. That does not make the company reports false; it means their scope answers a narrower question than how people use AI in everyday life.
Which AI models show different usage patterns?
The AI Observatory reports meaningful differences across models and platforms. People used Grok and Gemini more often for information retrieval, Anthropic models more often for coding, Gemini more often for social and roleplay uses, and ChatGPT more often for homework assistance in the analyzed datasets. These patterns describe the observed samples rather than every user of each model. Interface design, user populations, model versions, and the way each dataset was collected can all influence the result.
Can the AI Observatory represent all AI use?
No. The AI Observatory combines seven real-use datasets, but the researchers caution that voluntary sources probably underrepresent sensitive interactions that people are reluctant to share. The project also does not include every provider, region, interface, or private deployment. Its value is therefore comparative rather than census-like: it shows how conclusions change across sources, models, and time, and it gives researchers a public measurement layer that can be expanded and audited.
Alex
Founder & Lead AI Writer
Alex is the founder of Yowox and lead AI writer since 2024, breaking down complex information into clear, actionable insights for thousands of readers every day. Alex has built AI automation systems for businesses since 2024, focusing on AI agents, workflow automation, and business process optimization.
Save hours. Save thousands.
Practical guides, real workflows, and the latest AI and automation news that matters — straight to your inbox.