Anthropic opens 250,000 privacy-preserved Claude conversations to researchers
Anthropic will give three outside research groups access to aggregated Claude and Claude Code conversations under a privacy-preserving program. The pilot covers 250,000 conversations from April and May 2026.

TL;DR
- Three outside groups designed studies over roughly 250,000 Claude.ai and Claude Code conversations from April and May 2026; AnthropicAI's announcement says Anthropic collected the data and the groups analyzed the resulting outputs independently.
- Stanford's SALT Lab found that more than half of the sampled conversations involved consequential work, a finding AnthropicAI's study thread says was especially common for professional legal and financial guidance.
- Researchers never received chat transcripts: Anthropic's Insights system exposed aggregate categories and percentage shares, a form of platform transparency described by jackclarkSF.
- Oxford's study of user affect and METR's estimate of coding-agent productivity remain preliminary, while AnthropicAI's update says the company is weighing whether the privacy and review process can scale.
Anthropic's full pilot report says the partners may publish findings inconvenient for the company. The earlier Clio technical explainer describes the underlying system as an anonymized, aggregate view of usage patterns. The expression-of-interest form gives prospective collaborators a 5:00 PM PT, September 14 deadline.
The pilot
Anthropic says the sample mixed Claude.ai and Claude Code conversations. The three groups chose their own questions, while Anthropic ran collection through Insights, formerly Clio, and released the projects' aggregate data.
The partners were:
- Stanford University's Social and Language Technologies Lab, or SALT
- Oxford University's Human Information Processing Lab, or HIP
- METR, the frontier-model evaluation nonprofit
Anthropic limited its contractual review to privacy, policy-evasion information, company confidential material, and research accuracy; its pilot report says the teams remain free to publish inconvenient results. In jackclarkSF's explanation, Anthropic's cofounder framed the effort as opening platform telemetry to measurement outside the lab.
Facets, not transcripts
The most consequential detail is the analysis boundary: researchers can inspect distributions, but not individual classifications or conversations.
The workflow is compact:
- A researcher writes a question, such as what sort of guidance a person seeks.
- Claude answers it across every conversation in the study.
- The answers are grouped into categories, and researchers receive the categories and their shares of the sample.
Anthropic's method description says wording can materially change the categories Claude produces. Partners tested questions against the public WildChat dataset, but its casual, creative traffic did not always transfer cleanly to Claude usage. The privacy wall prevents a researcher from checking a questionable category against the underlying chats. jackclarkSF's thread describes the broader goal as a privacy-preserving form of platform transparency.
Stanford's consequential work
SALT studied the work people bring to Claude, the control they retain, and points of collaboration failure. Its findings complicate the familiar picture of AI as a home for low-accountability drafting.
- More than half of conversations delegated consequential tasks, meaning work that affects other people or is hard to reverse, according to the pilot report.
- Professional guidance, particularly legal and financial questions, had the strongest concentration of consequential use in the SALT results.
- In nearly three-quarters of conversations, people set the direction and adapted Claude's output rather than using it verbatim, the report says.
- Friction often improved results by making people clarify their intent, refine outputs, or stay engaged with the problem, per the researchers' summary.
Oxford and METR
HIP's early analysis paired conversational behavior with apparent user states. Warm Claude responses co-occurred with more positive users, refusals and disagreement with pushback, eccentricity with intellectual engagement, and helpfulness with satisfaction, according to the pilot report. The lab also found patterns of absorption, frustration, and enjoyment that resembled a separate study of ordinary web browsing.
METR is using Claude Code conversations to estimate productivity across model generations. Its preliminary method compares Claude's estimate of how long a task would take without AI against observed task time with different Claude models; the report says newer models appeared to create significant speedups and that Claude's time estimates correlated with completion times from an earlier developer study. Neither HIP nor METR has published its full writeup yet.
Voluntary access and removed clusters
Anthropic says the pilot was slow and resource-intensive because every external output passed privacy and legal review. The review process also changed the released record:
- A badly phrased question can yield misleading categories that cannot be audited against raw conversations.
- Anthropic removed or altered clusters that revealed how users bypassed safeguards, while retaining most categories describing attempted policy violations.
- Less than 5% of categories and conversations in each study were affected by those removals, according to the pilot appendix summary.
RazRazcle called out a governance limit in the arrangement: participation and access depend on a voluntary relationship with the lab, rather than recurring independent audit rights.
The next round
Anthropic is soliciting proposals for a possible second wave. The expression-of-interest form asks researchers to describe questions Insights can answer that are otherwise inaccessible, the societal value of the work, and the tool's limits; submissions close at 5:00 PM PT on September 14, 2026.
The company says it is still determining both the kinds of study its process can support and how many projects it can run concurrently.