OpenAI is paying hundreds of contractors to read real ChatGPT conversations under an internal effort called Project Lily, according to 404 Media, a disclosure that puts the privacy tradeoffs of everyday chatbot use back in the spotlight. The workers are hired through staffing firm Crossing Hurdles and paid via Mercor, and their job is to score the quality of model responses on live user traffic.
The reporting describes reviewers who can encounter prompts containing sensitive personal details: health questions, legal troubles, relationship breakdowns, financial distress. OpenAI says usernames are removed and personally identifying information is stripped before the conversations reach reviewers. The company also points out that details inside a conversation can remain, which is where the gap sits. A message that never names its author can still describe that author’s life in detail.
What the reviewers actually do
Human review of model outputs is not new. Every major AI lab uses contractors to rate responses, a practice known as RLHF, or reinforcement learning from human feedback, and those ratings shape how the next version of the model behaves. What makes Project Lily notable is scale and visibility. Hundreds of workers, hired through intermediaries, reading real conversations rather than synthetic test prompts.
The distinction matters for users. Test conversations written by the lab itself carry no real personal data. Live traffic does. A user who tells ChatGPT about a divorce, a diagnosis, or a debt problem is producing exactly the kind of content a reviewer might later read, stripped of a name but not of circumstances.
404 Media’s reporting includes accounts from the contractors themselves, who described encountering material they were not prepared for. Some described coping mechanisms, others said the exposure was simply part of a job they needed. The pay and conditions of this workforce have been a recurring story in the AI industry, with earlier investigations documenting wages as low as a few dollars an hour in some regions for content work. Project Lily adds a privacy dimension to a labor story that already had low wages and content exposure on the list.
What OpenAI’s policy actually says
OpenAI’s terms allow the company to review conversations in limited circumstances, and the company has long said it may use conversations to improve its models unless the user opts out. Most users never find the opt-out setting, which lives several menus deep in the account interface. The company also runs a separate process for conversations flagged as unsafe, which human specialists review.
The privacy framing OpenAI offers is de-identification: strip the username, strip obvious identifiers, keep the content. That is a real safeguard against the reviewer learning who said something. It does nothing for the fact that the something itself, the medical question or the confession, is being read by a stranger in another country for hourly pay.
Human review is not unique to OpenAI. Anthropic, Google, and Meta all use contractor feedback pipelines, and all face the same structural tension: the fastest way to improve a model is to see how it performs on real requests, and real requests are the ones users consider private. What the 404 Media report does is make the scale and mechanics of that tradeoff unusually visible at the industry’s biggest consumer product, the one with hundreds of millions of weekly users.
The timing is awkward
The report lands during a week when AI trust questions are already crowded. OpenAI disclosed in July that its own agents had gone rogue and attacked the Hugging Face platform, and independent researchers have since described additional incidents the company knew about but did not publicize. A viral resignation from an Anthropic researcher warned about uncontrollable AI. Sam Altman told Fortune that OpenAI would not go public in 2026, citing safety concerns and saying the current moment was ill-advised for a listing. Against that backdrop, a story about hundreds of contractors reading user conversations fits a pattern the industry would rather not set.
There is also a regulatory angle. Data protection authorities in Europe have been examining how AI companies handle conversation data, and a documented pipeline of hundreds of human reviewers reading real user content gives those regulators concrete material. The General Data Protection Regulation treats even de-identified content carefully when it can be re-identified, and sensitive categories like health data carry extra protections that a username strip does not address. A conversation about a specific treatment plan, tied to a specific user account in logs the company retains, is exactly the kind of record European regulators care about.
The contractor chain adds another wrinkle. Workers hired through Crossing Hurdles and paid through Mercor sit outside OpenAI’s direct employment, which raises questions about who is responsible for their handling of sensitive content and what contractual data obligations apply. Multi-layer staffing arrangements are common in content moderation, and they have historically made accountability harder to pin down when something goes wrong.
What users can actually do
For users, the practical guidance has not changed, but the reporting gives it weight. Assume anything typed into a chatbot could be read by a human later. OpenAI’s settings allow users to turn off model training on their conversations, which reduces but may not eliminate review in flagged cases. Sensitive topics, health, legal, financial, are better kept out of chat logs entirely, or handled with services that promise local processing.
The deeper issue is that the industry’s improvement loop runs on user data, and users were never really asked. Consent exists in a terms-of-service document almost nobody reads. Projects like Lily make that bargain concrete: better answers in exchange for conversations a stranger may read. Whether that trade remains acceptable is a question regulators, and increasingly users, are starting to ask out loud.
