404 Media has published internal materials describing an OpenAI programme codenamed Project Lily, under which hundreds of external contractors read real conversations between people and ChatGPT.
What the reviewers do
Contractors summarise what the user was trying to achieve, then grade the model's reply on a scale from 1 to 7. They flag specific failure modes, including robotic phrasing and excessive agreement with the user. The reporting is based on reviewer instructions, Slack messages, the scoring rubric and the user submissions themselves.
Pay is above $50 an hour, routed through intermediary firms rather than OpenAI directly. Some of the conversations reviewed contained sensitive personal information.
The consent gap
One source who worked on the programme was asked whether users realise their chats are reviewed this way. The answer was no: people do not imagine a contractor somewhere is analysing the conversation.
After publication, OpenAI pointed to a help page stating that humans may review content to improve model performance. The relevant setting, which sends chats into training, is enabled by default on the Free, Plus and Pro plans. Turning it off is the user's responsibility, and requires knowing it exists.
Why it matters
Human review is not a scandal in itself. It is how these systems are made to sound less robotic, and every major assistant does some version of it. What the reporting exposes is the distance between a defensible practice and what users actually understand about it.
The practical consequence is narrow and concrete. If you paste anything into a general assistant that you would not hand to a stranger — client data, medical details, unreleased work — the default settings are not on your side. Check the training toggle, or use a plan and product where the data handling is contractual rather than optional.