404 Media reported on 14 September 2026 that OpenAI runs an internal programme, Project Lily, in which hundreds of contract workers read real ChatGPT conversations and rate the model’s replies. The reviewers do not see usernames. They do see the prompt, often the surrounding conversation, and sometimes a short memory summary that hints at what the account has been used for. Their job is not the public safety queue. It is quality work: score answers on a one-to-seven scale, cut flattery, and stop the model talking as if it were a person.

That last part is not a small editorial preference. OpenAI has spent a year trying to pull ChatGPT back from the over-agreeable tone that got it into legal and reputational trouble. Project Lily is one of the human layers doing that work. Reviewers are hired through firms such as Crossing Hurdles and paid through Mercor. One North America-based reviewer told 404 Media the rate was more than $50 an hour.

OpenAI has said for years that authorised people and vendors may look at content to improve the product. The FAQ language has been on the site since at least 2023. What the reporting makes concrete is the scale, the fact that the conversations are live product traffic rather than synthetic test prompts, and how easy it is to miss the control that governs it.

The control is called “Improve the model for everyone.” On Free, Plus and Pro accounts it is on by default. On Business, Enterprise and Edu accounts it is off by default. Turning it off only applies to new threads. Temporary Chat is a separate mode that OpenAI says is not used to train the model. OpenAI also runs an automated privacy filter before a conversation reaches a person. The company itself says the filter can miss things. 404 Media found chats in the review pile where the user had asked ChatGPT to keep the contents private.

That is the story. The interesting part for a small business is not the scandal tone that will travel on social media. It is the gap between how SMBs actually use ChatGPT and how OpenAI has split its privacy defaults.

Most small-business AI work still sits on a personal plan

A five-person agency, a solo consultant, a trades firm that finally started drafting quotes with ChatGPT: these teams rarely start on Business. They start on Plus, or on the free tier, because that is the account the owner already had on their phone. Client emails get pasted in. Call notes get pasted in. A draft proposal, a price list, a competitor page, a messy spreadsheet export. The tool is treated like a private notebook that happens to talk back.

It is not a private notebook. It is a consumer product with a default that feeds some of that notebook into a human review pipeline. The chats are anonymised. They are not imaginary. A reviewer can still see a company name inside a brief, a margin inside a quote, a medical or legal detail that survived the filter, or the shape of a sales pipeline described in plain language.

There is a second mismatch. Plenty of owners believe “I asked it not to share this” is a control. It is not. It is a sentence inside a conversation that may later be sampled. The actual control is the Data Controls toggle, and it does not rewind old threads.

OpenAI’s paid business products already treat this differently. Business and Enterprise default the improvement setting off. That split is revealing. The company is selling privacy as a plan feature at the same time as it uses personal-plan traffic to tune tone. For a lab that needs to make ChatGPT less sycophantic, personal chats are a rich dataset. For a small firm that put a client’s strategy memo into Plus because Business felt like overkill, the same design looks like the cheap seat subsidising the product with its paperwork.

The privacy filter is doing a job it cannot fully do

Business language is full of identifiers that are not a passport number. “Our Seville hotel group,” “the Shopify store that does €40k a month,” “the dentist in Teatinos who wants the new booking flow.” A filter can strip an email address and still leave a paragraph that only one company in a city could have written. Reviewers also see a memory summary in some cases. That summary is there to help them understand the user. It also reconstructs context the anonymisation was supposed to blur.

None of this requires assuming bad faith. The stated goal is better answers and less creepy warmth. Anthropic and Google use human review as well. The uncomfortable part is the packaging. “Improve the model for everyone” sounds like a communal toggle, the kind of thing you leave on because you are not a conspiracy person. It does not sound like “a contractor may read this thread.” OpenAI did not immediately point 404 Media to a clear in-product disclosure. It later pointed to site copy. Site copy is not the same as a setting screen that says what will happen.

For SMBs the practical tension is simple. ChatGPT became useful exactly because people dropped real work into it. Real work is full of other people’s information: a client’s revenue, an employee’s sick-leave note, a supplier’s pricing. A consumer default that samples that work for human rating sits at odds with how the tool is now used.

Two products are living in one interface

There is ChatGPT the assistant, which small teams use like a junior hire who never sleeps. There is ChatGPT the training loop, which needs messy, current conversations to keep the assistant from sliding back into flattery. Project Lily belongs to the second product. The first product is what SMBs think they bought.

That split also explains why the story landed now. ChatGPT is no longer a novelty window. It is where proposals start, where support macros get drafted, where a founder thinks out loud about a hire. The more the tool sits in the middle of the business, the less the old consumer default fits. Business plans already admit that. Personal plans have not caught up, and personal plans are still where a large share of SMB usage lives.

The reporting will age. The underlying design choice will not, until the default and the disclosure catch up with the way small companies actually work.