An investigation has revealed that OpenAI is using human contractors under 'Project Lily' to review real user prompts and AI responses to train its models.
- OpenAI uses human contractors via 'Project Lily' to evaluate AI responses.
- Reviewers may see sensitive personal data within user prompts.
- Privacy filters used by OpenAI are not 100% foolproof.
A startling investigation by 404 Media has uncovered that OpenAI is employing hundreds of contract workers to read and review actual conversations held with ChatGPT. This process, conducted under the internal codename 'Project Lily', aims to refine the AI's accuracy and conversational tone.
The Privacy Gap in Project Lily
According to leaked internal documents and Slack communications, these human reviewers are tasked with assessing how well ChatGPT responds to specific user prompts. While usernames are reportedly masked, the actual content of the chats—which can include sensitive personal information, location details, and 'user memories'—remains visible to the contractors.
Why This Matters
BozokMedia analysis shows that this practice highlights a significant tension between rapid AI evolution and user data sovereignty. As AI models become more integrated into daily life, the 'black box' of data training becomes a major liability for users who assume their interactions are strictly machine-to-machine.
The illusion of total privacy in AI interactions is being shattered by the operational realities of model training.
The reviewers are not just checking for accuracy; they are specifically trained to reduce 'sycophancy'—the tendency of an AI to blindly agree with a user—and to ensure the model does not mimic human emotions or claim to have real-life experiences. They rate responses on a scale of 1 to 7.
Historical Background
Reinforcement Learning from Human Feedback (RLHF) has been the backbone of modern LLM development. However, the scale of human intervention has increased exponentially, moving from simple label-based tasks to deep, contextual reviews of complex human-AI dialogues.
Frequently Asked Questions
1. How can I prevent OpenAI from using my data for training?
You can opt-out by navigating to ChatGPT Settings > Data Controls and toggling off the 'Improve the model for everyone' option.
2. Is my identity safe during these reviews?
While OpenAI uses a privacy filter to redact personal info, they admit the filter can occasionally fail to catch sensitive details.