The next frontier in artificial intelligence isn't just building better models — it's finding where they break. A new paid research study is now recruiting professionals who can do exactly that: design challenging AI prompts that expose model limitations.
What This Study Is Actually Asking Participants to Do
The research trial is straightforward but demanding. Participants take a complex workflow from their own job or personal life and translate it into a prompt that requires both reasoning and real-world lookup. The goal isn't to create a simple question — it's to build a task that pushes the AI to its limits.
After crafting the initial prompt, participants run it in ChatGPT to identify where the model fails. Then comes the refinement stage: tweaking and adjusting until the prompt consistently breaks the system. Finally, each participant writes a clear grading rubric that a stranger could use to evaluate any AI's attempt at the task.
Why Breaking AI Models Has Become Valuable Work
As AI systems become more integrated into daily workflows, understanding their failure points has grown critical. Companies deploying AI need to know exactly where their tools fall short — before customers discover those flaws themselves.
This study sits at the intersection of quality assurance and AI development. By identifying people who can systematically expose model weaknesses, researchers hope to build a reliable bench of evaluators for ongoing work.
How the One-Hour Process Unfolds
The study is designed to be completed in about an hour, but that hour is intense. Participants must draw from their real-world experience — no generic prompts allowed. The task must reflect an actual workflow that requires genuine reasoning and external information lookup.
This requirement ensures the prompts are authentic and demanding, not artificial constructions that fail for trivial reasons. The refinement stage is where the real skill emerges: understanding why a model fails and adjusting the prompt to expose deeper limitations.
Who This Opportunity Is Best Suited For
Professionals who regularly use AI tools in complex, specialized work are ideal candidates. Think lawyers testing contract analysis, engineers evaluating code generation, or financial analysts probing data interpretation skills.
People who already spend time crafting detailed prompts in their daily work will find this study particularly natural. The ability to think systematically about AI behavior — not just use it — is the core skill being evaluated.
What the Research Team Is Looking For
The study's stated goal is to identify individuals suited for ongoing prompt engineering and evaluation work. This isn't a one-off task; it's a screening process for potential long-term collaboration.
Success in this trial could lead to continued paid opportunities in AI evaluation. The researchers are explicit that this initial trial helps them build a bench of skilled people for future projects.
Why the Grading Rubric Matters More Than the Prompt
The final deliverable — a grading rubric usable by strangers — is arguably the most important part of the study. A prompt that breaks an AI is only useful if someone can objectively evaluate whether another AI's attempt succeeds or fails.
This requirement separates casual AI users from systematic evaluators. Writing clear, unambiguous grading criteria requires deep understanding of both the task and the AI's likely failure modes.
Confirmed Details vs What Remains Unclear
Confirmed: The study is paid, involves about an hour of work, requires designing a challenging prompt, testing it in ChatGPT, refining it until failure, and writing a grading rubric. The goal is to identify people for ongoing evaluation work.
Unclear: The specific compensation amount, the number of participants being recruited, the timeline for the study, and the exact criteria for selection have not been publicly disclosed.
The Growing Field of AI Red Teaming and Evaluation
This study reflects a broader industry trend: the rise of AI red teaming as a professional discipline. Companies increasingly recognize that finding model weaknesses before deployment is cheaper and safer than discovering them after release.
Professional prompt engineers and AI evaluators are emerging as distinct career paths. Studies like this one serve as informal screening mechanisms for talent in this growing field.
What Professionals Should Consider Before Applying
Anyone interested in this study should think carefully about whether they can identify a genuinely complex workflow from their own experience. Generic tasks won't demonstrate the skill the researchers are looking for.
It's also worth considering whether you can articulate why a model fails — not just that it fails. The refinement process requires diagnosing AI behavior, which is a distinct skill from simply using AI tools.
What Happens After the Study
For participants who demonstrate exceptional skill, the likely outcome is ongoing prompt engineering and evaluation work. The researchers have framed this as a bench-building exercise, suggesting successful participants may be called upon for future paid projects.
For the broader AI industry, studies like this contribute to better understanding of model limitations — knowledge that ultimately leads to more robust and reliable AI systems.
Our Take
This study represents a smart approach to a real problem: AI evaluation needs people who understand both complex real-world tasks and AI behavior. The emphasis on real workflows rather than artificial puzzles is particularly promising.
The opportunity also highlights a growing truth about the AI economy — the people who understand model limitations are becoming as valuable as those who build the models. For professionals who already push AI tools to their limits daily, this could be a natural and lucrative side opportunity.
Frequently Asked Questions
What is the purpose of this AI prompt study?
The study aims to identify professionals who are exceptionally good at designing tasks that expose AI model limitations. Successful participants may be invited for ongoing prompt engineering and evaluation work.
How long does the study take?
Participants spend about an hour translating a complex workflow into a demanding prompt, testing it in ChatGPT, refining it until the model fails, and writing a grading rubric.
Who should apply for this paid AI research study?
Professionals who regularly use AI tools in complex, specialized work and can translate real-world workflows into demanding prompts are ideal candidates. The ability to systematically analyze AI failure points is essential.
What happens after completing the study?
Participants who demonstrate exceptional skill may be selected for ongoing prompt engineering and evaluation opportunities. The study is designed as a screening process to build a bench of skilled evaluators.