Vague Prompts Are the Real Problem, Not the Model — 50-Run Test Shows Prompt Quality Trumps Model Choice

A Reddit user ran an experiment to test the common claim that one AI model is smarter than another. They took ten common prompts and ran each one through ChatGPT 4, Claude Sonnet, and Gemini 1.5 Pro five times each — 150 outputs total.
What they found: the outputs were weirdly similar in quality. Not identical, but within the same tier. All three either gave something usable or all three gave "generic mush." They almost never disagreed on whether a prompt was answerable. The variable wasn't the model — it was the prompt.
Two prompts, different results
The same vague prompt produced identical bland output across models. For example:
"Write a cover letter for a marketing job"
All three returned the same kind of generic, applicable-to-anyone cover letter. People would call it a "ChatGPT cover letter" then try Claude and call it a "Claude cover letter" — same letter, different name.
But a specific prompt changed everything:
"Write a cover letter for a senior marketing role at a B2B SaaS company. I have 7 years of growth experience, mostly at Series A/B startups. The hiring manager is technical, ex-engineer. Avoid generic phrases like 'passionate about' or 'results-driven.' Use specific numbers from my background where it makes sense to invent plausible ones. Target 280 words."
All three returned something actually good. Different in style, but all useful.
Common pattern in complaints
The user reviewed dozens of "AI is so bad" complaints on Twitter and Reddit and noticed the same pattern: prompts like:
"Help me with my resume""Write a marketing plan""Explain quantum physics""Make this code better"
These prompts fail because they don't specify who you are, who it's for, what good looks like, or what to avoid. The model has to guess the most common version of that request — which is a generic template.
Mental model: prompt as brief
The key insight: stop thinking of it as "asking AI a question." Think of it as "writing a brief for an intern." A good brief tells the intern the audience, what success looks like, what to avoid, format, constraints, and at least one example of the kind of output you want.
Once the user started writing prompts like briefs, the model switching stopped. ChatGPT, Claude, and Gemini all got dramatically better — not because the models changed, but because the prompts changed.
If you're tempted to switch models because one gives bad results, try sharpening your prompt first. The model differences are real but much smaller than the prompt differences.
📖 Read the full source: r/ClaudeAI
👀 See Also

Claude Stealth Mode Directive for Autonomous AI Execution
A Reddit user shares a 'stealth mode' directive that forces Claude to operate silently and autonomously, delivering complete one-shot results without conversation output until work is complete.

Cut OpenClaw Boot Tokens 43% by Slimming Tool & Memory Files
Reduced boot tokens from ~9,457 to ~5,400 (43% drop) by converting TOOLS.md to an index, moving tool details to separate files, and implementing staged memory promotion.

Good AI-Assisted Development Happens at the Systems Level, Not the Task Level
A Reddit user explains how shifting from fixing AI agent output to designing constraints—like a linter rule that forces UI navigation—prevents entire classes of bugs permanently.
How I Proved My Own Bug Report Wrong: OpenClaw Telegram Debugging via apiRoot Proxy
OpenClaw's Telegram apiRoot can be pointed at a local proxy to log actual wire payloads. A dev retracted a bug report after learning that copied text was rendered, not reinserted.