AI Models Displayed Inconsistent Shopping Recommendations
Recent research reveals that generative AI models frequently shift product preferences and contradict their own prior outputs.
Updated on Sept. 30, 2026 in Artificial Intelligence

Live Poll
Do you trust AI-generated product recommendations for your own shopping and purchasing decisions?
Multiple studies have found that ChatGPT and Gemini struggle with output stability when responding to identical shopping queries. These reports highlight high rates of factual conflict and recommendation churn across both platforms in recent testing.
Why it matters
The lack of consistency and shared sourcing between major AI models suggests significant challenges for reliability in commerce-focused applications. This behavior creates uncertainty for users seeking stable, fact-based guidance from automated assistants.
Product.ai observed factual conflicts in 86% of 220 tested shopping questions, while an academic study found that ChatGPT and Gemini shared no common source domains in 76.7% of comparisons.
The players
ChatGPT
A generative AI platform developed by OpenAI that uses large language models to provide text-based responses.
Gemini
A suite of multimodal AI models developed by Google that integrates with its ecosystem of search and productivity tools.
Suff Digital
A research entity focused on analyzing consumer interaction patterns and consistency in automated AI responses.
Product.ai
A data analytics firm specializing in testing AI model accuracy and output reliability for shopping queries.
MentionBird
A market research firm that analyzes brand visibility and recommendation stability within large language model outputs.
The details
Researchers utilized provider APIs to simulate shopping interactions, asking 300 unique queries multiple times to track output variability. The high churn rates indicate that these models do not maintain a deterministic state when processing preference-based prompts. An academic study further noted that the models rely on disparate information ecosystems, sharing only 5.4% of source domains on average when answering the same questions.
Timeline
August 2026: Tracking showed Reddit citations in ChatGPT fell.
August 27, 2026: MentionBird published study results on recommendation churn.
September 16, 2026: Academic preprint published regarding source variation.
September 22, 2026: Product.ai published study results.
September 24, 2026: Suff Digital published study results.
The Tech Race
The observed instability in model outputs highlights the volatile nature of the information retrieval landscape across competing platforms. This divergence underscores a significant barrier to the widespread integration of AI as a reliable retail advisor.
Users currently cannot rely on AI models for consistent product comparisons, as both ChatGPT and Gemini frequently alter their recommendations on back-to-back queries. Until these models show greater output stability, individuals should treat AI shopping advice as preliminary and non-authoritative.
The takeaway
The lack of consensus between major models demonstrates that AI-driven product recommendations remain highly variable and prone to self-contradiction. Users should track the publication of upcoming model benchmarks to see if developers improve the deterministic quality of these shopping queries.
Further reading
For more on the current state of model reliability, see Artificial Intelligence.
Source note: This article includes information reported by TechRepublic.
Live Poll
Do you trust AI-generated product recommendations for your own shopping and purchasing decisions?










