Lantern Finds ChatGPT Shopping Results Mislead Brands
A new analysis by e-commerce analytics firm Lantern reveals that tracking individual AI shopping recommendations leads to misleading data, urging brands to measure long-term recommendation share.

E-commerce analytics firm Lantern has analyzed the volatility of generative AI shopping recommendations, warning marketers that tracking individual search queries provides a highly distorted view of brand visibility. While Adobe Analytics reported that traffic from generative AI tools to U.S. retail sites surged by 693.4 percent during the 2025 holiday season compared to the previous year, brands are struggling to measure their presence in these new search channels. Previously, Adobe found that 39 percent of surveyed U.S. consumers had already used generative AI for product research and recommendations.
To understand how these recommendations behave, Lantern collected data across approximately 3,300 AI shopping recommendation trends. This research involved 186 combinations of brands, prompts, and AI models, including OpenAI's ChatGPT, Google's Gemini, and Amazon's shopping assistant. The study revealed that while individual answers changed frequently, the broader distribution of recommendations remained highly consistent over time. Because generative models introduce natural variation into their responses, asking ChatGPT the same question ten times can yield ten different combinations of products.
For marketing practitioners, this volatility means that traditional search engine optimization tracking, which relies on static rankings, is no longer effective. Lantern CEO Andrew Lissimore advises brands to stop reacting to daily fluctuations or single screenshots of search results. Instead, marketers must define a baseline over several weeks and measure 'recommendation share'—the percentage of times a brand is recommended across hundreds of relevant conversations. This approach prevents companies from prematurely changing product content or redirecting marketing budgets based on isolated, non-representative query variations.
Furthermore, practitioners must analyze recommendation patterns by specific model rather than combining them into a single visibility score. OpenAI draws on merchant product data and public retail sources, while Google aggregates data from its own Shopping platform. Because these systems utilize different data sources, a brand might see its recommendation share rise in ChatGPT while remaining entirely stable in Gemini. Tracking these patterns over time allows brands to pinpoint exactly which consumer queries and AI models are driving their digital visibility.
This is our own summary of reporting by Unite.AI



