AI can't take over your holiday shopping quite yet, a new study suggests.
· Business Insider
Andriy Onufriyenko/Getty Images
Visit afnews.co.za for more information.
- Don't rely on LLMs to do your holiday shopping this year, a new study says.
- A test of four AI engines found that they had issues providing accurate information about products.
- Problems ranged from giving the wrong price to not finding the latest model.
AI isn't able to take over your holiday shopping yet, a new study suggests.
AI-powered shopping assistants frequently give conflicting answers about prices, product specifications, and which items are current, according to a study published this week by Product.ai, a startup that checks product claims against evidence.
The company tested the free and paid versions of ChatGPT, Claude, Gemini, and Perplexity using 220 shopping questions about products ranging from laptops and TVs to mattresses, sunscreen, and robot vacuums. It ran each question five times on each service, capturing 8,794 responses.
Eighty-six percent of the questions produced a repeatable factual conflict, Product.ai found. The study defined a conflict as a checkable disagreement — such as a different price, product model, or specification — that appeared in multiple responses.
"In short, it says that these LLMs aren't there yet when it comes to this end-to-end experience," said Dakota Nunley, Product.ai's head of search product.
AI has been creeping further into shopping, with both retailers and LLMs offering tools to find and buy things online. AI agents such as Meta's Muse or Instinct could do even more of the work for you.
Product.ai's study points to a more fundamental issue: AI still struggles to provide accurate information that humans or AI agents need.
"Perplexity is the only AI company relentlessly focused on achieving 100% accuracy, and we lead the industry in every measure of it," a Perplexity spokesperson said. Representatives for the other three LLMs did not respond to requests for comment.
The study found that 97% of head-to-head questions, such as asking the AI models to compare two products, produced a conflict. That was higher than the 75% conflict rate for straightforward factual or specification questions.
The models also struggled to provide users with the correct product prices. Of the 913 answers that Product.ai could verify, 85% matched the current or listed price.
There were differences between the LLMs. Gemini had the highest share of questions with what Product.ai classified as a costly error: 56% on its free tier and 54% on its paid tier.
The results using Claude improved with its paid version, with costly errors falling to 21% from 44%. Perplexity had the lowest costly-error rate, at 14% for its paid tier, followed by ChatGPT's paid version at 17%.
Product.ai also found that the tools sometimes contradicted themselves when asked the same question repeatedly. Gemini's free tier did so on 29% of questions, the highest rate in the study.
When the answers were wrong, though, the price was off by a median of $300, Product.ai said.
Such a big difference should be enough to give shoppers pause about relying on AI agents, Nunley said. "That's a big miss, right there," he said.
Instead of handing over payment details to agents or taking whatever AI agents say as gospel, Nunley said, users need to verify the information themselves. That can mean asking multiple AI engines and comparing the results — or simply checking the seller's websites themselves, he said.
"Use AI this upcoming holiday season as a discovery tool, as a thought partner," Nunley said. "But be wary of going hands-off at the moment."
Do you have a story idea about AI and shopping? Contact this reporter at [email protected] or via encrypted messaging app Signal at 808-854-4501. Use a personal email address, a nonwork WiFi network, and a nonwork device; here's our guide to sharing information securely.
Read the original article on Business Insider