Shelf Signal

Shopping Using AI Tools Across Different Product Categories

AI shopping shifts authority unevenly across product categories based on risk and complexity.

Reporter · · 13 min read
Cover illustration for “Shopping Using AI Tools Across Different Product Categories”
AI Shopping Behavior · September 17, 2026 · 13 min read · 2,932 words

AI shopping tools do not treat a bag of dog food and a diamond ring the same way, and that gap is the whole story. Delegation to an AI agent isn't a single switch that flips from "search" to "buy." It's a dial, and where a shopper sets that dial depends on the category: how consequential the purchase feels, how personal the fit is, and how well the product's attributes can be written down as data in the first place.

Most coverage of AI shopping treats this as one behavior: search faster, buy faster, regardless of what's in the cart. That's the intuitive read, and it's wrong. A shopper reordering paper towels and a shopper booking a honeymoon are using the same underlying technology, but they're handing over wildly different amounts of authority, and the AI is leaning on completely different signals to make the call.

IBM defines agentic commerce as AI systems that research, compare, and complete purchases on their own, not chatbots that just answer questions about a product. That distinction matters because scale has already caught up to the idea. OpenAI reported that ChatGPT processes 50 million shopping-related queries a day. Braze's Retail Customer Engagement Review projects agentic shopping adoption climbing from 19% of shoppers in 2025 to 46% by the end of 2026. Whatever a brand or a shopper thinks about this shift, it's already underway, and it's underway unevenly across categories.

This piece is for two readers at once: the shopper trying to figure out when it's smart to hand an AI the keys, and the brand trying to figure out what its product data needs to say for an agent to find it, trust it, and buy it. How agent-to-agent payments actually clear comes later. Here, the focus is the lens: category first, mechanics second.

How AI shopping tools move from assistant to autonomous buyer, and where most categories still sit

Diagram: The Delegation Ladder: Four Levels of AI Shopping Authority. Visualizes: Visualize a four-rung ladder showing how much authority a shopper hands to an AI, from bottom to top: (1) Assisted Discovery — AI surfaces options, human decides and…

There's a rough ladder to how much authority a shopper hands over. At the bottom sits assisted discovery: the AI surfaces options, a human decides, a human checks out. This is Amazon's Rufus assistant, or early ChatGPT shopping results, doing the legwork of narrowing choices without touching the transaction.

One step up is agentic shopping: the AI compares options, makes a recommendation, and can start the checkout process, but a human still has to confirm before money moves. Above that sits autonomous shopping, where an agent operates inside pre-approved boundaries, a budget cap, a brand preference, a delivery window, and executes without a human reviewing each step. At the top is agent-to-agent commerce: a buyer's agent negotiates directly with a merchant's agent, and no person is in the loop at the moment the transaction closes.

Right now, most real-world shopping is at the first two rungs. IBM's Institute for Business Value study found 45% of consumers already use AI for part of their buying journey, not the whole thing. That's a meaningful number, but it's a number about partial delegation, not full autonomy.

What decides where a given purchase lands on that ladder isn't the AI's capability. It's the shopper's tolerance for risk, how easily a bad decision can be reversed, how personal the fit needs to be, and how cleanly the category's attributes translate into structured data an agent can actually read. And there's a hard ceiling on all of this that has nothing to do with what the technology can do: Visa's "Earning Consumer Trust" report found 30% of consumers say they'd never let AI handle shopping or touch their payment information at all. That ceiling is psychological, not technical, and it moves depending on what's being bought.

Each category below sits in a different spot on that ladder, and the reasons are specific enough to learn from.

Everyday consumables: the category where full delegation is already normal

Toilet paper, dish soap, coffee pods, dog food. This is where full delegation already feels unremarkable, and the reasons are structural rather than emotional.

The shopper already knows the product. There's no discovery step, just reorder logic. Price and availability are the only variables that actually move the decision, and both are numbers a machine can read cleanly. Getting it wrong costs almost nothing, the wrong brand of paper towels is an annoyance, not a financial or emotional setback, so an agent doesn't need to model the risk of a bad outcome because there basically isn't one.

The typical agent behavior here: watch stock levels, trigger a reorder once a threshold is crossed, apply whatever loyalty rewards are available, and pick the lowest landed cost across retailers, all without asking the shopper to sign off. Deloitte's research offers a clean example of how far this stretches: a prompt like "plan a family dinner for six this weekend within a $100 budget" can send an agent across multiple retailers' inventories, applying rewards and coordinating delivery, without a human touching a single line item.

What the AI actually relies on here isn't reviews or brand copy. It's live inventory feeds, accurate pricing data, and structured SKU attributes, size, count, formulation, plus delivery-window data. For a brand, that means the agent is selecting almost entirely on catalog hygiene and price. A stale inventory feed or a missing attribute doesn't just hurt visibility, it loses the reorder outright to whichever competitor's feed is clean that day.

Subscription and auto-replenishment programs are really just this same logic pushed one step further: the agent is pre-authorized ahead of time and simply executes on a schedule.

Consumer electronics and appliances: where specs dominate but trust gates the final step

Electronics sit in a strange middle zone. Shoppers lean on AI heavily to do the comparison work, but they still want to be the one who pulls the trigger.

The price tag raises the stakes of getting it wrong, and unlike a bag of dog food, a laptop or a washing machine is attribute-rich: processor, resolution, battery life, compatibility, capacity. Those are exactly the kind of facts structured data carries well, and comparing a dozen SKUs across those fields is tedious for a person and trivial for a machine. But design preference, brand loyalty, and worry about the return process all pull the shopper back toward wanting one last look before paying.

Research on how these tools actually rank products suggests ChatGPT favors authoritative sources and review signals, while pricing competitiveness appears to factor into how tools like Copilot surface results. So the AI's judgment is built on structured specs, price positioning, and review sentiment, not on how a product is described in marketing copy.

What tends to happen is a kind of research gap: the shopper uses AI to compress what used to be days of tab-switching into a few minutes, arrives at a shortlist of two or three products, and then goes to the brand's own site or a retailer to confirm and buy. The AI is the shortlist engine. The human stays the final approver.

For a brand, missing or inconsistent specs are disqualifying at the shortlist stage, not just a minor visibility hit. If an agent can't confirm whether a soundbar is compatible with a given TV, it moves to the next brand whose data answers the question. That gap in accuracy between structured and unstructured content is well documented: AI tools handle structured spec data more reliably than paragraph-style marketing copy, which is a strong argument for treating spec sheets as more important than brand storytelling in this category. Payment authorization for high-ticket purchases is still a real friction point too, which is exactly why Visa's Intelligent Commerce program and Mastercard's Agent Pay exist: someone had to build the layer that lets an agent spend real money on something expensive without a human physically entering a card number.

Fashion and apparel: where personal fit and taste resist automation most

Apparel is the category where full agentic checkout is furthest away, and the reasons are almost all physical.

A size "M" fits differently from one brand to the next, and that variance simply isn't captured by size data alone. Style is contextual, tied to occasion, body shape, and what's already hanging in a shopper's closet, none of which fits neatly into a product feed. Return rates run high in this category, which means an agent has to model the logistics of sending something back, not just the logistics of buying it. And the attributes that matter most to a shopper, color, texture, drape, how a fabric moves, are exactly the ones that resist being turned into structured data.

AI does the narrowing well: filtering by size, price, category, matching a described aesthetic, flagging a retailer's return window. What it can't do reliably is predict whether a shopper is actually going to like how a specific jacket looks on their own body. Delegation stops at the shortlist, full stop.

Research tracking high-street fashion brands across AI shopping tools found a clear pattern: brands with consistent, specific positioning appeared in AI-generated recommendations, while brands leaning on vague, aspirational messaging often did not appear at all. COS, part of H&M Group, ranked highest for coherence across thousands of test prompts in that study. The lesson generalizes: in a category where the AI can't close the sale on its own, simply getting mentioned depends on whether a brand's product data and voice are specific enough for a model to describe accurately.

There's also a trust gap tied directly to returns. Visa's 2025 research found 79% of consumers interested in agentic commerce still worry about data privacy, and in a high-return category like apparel, that worry compounds: shoppers are afraid the agent will buy something that's a hassle to send back. Brands that publish return policies in clear, machine-readable formats remove one of the biggest psychological blocks to letting an agent shop for clothes at all.

Travel and experiences: a category that shows what full agentic orchestration looks like in practice

Travel is where agentic commerce moves from theory to reality, multiple vendors, multiple decisions, coordinated in one motion.

Booking a trip means flights, a hotel, ground transport, maybe activities, and stitching those together across separate vendors is precisely the kind of task agents are built for. Availability changes by the hour, so live inventory data isn't optional, it's the whole game: stale data means the agent books a room that's already gone. Comparing dozens of flight and hotel combinations across price, schedule, and rating is exhausting for a person and computationally nothing for a machine. Plenty of travelers would rather set the constraints once, under $2,000, direct flights, a 4-star hotel, near the beach, and let the agent go execute, rather than lose an evening to a booking site.

IBM names travel and ticketing directly as a core use case for agentic commerce, not a fringe example bolted onto retail. What the agent leans on here is real-time availability data, structured pricing, cancellation policy terms, aggregated reviews, and loyalty program details, not brand storytelling or marketing copy.

Delegation still has a ceiling, though. For a honeymoon, or a once-in-a-lifetime trip, shoppers want to verify the details themselves and often want to talk to an actual person at some point in the process.

Google's Agent-to-Agent protocol, launched in April 2025 with more than 50 supporting organizations at the start and growing to more than 150 including PayPal, Salesforce, and SAP, is the piece of infrastructure that makes this kind of multi-vendor orchestration possible. Instead of one agent trying to do everything, it hands sub-tasks off to specialists, an airline-booking agent here, a hotel-booking agent there. The broader lesson extends past travel: any purchase built from several coordinating decisions, an event plus a hotel plus transport, moves toward full delegation faster than a single standalone item ever will.

Health, supplements, and personal care: where regulatory signals and trust in source determine what an AI will recommend

Health and wellness purchases run on a different logic than anywhere else on this list, and it starts with caution built into the models themselves.

Claims in this category are regulated or actively disputed, so an AI trained to avoid liability will often hedge, or decline outright to recommend something for a specific health outcome. Shoppers in this space are frequently researching rather than buying, using AI heavily to check ingredients or drug interactions or read reviews long before any purchase intent exists. And personalization matters in a way that's hard to standardize: what works for one body, age, or condition can be wrong, even harmful, for another.

The signals an AI leans on here are third-party citations rather than brand copy. Tools in this space tend to prioritize third-party sources like community threads and expert-written content over branded product pages. Add review sentiment, ingredient transparency in the product data itself, and clearly written policy language, and that's the real evaluation stack. Trust architecture matters more in this category than any other on this list: a brand whose descriptions make vague claims, or that skips ingredient-level detail, is less likely to get cited, not because of a search ranking penalty, but because the model has no way to verify what's being claimed.

Delegation in this category stays lower than in consumables, even for something a shopper buys repeatedly, because "the right product" keeps shifting as a person's needs change. Shoppers use AI to research and narrow the field, then buy it themselves. For brands, the attributes that actually decide citation are structured ingredient data, visible third-party certifications, and accurate review data, not the marketing language wrapped around them.

How each major AI shopping surface evaluates products differently, regardless of category

Every AI shopping surface runs on its own sourcing logic, and that logic can override category in some cases.

ChatGPT draws heavily on external product data sources for its shopping carousels, and layers on weighting for authoritative list mentions, awards, and review volume, which means feed quality on the Google side is doing more work than most brands realize. Perplexity leans the opposite direction, favoring third-party citations from places like Reddit and expert blogs, so domain authority and earned press coverage carry more weight in its results than clean product-feed data. Amazon's Rufus draws mainly from review sentiment and Q&A activity inside Amazon's own ecosystem, with some outside web data folded in, which means a brand with no Amazon presence is functionally invisible to it. Microsoft's Copilot tracks price competitiveness through Bing as a central ranking input. Google's Gemini reads intent signals out of product highlights and local inventory data, weighting structured markup and freshness heavily.

A counterintuitive finding emerges from this: ranking well on Google doesn't mean an AI system will ever cite the product. Research has found that a large share of brands that rank well in Google search results aren't cited by AI systems at all, and a substantial majority of the URLs cited by ChatGPT, Perplexity, Copilot, and Google's AI Mode don't appear anywhere in Google's top 100 results for the same query. Traditional SEO and AI citation are, in practice, two different games.

The payoff for showing up in that second game is large. AI-referred shoppers convert at 15.9%, a rate that research suggests outpaces typical organic search traffic by a wide margin, nearly nine times higher. That's not a volume story. It's a value story: getting cited by an AI system is worth chasing even if it happens less often than a Google impression.

The category and the platform interact in specific ways too. In fashion, Perplexity's weighting toward editorial sources favors brands with real press coverage. In electronics, ChatGPT's reliance on Google Shopping data means that feed, not the brand's own website, is the actual lever. In consumables, Amazon's closed ecosystem means a brand absent from Amazon can be entirely missing from reorder queries, no matter how strong its own site is. And freshness cuts across every platform: AI-surfaced product pages tend to be noticeably newer than what shows up in traditional search results, so a stale product page gets penalized in AI discovery no matter which system is doing the discovering.

Diagram: AI Citation vs. Google Rank: Two Different Games. Visualizes: Show a stark before/after or split comparison contrasting two metrics: traditional Google organic search ranking versus AI citation rate.

What brands need in their product data for AI agents to find, cite, and convert their products

A growing share of consumers now start their product research with an AI assistant instead of a search engine, and most brand catalogs still aren't built for that reader. They're built for a person scrolling a page, not a model parsing a feed.

"Agent-ready" product data has a fairly specific shape. It means structured attributes that answer the exact questions an agent gets asked on a shopper's behalf: size, material, compatibility, current availability, delivery window, return policy, not just a product name and a price. It means ingredient or component-level detail in categories like health and personal care, where a vague claim is worse than no claim at all, because the model can't verify it and won't repeat it. It means pricing and inventory feeds that update in near real time, since an agent working from stale numbers either recommends something unavailable or gets quietly dropped from the comparison set the next time around. And it means return and cancellation terms written in plain, structured language rather than buried in a PDF, since that's one of the biggest levers on how much a shopper is willing to delegate in the first place.

None of this is really new advice dressed up for a new format. It's the old discipline of clean product data, applied to a reader that happens to be a model instead of a person, and enforced by the fact that the model won't guess at what a brand left out.

Sources

  1. Agentic Commerce: AI Shopping Agents Guide 2025
  2. What Is Agentic Commerce? | IBM
  3. arxiv.org
  4. alhena.ai
  5. shopify.com
  6. alhena.ai

More in AI Shopping Behavior