The user sends a photo, the bot identifies the product

Andrew Altair, Founder
The user sends a photo, the bot identifies the product

TL;DR: A user sends a photo in chat and aiSTAFF identifies or responds to the product, then matches it to your catalog based on price and availability, on any channel, from the same shared brain.

People send photos, not product codes

See how a real user asks for something. They rarely know the model name or SKU. They take a photo of a chair they saw at a friend's place, a screenshot of an item from an ad, or a picture of a broken part they need to replace and send a "Do you have this?" There is only a text bot stuck there. He can't see the photo, so he asks the user to describe what he's already shown, and the conversation stops.

aiSTAFF is reading the picture. A user submits an image and the bot names the product or answers a question about it, then goes directly to match your catalog. The friction of describing a visual thing with words disappears. If you want this on your channels, AI Agents Service installs it, and it's one of the features of the comprehensive aiSTAFF Platform.

What the view does in chat

Image understanding is not a trick cut out of a page. It is included in the same sales flow as the text question:

  • Identify the product. The bot recognizes what is in the photo and names it.
  • Compare Catalog. It searches your built-in catalog for this product or the closest equivalent, returning a card with price, discount and availability.
  • Answer it. Questions like “is it waterproof” or “what size is this” answered from your product data.
  • Keep selling. It offers related items, checks inventory before confirmation, and offers callbacks if intent is high.

So the photo is the starting point for the same conversation, the typed question will start, with the same guardrails. Catalog matching works on a hybrid search engine described in the Platform Hub and the match gate still works: if a photo looks nothing like your stock, the bot says so instead of inventing it.

It works on all channels

Vision is not limited to one surface. Users send photos most on WhatsApp and Instagram, where taking and sending is second nature, and aiSTAFF handles the image on both. Because the function is read from a single common brain, the same view behavior is displayed wherever the user can attach an image. Agent deployment on WhatsApp, where photo messages are most common, is covered by AI Agent on WhatsApp Business and the one brain design behind it is one brain on five channels.

Photo plus Georgian reading

Image and text work together. The user can send a photo with an inscription in Georgian, and the bot reads both: it recognizes the product from the image and responds to the inscription in Georgian. It recognizes the language from the text and responds in a similar way, so the Russian caption gets a Russian response. The catalog itself can be in English, and the user buys Georgian, which is a linguistic behavior explained by chatbot speaks fluent Georgian. The user never has to translate their own request.

Working example

A plumbing supply store receives a WhatsApp message at 8 p.m.: a photo of a corroded faucet cartridge with the caption in Georgian, "I need this part, do you have it?" A human clerk stares at the photo, trying to match it from memory, and probably asks the customer to sign in. aiSTAFF recognizes the type of cartridge from the image, searches the catalog, finds two compatible parts, and responds to both, showing the price and availability, in Georgian. He asks if the customer wants to be set aside for pickup and gets a phone number. The store opens the next day on a ready-made order tied to a specific part, from a photo that would confuse a text bot. The user solved the visual problem with the image, the way he wanted.

Here's a pattern: the user does the simple thing of sending a photo, and the bot does the hard thing of identifying and matching. It keeps the conversation moving instead of bringing it back to the customer, which is the same personal discipline described in A chatbot that doesn't sound like a bot.

Honest Limits

Vision is powerful in identifying a clear product photo and matching it to your catalog. Blurry, dark or ambiguous images can be difficult to locate, and in such cases the bot asks a clarifying question rather than making an incorrect guess. It also respects the match gate, so a photo you don't sell returns an honest "we don't carry that" instead of a forced match. And as everywhere at aiSTAFF, the result is product discovery plus delivery or call-out, not card payment. How much previous context the bot keeps during a photo conversation depends on your plan, which is covered in conversation memory tiers, and the tone per channel is set in per-channel volume control.

FAQ

Can a bot really read a photo sent by a user?

Yes. aiSTAFF reads the image, identifies the product or answers a question about it, then searches your catalog to match price and availability.

Which channels does photo messaging support?

Vision works wherever a user can attach an image, including WhatsApp and Instagram, where photo messaging is most common, all from the same shared brain.

What if the photo is unclear?

If the image is unclear or ambiguous, the bot asks a clarifying question instead of guessing. A photo of what you don't keep returns honest dissimilarity.

Can the user send a photo with a Georgian inscription?

Yes. The bot reads the image and caption together, recognizes the product and responds to the caption in the user's language, Georgian, Russian or English.