Without an EAN, image and name can suggest a candidate, but they are not enough to certify identity. Start by filtering on brand, model, or manufacturer reference when available. Then compare a normalized title, variant attributes, and several visual characteristics. Any divergence in size, color, quantity, version, or presentation must block automatic acceptance. The result should remain an explainable score accompanied by evidence, then be validated by a person or an official source. A generic image, an incomplete title, or old packaging requires review, even if the overall resemblance seems strong.
Why two nice-looking visuals can represent different products
Sites frequently reuse a family image across several sizes or colors. A brand may change its packaging without changing the contents, while a promotional bundle may use the single unit's photo. Cropping, compression, or reframing also affects pixel-by-pixel comparison. The image is therefore a signal, not an identifier.
GS1 publishes a specification dedicated to product images and their metadata. This standard shows the importance of an organized representation linked to the item, but the presence of a compliant image alone does not prove that two offers are identical. Keep the source URL, the observation date, the image's role, and, if the feed provides it, the associated variant reference.
The name poses a similar problem. "Black Runner Shoe" may omit the size, the generation, and the intended audience. "Shampoo 250 ml" and "Shampoo 2 × 250 ml" share almost every word, but not the marketed quantity. Before calculating a similarity, extract the attributes whose divergence changes the item.
Building an evidence file per candidate
A useful file separates source data from calculated signals.
| Family | Data to keep | Use | Limit |
|---|---|---|---|
| Provenance | URL, supplier, SKU, date | find the listing again | does not prove identity |
| Text | raw title, brand, MPN, description | extract model and variant | may be incomplete |
| Visual | URL, checksum, dimensions, role | spot a resemblance | may be generic |
| Variant | size, color, volume, pack | check for divergence | requires a normalized unit |
| Proof | brand listing, packaging, authorized registry | confirm | availability varies |

Keep the raw titles, then create a comparison version: harmonized case, neutralized punctuation, and normalized units where explicit. Don't strip out numbers or quantity terms, as they often distinguish variants. Also keep words like "mini," "maxi," "refill," "set," "pair," or their equivalent in the source language if their commercial meaning is relevant.
A two-stage procedure
1. Filtering by strong constraints
First discard impossible pairs. A different brand, a contradictory MPN, an incompatible size, or a diverging pack content is a blocker, unless official proof explains the difference. Missing fields are not matches: they simply reduce the available proof.
If a GTIN appears on only one side, validate its check digit and look for an authorized source linking it to the candidate. The method for validating a GTIN without false positives still applies: mathematical consistency does not confirm the product. If several variants improperly share the same code, apply the review of dangerous GTIN duplicates.
2. Ranking the remaining candidates
Then calculate distinct signals: title similarity, brand match, model, variant, visual characteristics, and source context. Don't just display an overall percentage. A person should be able to understand why the candidate is ranked first and which fields are missing.
A scoring system can be designed in-house, but no universal weighting is guaranteed. For example, a team might decide, as a design hypothesis, that an exact MPN is more discriminating than a close color, and that any contradictory pack quantity cancels the candidate. These rules must be tested on a labeled dataset, versioned, and revised based on errors. They should not be presented as a GS1 standard or as an already-active feature.
Comparing images without losing their context
A robust approach combines several observations. A perceptual hash can find the same visual cropped or recompressed. Descriptors can match shape or packaging. Text recognition can pick up a printed model number. Each of these signals can also be wrong: an identical background can dominate the image, staging can hide the variant, and tiny text can be misread.
Apply these caution rules:
1. Only download images whose use is authorized and keep the provenance. 2. Reject placeholders, logo-only images, and images too small to be informative. 3. Where possible, compare the main image and a secondary view, without exceeding the source's terms. 4. Detect identical images reused across several variants and reduce their discriminating power. 5. Never infer a pack from the number of objects visible if the text or data doesn't confirm it. 6. Present both visuals side by side during review.
An image can be useful for creating a candidate alert. It should not erase a structuring textual conflict.
Measuring quality with a ground-truth set
Before accepting matches at scale, build a set of confirmed examples and pairs that are deliberately close but different. Include packaging changes, similar sizes, similar colors, units versus packs, and generic images. Keep this set separate from the sample used to tune the rules.
Measure at minimum false matches and missed matches. In a purchasing context, a false positive can be more costly than a candidate left for review. The threshold should reflect this risk. It can also vary by evidence: a pair with an exact matching MPN and matching attributes doesn't carry the same case file as a pair based solely on two similar titles.
Log the human decision, the evidence, and the later outcome. A reasoned rejection improves the rule; a click with no rationale does not help understand the error. Once identity is confirmed, the offers can enter the comparison of suppliers on the same product.
Human validation checklist
- The brand matches or the discrepancy is officially explained.
- The model or MPN is not contradicted.
- Size, color, volume, and presentation have been compared.
- The pack unit count is proven textually or structurally.
- The images are neither generic nor placeholders.
- Provenance URLs and dates are kept.
- Missing attributes do not count as agreement.
- Every signal and every blocker is visible.
- The threshold comes from a tested, versioned dataset.
- Official proof is attached to sensitive validations.
What ArbitragePro+ can automate / what the seller must verify
The proposed score-rapprochement-image-nom specification could extract candidates, compare titles and visuals, apply variant blockers, and prepare an explainable review queue. It should never on its own promote a resemblance to a certain identity.
The seller must verify the brand, reference, variant, pack, image usage rights, and official proof before linking offers or making a purchasing decision.
Review product candidates
Open ArbitragePro+ search to review the available product matches and information.
Official sources
- GS1, «Product Image Specification Standard», https://www.gs1.org/standards/gs1-product-image-specification-standard/40 — accessed 2026-08-17.
- GS1, «Global Trade Item Number (GTIN)», https://www.gs1.org/standards/id-keys/gtin — accessed 2026-08-17.
- GS1 Support, «GS1 barcode commonly used for trade item identification», https://support.gs1.org/support/solutions/articles/43000734137-what-is-the-gs1-barcode-commonly-used-for-trade-item-identification- — accessed 2026-08-17.
