A supplier reliability score must summarize observed facts, not a commercial impression. Separately measure the feed's availability, freshness, completeness, stability, and the consistency of critical fields. Keep every raw measurement, make the weightings explicit, and impose blocking thresholds: an excellent response rate must never offset stale prices or unusable identifiers. The score then serves to prioritize a review, compare a source to itself over time, and decide whether an import can continue. It proves neither product profitability nor the supplier's overall quality.
A useful score starts with an operational question
"Is this supplier reliable?" is too vague. The real question is rather: "Can I use this feed today to prepare a purchasing decision, and with what level of human control?" This framing separates the catalog's technical quality, the offer's commercial quality, and the quality of the contractual relationship.
The score therefore applies to a precisely identified source: an API URL, an FTP file, a CSV export, or another documented feed. It must also specify the observed period. A rating calculated from a single successful download does not carry the same weight as a series of regular checks. The period, the frequency of observations, and any excluded incidents should remain visible next to the result.
Before scoring the source, keep at minimum: the time of every attempt, the technical result, the size or number of rows received, the dates present in the data, the number of rejected rows, and the exact rule for each rejection. To go deeper on the time dimension, the price and stock freshness check method complements this score without merging into it.
The five dimensions to measure separately
A robust grid avoids criteria that can't be audited. The following dimensions are working categories, not a universal standard:
| Dimension | Observation kept | Control question | Mistake to avoid |
|---|---|---|---|
| Availability | successful attempts, HTTP or network errors, duration | Can the feed be retrieved as expected? | counting an empty response as a full success |
| Freshness | file date, business date, collection date | Do price and stock correspond to a known period? | confusing download time with update time |
| Completeness | presence of required fields per row | Do usable rows contain the necessary data? | averaging critical and optional fields together |
| Stability | changes to columns, types, units, and formats | Does the data contract remain compatible? | penalizing an announced, correctly migrated change |
| Consistency | anomalies in price, stock, currency, GTIN, and variants | Do values respect the explicit rules? | turning an anomaly into a silent correction |
Item availability deserves a clean test of its own. A state like InStock expresses availability in Schema.org's vocabulary, but does not provide an exact quantity, a measurement date, or a reservation. The article on the difference between boolean stock and exact quantity helps define what the feed actually lets you claim.
Building a scale without manufacturing certainty
There is no official weighting valid for every seller. The scale must reflect the risk of the use case. For a fast-turnover purchase, stock freshness can be a blocker. For title enrichment, descriptive completeness can weigh more. Two teams can therefore arrive at different ratings from the same observations, provided they publish their method.
Here is an entirely hypothetical, reproducible example. Assumptions: a team allocates 100 points total, split between availability 20, freshness 30, completeness 20, stability 15, and consistency 15. Each dimension gets a ratio between 0 and 1 defined in a versioned rules dictionary. If the observed ratios are respectively 0.95, 0.70, 0.90, 1.00, and 0.80, the calculation is:
20 × 0.95 + 30 × 0.70 + 20 × 0.90 + 15 × 1.00 + 15 × 0.80 = 85.
This result of 85 is neither a market value nor a purchasing recommendation. It only demonstrates the formula with fictional assumptions. The breakdown mainly reveals that freshness is the weak point. That's more useful than an overall green color.
Then add veto rules. For example, in your internal policy, a file with no currency or no verifiable date can be placed in review, even if the weighted average stays high. A veto is not a penalty against the supplier: it's a safeguard against the mathematical offsetting of a critical flaw.
Reproducible seven-step procedure
1. Define the source and the use case. Identify the feed, the authorized account, the channel, the expected frequency, and the decision the data must support. 2. Write the rules dictionary. For each criterion, state the field, the rule, the unit, the tolerance, and the expected response. An unknown rule stays "not evaluated," never "compliant." 3. Collect the raw evidence. Log timestamps, response codes, useful headers, file checksum, volumes, and rejections without substituting the final score for them. 4. Calculate each dimension. Use only observations from the displayed period. Keep the numerator, the denominator, and the exclusions. 5. Apply weightings and vetoes. Version the scale. Any change should be able to explain why two successive calculations differ. 6. Compare the source to its own history. A sudden drop is often more informative than a comparison between two suppliers with different contracts. 7. Decide and document. Authorize the feed, restrict it to a test batch, request a correction, or suspend the import. The human decision should cite the main evidence.
RFC 9110 notably defines HTTP statuses and the Last-Modified and ETag validator fields. These elements can document a retrieval, but they alone do not certify the business date of a price. A supplier could republish today a file containing old data. You therefore need to keep both the transport metadata and the catalog's internal dates.
Interpreting the trend rather than enshrining a single rating
An isolated score gives a snapshot. A time series lets you see a behavior change: growing delays, rising rejections, alternating patterns, or recovery after an incident. Display the overall rating, but also the components and the events that explain the movement.
Avoid absolute ranking of suppliers when scopes differ. A small catalog updated daily and a very large base refreshed in segments don't promise the same service. Compare each source first to its own documented contract, then to its own previous periods. A cross-comparison is only relevant if the criteria, windows, and required fields are identical.
The score also does not evaluate usage rights, solvency, product compliance, actual delivery times, or final profitability. Those checks rely on different evidence. Bundling them into a single rating would make causes invisible and corrections imprecise.
What ArbitragePro+ can automate / what the seller must verify
The score-sante-source function here is a proposed specification, not a feature declared as available. It could collect import results, calculate the dimensions according to a versioned scale, display missing observations, and flag a veto. It should always allow going back to the raw measurements and distinguishing "zero," "unknown," and "not applicable."
The seller must verify the feed's contract, access rights, the meaning of the fields, the relevance of the weightings, and the final decision. They also confirm that the observed period is representative. Before scaling up a source, a supplier canary batch lets you test the rules on a limited, reversible scope.
Checklist before using the score
- The source, channel, and period are named.
- Each dimension traces back to a consultable raw measurement.
- Critical fields and vetoes are written before the calculation.
- Weightings are dated, versioned, and justified.
- Missing values are not turned into arbitrary zeros.
- A collection date is not presented as a business update date.
- The overall rating comes with its components.
- The human decision and its rationale remain traceable.
- The score is not used as proof of profitability or legal compliance.
Monitor source signals
Open source health in ArbitragePro+
Official sources
- IETF, RFC 9110 — HTTP Semantics: https://www.rfc-editor.org/rfc/rfc9110.html (accessed 2026-08-17).
- Schema.org, ItemAvailability: https://schema.org/ItemAvailability (accessed 2026-08-17).
- Schema.org, InStock: https://schema.org/InStock (accessed 2026-08-17).
