Source ratings
How each provider actually performs, pair by pair. Measured by us, published as measured, never declared by the provider.
The public rating is not built. The measurement behind it is, and it is running right now: the engine scores every source continuously and routes on those scores. What is missing is publishing them, and publishing a rating badly is worse than not publishing one.
Every source already carries a score built from measurement rather than from a listing: how often it delivers, how quickly, how fresh its data is, and how often it disagrees with the rest. That score decides where each call goes. It is why providers cannot buy position, and why a source that starts degrading loses traffic within minutes rather than after someone files a ticket.
Why pair by pair
A single number per provider is the wrong shape. A source can be excellent on one pair and absent on another, and averaging those hides the only thing a reader wants to know. Coverage is a property of the pair rather than of the provider, so the rating has to be too.
The engine already works this way internally. An arm is a source and a type together, never a source alone, which is what stops a provider's record on a pair it serves badly from polluting the pair it serves well.
What makes it hard to publish
A rating is read as a verdict, so it has to survive being read that way.
Sample size first. A provider with forty calls behind it and one with forty thousand cannot carry the same number without an interval around it. A rate reported as a single number is not a measurement, and a leaderboard of point estimates would be a ranking of who happened to get lucky in a small window.
Then the meaning of disagreement. A source that diverges from the others may be the only one that is right, and a rating treating agreement as correctness would quietly punish exactly that. Where a reference exists we can distinguish the two. Where none does, disagreement is genuinely ambiguous and a published number would be asserting otherwise.
Then the feedback loop. Ratings drive routing, routing drives traffic, and traffic drives the evidence the next rating is built from. A provider ranked low gets fewer calls and therefore fewer chances to prove it has recovered. The engine already spends a fraction of its calls exploring sources it does not currently favour, precisely so a healed source is not starved, and a public rating has to reflect that or it will freeze the order it describes.
What it would report
Delivery rate and latency at the tail rather than the mean, since the mean hides the calls that hurt. Freshness as measured from the source's own timestamps, with a null where a source does not stamp its data, because reporting fetch time as data time is the dishonesty this whole product exists to remove. Divergence against the network. And the sample behind each figure, with its interval.
Not a single composite score. Aggregating those into one number would be us inventing a metric and presenting it as a fact about the provider. Metera does not build composites and attribute them to someone else.
What providers agree to
Consent to be measured and published is a condition of joining rather than a setting. It is asked plainly during the application and cannot be skipped, because the published measurement is the product. A source that will not be measured in public cannot be listed.
Until we can publish this without it being read as a claim we cannot support, it stays unpublished. The measurement is already doing its real job, which is deciding where your calls go.