# 402 Score - Scoring Methodology (public document) Market402's credibility as a verification layer depends on the transparency of its methodology. This document is public; every change to the scoring logic is recorded here with a version and date. ## Principles 1. **Measured only.** On-chain volume and self-reported metrics are never used (both can be gamed). Every score derives from facts observed by Market402's own probes. 2. **Reproducible.** Every score can be traced back to the underlying observation records (timestamp, request, response code). 3. **Conflict of interest disclosed.** Services operated by our own operator (Mart402) are flagged `operator-provided` and scored by the same automated rules as everyone. Manual score adjustments are prohibited for all services. ## v0 - four axes (measurable with unpaid probes) | Axis | How it is measured | Points | |---|---|---| | Availability | Response success rate of weekly probes (402/200 = alive; timeout/5xx = failure), last 4 rounds | 40 | | Latency | p50 time to the 402 challenge. <500ms = full points, >5s = 0, linear in between | 20 | | Spec compliance | x402 schema of the 402 body (x402Version, accepts, payTo, amount, network) | 20 | | Price integrity | Listed price vs the live 402 challenge price; stability of the price history | 20 | - Displayed as 0-100 plus a grade (A >=85 / B >=70 / C >=50 / D <50) and the last-observed timestamp. ### v0.1 revision (2026-07-21) Spec compliance (20) split into three parts: base schema 10 + network name format 5 (v1 requires a plain name like `base`; CAIP-2 `eip155:*` is a v2 form that official v1 SDK clients cannot pay) + presence of the EIP-712 `extra` block 5. ## Probe conduct (v0.2, revised 2026-07-26) Revision note: unpaid probe frequency raised from "1-3 rounds/week" to the tiers below, to improve badge freshness and time-series density. Per-request behavior is unchanged (no payment attached, one request per probe, robots respected). - **Full universe** (every resource in the public Bazaar catalog plus self-submitted services): probed **weekly**. - **Index head** (Verified, opted-in, top-500 by score, and self-submitted services): probed **daily** (once per day per service - well below common uptime-monitor rates in this ecosystem). - Same-operator resources are probed serially, never in parallel, so providers with many endpoints see at most one concurrent request from us. - A weekly **census snapshot** of the full catalog (counts by network, method, price band; entries and exits) is retained as a public time series and feeds the monthly reports at /reports. - Paid delivery probes remain **at most monthly per service, <=$0.05 per probe**, opt-in prioritized, with a rotating monthly sample of the qualifying pool (seeded by calendar month, so the sample is reproducible), under a **$3/month hard budget cap**. - Explicit User-Agent: `Market402Probe/0.1 (+https://market402.com; verification probe)` (paid probes: `Market402PaidProbe/...`). - Unpaid probes attach no payment. A healthy x402 service returns 402 before any side effect; a design that performs side effects before the 402 is itself recorded as a compliance deduction. - Frequency is limited (1-3 rounds/week, a few requests per service). robots.txt and explicit probe opt-outs are respected; opted-out services are shown as "unverified". - Paid probes run under preset budget caps and identify themselves as test calls in the User-Agent. ## 402 Verified badge (2026-07-21) Three tiers, decided only by Market402's measurements - never by self-reporting: Criteria (v0.3, 2026-07-26 - delivery evidence required): - **verified**: 402 Score >=70, x402 spec compliant, >=5 observation rounds spanning >=14 days, price integrity >=90% (where measured), **and a successful paid delivery to the operator within the last 60 days** (we bought a listed endpoint at list price and a valid result was delivered; evidence bundle retained). - **verified-plus**: 402 Score >=85, correct v1 network name, EIP-712 extra present, price integrity >=95%, >=8 rounds spanning >=30 days, **and the operator's last 2 paid deliveries both succeeded**. - **none**: everything else (including compliant services not yet old enough or not yet delivery-proven). Notes on v0.3: - Delivery evidence is **operator-level** (we purchase one qualifying endpoint per operator per month) and **expires after 60 days** - the badge is continuously re-earned, never permanent. - We deliberately do **not** cap the number of badges. The criteria themselves bound it: a badge requires a recent real purchase, so the badge set can never exceed the operators our paid probes actually cover. - Transition: resources that held the badge under v0.2 on 2026-07-26 are exempt from the observation-span requirement until 2026-09-01 (the delivery requirement is NOT waived); they are probed with priority during the window. ### Badge criteria v0.4 (published 2026-08-11, effective 2026-09-01) Why: edge monetization products (e.g. Cloudflare's Monetization Gateway) will grow the x402 endpoint population quickly. We are hardening the badge BEFORE that influx, in a form every current holder already satisfies - tightening the gate is fair when nobody is retroactively pushed through it. 21 days of notice; effective the same day the v0.2 transition list expires. - **verified** (integrity tier): 402 Score >=80, x402 spec compliant, >=8 observation rounds spanning >=14 days, **price integrity 1.0 where measured** (any observed price mismatch in the window blocks the badge), and a successful paid delivery to the operator within the last 60 days. Principle: the honesty axes (compliance, price, delivery) must be perfect; the performance axes (latency, availability) may be imperfect - an honest but slow operator can be verified. Impact at publication: all 12 current verified operators already satisfy v0.4 (zero demotions by design). - **verified-plus** (excellence tier): 402 Score of 100, >=20 rounds spanning >=30 days, price integrity 1.0, x402 v2 declared, observed compliant from both EU and US vantages, and the operator's last 2 paid deliveries both succeeded. This is where "perfect score only" lives: performance-based exclusivity belongs in the excellence tier, not in the base trust signal. (Span threshold corrected 21 -> 30 days on 2026-08-21, before the effective date - see the correction note below.) - Anti-inflation mechanism, stated explicitly: badge population is bounded by real purchases (delivery evidence expires after 60 days and our paid probe budget has a hard monthly cap), not by listing volume. If the endpoint population grows 10x, the verified share falls - it cannot be inflated by listings. - Correction (2026-08-21, before the effective date): the 2026-08-11 publication of this section claimed that observation spans beyond 21 days were "structurally impossible" due to the results retention window, and lowered the verified-plus span requirement to 21 days on that basis. That diagnosis was wrong: the 21-day ceiling observed on 2026-08-11 was an artifact of the results file's age after an archival rotation, not a structural cap. Live data on 2026-08-21 shows spans reaching 31 days, and the v0.3 verified-plus tier has since activated naturally (6 operators). The v0.4 verified-plus span requirement is therefore restored to >=30 days, effective as originally scheduled. We publish this correction rather than silently editing history: versioned rules mean versioned mistakes. ### Purchase-testable scope (v0.4.2, published 2026-08-22, effective 2026-09-01) "The verified ones were bought with real USDC" must never claim more than our purchase-test regime actually reaches. Delivery evidence is gathered under hard budget guardrails (max $0.05 per test purchase), which creates a decoy vector: list one cheap deliverable endpoint, earn operator-level delivery evidence, and let expensive endpoints - which we can never buy - ride along as "verified". Two rules close this, both effective 2026-09-01: - **(a) Verified Endpoint requires a purchase-testable price.** An endpoint can hold Verified Endpoint status only if its observed live price is known and within the test-purchase cap ($0.05). Endpoints priced above the cap, or whose price cannot be established, show their measured score and compliance as before but never the verified claim. Declaring your price (via Bazaar listing or /submit) is therefore in your interest. - **(b) The operator emblem requires majority coverage.** The 402 Verified operator emblem (and its ranking tier) additionally requires that a **majority (>50%) of the operator's listed endpoints are Verified Endpoints**. One sentence: the emblem means "most of this catalog is in the purchase-verified class." A catalog dominated by untestable listings cannot be summarized by it - however good the testable minority looks. Threshold note: 75% was considered and rejected by impact simulation (it would strip the emblem from honest operators whose typical price sits at the cap boundary, correlating the emblem with cheapness rather than honesty; majority coverage blocks decoy structures - which sit near 0-10% - with minimal collateral). - Impact (full effective-day rehearsal on live data, 2026-08-22, all v0.4 rules combined): Verified Endpoints 1,343 -> 689; operator emblems 22 -> 11; Verified+ emblems 6 -> 4 (the two departures are driven by the published v0.4 criteria - v2 declaration and majority coverage - not by discretion). Everyone affected retains score, grade, compliance display and the honest intermediate "verified endpoints: k/n" form. Recovery paths: declare prices, add testable products, or - future option - operator-funded/owner-approved high-value test purchases may extend the testable band case by case. ### Operator emblem vs verified endpoints (v0.4.1, effective 2026-08-21) The badge is and always was **endpoint-scoped evidence**: we probe and buy specific endpoints, so verification attaches to those endpoints. What changed is the operator-level presentation, after observing operators with large listing sets where many endpoints are dead or non-compliant while a subset is genuinely verified (portfolio median grade C/D next to a "Verified" chip - a mixed signal we will not show). - **Verified Endpoint** (endpoint level, unchanged in substance): an endpoint that individually meets the verified criteria. Shown in search results and machine surfaces (index.json, verified.json) at endpoint granularity. This is the buyer's actionable signal. - **402 Verified operator emblem** (operator level): shown only when the operator holds verified endpoints AND the operator's whole-portfolio median score is grade A (>=85). Operators below that show the honest intermediate form "verified endpoints: k/n" instead, rank in the standard tier (not the Verified tier), and the embeddable SVG reads "VERIFIED ENDPOINTS" rather than "402 VERIFIED". - Rationale: an operator-level emblem is a portfolio-wide reputational summary; granting it on the strength of a subset while the median portfolio grade is C/D would claim more than we measured. No endpoint loses its verification; nothing is revoked - the aggregation is simply no longer allowed to overstate. Side effect by design: listing hygiene (removing dead endpoints) now directly improves an operator's standing. - At publication this affects 3 operators (portfolio medians 40.0-53.4 with 1-128 verified endpoints each); all 12 grade-A verified operators are unaffected. Distribution: the index page badge column, `GET /badge/{domain}` (live verification JSON), `GET /badge/{domain}.svg` (embeddable image), and `GET /verified.json` (an agent-facing allowlist). ### Displaying the badge (verified operators) Embed the badge on your site so visitors can verify your status with one click: ```html 402 Verified - measured by Market402 ``` Replace YOUR-DOMAIN with your operator domain (e.g. `example.com`). The image is generated live from the current scoring run: if the badge lapses (delivery evidence expires after 60 days, or checks fail), the URL returns 404 and the image stops rendering - the badge cannot be displayed stale. The link target returns the machine-readable verification (badge tier, score, methodology), so anyone can audit the claim. Badge benefit (2026-07-26): Verified holders are **guaranteed a monthly paid delivery probe** - their listing carries measured "we actually bought it" evidence at no cost to them. The benefit is earned by measurement; it cannot be purchased. Monetization policy: listing is free (neutrality). Any future paid offering will be of the form "guaranteed monthly re-verification" or "featured placement" - the badge itself is never for sale, because measured trust is the product. ## v1 addendum: the delivery axis (paid probes - rules published 2026-07-23) **Rules first, operation second.** We publish the measuring stick before we start using it. ### Definition - **delivery**: for requests actually paid via x402, the rate at which (a) HTTP 200 is returned and (b) the response is valid for the service's category (non-empty, type-consistent, passes hallucination checks). - Measurement: at most monthly per service, <=$0.05 per probe, read-only idempotent endpoints only. An evidence bundle (payment tx, receipt id, response hash, latency) is retained so every result is auditable. ### Scoring integration (v1 weights - only for services with delivery data) | Axis | v0 | v1 | |---|---|---| | Availability | 40 | 30 | | Latency | 20 | 15 | | Spec compliance | 20 | 20 | | Price integrity | 20 | 15 | | **delivery** | - | **20** | - **Scope (v1.1, revised 2026-07-26).** A paid probe is an ordinary customer purchase of a publicly listed API at its listed price. Like unpaid probes, paid probes therefore cover **every qualifying listed service** (read-only GET, <=$0.05, idempotent, valid input constructible from the listed schema); an explicit probe opt-out is always respected. Coverage tiers: - **Opted-in sellers** (`paid_probe_optin: true` via `POST /submit`, roster at `/optin.json`): guaranteed every month, probed first. The guarantee applies to **qualifying** endpoints; the `/submit` response reports your qualification status with a machine-readable reason and fix hint. If your listed schema is not probeable (non-GET, or placeholder example values), you may declare a safe `sample_input` ({method, queryParams, body}) in your opt-in - we will probe with your declared request instead. The per-probe price cap and the side-effect screen are never waived by declaration. - **402 Verified holders**: probed at the highest cadence the paid-probe budget supports - currently every month - as a badge benefit, earned by measurement and never for sale. (If the verified set ever outgrows monthly capacity, the cadence stated here will be updated rather than silently degraded.) - **All other qualifying services**: monthly rotation covering the full qualifying pool, processed in small daily batches to keep load polite. When the delivery axis is published, the observation count n is always shown next to it. (v1.0 said probes ran only for opted-in sellers; superseded because buying a listed public API at list price requires no special consent, and full coverage makes the axis fair rather than selective.) - **Disqualification.** Two consecutive observations of "payment settled, result not delivered" set delivery to 0 and revoke the 402 Verified badge. The listing is annotated with the observed facts (a record, not an accusation). The annotation is removed once a re-verification confirms recovery. - **Dispute procedure.** On a non-delivery observation: preserve the evidence bundle -> machine-readable notification to the service -> re-probe after 48 hours -> if still failing, apply the rule above. The procedure is published here precisely so that it cannot be applied arbitrarily. ### Self-assessment vs measurement (important) Passing the free self-test at `market402.com/selftest`, or the `instant_check` returned by `POST /submit`, has **no effect** on the delivery axis or the 402 Verified badge. Badges and scores derive solely from Market402's own scheduled probes. ## Operator-provided listings (v0.3.1, added 2026-07-26) Market402 is operated by the same company that runs Mart402 (mart402.com and its sandbox mart402.dev). Policy for those "operator-provided" listings: 1. **Same rules.** Operator-provided endpoints are discovered, probed, scored and ranked by the same automated pipeline as every other listing. No special treatment exists in the code, in either direction. 2. **Machine-readable disclosure.** Every operator-provided row carries `"operator_provided": true` in /index.json, /operators.json and /search.json, and an "operator" chip on human pages. 3. **No self-awarded badge.** 402 Verified requires delivery evidence from real paid purchases. For operator-provided services that evidence would be circular (the operator would be paying itself), so operator-provided listings are **not eligible** for the Verified badge regardless of their measured numbers. Their scores remain published, and anyone can reproduce them by probing the endpoints directly. A consequence: operator-provided rows never benefit from verified-first ranking. Reason for the change (2026-07-26): buyer-side experiments showed agents rely on this index when selecting services. Completeness requires listing our own endpoints; honesty requires that the one circular evidence path (self-paid delivery) can never mint a badge. This rule will be revisited if third-party attestation of delivery (e.g. on-chain attestations by independent verifiers) becomes available. ## Price integrity sources (v0.4, added 2026-07-26) Price integrity compares the amount actually demanded in the live 402 challenge against a published reference price. Reference sources, in priority order: 1. The listing's own `accepts[].amount` as published in the x402 Bazaar (unchanged - this has been the source since v0). 2. NEW: a price the operator declares at submission time. `POST /submit` now accepts an optional `declared_price_usd` (USD per call, the base price an unpaid probe should be quoted). The latest declaration per resource is used for every subsequent probe round. A mismatch against either source lowers the score in exactly the same way. Declaring a price is optional; resources with no reference price from either source keep `price_integrity: null` (unmeasured, 0 of 20 points - see "Score axes" below). Changelog reason (2026-07-26): self-submitted listings previously had no measurable reference price, which structurally capped them at 80 points. Declared prices give every self-submitted operator the same published-vs-actual honesty measurement that Bazaar-listed operators already face. ## Price integrity v0.4.1: freshness and price epochs (2026-08-15) Two defects surfaced when an operator legitimately repriced (discovered via mart402.com's own 2026-08-10 repricing; both rules apply identically to every operator in the index): 1. FRESHNESS. The Bazaar listing amount could be a stale crawl that froze a superseded price, and it always took precedence over a newer operator declaration - so an honest repricing was recorded as a daily mismatch forever. New rule: the reference price is the MOST RECENTLY UPDATED of (a) the Bazaar listing amount, timestamped by the listing's lastUpdated, and (b) the operator's declared price, timestamped by its submission. Probe records carry listed_source so the choice is auditable. 2. PRICE EPOCHS. price_integrity was a lifetime average over all rounds, so one legitimate price change made a perfect score (and the v0.4 verified-plus requirement of price_integrity == 1.0) permanently unreachable. New rule: for resources with declared prices, integrity is computed over the CURRENT PRICE EPOCH - rounds observed since the most recent declaration whose VALUE differs from the previous declaration. Within the epoch each round is judged by comparing the live 402 amount observed in that round against the epoch's declared amount (stored live_amount, re-evaluated at scoring time - this also retroactively corrects rounds that were recorded against a stale reference). Re-declaring the SAME value does not start a new epoch, so mismatch history cannot be washed away by repeated declarations. Resources with no declarations keep the lifetime calculation unchanged. Incentive note: an operator who changes the live price without declaring it keeps accumulating mismatches until they publish the change - exactly the behavior the index wants to reward: say your price out loud, then charge it. ## Self-submission hygiene v0.4.2 (2026-09-01) A submission flood (9,500+ accepted posts from 2 projects re-submitting the same endpoints in loops, 2026-08-24..31) surfaced two ingestion rules. Both were designed to leave every good-faith use untouched: 1. IDEMPOTENT SUBMISSIONS. A POST /submit identical to one accepted for the same resource within the last 24h (same declared price, same sample_input) still runs the instant self-check and returns full feedback - the developer iteration loop is preserved exactly - but is not persisted again ("already_listed": true in the response). A submission that CHANGES the declared price or sample_input is always persisted (price epochs depend on declared changes being recorded). 2. SELF-SUBMISSION PROBE CAP. Free scheduled probes cover at most 20 self-submitted resources per domain (oldest submissions first, so an operator's canonical catalog wins over later floods). Resources listed in the x402 Bazaar are exempt - the full-universe census stays complete. This bounds the index's free-probe budget against submission flooding without touching any Bazaar-listed operator. ## Score axes visibility (v0.4, added 2026-07-26) Every resource row in /index.json now carries `score_axes`: `{"measured": [...], "unmeasured": [...]}` over the four axes (availability, latency, spec_compliance, price_integrity). An axis in `unmeasured` contributed 0 points because it could not be measured - not because the service failed it. Agents comparing scores should prefer rows whose axes are fully measured, or compare only over shared measured axes. ## Delivery watch (v1.2, added 2026-07-26) A paid probe that settles payment but receives no valid result ("paid, not delivered") is the single worst buyer outcome, so it gets its own flag, separate from the score (scores remain pure measurements; we do not mix penalty points into them). - Entry: an operator has at least one paid-not-delivered event in the last 60 days, with no successful paid delivery after it. - Cure (automatic): a later successful paid delivery, or 60 days elapsing, or the operator fixes the endpoint and a later paid probe succeeds. Entry and cure are both decided only by recorded probe evidence - there is no manual judgement in either direction. - Effects: 1. `delivery_watch: true` (with evidence: date, resource, amount) on the operator in /operators.json and on each of its rows in /index.json and /search.json. 2. Ranking: pages sort Verified > clean > delivery-watch. A watch-flagged operator never outranks a clean operator regardless of score. 3. Badge: watch implies no current successful delivery, so the operator is ineligible for 402 Verified until cured (this restates the v0.3 delivery requirement explicitly). ## Delivery watch v1.3: settlement-verified "paid" + two classes (added 2026-07-31) Reason for the change: on-chain reconciliation of our July 2026 paid probes showed that a payment-response header alone does not prove settlement. Of the events flagged "paid, not delivered" under v1.2, most had no matching USDC transfer on Base at all: the probe handed over a signed payment authorization (x402 exact scheme) but the operator never executed it. Treating those as "paid" was overstated. v1.3 makes "paid" evidence-based and splits the watch into two classes. Entry and cure remain fully automatic (recorded evidence only, no manual judgement). - **"Paid" definition (tightened)**: a probe counts as paid only when its X-PAYMENT-RESPONSE decodes to success with a transaction hash. The hash is stored in the public spend ledger, so every "paid" claim we publish can be verified on-chain by anyone. A payment header without verifiable settlement is recorded as `payment_attempted`, never as paid. - **Class 1 - paid-not-delivered**: settlement confirmed on-chain, no valid result, no later successful paid delivery. (v1.2 semantics with stricter evidence.) A refund observed on-chain is annotated (`refunded: true`) but is NOT a cure: the buyer got the money back but never the service. Refunding is still better buyer treatment than keeping the funds, and the annotation makes that visible. - **Class 2 - payment-broken** (new): the probe handed over a signed payment authorization but the operator never settled it, and no later paid delivery succeeded. No funds move, but the purchase path is non-functional - for an agent that is worse than being untested. Cure: a later settled, delivered paid probe. - Effects: both classes carry `delivery_watch: true` plus `watch_class: "paid-not-delivered" | "payment-broken"` with evidence, and both sort in the bottom tier exactly as v1.2. From v1.3 this ordering also applies to the `resources` array of /index.json (previously operators and /search only), so machine readers that take the array order as ranking see the same rule as the list pages. - Historical records (before 2026-07-31) predate settlement-hash recording; they are classified via an on-chain reconciliation sidecar (ledger entries matched against the probe wallet's public USDC transfers). Original probe records are never rewritten. ## Delivery watch: path-template exclusion (clarified 2026-08-03) Path-template resources (URLs containing `:param` or `{param}`) are already excluded from scoring as unprobeable (v1.3). The same exclusion applies to the delivery watch: if our paid-probe harness ever pays such a resource by mistake, the non-delivery is our measurement artifact - an unfilled template is not valid input - and it is not evidence against the operator. The spend still appears in our published spending ledger; only the operator flag is excluded. (Clarified after one real case: $0.01 paid to a `/:domain` template that returned 400. The harness now refuses to pay template URLs at all.) ## Ranking tiers and tie-breaks (v1.3, added 2026-07-26) Updated 2026-08-21 (v0.4.1): tier placement for badge holders additionally requires the operator emblem qualification (grade-A portfolio); operators with verified endpoints but a sub-A portfolio rank in the standard tier. Updated 2026-08-22: **Verified+ ranks above Verified as its own top tier** (the excellence tier leads the ordering, at both operator and endpoint granularity). List pages (/ and /search) and /operators.json rank operators in tiers: 1. **402 Verified** (badge holders; Verified+ first, then Verified). 2. **Operator-provided, criteria met**: operator-provided listings (see "Operator-provided listings") are never awarded the badge, but the same automated badge algorithm is still evaluated on them, shadow-fashion, with no other differences. If it passes, the listing ranks directly below all badge holders and above listings without the badge, marked `verified_criteria_met` (machine-readable) instead of a badge. The delivery evidence used for this evaluation comes from the operator's own disclosed calibration purchases (real on-chain settlements, published in the spend ledger), which is exactly why it earns a rank position but never the badge. 3. **Clean**: everything else. 4. **Delivery watch** (see "Delivery watch (v1.2)"). Within a tier: score median, then number of verified endpoints, then number of endpoints. The verified-endpoint tie-break (new in v1.3) keeps an operator with 7 verified endpoints above one with 1 verified endpoint at equal score; total endpoint breadth alone no longer wins ties in the Verified tier. Transitional badges: badges whose observation-period requirement is satisfied only by the published 2026-09-01 transition roster are marked `badge_transitional: true` and annotated "transitional until 2026-09-01" on list pages. The delivery requirement is never waived. Path-template listings: some catalogs list route templates (URLs containing `:param` or `{param}` segments) rather than callable URLs. Probing them can only produce artifacts (typically 404), so from v1.3 they are marked `unprobeable: "path_template"` and excluded from operator score aggregates (median, range). They remain visible as endpoints. ## Probe vantages (v1.4, added 2026-07-27) Market402 probes from more than one network location ("vantage") to report latency as agents in different regions would experience it. Vantages: - **eu-central** (Falkenstein, DE): the primary and historical vantage. Every score and every historical observation is measured from here. - **us-east** (Ashburn, US): a secondary vantage added 2026-07-27, head-probe only (no payment), for supplementary US-perspective latency. Rules that protect the time-series (the point of a verified index): 1. **Scoring uses only the eu-central vantage.** The 402 Score, badges, and all historical continuity derive from eu-central observations exclusively. Adding a second vantage never changes a score. This keeps the time-series unbroken. 2. **US latency is supplementary, not scored.** Resources carry an optional `latency_p50_ms_us` field alongside `latency_p50_ms` (which stays eu-central). Agents can compare US-vs-EU latency; neither is a scoring input beyond the existing eu-central latency axis. 3. **Every observation records its `vantage`.** Machine-readable, so the source of each latency figure is auditable. 4. If the eu-central box is ever lost and the index is promoted to us-east (see operational recovery), the vantage change is announced here with a date, because a verifier that quietly changes how it measures is not a verifier. ## Score history (v1.5, added 2026-07-31) Why: agents cannot taste-test stability. A current score answers "is it good now"; history answers "will it stay good". We publish both the derived verdicts and the raw series they were derived from, so the verdicts are verifiable. Data: - One snapshot per ISO week per operator: the operator's `score_median` and `badge_best` at the first scoring run of that week (normally Monday's full census; the census that scores are computed from). Later runs in the same week never overwrite a recorded snapshot. - Snapshots are retained permanently. The published window is the most recent 12 weeks (matching the 60-day delivery-watch and badge-expiry constants). - Backfill note: weeks 2026-W30 and 2026-W31 were seeded from retained weekly backups when this section was added; earlier weekly scores were not archived per-operator and are not reconstructed (we publish what was measured, not recomputed). Published at `GET /op/{operator}.json`: - `series`: `[["2026-W30", 95.0], ...]` - the raw weekly points (score_median). - `score_now`, `median_12w`, `min_12w`, `max_12w`, `observations`. - `trend` (derived): least-squares slope of score over the window, in points per week. `improving` if slope > +0.5, `degrading` if slope < -0.5, else `stable`. Requires >= 4 weekly observations; otherwise `insufficient_history` (we do not guess). - `volatility` (derived): population standard deviation of the window. `low` < 2.0, `medium` < 5.0, else `high`. Requires >= 4 observations, otherwise `null`. Thresholds are part of these rules; changing them requires a version bump here. The human page (`/op/{operator}`) renders the same series as an inline sparkline and a table - same data, no separate source. ## Network naming and protocol version (v1.6, updated 2026-07-31) The original naming check (2026-07-21) awarded the naming point only to v1-style plain names ("base"), because hybrid responses - a v1 envelope carrying CAIP-2 identifiers - broke payment SDKs at the time (measured on our own shop). The ecosystem has since converged on x402 version 2, where CAIP-2 identifiers (eip155:8453) are the standard; a rule frozen in the v1 era would now penalize the majority for being current. From v1.6 the check is **consistency between the declared protocol version and the network naming form**: - `x402Version: 1` with a plain network name ("base", "base-sepolia"): pass - `x402Version: 2` with a CAIP-2 identifier ("eip155:8453", "solana:..."): pass - Mixed forms (v1 envelope + CAIP-2, or v2 envelope + plain name): fail - this is the original failure mode that made services unpayable. Effect: services correctly declaring v2 with CAIP-2 now receive the naming point they were previously denied; scores may step up at their next round. For machine compatibility the observation field keeps its historical name (`spec_network_named`) and the monthly report keeps `v1_network_name_ok_rate`; both now measure version-consistent naming as defined here. ## Extended spec observations (v1.7, added 2026-07-31) We cross-referenced our checks against public x402 conformance checklists (notably x402-list's 14-point list) and adopted the axes we lacked that bear directly on whether an agent can actually pay. From v1.7 every probed 402 records six additional observations, published per resource in index.json: - `x402_version`: the protocol version the service declares (census statistic; the monthly report publishes the v2 adoption rate). - `spec_payto_valid`: `payTo` parses as a valid address for the network family (EVM: 0x + 40 hex; Solana: base58). A malformed payTo is unpayable. - `spec_asset_declared`: an asset identifier is present. - `spec_scheme_declared`: a payment scheme is declared. - `spec_v2_header`: for services declaring v2, whether the 402 carries a decodable `PAYMENT-REQUIRED` header with accepts[]. The header is v2's primary transport - current SDKs treat the body as a v1 fallback only - so a v2 declaration without the header is unpayable by standard clients (failure mode measured on our own shop during its v2 migration). - `spec_https`: the resource is served over HTTPS. Scoring: these are **observations only in v1.7** - they do not move the 402 Score yet. Following our practice of validating new judgments against a known answer set before they affect scores, they will be considered for scoring in a future version bump after at least one full weekly census of measurement. Deliberately not adopted: an absolute "price sanity range". Our price integrity axis cross-checks live prices against listed and self-declared prices, which catches manipulation without penalizing legitimately premium services by an arbitrary threshold. Observation fields added 2026-08-11 (observation only, no scoring change), in response to Cloudflare's Monetization Gateway announcement and the Agents SDK x402 documentation (upto scheme, multi-network support): - scheme_upto_offered: whether any accepts[] entry declares the "upto" (max-authorization, metered) scheme - networks_offered: the set of networks declared across accepts[] (up to 6) Purpose: measure how fast edge-gated monetization and non-Base networks spread through the 402 economy. Any scoring implication follows the same rule as all v1.7 observations: at least one full weekly census of measurement first, then a versioned decision. ## Traction observation (added 2026-07-31) Market402 records inbound USDC transfers to the payTo addresses that services publicly declare in their 402 responses (both are public: the declaration and the ledger). Rules, stated up front: 1. **Traction never enters the 402 Score, badges or rankings.** On-chain volume can be self-dealt, which is exactly why this index scores measured behavior instead (see the front page: "No on-chain volume - it can be gamed"). Collecting the data does not change that promise. 2. **Published externally: monthly market aggregates only** - total observed volume, active payment addresses, median settlement, growth. Per-provider figures are not published. 3. **Figures are an upper bound, not verified revenue**: gross inbound to a declared address may include non-x402 activity, and shared addresses are flagged as ambiguous. Aggregates always carry this caveat. 4. Collection window starts 2025-05 (the x402 protocol's public launch); transfers before the protocol existed are by definition not x402 and are not collected. ## Data license (added 2026-07-31) All data Market402 publishes (index.json, operators.json, search results, badge JSON, monthly reports, and the published score-history window) is licensed under Creative Commons Attribution 4.0 (CC BY 4.0): https://creativecommons.org/licenses/by/4.0/ You may copy, redistribute and build on it, including commercially, provided you attribute "Market402 (market402.com)" with a link. Note what the license does and does not cover: - It covers the published snapshots. Market402 retains the full measurement history internally; the published score-history window is the most recent 12 weeks. A copy of today's data ages; the live index does not. - It does not make copied data "verified by Market402". Verification claims (badges, scores) are only valid when served live from market402.com - a badge image 404s when verification lapses, and copies obviously cannot.