CardFax

CardFax Research · Precision Series · No. 001

The Repeatable Instrument

Can AI card grading tell a 9.1 from a 9.8? We rescanned the same graded cards 25 times, in 25 different positions, and measured the answer.

CF-PS-2026-001v1.0October 4, 2026MOSES 2.18

A decimal grade is a measurement claim. Telling a 9.3 from a 9.7 only means something if the instrument, handed the same card twice, says the same thing twice. So we tested exactly that — with the thresholds published before the data existed, and every miss reported next to every pass. Five registry-locked cards. Five rescans each, deliberately shifted, slid, and rotated. One hundred twenty-three measurements. One accidental drop that became the most useful part of the study.

Section · 01

Why repeatability comes before accuracy

Every grading company publishes grades. None publish the precision of the instrument producing them. Yet precision decides what a decimal means: if a grader's reading of the same card wobbles by ±0.4 between scans, "9.3 versus 9.7" is a coin flip wearing a decimal point. If it wobbles by ±0.03, the distinction is a measurement. Industrial metrology settles this with a repeatability study — the same object, measured repeatedly under varied conditions — and the spread of the results, the standard deviation σ, is the instrument's noise floor. You cannot resolve differences smaller than your own noise. This study establishes that floor for MOSES, axis by axis.

Pre-registered before data collection

Hypothesis: per axis, σ across five varied-placement rescans of the same card is at most 0.10. Failure bar: σ of 0.25 or more means fractional claims are not defensible on that axis. Acceptance: every axis mean σ ≤ 0.10 and no card above 0.20. These numbers were fixed in the registered protocol before a single scan was processed, so the goalposts provably never moved.

Section · 02

The instrument and the cards

MOSES grades each axis with independent machinery — measurements first, deductions from the measured math, never a model guessing a final number from a photo. Corners run layered analysis per corner: a trained detector, a reference-library match, an AI examiner for whitening and fray, a silhouette layer that traces the physical outline to the degree, a lift scan for bends and curls, and reconciliation rules. Engines were pinned for the study at MOSES 2.18.0 and Edge Engine 3.1.0; every read logged its layer outputs to an append-only record, which is what makes the analysis below auditable.

CardSerialGradeC / K / E / SWhy it is in the study
Shohei Ohtani — Topps Now 2026220737889.9 / 9.4 / 8.0 / 10Matte; edge-limited 8
Luka Dončić — Donruss Optic 20250199483910 / 9.5 / 9.4 / 10Holographic front; two borderline axes
Jayden Daniels — Panini Instant 2025889636689.8 / 9.3 / 9.8 / 8.8Dark-art front; lowest corner in the set
Aaron Donald — Topps Now 2026225688299.8 / 9.4 / 9.8 / 9.4Matte control
Jared McCain — Donruss Optic 2025350280199.6 / 9.4 / 9.5 / 10Second holographic front

Five rescans per card, front and back, production scanner at production settings. Placement varied by protocol each round: centered, shifted inward, shifted outward, visibly rotated 5–10°, then free placement. All records registry-locked: the published grades cannot change during the study.

Section · 03

Result one — corners repeat within hundredths

Chart of corner subgrades across five rescans per card: four cards cluster within a few hundredths of a point; one holographic card shows a single rogue 8.0 read
Figure 1. Corner subgrade, five rescans per card. Four cards cluster within a few hundredths no matter how they sit on the glass. The red dot is a single rogue read — examined in Section 04.

On matte and dark-art stock the corner engine behaved like a measurement instrument: σ between 0.019 and 0.036 across Ohtani, Donald, Daniels and McCain — a pass against the 0.10 bar with three times the margin. The set was not easy: it includes two 8s, corners sitting at the 9.3–9.4 borderline, and one of the two holographic cards (McCain) posted the tightest spread of the entire study at σ 0.019. Precision near ±0.03 means consecutive fractional grades, 0.1 apart, are separated by more than three noise-widths — and a 9.1-versus-9.8 comparison by more than twenty. For corner condition on standard stock, decimal distinction is not an aspiration. As of this study it is a measured property of the system.

Section · 04

Result two — anatomy of a rogue read

Luka's third round produced the study's one large error, and the stored layer outputs let us dissect it exactly. The front bottom-left corner — a busy holographic region — read 8.5 on that round, against 9.1 / 9.5 / 9.5 on its other rounds. One read, a full point below its own neighborhood. The grading math then did what it is designed to do with genuine damage: a single corner below 9.0 floors the whole corner subgrade, so one 8.5 pulled a card averaging 9.5 down to a run-level 8.0.

The amplification arithmetic: an examiner error of −1.0 on one of eight corners became −1.5 on the subgrade, because crossing the 9.0 boundary engages the floor rule. The floor rule is correct — it is what stops a crushed corner from hiding inside an average — but it converts rare read noise into rare large errors. Measured frequency: one rogue read in roughly 200 corner reads, confined to busy holographic regions. The other holographic card never produced one.

Section · 05

Result three — edges flip with placement

Chart of edge subgrades across five rescans per card: readings split between the 9.3 level and the 9.8 level rather than scattering
Figure 2. Edge subgrade, five rescans per card. The readings do not scatter — they bifurcate between two levels.

Edges failed both pre-registered bars (mean σ 0.206, worst 0.271) — and the shape of the failure is the finding. Edge readings are quantized: the engine produces discrete levels, and a placement change alone flips which level a borderline edge lands on. Every card exhibited the flip; none scattered randomly. A placement change can move an edge subgrade by 0.5. Because all 25 runs stored their per-edge layer outputs, the diagnosis set already exists, and this exact study re-runs afterward as the before-and-after proof. Until then, edge decimals should be read as band-resolution, not fractional-resolution.

Section · 06

The drop — an accidental blind test

Between rounds four and five, the study owner dropped the Luka card while rearranging the scanner bed, bending a corner. Instead of contaminating the study, this created something a repeatability protocol normally cannot have: a genuine physical change, timestamped mid-series, with the location deliberately withheld from the analysis. The protocol was amended on the record before analysis: Luka's stability statistics use rounds one through four; round five became a blind sensitivity probe.

Round five arrived doubly scrambled — the cards had been rearranged across bed quadrants, and the Luka card sat rotated 180°. Both were detected from image evidence alone, by matching every extracted card against its stored registry scan in both orientations. An initial corner-level call made before the rotation was discovered was withdrawn, on the record. Then the corrected numbers said this:

Physical cornerBaseline (rounds 1–4)Post-dropShift
Top-left (front view) — front face9.509.4−0.10
Same physical corner — back face9.389.3−0.08
All six other corner-faces——+0.02 … +0.25
Four control cards, corner subgrade——within ±0.07

The owner then confirmed in hand: a bent top corner on the back — a bend with no crease. The engine had localized the change: the only corner to drop, independently on both faces of the card, in an otherwise stable session. But it priced the bend at roughly a tenth of a point, and the owner's grading standard, stated for the record during this study, is that a bent corner, creased or not, must never grade in the 9 range. Detection: yes. Severity: wrong by about two grade bands on this finish.

The mechanism, from the engine's own telemetry: the corner-lift layer — the component built to catch bends — reported for every corner of this card: not measurable: too little bare card stock at the tip (dark/printed corner). And the examiner's primary damage cues, whitening and fray, are nearly invisible on holographic stock, where a bend does not whiten. The one defect this card suffered fell exactly between two layers' blind spots.

Section · 07

What changed the same day

  • MOSES 2.18.1 — lift telemetry. Every production corner read now permanently records whether the bend detector could measure, alongside the outline verdict. Within weeks, the blind-spot rate by finish is a measured number rather than an estimate, and the bend-detection fix will be proposed with data attached. Grading behavior: unchanged, verified by re-reading study crops against their baselines.
  • The grade evidence chain. Every grade event on the registry now appends an immutable, hash-chained evidence record — each entry carries the fingerprint of the one before it, back to a genesis record, and the ledger is append-only at the database privilege layer. Editing grading history would visibly break every link that follows.
  • The canary set. These five cards and their measured tolerances become a permanent regression gate: future engine versions must reproduce them within tolerance before activation. Drift stops being a risk managed by vigilance and becomes a property blocked by process.

Section · 08

Limitations

  • Five cards, one scanner. Sufficient for an instrument noise floor; not for population statistics. The main study (~25 cards, weighted toward holographic fronts and sub-9 corners) addresses breadth.
  • Centering was uninformative here: the standalone centering engine scored every study card 10.0 on every run — a ceiling effect on near-centered cards, not evidence of precision. The ratio-based decimal centering path joins the next study.
  • Surface was deferred pending a read-only wrapper for the etched-map engine.
  • Repeatability is not accuracy. This study says the instrument returns the same number; the accuracy and discrimination phases — against expert-confirmed boundary cards — say whether it is the right number. Those are the next phases of the validation program.

Section · 09

Raw axis values

CardCorners, rounds 1–5Edges, rounds 1–5
Ohtani 22073789.49 · 9.46 · (fail-closed) · 9.51 · 9.569.3 · 9.3 · 9.8 · 9.9 · 9.3
Luka 01994839.45 · 9.48 · 8.0 · 9.55 ‖ 9.59.3 · 9.3 · 9.3 · 9.8 ‖ 9.8
Daniels 88963669.56 · 9.53 · 9.55 · 9.47 · 9.549.3 · 9.8 · 9.8 · 9.3 · 9.3
Donald 22568829.58 · 9.55 · 9.62 · 9.60 · 9.609.3 · 9.3 · 9.8 · 9.3 · 9.3
McCain 35028019.49 · 9.46 · 9.45 · 9.47 · 9.509.6 · 9.8 · 9.8 · 9.9 · 9.8

‖ separates the pre-drop baseline from the post-drop round (documented material-change event). One corner run failed closed when the vision provider errored after retry and fallback — the engine refused to invent a grade; the run is recorded, not imputed. One round's card rearrangement and one 180° placement were detected and corrected from image evidence; the incident log lives in the registered protocol and results documents.

Section · 10

Conclusions

  1. On standard card stock, MOSES's corner measurement repeats at σ 0.019–0.036 under deliberately adversarial placement — the precision a fractional-grade claim requires, demonstrated under pre-registered acceptance criteria.
  2. The two failures are localized and instrumented: a 0.5% holographic rogue-read rate amplified by the sub-9.0 floor, and a placement-driven edge-level flip of 0.5 — each with a named engineering path and a built-in before-and-after test.
  3. An accidental blind test exposed a structural blind spot — bends on dark printed corner stock — and the telemetry to size it shipped the same day.
  4. The deeper claim is not that the engine is perfect. It is that the grading system measures itself, publishes the measurement, and is structurally prevented from drifting — pinned versions, canary cards, an append-only evidence chain, and corrections made in the open.

Frequently Asked Questions

Can AI card grading really tell a 9.1 from a 9.8?

For corners on standard card stock, yes — and this study is the evidence. Rescanning the same cards in different positions, MOSES reproduced its corner subgrade within ±0.02 to ±0.04. At that precision a 9.1 and a 9.8 are separated by many times the instrument's own noise. Edges are not there yet: placement alone moved edge readings by half a point, so edge decimals should be read as bands until that engine is hardened.

What does repeatability mean in card grading, and why does it matter?

Repeatability asks: if the same physical card is scanned again — shifted, rotated, placed differently — does the system return the same grade? It is the foundation under any decimal claim. A grader whose reading of the same card wobbles by half a point cannot meaningfully distinguish fractional grades, no matter how confident the label looks.

What is a pre-registered study?

The hypothesis, the cards, the pass and fail thresholds, and the analysis rules were published before any data was collected, and every correction made during analysis was recorded in the open. That means the acceptance bars provably never moved to flatter the results — which is also why the failures are in this report alongside the passes.

What did the study find wrong with the engine?

Two things, each with a mechanism. First, on busy holographic corners the AI examiner produced one rogue read in roughly 200 corner reads, and the sub-9.0 floor rule amplified that single read into a 1.5-point swing. Second, edge readings snapped between two levels (9.3 and 9.8) depending on placement alone. Both now have named engineering fixes, and this exact study re-runs afterward as the before-and-after proof.

What happened with the dropped card?

Mid-study, one card was accidentally dropped and a back corner was bent — and the damaged corner's location was deliberately withheld from the analysis, making it a blind sensitivity test. The engine flagged the correct physical corner on both faces of the card, but priced the bend at only about a tenth of a point, because the bend-measurement layer cannot run on dark printed corner tips and a bend on holographic stock shows almost no whitening. Telemetry to measure that blind spot across all production grading shipped the same day.

How is this different from how other grading companies describe accuracy?

Grading services publish grades; this study publishes the measurement precision of the instrument that produces them — with raw run values, pre-registered thresholds, incidents, and corrections included. Every grade referenced is registered on the CardFax, and every grade event in the system is now recorded in an append-only, hash-chained evidence ledger.

Every grade in this study

Measured by MOSES. Registered on the CardFax.

Grade your first card

CF-PS-2026-001 · v1.0 · October 4, 2026 · Study PILOT-1, pre-registered 2026-10-04 before data collection; engines pinned MOSES 2.18.0 / Edge Engine 3.1.0; remediation v2.18.1 deployed the same day. Measurements are published; activation thresholds are proprietary. Related reading: How AI card grading caught a damaged corner · How to grade card corners.