Transparency report
How Motus Engine scores — and where it is not ranking-safe yet.
Every figure on this page is computed live from the scoring corpus. It contains no athlete names, no video and no personal data — only aggregate statistics.
engine motus-1.2.0 · generated 2026-08-25 01:52 UTC
30
Runs scored
1
Countries
2
Athletes
2
Forms covered
Reliability
0 runs carry an official judge label. The question that matters is not whether the engine is perfect, but whether it disagrees with a judge less than two judges disagree with each other.
Engine vs judge (MAE)
—
Average absolute gap between the engine total and the official judge total.
Engine vs judge (r)
—
Correlation across labelled runs. Higher means the engine ranks athletes like judges do.
Judge vs judge (MAE)
—
Human baseline from 0 judge pairs on the same run.
Provisional runs
43.3%
Runs whose capture quality is below the trust gate; these are flagged, never silently scored.
Fairness status
Unaudited
No judge-labelled runs yet on this corpus. Scores are engine estimates only and must not be used for placement.
Internally the engine is audited for score differences across gender, age division and country using Welch tests with Benjamini–Hochberg correction. A dimension only becomes ranking-safe once no significant unexplained gap remains.
Audit trail
- Every score is versioned: 10 score records are stored append-only with the engine version and scoring configuration that produced them.
- Official results can be frozen — 0 runs are locked so no later re-score can overwrite them.
- Capture quality is measured per run; low-quality runs are marked provisional rather than scored silently.
- Re-scores never delete history; the previous value stays readable for any audit.
How the score is computed
- 1. Pose capture. 17 body keypoints are tracked per frame from one or more camera seats, with per-keypoint confidence.
- 2. Quality gate. Confidence, jitter and occlusion form a data-quality index. Below the threshold the run is provisional.
- 3. Movement segmentation. The sequence is split into techniques using limb elevation and velocity, then matched to the form.
- 4. Scoring. Accuracy and presentation components are computed from joint angles, balance, symmetry and rhythm under a pinned scoring configuration.
- 5. Panel consensus. With multiple camera seats, unusable seats are excluded and the remaining seats are combined, with spread reported alongside the total.
Questions about methodology or an audit request: support@olympico.app
