How We Validate

Apex makes the call first.
Then we open the lab report.

That order is the test. Reverse it and the result means nothing.

Apex Motion's assessment was developed by comparing what it found against independent 3D motion-capture and biomechanics reports it was never allowed to see first.

Every run on this page was produced from video and intake alone, locked to a non-editable record, and only then scored.

“If Apex got it wrong, the wrong answer stays on the record.”
Every miss on this page is an output we left exactly as it was generated
The Benchmark

Why this benchmark matters.

Apex is benchmarked against independent 3D motion-capture and biomechanics reports. Before the results mean anything, it's worth knowing what that class of testing actually is.

What it is

3D motion capture is a laboratory biomechanics method used to analyse movement in three dimensions. Multi-camera systems capture fast athletic movement and allow researchers and high-performance professionals to quantify movement that cannot be directly measured from a single ordinary video angle.

Why it's the right bar

This class of technology is used in elite sport, biomechanics research, high-performance training environments and professional baseball. It runs on specialised multi-camera laboratory infrastructure that most pitchers do not have routine access to — which is exactly why it makes a demanding benchmark.

How a laboratory benchmark is produced

Step 01

Multiple camera views

Synchronised cameras record the same delivery from several angles at once.

Step 02

3D reconstruction

Those views are combined into a three-dimensional model of the movement.

Step 03

Biomechanical measurement

Joint angles and segment speeds are quantified — values a single video angle cannot produce.

Step 04

Actionable report

The findings are written up as prioritised, actionable conclusions about the athlete.

So we hid that information from Apex.

FirstApex made its assessment from video and intake alone.

ThenThe findings were locked to a non-editable, timestamped record.

Only thenThe independent biomechanics benchmark was opened.

How closely could a structured video-based assessment reproduce the actionable findings contained in high-end biomechanics testing?

Side by Side

What Apex called.
What the benchmark said.

Four representative comparisons, anonymised. One clean agreement, one outright failure, one deliberate withholding, and one recovery on the frozen system.

Lead leg braking flagged as a primary limiter, from video and intake alone.

The independent biomechanics report identified a lead-leg braking deficit of its own accord.

Agreement Recovered and held on the frozen production baseline.

A large, visually dramatic torso rotation read as fast rotational output.

Measured rotational output was among that athlete's weakest features — the opposite direction.

Wrong direction Pre-revision system. This failure is what triggered the rebuild.

No rotational-velocity call made. Video cannot establish segment velocity, so the claim is blocked rather than estimated.

The report contains a directly measured value.

Withheld Scored as its own outcome — never quietly counted as a correct answer.

On the frozen system, the two athletes whose rotation had been read backwards now return “not assessable”.

Same benchmark reports, re-scored against the same pre-written rules.

Recovered Neither control regressed in the frozen confirmation suite.

Athletes are not identified, and no third-party report is reproduced. These are the comparison outcomes as they were recorded in the scoring ledger.

Where It Came From

Apex didn't start with an empty prompt.

The assessment framework was built on an existing body of knowledge first, then tested against independent measurement — not improvised and rationalised afterwards.

The evidence foundation

What informed the framework

The eight-category framework and the reasoning behind it were built from an extensive evidence library spanning pitching biomechanics, mechanics, physical preparation, movement and training — drawing on peer-reviewed research, biomechanics literature, case studies, expert educational material and applied pitching-development resources.

Those source types do not carry equal authority, and we don't treat them as if they do. Peer-reviewed biomechanics research and a coaching video are both useful inputs to how the framework reasons; neither one, on its own, licenses a claim about a specific athlete. What an individual assessment is permitted to state is decided separately — see below.

The governance layer

What Apex is allowed to claim

Governance is a different job from research. The evidence library helps Apex reason about pitching in general. Governance decides what Apex may actually say about you, given only the evidence your own submission contains.

  • Research shapes how the system thinks about a delivery.
  • Governance gates what any single assessment may assert.
  • If your footage can't support a finding, the framework's general knowledge does not fill the gap — the category returns not assessable.
How the system got to production
  1. 01Evidence foundationExisting biomechanics, training and movement literature and applied resources.
  2. 02Assessment frameworkEight categories defined, with the reasoning structure behind each one.
  3. 03Early blind testingAssessments locked before independent benchmark reports were opened.
  4. 04Documented failuresSystemic error patterns named and recorded rather than smoothed over.
  5. 05Governance rulesWhole classes of unsupported output made structurally impossible.
  6. 06Targeted revisionsNarrow, line-by-line changes aimed at the named failures.
  7. 07Frozen confirmationCriteria written and frozen first, then the system tested against them.
  8. 08Production lockConfiguration locked 2026-08-14 with cryptographic hash verification.
Method

How a comparison actually runs.

Five steps, in this order. The order is the entire test — it is what makes it impossible for the answer to have been influenced by already knowing the answer.

01 — Input

Video and context go in

Slow-motion pitching video plus the athlete's intake — strength numbers, mobility findings, injury history. Nothing from a lab.

What a coach can actually get.
02 — Apex call

Apex makes the call

The assessment is produced from that material alone. The session generating it has no access to the athlete's biomechanics report.

No peeking.
03 — Lock

The call is locked

Saved to a non-editable, timestamped record — 24 of them carrying a cryptographic file fingerprint. From this moment it cannot be changed.

The answer is sealed.
04 — Reveal

The benchmark is opened

Only now is the independent 3D motion-capture and biomechanics data revealed to the scoring session.

Now the benchmark opens.
05 — Score

Both are scored

Compared under scoring rules written down before any athlete was opened. Agreements and misses are recorded with equal weight.

Graded using rules that were already written.

What that produced

Four different kinds of proof: scale and method, cohort, mechanism, and verification. Each carries its own denominator, and the denominators are never swapped.

The Cohort

13 athletes.
58 blind assessment runs.

Athletes are counted once. Runs are counted separately. Those are different numbers and we keep them apart.

13 distinct athletes

produced

58 blind assessment runs
38 first-exposure runs — the athlete's benchmark had never been compared before 20 re-tests — regression, confirmation and baseline runs

A re-test is another run, not another person. Thirteen people produced fifty-eight runs.

24 of the 58 carry a recorded cryptographic file fingerprint. A note on wording: there are 58 runs and 57 benchmark comparisons — one blind Core v1 run was produced but never compared, because that athlete's benchmark was not supplied in the scoring session. We report runs.

The Question Everyone Asks

How close did Apex get to the benchmark?

There isn't one honest blended accuracy number — so we show the comparisons that actually matter instead. Different rounds tested different scopes, categories and scoring rules; averaging them would produce a figure that sounds precise and means nothing. What follows is the progression as it was recorded, strongest evidence first.

01Where we started

The first properly controlled test.

Seven athletes went through the system twice — once with the safety layer, once without — with the benchmark data sealed until both were locked. It was scored at the strictest possible standard, where only an exact category match counts and “close” scores zero.

Real signal, well short of good enough. That result is what triggered the rebuild.

Pre-revision system — retired 44 / 99 Exact feature-level matches, pre-revision system · 7-athlete controlled retest · 99 scoreable feature-level decisions. Not current Apex performance.
02What broke

The pattern was worse than the average.

The most-used primary archetype was confirmed by the benchmark in only 2 of 5 cases. And a specific optical illusion kept repeating: a large, visually dramatic torso rotation was read as fast rotation, when the measurement showed it was one of that athlete's worst features. Every configuration willing to state a direction got it backwards.

A confident, specific, wrong answer is worse than no answer. That is the failure mode the rebuild was aimed at.

Pre-revision system — retired 2 / 5 Most-used primary archetype, confirmed by the benchmark · pre-revision system · 5 tested instances. Not current Apex performance.
03What we rebuilt

Three classes of failure, closed at the source.

  • Amplitude read as velocity. Rotational output can no longer be inferred from how dramatic a rotation looks. Where direction cannot be established from video, the system returns “not assessable”.
  • Over-selection of a favourite archetype. The primary finding that failed most often can no longer be selected on thin evidence.
  • Claims the evidence could not support. Projected velocity, future ceilings, numeric grades and asserted root causes became structurally impossible to emit rather than discouraged.
3 Surgical revisions, each diffed line by line against the source files before the final freeze. Full amendment ledger — including the fixes that failed — is in the technical details.
04After the rebuild

What the frozen production baseline did.

The final configuration was locked with cryptographic hashes. All 13 athletes were then run through that exact configuration. On the two movement categories the rebuild targeted — the ones with the strongest benchmark detail — it recorded:

0Misses
0False positives
0Unsupported predictions

Scope: frozen production baseline · Lead Leg Braking + Rotational Output · 2 of 8 formal categories · across the 13-athlete baseline. This result does not describe the other six categories.

Both lead-leg findings the benchmark had genuinely confirmed were recovered. Both athletes whose rotation reading had been backwards now return “not assessable” instead of a confident wrong answer.

05The remaining gap

Where the evidence still stops.

Only one athlete was genuinely held out from earlier system-development exposure — designated by a documented random draw, using a recorded seed, before any athlete's identity, video or benchmark data had been looked at. That athlete had a real lead-leg braking limiter, ranked fourth in their own benchmark report's development priorities. The governed system withheld rather than calling it: correct by its own evidence rules, and still a finding the athlete did not receive.

We publish this because a validation page without it isn't a validation page.

Held out 1 Athlete truly held out from previous development exposure — and that case produced a miss.
Failure → Fix → Retest

We found the failures ourselves.

None of this was discovered by a customer. Each of these is documented in the same ledger as every result above.

Wrong-direction rotational read

Amplitude is not velocity.

  • Observed. A dramatic-looking torso rotation was read as fast rotation, in 3 of 4 scored athletes.
  • Cause identified. Visual amplitude was being treated as a proxy for a quantity video cannot measure.
  • Rebuilt. Rotational handling was reconstructed; direction can no longer be asserted from amplitude.
  • Retested. No recurrence in the confirmation suite, in either athlete who originally produced it.
Unsupported claim generation

Answers the evidence could not carry.

  • Observed. Without governance, 6 of 6 unseen athletes received at least one unsupported claim.
  • Cause identified. Nothing structurally prevented the system from producing predictions it could not support.
  • Rebuilt. A fixed, hashed governance layer was introduced, blocking whole classes of output.
  • Retested. 0 of 6 with governance, in each of three independently scored configurations.
Conservative withholding

The cost of not guessing.

  • Observed. The one truly held-out athlete had a real lead-leg limiter that was withheld rather than called.
  • Cause identified. The evidence threshold that blocks confident wrong answers also blocks some correct ones.
  • Recorded, not patched. Loosening the threshold to catch it would reopen the failure mode above.
  • Measured on both sides. Correct withholding and information lost to over-caution are logged in the same ledger.
Governance

A system that knows when it isn't allowed to answer.

General-purpose AI tends to answer even when evidence is thin. Apex is designed to withhold claims when the available evidence cannot support them.

6 of 6 Without governance
0 of 6 With governance

Scope: N = 6 unseen athletes / independently scored configurations. Each of the six produced at least one unsupported claim with governance off; none did with governance on.

What governance blocks

Why it holds

  • It is fixed. The same rules run for every athlete. There is no per-athlete tuning and no way to loosen a rule because a particular result would look better.
  • It is verified. The rule set carries a cryptographic hash. Any change requires a new version, a verified diff and a fresh confirmation suite.
  • It fails closed. When evidence is insufficient the system withholds. “Not assessable” is tracked as its own outcome, so it can never be quietly confused with a correct answer.

And what it costs

  • It withholds real findings too. In at least 4 of 7 early tests, governance withheld the single highest-value correct finding in the report.
  • In one case it introduced an error that the ungoverned read had avoided.
  • We do not oversimplify this. Our own record says it plainly: governance does not always help. Both sides are logged in the same ledger.

This is reliability engineering, not fine print. A system that never says “I can't tell from this” isn't more accurate — it is guessing more often and hiding it better.

Confirmation

Final system.
Frozen before the test.

A pass rate is only as good as the protocol that produced it. The protocol comes first here, because it is the reason the score counts.

Control 01

System frozen

The candidate prompt and governance rules were locked and cryptographically hashed before the first case ran.

Control 02

Pass criteria written first

Each of the nine cases had its success condition defined in advance, including which behaviours were not allowed to regress.

Control 03

No edits between cases

A failure in case 3 could not be patched before case 4. That is the difference between a test and a rehearsal.

Control 04

No scoring until all cases locked

Results were withheld until every case was complete, so no interim result could influence how a later one was read.

9 / 9 Final confirmation cases passed

Confirmation cases, not nine unique production athletes. These are re-tests of athletes already in the programme, using their original blind inputs. We do not convert 9 of 9 into a percentage.

3 cases False positives on the lead-leg finding are gone
2 cases Genuinely correct lead-leg calls recovered, not destroyed
2 cases The rotation-reading failure, targeted directly
2 cases Controls that simply had to not get worse

All nine passed. Both target failure modes were fixed. Neither control regressed. Zero prohibited outputs across all nine. It is confirmation evidence — proof the fixes hold — not proof the system generalises to a large new population.

Category Scope

Two of eight categories were formally scored.

A category in the Apex framework is not the same thing as a category included in a frozen scoring pass. We are not going to show you eight scores when we formally scored two.

CodeCategoryFormal category scoringFinding-level evidence
LLB Lead Leg Braking Formally category-scored Finding-level evidence
RO Rotational Output Formally category-scored Finding-level evidence
HIR Hip Internal Rotation Not formally category-scored Finding-level evidence
FA Force Acceptance Not formally category-scored Finding-level evidence
FT Force Transfer Not formally category-scored Finding-level evidence
HSS Hip-Shoulder Separation Not formally category-scored Finding-level evidence
AA Arm Action Not formally category-scored Finding-level evidence
AO Arm Output Not formally category-scored Finding-level evidence
Formally category-scored in the frozen pass — 2 of 8 Finding-level supporting evidence only Not formally category-scored

The other six categories are live parts of the assessment framework and carry finding-level evidence. They did not receive frozen, category-level scoring, so nothing on this page claims they did.

Human Review

Where people enter the process.

Apex Motion is a governed system, not an unsupervised one. It matters where the humans are — and just as much, where they are not.

Where human review applies

  • In the production workflow, when governance flags a serious or major issue on a real athlete's assessment.
  • Hands-on review of the athlete's own footage by people with athletic, baseball and movement backgrounds.
  • As an escalation and quality-control step, on cases the system itself has marked as needing another look.
  • In deciding whether a system change is warranted, through a documented revision and re-test cycle.

Where it does not

  • Not inside the blind benchmark workflow — a locked assessment is never revised after the benchmark is opened, including when it was wrong.
  • Not as a way to improve a benchmark result after seeing the answer.
  • Not to overrule an evidence boundary and turn a “not assessable” into a confident claim.
  • Not as silent editing — configuration changes require a new version, a verified diff and a re-run confirmation suite.

These are two separate workflows. The benchmark workflow exists to measure the system honestly, so nothing in it may be revised after the comparison. The production workflow exists to serve a real athlete well, so escalation to human review is available there. Keeping them apart is what makes the benchmark meaningful.

The Claim

What the evidence supports.

Supported

  • Video-based assessment, benchmarked against 3D motion capture. 58 blind assessment runs across 13 distinct athletes.
  • Blind findings were locked before benchmark reveal — to a non-editable, timestamped record, 24 of them carrying a cryptographic file fingerprint.
  • Frozen formal scoring covers Lead Leg Braking and Rotational Output at category level, across the 13-athlete baseline: no misses, no false positives, no unsupported predictions.
  • Governance materially reduced unsupported output — from 6 of 6 unseen athletes to 0 of 6, N = 6.
  • The final confirmation suite passed 9 of 9 cases under a protocol frozen and hashed before the first case ran.
  • A documented failure, a documented fix, and a documented re-test — including the fixes that failed.

How to say it

  • Video-based assessment, benchmarked against 3D motion capture.
  • Benchmarked against independent 3D motion-capture and biomechanics reports.
  • Built to test how closely a structured video-based assessment could reproduce the actionable findings contained in high-end biomechanics testing.
  • “Not assessable” is a real answer. Where footage can't support a call, the category is marked Not Assessable rather than guessed at — and that outcome is scored separately, so it is never quietly counted as correct.
Limits

Where the evidence stops.

These are real constraints on what the evidence supports, not a disclaimer block. If one of them matters to your decision, it should.

Scope of the evidence

  • The validation cohort contains 13 distinct athletes. The formal unseen cohort was N = 6. All rates are directional evidence, not statistically tight estimates.
  • Formal category-level scoring covers 2 of 8 categories. The other six carry finding-level evidence only.
  • Only one athlete was truly held out from earlier system-development exposure — and that case produced a miss.

Limits of the method

  • Video is observational, not laboratory measurement. Apex does not produce joint angles, segment velocities, hip-shoulder separation in degrees or projected velocity gains. Those are hard-blocked output categories, not features awaiting a release.
  • One global accuracy percentage is not methodologically supported. Rounds used different scoring methods and denominators; none may be averaged into a single figure, and we do not publish one.
  • “Not assessable” is an intentional outcome. When evidence is insufficient the system withholds — which sometimes means a real finding is not delivered.

The full limitation register — benchmark format heterogeneity, era gaps, documentation gaps, procedural deviations and amendment failures — is published in the technical details below. Nothing has been removed from it; it is staged there because it is engineering detail rather than buying information.

Traceability

Ask us where any number came from.

Every published validation figure traces to a named locked source record — file, sheet and row. If you want to see the working, ask.

View Technical Validation Details

Provenance — where each published number comes from

Published numberWhere it comes fromHow it is calculated
58Blind-Output Lock Registry and phase records14 + 24 + 7 + 9 + 4 locked runs
13Validation closeout, distinct-athlete count7 + 2 + 4, with no double counting
6 of 6 → 0 of 6Prohibited-output rate, unseen formal cohortAthletes with at least one violation, out of 6
9 / 9Final confirmation suiteCases passed, out of 9 frozen cases
0 missesFrozen production baselinePer-athlete rows, Lead Leg Braking + Rotational Output
Retired44 of 99Phase-3 pooled view, pre-revision system — superseded, not current performanceExact matches, out of scoreable feature-level decisions
Retired2 of 5Phase-3 archetype accuracy, pre-revision system — superseded, not current performanceConfirmed, out of tested instances

Production configuration locked 2026-08-14 with cryptographic hash verification. Any future analytical change requires a new version, a verified diff and a fresh confirmation suite before it can be authorised.

The percentages we can publish, individually

Rows marked Retired describe the pre-revision system. They are published for traceability and do not describe current Apex performance. The frozen production baseline is the current result.

The frozen production system has no publishable percentage, and that is a deliberate feature of the record rather than an omission. Its scored pass produced counts, not a rate; the counts that reconcile exactly are all zeros; and our own validation record prohibits building a blended cross-round figure. A percentage here would be manufactured, so there isn't one. Every publishable percentage below describes the system we replaced.

ValueFractionWhat it measuresScope and denominator
Retired44.4% 44 of 99 Exact-match rate at the individual-finding level Pre-revision system, ungoverned, 7-athlete controlled retest, Full-7 cohort view · 99 scoreable feature-level decisions
Retired42.3% 33 of 78 Same measure, same-era subset Excludes the two athletes whose benchmark data is years older than their assessment · 78 scoreable feature-level decisions
47.5% 28 of 59 Same measure, clean-input subset Excludes the two athletes with documented blind-test contamination exposure · 59 scoreable feature-level decisions
40% 2 of 5 Accuracy of the most frequently selected primary finding Pre-revision system, the 5 athletes on whom that archetype was tested · 5 tested instances
25% 1 of 4 Same finding, ungoverned, on the next unseen cohort 4 resolvable athletes; 2 excluded because their benchmark format could not grade it · 4 resolvable athletes of 6

These are separate measurements with separate denominators. They are never averaged together, and no cross-round global accuracy percentage exists.

Development phases, reported separately

Phases are reported separately by design — each used its own scoring method and denominator. The progression itself is the evidence; blending it into one figure would destroy it.

  • 2026-05 → 2026-07 Early development

    Low accuracy under an informal rubric. Superseded — these figures are not publishable.

  • 2026-07-19 → 2026-07-20 Phase 3 — controlled retest

    Ungoverned system matched the benchmark exactly on 44 of 99 scoreable decisions. Four systemic failure modes documented.

  • 2026-07-20 → 2026-07-21 Revision design

    12 amendments drafted, frozen and hashed. Not authorised for operational or commercial use during the round.

  • 2026-07-21 → 2026-08-14 Phase 6 — unseen four-condition round

    Unsupported claims: 6 of 6 athletes ungoverned, 0 of 6 governed. Two new failure modes found.

  • 2026-08-15 Amendment regression

    8 fixes tested against pre-registered criteria: 4 passed, 4 failed or were confounded. All published.

  • 2026-08-17 → production lock Targeted revision + frozen final confirmation

    9 of 9 confirmation cases passed. Candidate frozen and hashed before the first case ran.

  • 2026-08-14T17:43:25Z Production lock

    Prompt and governance locked with SHA-256 hashes. Substantive diff vs. the tested artifacts: zero.

  • 2026-08-15 Apex Core v1 frozen baseline

    6 direct hits, 0 misses, 0 false positives — Lead Leg Braking and Rotational Output only, across 13 athletes.

Documented misses, in full

  • Amplitude is not velocityA large, visually dramatic torso rotation was read as fast rotation; the measurement said it was one of the athlete's worst features. Every configuration willing to state a direction got it backwards, in 3 of 4 scored athletes. This is the failure the final revision was built to fix, and it did not recur in the confirmation suite.
  • A real limiter went unconveyedThe one held-out athlete had a genuine lead-leg braking limiter ranked fourth in their benchmark report. The locked system withheld rather than calling it. Separately, a hip-mobility limiter flagged as the single largest deviation in one report was underweighted or withheld by all four configurations tested.
  • A good composite can hide a bad featureA correctly-called above-average build-up composite contained a feature scoring 6 out of 100. A 32-degree trunk lateral-tilt outlier was surfaced by none of the four configurations.
  • Half of our own fixes failed on first testOf eight proposed amendments, four passed and four failed or could not be cleanly tested. Two were confounded by an unrelated pitch-count gate; one feature turned out never to have been implemented at all; one was declined for production. All four are on the record as failures.

Full limitation register

  • Sample size13 distinct athletes; the formal unseen cohort was N = 6; one realised held-out athlete. All rates are directional evidence, not statistically tight estimates.
  • No overall accuracy figure existsScoring methods and denominators differ across rounds by design. The locked source states a cross-round global accuracy percentage “DOES NOT EXIST” and “must never be created”.
  • Category scopeThe frozen validation pass scored 2 of the 8 framework categories at full category level. The other 6 have finding-level evidence only.
  • Benchmark format heterogeneityGround truth came from several independent 3D motion-capture and biomechanics providers plus narrative laboratory reports, which differ in metrics, normalisation and whether population reference bands are printed at all.
  • Some benchmarks cannot grade a claimOne formal-cohort export carried no population reference bands, so above/below-average claims were logged Not Determinable rather than scored. One pilot athlete's benchmark images were cropped.
  • Era gapsOne athlete's benchmark predates their assessment by roughly five years. Another has two sessions six weeks apart with real test-retest variability.
  • Validation footage is not production footageLegacy validation footage is explicitly not representative of the standardised submissions production expects. Video quality was logged as methodological context and no locked result was re-weighted because of it.
  • Older intakes lack current fieldsOlder intake packets lack fields the current Master Intake Form captures — one documented miss (Hip Internal Rotation) traces directly to unstructured intake wording rather than to the assessment logic.
  • Documentation gapsSeveral historical and condition files lack input-hash manifests or precise UTC lock timestamps. Only 2 of 9 final confirmation reports self-document input hashes.
  • Procedural deviations, logged not curedThe two pilot athletes ran in fixed rather than randomised condition order, and were opened before any held-out subset was designated. Both were permanently excluded from the formal held-out subset rather than retroactively reclassified.
  • Amendment failures4 of 8 amendments failed or were confounded on first regression. Two were confounded by an unrelated pitch-count gate; one feature was never implemented; one was declined for v1.0. All four are published as failures.
  • Held-out subset was under-filledTwo held-out slots were designated; only one was filled, because the cohort closed at six athletes when eligible athletes with usable benchmark data ran out.
  • Video observation is not laboratory measurementApex reads observational evidence from video plus intake. It does not measure joint angles or segment velocities. Claims that require direct measurement — exact hip-shoulder separation in degrees, rotational velocity in deg/s, projected velocity gain — are categorically blocked, not estimated.
  • Computer vision is not production-approved for pitchingZero pitching CV measurements are production-approved. The most promising metric (throwing elbow flexion at ball release, mean absolute difference 5.44 deg) remains RESEARCH_ONLY for insufficient N and replication. Mobility CV is authorised only as a lower-authority supporting layer that cannot change or create a grade.

Source figures, as generated from the validation package

These are the charts produced directly from the locked validation package. They are retained here unchanged so that every figure on this page can be checked against its source rendering.

Cohort and run composition
Bar chart of blind assessment runs by group totalling 58, with three excluded groups shown separately in a different colour.
Where the 58 comes from — and what we chose not to count.
Two stacked bars comparing 13 athletes with 58 runs, segmented by testing phase.
13 people, 58 runs. Two different numbers.
Results, governance and confirmation
Before: 44 of 99 exact matches and 2 of 5 confirmed findings for the pre-revision system. After: zero misses, zero false positives and zero unsupported predictions for the locked system across two categories.
Exact agreement, before and after. Two panels, two denominators, never combined.
Prohibited outputs appeared in 6 of 6 athletes without governance and 0 of 6 with each of three governed configurations.
N = 6 unseen formal cohort. Each governed configuration was scored independently.
Nine confirmation cases across four target groups, all passed.
Nine confirmation cases across four target groups, all passed.
Category coverage, amendments and phase progression
Eight framework categories, with Lead Leg Braking and Rotational Output marked as formally scored and the other six as finding-level evidence only.
The frozen pass covered the two categories the rebuild targeted.
Of eight proposed amendments, four passed and four failed or were confounded.
Our own failure rate on our own fixes, tested against criteria written before the test.
Eight-phase development timeline from early development through production lock and the frozen Core v1 baseline.
Eight phases, reported separately. Each used its own scoring method and denominator.

Want the assessment itself?

See the eight-category framework, the sample Blueprint, and what lands in your app — then start when you're ready.

Start Your Assessment See What You Get