That order is the test. Reverse it and the result means nothing.
Apex Motion's assessment was developed by comparing what it found against independent 3D motion-capture and biomechanics reports it was never allowed to see first.
Every run on this page was produced from video and intake alone, locked to a non-editable record, and only then scored.
“If Apex got it wrong, the wrong answer stays on the record.”Every miss on this page is an output we left exactly as it was generated
Apex is benchmarked against independent 3D motion-capture and biomechanics reports. Before the results mean anything, it's worth knowing what that class of testing actually is.
3D motion capture is a laboratory biomechanics method used to analyse movement in three dimensions. Multi-camera systems capture fast athletic movement and allow researchers and high-performance professionals to quantify movement that cannot be directly measured from a single ordinary video angle.
This class of technology is used in elite sport, biomechanics research, high-performance training environments and professional baseball. It runs on specialised multi-camera laboratory infrastructure that most pitchers do not have routine access to — which is exactly why it makes a demanding benchmark.
How a laboratory benchmark is produced
Synchronised cameras record the same delivery from several angles at once.
Those views are combined into a three-dimensional model of the movement.
Joint angles and segment speeds are quantified — values a single video angle cannot produce.
The findings are written up as prioritised, actionable conclusions about the athlete.
So we hid that information from Apex.
FirstApex made its assessment from video and intake alone.
ThenThe findings were locked to a non-editable, timestamped record.
Only thenThe independent biomechanics benchmark was opened.
How closely could a structured video-based assessment reproduce the actionable findings contained in high-end biomechanics testing?
Four representative comparisons, anonymised. One clean agreement, one outright failure, one deliberate withholding, and one recovery on the frozen system.
Lead leg braking flagged as a primary limiter, from video and intake alone.
The independent biomechanics report identified a lead-leg braking deficit of its own accord.
A large, visually dramatic torso rotation read as fast rotational output.
Measured rotational output was among that athlete's weakest features — the opposite direction.
No rotational-velocity call made. Video cannot establish segment velocity, so the claim is blocked rather than estimated.
The report contains a directly measured value.
On the frozen system, the two athletes whose rotation had been read backwards now return “not assessable”.
Same benchmark reports, re-scored against the same pre-written rules.
Athletes are not identified, and no third-party report is reproduced. These are the comparison outcomes as they were recorded in the scoring ledger.
The assessment framework was built on an existing body of knowledge first, then tested against independent measurement — not improvised and rationalised afterwards.
The eight-category framework and the reasoning behind it were built from an extensive evidence library spanning pitching biomechanics, mechanics, physical preparation, movement and training — drawing on peer-reviewed research, biomechanics literature, case studies, expert educational material and applied pitching-development resources.
Those source types do not carry equal authority, and we don't treat them as if they do. Peer-reviewed biomechanics research and a coaching video are both useful inputs to how the framework reasons; neither one, on its own, licenses a claim about a specific athlete. What an individual assessment is permitted to state is decided separately — see below.
Governance is a different job from research. The evidence library helps Apex reason about pitching in general. Governance decides what Apex may actually say about you, given only the evidence your own submission contains.
Five steps, in this order. The order is the entire test — it is what makes it impossible for the answer to have been influenced by already knowing the answer.
Slow-motion pitching video plus the athlete's intake — strength numbers, mobility findings, injury history. Nothing from a lab.
What a coach can actually get.The assessment is produced from that material alone. The session generating it has no access to the athlete's biomechanics report.
No peeking.Saved to a non-editable, timestamped record — 24 of them carrying a cryptographic file fingerprint. From this moment it cannot be changed.
The answer is sealed.Only now is the independent 3D motion-capture and biomechanics data revealed to the scoring session.
Now the benchmark opens.Compared under scoring rules written down before any athlete was opened. Agreements and misses are recorded with equal weight.
Graded using rules that were already written.What that produced
Four different kinds of proof: scale and method, cohort, mechanism, and verification. Each carries its own denominator, and the denominators are never swapped.
Athletes are counted once. Runs are counted separately. Those are different numbers and we keep them apart.
produced
A re-test is another run, not another person. Thirteen people produced fifty-eight runs.
24 of the 58 carry a recorded cryptographic file fingerprint. A note on wording: there are 58 runs and 57 benchmark comparisons — one blind Core v1 run was produced but never compared, because that athlete's benchmark was not supplied in the scoring session. We report runs.
There isn't one honest blended accuracy number — so we show the comparisons that actually matter instead. Different rounds tested different scopes, categories and scoring rules; averaging them would produce a figure that sounds precise and means nothing. What follows is the progression as it was recorded, strongest evidence first.
Seven athletes went through the system twice — once with the safety layer, once without — with the benchmark data sealed until both were locked. It was scored at the strictest possible standard, where only an exact category match counts and “close” scores zero.
Real signal, well short of good enough. That result is what triggered the rebuild.
The most-used primary archetype was confirmed by the benchmark in only 2 of 5 cases. And a specific optical illusion kept repeating: a large, visually dramatic torso rotation was read as fast rotation, when the measurement showed it was one of that athlete's worst features. Every configuration willing to state a direction got it backwards.
A confident, specific, wrong answer is worse than no answer. That is the failure mode the rebuild was aimed at.
The final configuration was locked with cryptographic hashes. All 13 athletes were then run through that exact configuration. On the two movement categories the rebuild targeted — the ones with the strongest benchmark detail — it recorded:
Scope: frozen production baseline · Lead Leg Braking + Rotational Output · 2 of 8 formal categories · across the 13-athlete baseline. This result does not describe the other six categories.
Both lead-leg findings the benchmark had genuinely confirmed were recovered. Both athletes whose rotation reading had been backwards now return “not assessable” instead of a confident wrong answer.
Only one athlete was genuinely held out from earlier system-development exposure — designated by a documented random draw, using a recorded seed, before any athlete's identity, video or benchmark data had been looked at. That athlete had a real lead-leg braking limiter, ranked fourth in their own benchmark report's development priorities. The governed system withheld rather than calling it: correct by its own evidence rules, and still a finding the athlete did not receive.
We publish this because a validation page without it isn't a validation page.
None of this was discovered by a customer. Each of these is documented in the same ledger as every result above.
General-purpose AI tends to answer even when evidence is thin. Apex is designed to withhold claims when the available evidence cannot support them.
Scope: N = 6 unseen athletes / independently scored configurations. Each of the six produced at least one unsupported claim with governance off; none did with governance on.
What governance blocks
This is reliability engineering, not fine print. A system that never says “I can't tell from this” isn't more accurate — it is guessing more often and hiding it better.
A pass rate is only as good as the protocol that produced it. The protocol comes first here, because it is the reason the score counts.
The candidate prompt and governance rules were locked and cryptographically hashed before the first case ran.
Each of the nine cases had its success condition defined in advance, including which behaviours were not allowed to regress.
A failure in case 3 could not be patched before case 4. That is the difference between a test and a rehearsal.
Results were withheld until every case was complete, so no interim result could influence how a later one was read.
Confirmation cases, not nine unique production athletes. These are re-tests of athletes already in the programme, using their original blind inputs. We do not convert 9 of 9 into a percentage.
All nine passed. Both target failure modes were fixed. Neither control regressed. Zero prohibited outputs across all nine. It is confirmation evidence — proof the fixes hold — not proof the system generalises to a large new population.
A category in the Apex framework is not the same thing as a category included in a frozen scoring pass. We are not going to show you eight scores when we formally scored two.
The other six categories are live parts of the assessment framework and carry finding-level evidence. They did not receive frozen, category-level scoring, so nothing on this page claims they did.
Apex Motion is a governed system, not an unsupervised one. It matters where the humans are — and just as much, where they are not.
These are two separate workflows. The benchmark workflow exists to measure the system honestly, so nothing in it may be revised after the comparison. The production workflow exists to serve a real athlete well, so escalation to human review is available there. Keeping them apart is what makes the benchmark meaningful.
These are real constraints on what the evidence supports, not a disclaimer block. If one of them matters to your decision, it should.
The full limitation register — benchmark format heterogeneity, era gaps, documentation gaps, procedural deviations and amendment failures — is published in the technical details below. Nothing has been removed from it; it is staged there because it is engineering detail rather than buying information.
Ask us where any number came from.
Every published validation figure traces to a named locked source record — file, sheet and row. If you want to see the working, ask.
| Published number | Where it comes from | How it is calculated |
|---|---|---|
| 58 | Blind-Output Lock Registry and phase records | 14 + 24 + 7 + 9 + 4 locked runs |
| 13 | Validation closeout, distinct-athlete count | 7 + 2 + 4, with no double counting |
| 6 of 6 → 0 of 6 | Prohibited-output rate, unseen formal cohort | Athletes with at least one violation, out of 6 |
| 9 / 9 | Final confirmation suite | Cases passed, out of 9 frozen cases |
| 0 misses | Frozen production baseline | Per-athlete rows, Lead Leg Braking + Rotational Output |
| Retired44 of 99 | Phase-3 pooled view, pre-revision system — superseded, not current performance | Exact matches, out of scoreable feature-level decisions |
| Retired2 of 5 | Phase-3 archetype accuracy, pre-revision system — superseded, not current performance | Confirmed, out of tested instances |
Production configuration locked 2026-08-14 with cryptographic hash verification. Any future analytical change requires a new version, a verified diff and a fresh confirmation suite before it can be authorised.
The frozen production system has no publishable percentage, and that is a deliberate feature of the record rather than an omission. Its scored pass produced counts, not a rate; the counts that reconcile exactly are all zeros; and our own validation record prohibits building a blended cross-round figure. A percentage here would be manufactured, so there isn't one. Every publishable percentage below describes the system we replaced.
| Value | Fraction | What it measures | Scope and denominator |
|---|---|---|---|
| Retired44.4% | 44 of 99 | Exact-match rate at the individual-finding level | Pre-revision system, ungoverned, 7-athlete controlled retest, Full-7 cohort view · 99 scoreable feature-level decisions |
| Retired42.3% | 33 of 78 | Same measure, same-era subset | Excludes the two athletes whose benchmark data is years older than their assessment · 78 scoreable feature-level decisions |
| 47.5% | 28 of 59 | Same measure, clean-input subset | Excludes the two athletes with documented blind-test contamination exposure · 59 scoreable feature-level decisions |
| 40% | 2 of 5 | Accuracy of the most frequently selected primary finding | Pre-revision system, the 5 athletes on whom that archetype was tested · 5 tested instances |
| 25% | 1 of 4 | Same finding, ungoverned, on the next unseen cohort | 4 resolvable athletes; 2 excluded because their benchmark format could not grade it · 4 resolvable athletes of 6 |
These are separate measurements with separate denominators. They are never averaged together, and no cross-round global accuracy percentage exists.
Phases are reported separately by design — each used its own scoring method and denominator. The progression itself is the evidence; blending it into one figure would destroy it.
Low accuracy under an informal rubric. Superseded — these figures are not publishable.
Ungoverned system matched the benchmark exactly on 44 of 99 scoreable decisions. Four systemic failure modes documented.
12 amendments drafted, frozen and hashed. Not authorised for operational or commercial use during the round.
Unsupported claims: 6 of 6 athletes ungoverned, 0 of 6 governed. Two new failure modes found.
8 fixes tested against pre-registered criteria: 4 passed, 4 failed or were confounded. All published.
9 of 9 confirmation cases passed. Candidate frozen and hashed before the first case ran.
Prompt and governance locked with SHA-256 hashes. Substantive diff vs. the tested artifacts: zero.
6 direct hits, 0 misses, 0 false positives — Lead Leg Braking and Rotational Output only, across 13 athletes.
These are the charts produced directly from the locked validation package. They are retained here unchanged so that every figure on this page can be checked against its source rendering.
See the eight-category framework, the sample Blueprint, and what lands in your app — then start when you're ready.