Cost calculator
What is this actually worth to your practice?
Put in your own numbers and see what typing assessment data by hand actually costs you — in money, in hours, and in wrong figures that reach a client. Every number behind the arithmetic is listed at the bottom, with where it came from.
At 250 candidates a year, typing them by hand puts about 95 wrong numbers into your data. Nothing marks which ones.
AeroPoint
the shipped extractorBuilt yourself, with AI
——
At this volume, extraction alone probably will not pay for itself.
Below roughly 150 candidates a year the hand-typing it replaces costs less than the software does, and you should hear that from me rather than work it out afterwards. Where it can still make sense at this size is if extraction is not the only thing being done by hand — if the reports that go out to your clients are also assembled manually, or the same numbers get retyped into a second document, a summary, or a comparison across candidates, then the case is built on all of that together rather than on extraction by itself. Worth a conversation either way; the answer might be no.
What if you get hit by a bus?
It is a fair question to ask a business that is one person, and there is a boring answer before the dramatic one: nothing about this runs on my machine. The software is installed on yours, your files never leave it, and the data it produces is yours in a format you can open without me. If I disappeared tomorrow you would still have every dataset it has ever produced. For continuity beyond that — source-code escrow, or a clause releasing the code if I can no longer maintain it — ask, and we will write it into your agreement.
You do not have to take any of this on trust. There is a free trial — install it, verify your email, and it unlocks for 24 hours or 50 assessments on that machine. No payment details, nothing to cancel. Run it against your own files and see what comes back before any of these numbers matter. Download it here.
These figures are a starting estimate, not a quote. Every practice is priced after an actual conversation — how many instruments you use, which report formats need rebuilding, what your files look like, and how much of the work around the extraction you want handled. The numbers above exist so you can see the shape of it before that call, not to stand in for one.
No hidden constants
Where every number comes from
A calculator that hides its inputs is a sales toy, so here are all of them. Most are measured on real assessment files. Three are assumptions, and those are labelled as assumptions rather than buried.
| Constant | Value | Where it comes from |
|---|---|---|
| Values in a full 8-test battery | 76 | Measured. A full battery is the eight instruments listed on the home page, and it holds exactly 76 stored values — minimum, median and maximum are all 76, across 51 full batteries in one suite and 171 in another. |
| AeroPoint time per battery | 0.417 s | Measured. Full 115-candidate suite run in parallel, then normalised to a full 8-instrument battery. Batteries range from 1 to 8 instruments, so a raw per-candidate average would not be comparable. |
| AeroPoint accuracy | 76 / 76 | Measured. Against a hand-verified gold standard. The harness is not one battery: it re-checks 6,857 values across 115 candidates on every build, and a second suite covers 25,270 values across 444. |
| AI reading accuracy | 99.8% | Measured at scale. 1,233 of 1,235 printed values read correctly by Claude Haiku 4.5 — the cheapest model tested — across 42 candidates and three instruments (CPI 988/988, Strong/Skills 126/126, cognitive and critical-thinking 119/121). Earlier smaller runs agree: 75/75 on Claude Sonnet 5, and 24/24 MBTI values across four different editions of that report. Where a number is printed on the page, reading it is a solved problem, even cheaply. Both failures were on the same chart: a bar with the value in small grey type at its end. One misread 95 as 85 where the digits overlap the bar edge. The other read every digit correctly and attached it to the wrong label — on that chart the label column and the bars are vertically offset and the correspondence is carried by colour. Neither error is detectable in the output. |
| AI coverage | 43 / 76 | Measured. Given the documents but not the conventions, a frontier model attempted 43 of the 76 values and never went looking for the other 33 — nothing in the reports says they exist. This is separate from getting them wrong, and it is the larger of the two problems. |
| AI accuracy on what it attempted | 38 / 43 | Measured. 88.4% on the first pass. All five failures were schema errors; none were misreadings. |
| Input tokens per battery | 64,600 | Measured. 24,600 for planning plus 40,000 for reading twelve rendered pages at about 2,800 tokens each. |
| Values no reader can recover | 3 per WG‑II | Measured, and the clearest case on this page. On the older Watson-Glaser II report the three subscale percentiles are printed as unlabelled bars with no axis, no scale and no number anywhere in the file. AeroPoint recovers them by measuring bar geometry against a calibrated box. A reader of any price returns the one printed overall percentile and a band word from the prose — the other three values are not in the document to be read. |
| Hourly rate | $90 assumed | An assumption, and a deliberately conservative one. In most practices the person retyping these figures is an associate or an administrator rather than the psychologist, and the loaded cost of that hour is nearer $60–90. If a clinician is doing it, put their rate in — every saving on this page roughly doubles. |
| Minutes per battery, by hand | 15 assumed | An assumption. Covers transcription and checking for one candidate. Adjust it above. |
| By-hand accuracy | 99.9% small sample | From a real but small sample: 3 transcription errors across 8 candidates, found early in this project — 3 mistakes in roughly 4,800 numbers, which is nearer 99.94%. Rounded down to 99.9% here, which is generous to nobody in particular and keeps the number readable. Small sample: treat it as indicative rather than settled, and change it above to see how little it takes to matter. |
How the AI accuracy figure is built — and why the model you pick barely moves it
The estimate is coverage × accuracy-on-attempted × reading accuracy, applied to 76 values. All three were measured, and they behave very differently.
Reading is solved. Two model tiers were tested on the hardest part of the battery — CPI scale values that exist only as pixels on a chart. Claude Haiku 4.5, a small cheap model, scored 125 / 125. Claude Sonnet 5 scored 75 / 75. Nothing suggests a more expensive model reads these pages better, so every model in the dropdown is credited with the same reading accuracy.
Schema is not, and it does not improve with spend. A frontier model given the documents but not the conventions scored 88.4%. Every one of its errors was a decision that is not written anywhere in the document: whether an MBTI preference score is stored positive or negative, and that one CPI scale is stored as 100 − the printed value. Running the same schema questions past a much cheaper model reproduced the identical failures — the same inverted signs, the same untransformed scale. That is the evidence for treating schema accuracy as a property of the documents rather than of the model, including for the models below that could not be tested directly.
There is a cheaper strategy that is worse still: skipping the rendered pages and reading the PDF text layer instead. On this battery that returns a clean, plausible, correctly formatted answer for all 26 CPI scales and is wrong on every one of them — thirteen consecutive scales come back as the same number. Nothing errors and nothing is flagged. It is also the cheapest thing to build, which is why it is the one a competitor reaches for first.
Model prices, speeds, and which ones were actually tested
| Model | Input $/M | Output $/M | Status here |
|---|
Models marked estimated were not tested. Their prices are the vendor’s published list rates, their speed is inferred from their tier, and they are credited with the same accuracy as the tested models — this page does not claim a competitor’s model reads worse. The schema ceiling applies to them for the reason given above: the missing information is absent from the document, so no model recovers it.
What the “catches its own errors” row actually means
It is the difference between an accuracy number and a checkable accuracy number, and it is the one row that does not move when you change the inputs above.
A person working by hand and an AI pipeline both produce a spreadsheet with no indication of which cells are wrong. The AI pipeline reached its answer in one pass with nothing to compare against — in production there is no answer key, so it cannot know it needs a second look and does not take one. AeroPoint’s figure is checked against a hand-verified corpus and re-checked automatically on every build, which is a different kind of claim rather than just a better number.
Be sceptical of the 100%. It is 61 of 61 on a corpus that has been hand-verified over two years, on the eight instruments listed on the home page. A new instrument, or a report layout that has changed, starts at zero until it is built and verified. That work is the engagement.
“Data stays with you” — the caveat in full
By hand and with AeroPoint, candidate files never leave the machine they are already on. A build-your-own AI pipeline transmits candidate reports to a third-party provider, which is a different question for a practice holding assessment data on identified people.
Zero-retention and no-training terms are genuinely available on enterprise agreements with the major providers, and a practice that negotiates them has materially reduced the exposure. But that is a contract to negotiate, verify, and keep current as terms change — it is not how an API key behaves out of the box, and it is worth establishing before candidate files start moving rather than after.
Run on 6 September 2026 against one real eight-instrument battery, with follow-up model tests on 7 September 2026. Figures describe that corpus. If you want to see the method rather than the summary, ask — it is more convincing than this page is.