100-task capstone
Accounting Capstone
Hybrid scoring · 17 published model runs
View benchmark →Independent financial AI evaluation
Kaelum Bench evaluates accounting, reporting, and corporate finance with expert-designed tasks and signed evidence attached to every published model run.
v0.2 · Public CalibrationUpdated August 14, 2026Signed public release
Top models this release
v0.2 · signed release
Mean of nine benchmark scores · All 15 models →
9 Published benchmarks
of 10 defined
580 Tasks represented
across both professional directions
164 Verifiable run receipts
132 deterministic · 32 hybrid
Latest benchmark release
v0.2 publishes the focused accounting and corporate-finance benchmarks alongside the first integrated 100-task capstone.
v0.2 · Signed public release
Explore focused 60-task benchmarks and the published Accounting Capstone. Every observed score remains connected to methodology, configuration, cost, token use, and exact run evidence.
Open all benchmarks →100-task capstone
Hybrid scoring · 17 published model runs
View benchmark →Next 100-task capstone
Benchmark definition prepared; calibration pending
Preview scope →Choose your depth
Scan every published benchmark, focused suite, and 100-task capstone in one consistent overview.
Explore benchmarks →Select models and compare observed benchmark performance and recorded resources across the release.
Compare models →See what is tested, how points are awarded, and which limits matter when reading a score.
Understand the method →Transparent by design
Definitions, task counts, scoring modes, and tested scope remain visible.
Provider route, runtime, tokens, cost, and validity stay attached to each run.
Versioned data and authenticated manifests preserve the exact public release.