Independent financial AI evaluation

Which AI models can actually perform professional financial work?

Kaelum Bench evaluates accounting, reporting, and corporate finance with expert-designed tasks and signed evidence attached to every published model run.

v0.2 · Public CalibrationUpdated August 14, 2026Signed public release

Top models this release

v0.2 · signed release

  1. Kimi K2.688.9
  2. Kimi K388.9
  3. Claude Opus 588.5
  4. GPT-5.6 Terra88.4
  5. Claude Sonnet 5 (Max)86.1

Mean of nine benchmark scores · All 15 models →

Current release metrics

9 Published benchmarks

of 10 defined

580 Tasks represented

across both professional directions

164 Verifiable run receipts

132 deterministic · 32 hybrid

Latest benchmark release

The current calibration, clearly published.

v0.2 publishes the focused accounting and corporate-finance benchmarks alongside the first integrated 100-task capstone.

v0.2 · Signed public release

Accounting and corporate finance, evaluated as professional work.

Explore focused 60-task benchmarks and the published Accounting Capstone. Every observed score remains connected to methodology, configuration, cost, token use, and exact run evidence.

Open all benchmarks →

100-task capstone

Accounting Capstone

Hybrid scoring · 17 published model runs

View benchmark

Next 100-task capstone

Corporate Finance Capstone

Benchmark definition prepared; calibration pending

Preview scope

Transparent by design

Evidence before comparison.

Task and scoring evidence

Definitions, task counts, scoring modes, and tested scope remain visible.

Execution and cost evidence

Provider route, runtime, tokens, cost, and validity stay attached to each run.

Signed publication

Versioned data and authenticated manifests preserve the exact public release.