What is being tested?
Kaelum Bench evaluates structured financial work: calculations, accounting-rule application, analytical decisions, and integrated cases. Published suite titles describe tested scope; they are not hidden capability subscores.
Accounting, Reporting & Financial Analysis
5 published suites · Accounting knowledge, HGB and IFRS reporting, financial analysis, and an integrated accounting capstone.
- Accounting Knowledge
- HGB Reporting
- IFRS Reporting
- Financial Analysis
- Accounting Capstone
Corporate Finance & Valuation
4 published suites · Capital budgeting, financing, valuation, transactions, and an integrated corporate finance capstone.
- Capital Budgeting
- Cost of Capital & Financing
- Valuation & Financial Modeling
- M&A, LBO & Capital Allocation
- Calculate
- Derive exact financial values, reconciliations, and structured fields.
- Apply rules
- Use accounting standards and stated assumptions correctly.
- Analyze & decide
- Interpret performance, financing, valuation, or transaction evidence.
- Integrate
- Carry earlier results through a longer multi-step capstone case.
From task to published result
The complete evaluation contract is fixed before execution and remains connected to the resulting public evidence.
- Task contractCase inputs, question, required output, tools, and available points.
- Recorded runExact model, provider route, runtime, reasoning, and tool setup.
- ScoringPredefined exact checks, plus a disclosed judge only where required.
- Run receiptScore components, cost, usage, configuration, and judge evidence.
- Observed resultA percentage for this benchmark version and recorded configuration.
Scoring
Benchmark pages show only the scoring mode and components that were actually used.
| Mode | How it works | Current release |
|---|---|---|
| Deterministic | Exact fields, calculations, and structured outputs are checked using predefined rules. Deterministic score components can be recomputed exactly from their authenticated inputs.Earned deterministic points ÷ available points | 132 runs · no AI judge |
| Hybrid | Exact checks are combined with rubric assessment where deterministic checks are insufficient. A model-judged component records one stored judge outcome.Deterministic + rubric ± published modifier | 32 runs · judge disclosed |
Comparing results
A published score is evidence for a specific recorded run, not a statistically certain universal ranking.
- Configuration-specificModel, route, runtime, reasoning, and tools stay attached.
- Missing is not zeroUnpublished evidence remains visibly unavailable.
- No universal scoreBenchmarks are not merged into an opaque intelligence measure.
Evidence and limitations
The release states both what the recorded evidence supports and what it does not establish.
Supported
- Versioned benchmark and recorded configuration
- Observed scores and published score components
- Recorded cost, usage, execution, and judge identity
- Signed sanitized release artifacts
Not established
- Standardized best-effort model performance
- Statistical certainty from repeated runs
- Model-judge repeatability qualification for the current public release
- A guarantee that a fresh judge call reproduces a stored hybrid outcome
- Freedom from conceptual contamination
- Spreadsheet or workbook-artifact performance
Release integrity
Ed25519 authenticates the exact public manifest and can expose publication tampering. The signature authenticates the public projection; it does not make a judge decision repeatable or turn a provisional result into a verified result.
- Signature
- Ed25519
- Inventory
- 164 completed run receipts
- Status
- Signature material published
Status glossary
These terms keep the same meaning across benchmarks, comparison views, and run receipts.
- Provisional
- Published calibration evidence that has not reached the future verified standard.
- Verified
- Reserved for a future result that satisfies the published verification standard.
- Hybrid
- Scoring that combines deterministic checks with rubric-based AI assessment.
- Deterministic
- Scoring produced entirely by predefined exact checks without an AI judge.
- Configuration-specific
- Evidence tied to the recorded model, provider route, runtime, reasoning, and tools.
- Signed release
- A release whose exact manifest bytes are authenticated with a published key.