Methodology

How Kaelum Bench works.

This page explains what the benchmarks test, how answers are scored, how results may be compared, and where the published evidence stops.

What is being tested?

Kaelum Bench evaluates structured financial work: calculations, accounting-rule application, analytical decisions, and integrated cases. Published suite titles describe tested scope; they are not hidden capability subscores.

Accounting, Reporting & Financial Analysis

5 published suites · Accounting knowledge, HGB and IFRS reporting, financial analysis, and an integrated accounting capstone.

  • Accounting Knowledge
  • HGB Reporting
  • IFRS Reporting
  • Financial Analysis
  • Accounting Capstone

Corporate Finance & Valuation

4 published suites · Capital budgeting, financing, valuation, transactions, and an integrated corporate finance capstone.

  • Capital Budgeting
  • Cost of Capital & Financing
  • Valuation & Financial Modeling
  • M&A, LBO & Capital Allocation
Calculate
Derive exact financial values, reconciliations, and structured fields.
Apply rules
Use accounting standards and stated assumptions correctly.
Analyze & decide
Interpret performance, financing, valuation, or transaction evidence.
Integrate
Carry earlier results through a longer multi-step capstone case.

From task to published result

The complete evaluation contract is fixed before execution and remains connected to the resulting public evidence.

  1. Task contractCase inputs, question, required output, tools, and available points.
  2. Recorded runExact model, provider route, runtime, reasoning, and tool setup.
  3. ScoringPredefined exact checks, plus a disclosed judge only where required.
  4. Run receiptScore components, cost, usage, configuration, and judge evidence.
  5. Observed resultA percentage for this benchmark version and recorded configuration.

Scoring

Benchmark pages show only the scoring mode and components that were actually used.

ModeHow it worksCurrent release
DeterministicExact fields, calculations, and structured outputs are checked using predefined rules. Deterministic score components can be recomputed exactly from their authenticated inputs.Earned deterministic points ÷ available points132 runs · no AI judge
HybridExact checks are combined with rubric assessment where deterministic checks are insufficient. A model-judged component records one stored judge outcome.Deterministic + rubric ± published modifier32 runs · judge disclosed

Comparing results

A published score is evidence for a specific recorded run, not a statistically certain universal ranking.

  • Configuration-specificModel, route, runtime, reasoning, and tools stay attached.
  • Missing is not zeroUnpublished evidence remains visibly unavailable.
  • No universal scoreBenchmarks are not merged into an opaque intelligence measure.

Evidence and limitations

The release states both what the recorded evidence supports and what it does not establish.

Supported

  • Versioned benchmark and recorded configuration
  • Observed scores and published score components
  • Recorded cost, usage, execution, and judge identity
  • Signed sanitized release artifacts

Not established

  • Standardized best-effort model performance
  • Statistical certainty from repeated runs
  • Model-judge repeatability qualification for the current public release
  • A guarantee that a fresh judge call reproduces a stored hybrid outcome
  • Freedom from conceptual contamination
  • Spreadsheet or workbook-artifact performance

Release integrity

Ed25519 authenticates the exact public manifest and can expose publication tampering. The signature authenticates the public projection; it does not make a judge decision repeatable or turn a provisional result into a verified result.

Signature
Ed25519
Inventory
164 completed run receipts
Status
Signature material published

Status glossary

These terms keep the same meaning across benchmarks, comparison views, and run receipts.

Provisional
Published calibration evidence that has not reached the future verified standard.
Verified
Reserved for a future result that satisfies the published verification standard.
Hybrid
Scoring that combines deterministic checks with rubric-based AI assessment.
Deterministic
Scoring produced entirely by predefined exact checks without an AI judge.
Configuration-specific
Evidence tied to the recorded model, provider route, runtime, reasoning, and tools.
Signed release
A release whose exact manifest bytes are authenticated with a published key.