Kaelum BenchBack to calibration results

Public Calibration

Accounting Capstone

Kaelum Bench · FMB-CAP-ACC-03

This calibration view covers the concrete 100-task suite only. Its provisional 1000-point hybrid score combines deterministic and rubric components, applied modifiers, and an independent AI judge. It exposes only sanitized score, execution, token, cost, and judge-identity aggregates. It contains no prompts, responses, expected solutions, rubrics, task IDs, provider generation IDs, keys, rationales, digests, or private file paths.

Provisional · Hybrid

Accounting Capstone · 100-task sub-benchmark

Calibration results

Sanitized lifecycle, coverage, score, token, and cost aggregates for FMB-CAP-ACC-03. Results are provisional and specific to the recorded model configuration.

Public Calibration
Sub-benchmark
Accounting Capstone · FMB-CAP-ACC-03
Models
1 calibration run
Coverage
100 tasks per model
Deterministic
800 deterministic points
AI rubric
200 AI-judged rubric points

Execution detail

Model runs

Provisional hybrid results combine deterministic scoring, an independent AI judge, and applied modifiers.

Provisional · Configuration-specific calibration · Single runs; no statistical ranking.

OpenRouter (moonshotai/int4) · max reasoning

Kimi K3 — open run evidence

moonshotai/kimi-k3

Judge openai/gpt-5.6-sol · OpenRouter (Azure EU)

StatusProvisional · Hybrid
Progress

100 / 100

Final score

95.3%

952.8 / 1000

769.0 / 800 deterministic

192.0 / 200 AI rubric

-8.2 modifier adjustment

Valid

100 valid answers

0 invalid JSON responses

AI judge

100 judge assessments

3 judge retries

Tokens

493,806 total tokens

Cost

$4.33 total cost